Bar Zik

AI architecture consultancy

Most AI systems don’t fail on the model.
They fail on the architecture around it.

I help engineering teams take LLM systems from convincing prototype to something that holds up in production — with retrieval that works on real queries, evaluation you can trust, and costs that survive contact with scale.

Engagement start
2–3 weeks
Typical first deliverable
10 days
Model & vendor stance
Independent

Where I help

You probably recognise at least one of these.

These are the situations that bring teams here. They are all architecture problems wearing a model-shaped disguise.

  • A demo that impressed everyone and then stalled on the path to production
  • Retrieval that works on the happy path and falls apart on real user queries
  • An agent that works until it does not, with no way to tell which is happening
  • Inference costs scaling faster than the value the feature delivers
  • No reliable way to know whether last week’s prompt change helped or hurt
  • A team that can ship features but has no one to own the system design

Services

Six ways engagements usually start.

Most work is some combination of these. The shape gets decided in the first conversation, once I understand what is actually blocking you.

01

Architecture review

A structured audit of an existing AI system — retrieval, prompting, orchestration, evaluation, cost. You get a written findings document with ranked, specific changes, not a slide deck.

02

System design

Greenfield design for retrieval, agentic workflows, and tool orchestration. Component boundaries, data flow, failure modes, and the build sequence your team can execute against.

03

Evaluation engineering

The part most teams skip. Task-specific eval sets, graders, and regression harnesses so you can tell whether a change actually improved anything before it reaches users.

04

Cost & latency work

Model routing, caching strategy, context budgeting, and batching. Usually the difference between a demo that impresses and a product with defensible unit economics.

05

Fractional AI architect

Ongoing design authority for teams without a senior AI specialist in-house. Design reviews, technical direction, and hands-on unblocking on a recurring cadence.

06

Team enablement

Working sessions that transfer judgment, not vocabulary. Your engineers leave able to make the next set of architectural calls without me in the room.

Full service detail and engagement formats

How I think about it

Four positions that shape every engagement.

01

Evaluation before architecture

You cannot improve what you cannot measure. Every engagement starts by establishing how we will know the system got better.

02

The boring parts decide it

Chunking, retrieval quality, and context construction determine outcomes far more often than model choice. That is where the work goes.

03

Build for the failure case

Production AI systems are defined by how they behave when the model is wrong. Fallbacks and guardrails are architecture, not polish.

04

Leave the team stronger

A consultancy that makes itself indispensable has failed. The deliverable is capability your team keeps.

Next step

Have a system that needs a second set of eyes?

Tell me what you are building and where it is stuck. If I am not the right fit, I will say so and point you somewhere better.