Bar Zik

Services

What I actually do.

Six areas of work. Most engagements combine two or three of them — the exact shape gets decided once I understand what is blocking you, not before.


01

Architecture review

An independent audit of a system that already exists but is not behaving the way you need it to.

I work through your retrieval pipeline, prompt construction, orchestration logic, evaluation setup, and cost profile — reading code, running the system against real inputs, and talking to the engineers who built it. The output is a written findings document: what is wrong, why it is wrong, what to change, and in what order.

  • Retrieval quality analysis against real user queries, not curated examples
  • Context construction and token budget audit
  • Failure-mode mapping, including silent failures
  • Ranked remediation plan with effort estimates
02

System design

Greenfield architecture for retrieval systems, agentic workflows, and tool orchestration.

For teams starting something new, or rebuilding after a prototype has hit its ceiling. I design the component boundaries, data flow, state management, and failure handling, then sequence it into a build plan your team can execute without me. Design decisions come with the reasoning attached, so your engineers can revise them intelligently as things change.

  • Component and data-flow architecture with explicit boundaries
  • Retrieval strategy: indexing, chunking, ranking, and freshness
  • Agent and tool-orchestration design, including error recovery
  • Sequenced build plan with milestone definitions
03

Evaluation engineering

The discipline that separates teams who improve their system from teams who just change it.

Most teams ship prompt and retrieval changes on vibes because building real evaluation feels like a detour. It is not — it is the thing that makes every subsequent change cheap and safe. I build task-specific eval sets from your actual traffic, design graders that correlate with what your users care about, and wire it into CI so regressions surface before release.

  • Eval sets built from production traffic and real edge cases
  • Grader design: rubric, model-graded, and programmatic
  • Regression harness integrated into your existing CI
  • Dashboards your PM can read without a translation layer
04

Cost & latency engineering

Making the unit economics work without giving up the quality you shipped for.

Inference cost and response latency are architectural properties, not billing problems. I work through model routing, prompt caching, context budgeting, batching, and speculative strategies — measuring quality at every step so savings do not quietly become regressions. Teams typically find that most of the spend is going somewhere they had not looked.

  • Per-request cost and latency attribution
  • Model routing and tiering strategy
  • Caching architecture, including prompt and semantic caching
  • Quality-guarded optimisation with before/after evaluation
05

Fractional AI architect

Ongoing design authority for teams that need senior AI judgment but not a full-time hire.

A recurring engagement — typically a fixed number of days per month — where I hold architectural direction alongside your team. Design reviews, technical decisions, unblocking hard problems, and reviewing PRs that touch the AI layer. Useful when you have capable engineers who are new to production LLM work and need someone to check their reasoning.

  • Recurring design reviews and architectural decision records
  • Hands-on unblocking of specific hard problems
  • Code review on the AI layer
  • Hiring support: scoping roles and technical interviewing
06

Team enablement

Transferring judgment, not vocabulary.

Intensive working sessions built around your codebase and your problems, rather than generic training material. We work through the decisions your team is actually facing, and I explain the reasoning behind each call as we make it. The measure of success is that the next set of architectural decisions gets made well without me.

  • Sessions built on your codebase, not slides
  • Decision frameworks your team can apply independently
  • Written architectural guidance that outlasts the engagement
  • Optional follow-up reviews as the team applies it

Engagement formats

Three shapes the work takes.

Pricing is fixed and quoted up front for diagnostic and design work, and monthly for retainers. No hourly billing — it rewards the wrong things.

1–2 weeks

Diagnostic

A focused review of an existing system, ending in a written findings document and a working session to walk through it. The most common way engagements start.

Best when you know something is wrong but not what.

3–6 weeks

Design engagement

Full architecture design or a substantial rebuild, delivered as documented decisions plus a sequenced build plan. Includes working sessions with your engineers throughout.

Best for greenfield builds or post-prototype rebuilds.

Ongoing

Retainer

A set number of days per month holding architectural direction alongside your team, with a standing review cadence and availability for urgent problems.

Best when the AI layer is core and permanent.

Next step

Not sure which of these you need?

That is a normal place to start. Describe the situation and I will tell you which shape of engagement fits — or whether you need one at all.