AI Architecture Consulting
AI architecture consulting for LLM and agentic systems
Senior architects design and review the architecture behind AI products - model strategy, retrieval, orchestration, evaluation and cost - and stay to help build it.
For product and engineering teams building LLM, RAG or agentic systems.
- Architecture-first
- Vendor-neutral
- Practitioners who build
Powered by Comtau Inc.
Where AI prototypes break
Demos are easy. Production is where the architecture shows
Most AI products don’t fail because the model is wrong. They fail because nobody designed the system around it.
Model and vendor lock-in
Architecture decisions get made implicitly, by whichever API was fastest to integrate first.
Swapping providers later means rewriting half the system.
Fixed at04 Model gateway
Retrieval that degrades quietly
RAG pipelines are tuned once at launch and never revisited as data grows.
Answer quality drifts down and nobody notices until a customer does.
Fixed at03 Retrieval & RAG
No way to measure quality
There is no evaluation process, so nobody can say if a change made things better.
Every release is a guess dressed up as a deploy.
Fixed atEvaluation rail
Inference cost outrunning the business case
Cost was never designed in, only discovered on the first real invoice.
Usage growth starts to look like a threat instead of a win.
Fixed atObservability & cost rail
AI-generated code nobody fully owns
Code lands faster than anyone can review or understand it.
Technical debt accumulates faster than the team can see it.
Fixed atArchitecture decision records
Safety and governance bolted on late
Guardrails get added after an incident, not designed in from the start.
A single bad output becomes a company-wide fire drill.
Fixed atSecurity & governance rail
Reference AI system architecture
What we design and review, layer by layer
An AI system is a set of architecture decisions, not one model call. Each layer below is a decision Comtau makes explicit and documents.
Product & UX
Decision we make explicit
Where AI sits in the workflow, what the user sees when it is wrong, and where a human approves.
Artefact
Interaction and fallback design
Decision we make explicit
- Where AI sits in the workflow, what the user sees when it is wrong, and where a human approves.
- Artefact
- Interaction and fallback design
Orchestration & agents
Decision we make explicit
How agents, tools, memory and human-in-the-loop steps are composed.
Artefact
Orchestration design · Tool & agent boundaries
Decision we make explicit
- How agents, tools, memory and human-in-the-loop steps are composed.
- Artefact
- Orchestration design · Tool & agent boundaries
Retrieval & RAG
Decision we make explicit
How context is found, ranked and handed to the model for each request.
Artefact
Retrieval architecture
Decision we make explicit
- How context is found, ranked and handed to the model for each request.
- Artefact
- Retrieval architecture
Model gateway
Decision we make explicit
Which models and providers the system relies on, and how it fails over.
Artefact
Model selection criteria · Vendor & fallback strategy
Decision we make explicit
- Which models and providers the system relies on, and how it fails over.
- Artefact
- Model selection criteria · Vendor & fallback strategy
Data & pipelines
Decision we make explicit
How data is ingested, embedded, stored and kept current.
Artefact
Data pipeline design
Decision we make explicit
- How data is ingested, embedded, stored and kept current.
- Artefact
- Data pipeline design
Platform
Decision we make explicit
Where the system runs, how tenants are isolated, and how changes ship safely.
Artefact
Deployment and platform design
Decision we make explicit
- Where the system runs, how tenants are isolated, and how changes ship safely.
- Artefact
- Deployment and platform design
Across every layer
Evaluation & guardrails
How output quality and safety are measured and enforced before release.
Artefact: Evaluation plan · Guardrail design
Observability & cost
How the system is monitored in production and kept inside budget.
Artefact: Observability plan · Cost model
Security & governance
Access control, data handling and compliance across the AI stack.
Artefact: Security review · Governance framework
When teams bring us in
Common questions that turn into an engagement
“We’re evaluating three model vendors and can’t tell which one actually fits.”
A vendor and model strategy tied to your data, cost and latency constraints.
“Our RAG quality plateaued and we don’t know why.”
A retrieval architecture review, from ingestion to ranking.
“The agent prototype works in a demo. We need it to work in production.”
A production-grade design: orchestration, evaluation, guardrails, cost.
“We’re adding AI to an existing platform and don’t want to re-architect everything.”
An integration design that respects what already works.
Engagement types
Review, design, or hands-on build
The three connect: a review can lead to a design, and a design can carry straight into implementation with the same team.
01 · Assess
Architecture Review
You have an AI system or a plan for one, and need an independent view of it.
An independent review of an existing or planned AI system architecture, with findings and prioritised recommendations.
- Current-state assessment
- Risk list
02 · Design
Architecture Design
You know what the product must do and need the system designed before it is built.
A target reference architecture and the decisions behind it - model strategy, data, orchestration, evaluation and cost - recorded as ADRs your team can build from.
- Reference architecture
- Architecture decision records
03 · Execute
Implementation Guidance
The design is agreed and you want the people who made it involved in the build.
Hands-on support carrying the design into a working system, through Comtau’s own build process.
- Implementation roadmap
- Build handoff
Looking for a review of a non-AI system? See Software Architecture Review
Building on AI you can’t yet explain?
Tell us what the system needs to do. We’ll show you the architecture decisions that matter before you build the wrong thing twice.
Direct conversation with a senior CTO/architect. We’ll review your situation and determine the useful next step.
How we work
Assess. Design. Execute.
Every stage leaves an artefact your team keeps, whether or not Comtau builds the system.
Assess
Understand the current prototype or system, the data available, and what actually constrains it.
Design
Decide the target architecture layer by layer, and record why.
Execute
Carry the design into a working system, hands-on, or hand it to your team to build.
Architecture you can read, not just a diagram
Every design ends in documents your team keeps. The decision record is the core of it: one per significant decision, tied to a layer of the architecture.
Reference architecture
The layered target system, with components and boundaries.
Architecture decision records
One record per significant decision, with options and consequences.
Evaluation plan
How quality and safety are measured before each release.
Roadmap
The order in which the architecture is built or changed.
Risk list
What could break, where, and what reduces it.
Architecture Decision Record
Illustrative format · not client material
ADR-0XX · Route requests through a model gateway with provider fallback
Accepted- Layer
- 04 Model gateway
- Rails touched
- Evaluation · Cost
- Owner
- Architect of record
Context
Why the decision is needed now: the constraint, the requirement, what breaks if nothing changes.
Options
- ASingle provider, called directly
- BGateway with routing and fallback
- CSelf-hosted open-weight model
Decision
The chosen option and the reasoning, in terms of your data, cost and latency constraints.
Consequences
- gainWhat becomes easier
- costWhat becomes harder or must be maintained
Evaluation gate
The check that must pass before the change ships.
- Architecture-first: decisions are made and documented before code, not discovered after an incident.
- Practitioners who build: the people who design the architecture can also implement it.
- Vendor-neutral: recommendations follow your constraints, not a partnership with a model provider.
From operating experience
We build with AI agents ourselves, under explicit controls
Comtau builds its own software with AI coding agents inside its own system for planning, running and verifying the work. The controls that make that safe are the same ones a production AI system needs, so we design them from practice, not from a vendor diagram.
- Bounded scope
- Inside ComtauEvery agent task states what it may change and what is out of bounds.
- In your systemAgents and tools get explicit permissions and boundaries, not open-ended access.
- Stop and escalate
- Inside ComtauWork stops on a failed check, scope growth or a missing approval, and a person decides.
- In your systemEscalation paths to a person are designed in for the cases the system should not decide alone.
- Traceable runs
- Inside ComtauEach run records the prompt it was given, the model used, an event trail and its token and cost figures.
- In your systemEvery AI action is traceable: inputs, model version, tool calls and cost.
- Verification separate from generation
- Inside ComtauOutput is checked by tests and by independent read-only reviewers, never signed off by the agent that produced it.
- In your systemEvaluation and guardrails sit outside the component that generates the output.
- Human-owned irreversible actions
- Inside ComtauArchitecture changes, new dependencies and releases need explicit human approval.
- In your systemActions that cannot be undone sit behind an approval step, by design.
This describes how Comtau engineers its own software. It is not a product we sell; it is the operating experience behind the architecture work.
A short note on roles and cost
AI architect, AI engineer, or ML engineer?
- AI architect
- Decides the system’s structure: which models, how data flows, how quality is evaluated, and where the risk sits.
- AI engineer
- Builds the product inside that structure: orchestration, retrieval, integrations.
- ML engineer
- Builds, trains and tunes the models and the pipelines around them.
Comtau does both - the architecture and, where useful, the build.
Cost follows scope, not a rate card
What drives the scope
- How many layers need a decision
- How deep the review goes
- Whether the engagement ends at a design or carries into implementation
Related
Choose the right starting point
AI MVP → Production Readiness
Already have an AI prototype or MVP? Start with a readiness audit instead of a from-scratch design.
See AI Production ReadinessSoftware Architecture Consulting
Building or fixing a non-AI system? See our general architecture consulting.
See Software Architecture Consulting
FAQ
AI architecture consulting questions
What does AI architecture consulting include?
Model and vendor strategy, retrieval and data architecture, agent orchestration, evaluation and guardrails, observability, and security and cost. Each area is a decision we make explicit and document, not a checkbox.
Do you only review, or do you also design and build?
Both. A review can stand alone, or it can lead into a target architecture and, from there, hands-on implementation with the same team - assess, design, execute.
How is this different from AI Production Readiness?
This page is for designing or reviewing an architecture. If you already have an AI MVP and want it audited and hardened for production, see AI Production Readiness.
Are you tied to a particular model or vendor?
No. Recommendations follow your data, cost and latency constraints, not a partnership with a model provider. Vendor and model choices are treated as an architecture decision, not a default.
Who owns the architecture and documentation afterwards?
You do. Architecture decision records, the reference architecture and the roadmap are written to be used by your team, with or without Comtau continuing.
What does an engagement cost?
Cost follows scope: how many layers need a decision, how deep the review goes, and whether the engagement ends at a design or carries into implementation. There is no fixed rate card - each engagement is scoped individually.
How is this different from general software architecture consulting?
AI systems add decisions that traditional architecture work doesn’t cover: model behaviour, retrieval quality, evaluation and inference cost. For architecture work without those AI-specific concerns, see Software Architecture Consulting.
Design the AI architecture once, properly
Tell us what you’re building. We’ll help you work out whether you need a review, a design, or hands-on implementation - and what that would produce.
Direct conversation with a senior CTO/architect. We’ll review your situation and determine the useful next step.
What happens next
- 01You describe the system and where it’s stuck
- 02We identify the architecture decisions that matter most
- 03We define the useful next engagement
Talk to a senior CTO/architect
Book a call