AI Systems Architecture Blog — Agentic AI, RAG Pipelines, Multi-Agent Orchestration
On enterprise AI implementation — agentic AI deployment, RAG pipelines, multi-agent orchestration, and what it actually takes to ship production AI systems at enterprise scale.
Where Governance Can Actually Reach
Agent memory isn't one thing — it's four distinct surfaces, each carrying different privileges, each failing differently, each needing its control in a different place. Here's the map: every door, and where the lock goes.
The Prompt Is No Longer Where the Behavior Lives
On July 28, the protocol connecting AI agents to their tools stopped tracking who was talking — on purpose. It wasn't an isolated decision. Every layer of the stack is making the same move, and your audit surface hasn't followed.
Your Agents Remember Things Nobody Reviewed
Your system prompt went through security review. Your prompt library is versioned. The memory store now driving your agents' behavior went through neither — and your name is on the output either way.
The Two Failure Modes Nobody Formalizes
Every AI team can tell you what happens when their system works. Ask what happens when it fails silently, and the room goes quiet. Two disciplines close that gap — and almost no one has named them.
Your Corpus Is Lying to Your Model
When an AI system gives a confidently wrong answer, everyone audits the model. Almost no one audits what you fed it. The failure is usually in the corpus — and corpus quality is measurable.
Before You Build
The most expensive AI mistake happens before a single model is deployed — in the gap between approving the budget and asking whether the organization can hold what it's about to build.
The Comprehension Gap
The most expensive gap in enterprise technology isn't compute. It's the distance between 'we deployed AI' and 'we understand what we deployed' — and it only shows up when something fails.
Five States, Not Two — Why Most AI Tools Treat Their Users Like an Afterthought
Most builders design for success and bolt on an error page. I designed five explicit states. Here's why the failures are the product.
We Published Our AI Rubric Calibration Data. Everyone Told Us Not To.
The AI judge scored 60% of responses at Level 4. Manual review put the real number at 15%. We published the gap.
Our AI Judge Went Down. Nobody Noticed.
The system didn't crash. It succeeded at writing to the wrong place — for weeks. Here's what I built after.