Why “Chat With Your Docs” Pilots Keep Stalling

A "chat with your docs" pilot that reaches month six without a production date is already costing more than it looks. Engineering hours burn against a demo nobody will approve. The business sponsor stops showing up to the standup. And the vendor keeps suggesting the next model version will improve retrieval.

The uncomfortable part is that the demo probably worked. A sales engineer pointed a hosted model at a folder of PDFs, asked three canned questions, and got clean answers. That's the trap. Everything that makes a retrieval-augmented generation system fail in production sits in the plumbing between those PDFs and the model, and that plumbing is where most pilots never do the real work.

Before the Pilot, the Demo Hides the Hard Parts

The demo answers questions about documents the vendor hand-picked. The pilot has to answer questions about everything else. Those are two different problems, and the distance between them is where budgets disappear.

RAG is a system, not a feature. Ingestion, chunking, embedding, indexing, retrieval, re-ranking, generation, evaluation, and re-embedding each sit in the critical path, and each one can silently break the next. The demo skips almost all of it.

Before you green-light a pilot, write what success looks like in a sentence a business owner would sign. "Users can ask questions and get answers" is not that sentence. "Claims adjusters can resolve coverage questions without opening a second system, with citations a supervisor can verify" is closer. If you can't name the user, the task, and the verifiable outcome, there's no finish line to run toward.

During Build, the Pipeline Is Where the Work Lives

The bulk of engineering effort in a serious RAG build goes to the retrieval layer, not prompt tuning, and the retrieval layer is unforgiving. One industry analysis of production programs reported that a large share of enterprise RAG projects fail to reach reliable production use, with the majority of those failures traced to retrieval rather than the model itself. Swapping to a newer frontier model does little when the system is handing it the wrong three paragraphs.

That shift from subscription to engineered system is why the cstm.ai launch announcement from DEV.co frames custom AI as a software engineering practice rather than a model-selection exercise. The decisions that matter show up in unglamorous places:

  • Chunking strategy. How a document gets split determines what the retriever can even find. Fixed-size windows are easy and often wrong. Semantic or structure-aware chunking is harder and usually better.
  • Metadata and permissions. Every chunk needs to carry the access controls of its source. Skip this and the pilot leaks confidential material to anyone with a login.
  • Re-ranking. A first-pass vector search returns plausible neighbors. A re-ranker decides which of them actually answer the question. Pilots that skip this step hallucinate with confidence.
  • Evaluation harness. You need a labeled set of real questions and acceptable answers before the model goes near a user. Without it, "it feels better" becomes the release criterion.
  • Re-embedding. Documents change. If the index doesn't change with them, answers go stale within weeks.

None of this is research. It's software engineering, done against a messy corpus, with a model that will cheerfully invent a policy number when retrieval returns garbage.

At Launch, Governance Decides Whether It Survives

A pilot that works in a sandbox meets its real adversary at launch: the people whose job is to say no. Security, legal, compliance, and the data owners who were not in the kickoff. They ask about training data, retention, audit trails, model residency, and what the system does when it's wrong. If the answers are being drafted the week before go-live, the launch slips.

The NIST AI Risk Management Framework and its generative AI profile name confabulation, data leakage, and insecure outputs as first-class risks to be managed with concrete controls, rather than left to a model's good behavior. Build those controls in as design inputs from week one, and the launch review becomes a conversation. Leave them for paperwork, and it's a wall.

After Launch, the Pipeline Becomes the Product

The teams that get past pilot stop talking about "the chatbot" and start talking about the pipeline. Retrieval quality is monitored the way uptime is monitored. New document types trigger re-chunking and re-embedding jobs. Prompt changes go through code review.

A regression suite of real user questions runs on every deploy. Hallucinations get logged, categorized, and fed back as fixes to retrieval or guardrails, rather than shrugged off as the model being weird.

This is where the economics change. A hosted assistant with a monthly seat fee is a product someone else owns. A RAG system wired into your corpus, your identity provider, your workflow tools, and your evaluation harness is infrastructure you own. It's also the thing competitors can't copy by signing the same vendor contract.

The teams shipping useful internal AI didn't pick the best model. They stopped treating RAG as a feature and started treating it as the product.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *