It was mid-2025. Every company seemed to be building its own chatbot, and the Model Context Protocol was still new enough that most people had not heard of it. We were putting a generative AI chatbot into an existing hospital operations platform, so that staff could navigate the system and ask about operational data in plain language instead of clicking through it.
The hard part was never the model. It was getting the model to read live data without either copying that data somewhere else or handing it more access than the person asking deserved.
Three approaches that did not hold up
RAG
The first instinct was retrieval-augmented generation: embed the existing data, put it in a vector database, let the model pull context as it needs it. It is a good pattern for documents that sit still. Ours did not. Operational data changes by the minute, so RAG meant duplicating large volumes of it and then owning a synchronization problem forever. The answer would always be a little bit stale, and “a little bit stale” is not a thing you want in a hospital.
The existing APIs
Next we looked at the APIs the application’s own frontend already called. This felt right, because those APIs already enforced permissions. Ask through them and a user cannot see anything they could not see in the UI. That property is worth a lot and we kept it.
The problem was coverage. Plenty of the data the chatbot needed was not exposed by any endpoint, and every gap meant writing a new one. The chatbot’s capabilities were capped by however much backend work we were willing to do that week.
Direct database access
So we considered going straight to the database, on the theory that everything is in there somewhere. It is, and that is the trap. A schema is not an interface. Without the application’s logic layer, rows lose the meaning the application gives them: which timestamp is the one that counts, which records are soft-deleted, what a status value actually implies. And the permission model we had just been handed for free was gone.
The thing that worked
What we shipped was a small server sitting between the chatbot and the backend APIs, passing the user’s own auth token through on every call. The model got a usable set of operations, and the access check stayed where it already worked: in the application, as that user. People could ask questions of live data and nobody had to reason about whether the AI had too much access, because it had exactly theirs.
It worked, and it was still annoying. Every new capability was a new endpoint or a change to an old one. Each iteration was more code to write, review, and maintain, and all of it was glue. The interesting work was not the glue.
What that turned into
The realization was that the glue is the same every time. Every team wiring an agent to a real system writes a version of that proxy, and then writes it again for the next system, and then copies the config onto a second laptop along with every secret in it.
That is the layer Infragate builds. Not a service that generates backends for you, which is where this started and is not where it landed, but tooling for the part everyone was hand-rolling:
- CAPA declares what an agent is allowed to do in one file per repo: skills, tools, MCP servers, sub-agents. One command installs it, and a teammate gets the same setup without a walkthrough.
- Lanyard is the auth-passthrough idea from that first proxy, as a product. One gateway URL, each person signed in to each server as themselves, one revocable key per agent.
- ShareCube is for what the agent produces, so a document ends up at a URL a colleague can comment on instead of scrolling away in a chat.
All of it speaks MCP, which was the bet in 2025 and looks like the right one now: an open protocol beats a proprietary integration surface, because the integrations outlive whichever client is fashionable.
The rest of this blog is that work as it ships.