LLM Integration: Adding Language Models to Your Existing Product
How we add language model features to an existing product: swappable models, retrieval over your data, structured outputs, evals in CI and production safeguards.
You already have a product, a backend and users, and you want to add a feature that reads or writes language. It might summarize a case file, answer questions from your help center, turn inbound emails into structured tickets, or draft a reply for a support agent to approve. LLM integration is the engineering that turns a feature like that into a dependable part of your system instead of a demo wired to an API key.
A finished integration has a clear contract with the rest of your code, and you can change the model without a rewrite. Every prompt change gets tested against real examples before it ships. You can see cost and latency per feature, failures degrade gracefully, and personal data only goes where you decided it should.
Where LLM integration fits in your system
Start by checking that the feature needs a model at all. Exact lookups, calculations and validations belong in ordinary code. Our overview of AI app development goes through which problems suit a model and which don’t.
When a model is the right tool, we treat it like any other external dependency, a payment provider for example. It’s called from your backend, never directly from a browser or mobile app, and the credentials stay on the server. Most integrations end up in one of three shapes:
- Streamed, interactive features for drafting, rewriting and question answering, where the user watches the output arrive.
- Background jobs that extract, classify and summarize documents through the queue your backend already uses, with results written back to your database.
- Event-driven enrichment, where a new record triggers tagging, routing or indexing for search.
The integration usually lives in a small module with a narrow interface, for instance a function that takes a support email and returns a typed ticket. Prompts and provider details stay inside that module, and the rest of your code never sees them. We build these modules in whatever your backend runs on; the systems around them are covered in our article on backend development.
Choosing a model and keeping it replaceable
We choose models with your own examples, not leaderboards. Candidates are compared on quality against your evaluation set, and on latency, cost per task, context size, data-handling terms, regional availability and reliability. Sometimes the answer is two models, a small, fast one for routine requests and a larger one for the hard cases, with routing between them.
Hosted APIs are the usual starting point, since they need no infrastructure. Running an open-weight model yourself makes sense when data can’t leave your environment, when volume is steady enough to keep GPUs busy, or when you need full control over the model. The catch is that serving, scaling and upgrades become your job.
The provider sits behind an internal interface defined around your tasks rather than a generic chat wrapper. Providers differ in how they handle structured output, tool calling and caching, and a lowest-common-denominator abstraction throws those differences away. Where a provider offers pinned model versions, we pin them. A model upgrade then becomes a tested change, like any other dependency upgrade.
Retrieval over your own data
Most useful features need knowledge the model doesn’t have, like your documentation, your policies or the user’s own records. Retrieval-augmented generation supplies it at request time, and answer quality depends more on the retrieval than on the model. The parts that need the most care:

- Documents need proper parsing, tables and headings included. We split them along their structure instead of at fixed lengths and store each chunk with metadata such as source, date, tenant and access rights.
- Search combines keyword and vector search, since each finds what the other misses, and then reranks the candidates. Vector support in the database you already run, such as pgvector for PostgreSQL, is often enough before a dedicated vector store is justified.
- Results are filtered by what the user is allowed to see before anything reaches the model. Asking the model to withhold restricted content isn’t access control.
- Re-indexing happens when source content changes, triggered by events rather than by occasional full rebuilds.
We evaluate retrieval separately, checking whether the right passages come back for each test question. For structured data such as orders, balances or appointments, we skip embeddings and let the model query your existing APIs.
Structured outputs and tool calling
If the result feeds your code, free text is the wrong format. We ask for output that matches a JSON schema, using the provider’s structured output mode where it exists, and we validate it anyway. The first check is against the schema. The second is against business rules, such as dates in range, totals that add up and categories that exist. When validation fails, the call is retried with the error attached, and if the retry fails too, the item goes to a fallback or a human review queue.

Tool calling lets the model request an action, like looking up an order or checking availability. Your code executes it with the user’s permissions and the same validation as any other request. Everything the model reads should be treated as untrusted, because text inside a document or email can try to redirect it, and model output alone must never authorize a sensitive operation. Once a feature starts choosing its own sequence of actions, it’s an agent, and our article on AI agent development covers the extra controls that requires.
Prompts, evals and regression tests
We treat prompts as code. They live in your repository as templates with typed variables, go through code review, and ship together with the model and parameters they were tested with. Every production call logs its prompt version, so you can trace any output back to the configuration that produced it.
The evaluation set is built from real, anonymized inputs, with expected results agreed with your domain experts. Deterministic checks score schema validity and exact fields. Open-ended answers are scored by a grading model against a rubric, but only once its judgments match those of human reviewers.
The suite runs in continuous integration whenever a prompt, model or retrieval setting changes, and a drop below the agreed thresholds blocks the release. Outputs vary between runs, so the important cases run several times and we look at the spread instead of a single result.
Running it in production
Caching
Exact-match caching answers repeated requests, such as the same document summarized twice, without a new call. Several providers can also cache a long shared prompt prefix, so it pays to put stable instructions and reference material first. We only use semantic caching, which reuses answers to similar questions, where a slightly wrong reuse would be harmless.
Rate limits and fallbacks
Providers enforce quotas, and now and then they slow down or fail. Every call gets a timeout, retries with exponential backoff and jitter, and a circuit breaker. Background work runs through queues with controlled concurrency, so a batch job can’t starve interactive users. When the primary model is unavailable, a fallback model that passed the same evaluation set takes over. If both fail, the product offers a clear path without AI rather than an error.
Personal data
A task gets only the fields it needs. Where the task allows it, identifiers are redacted or pseudonymized before the call. We pick provider settings and regions that match your obligations, and logs and traces get the same redaction and retention limits.
Observability
Each request leaves a trace with the prompt version, model, retrieved sources, tokens, latency, cost, validation result and user feedback. Per-feature dashboards show quality signals, spend and errors, and spikes trigger alerts.
How an engagement works
We start small, with one feature, one success metric and access to real examples. A business analyst works with your domain experts to define correct behavior and build the first evaluation set. An engineer prototypes against it, and we go through the results together before anyone commits to production work.
The build team is usually one or more backend engineers with QA, plus a designer if the feature has a user interface. You get regular demos and written reports. At handover, everything is in your repository: the module, the prompts, the evaluation suite wired into continuous integration, the dashboards and a runbook.
Frequently asked questions
Can you work with the stack we already have?
Yes. The integration sits behind a narrow interface in your backend, whatever language and framework it uses, and it follows your existing conventions for queues, configuration and deployment.
Do we need a vector database?
Not necessarily. Plenty of products start with vector search inside their existing database, or with good keyword search. A dedicated vector store earns its place at larger scale or when the search needs are specialized.
What happens when a provider retires a model?
Since models are pinned and wrapped behind an interface, the replacement is a tested change. We run the evaluation suite on candidates, adjust prompts where needed and roll out behind a flag.
Will the provider train on our data?
It depends on the provider, the plan and the settings. We go through the terms with you and configure retention and training options to match. If your requirements rule out third-party processing, we self-host a model.
How do we keep costs under control?
Give each feature a budget, route easy requests to smaller models, keep context lean, cache where it’s safe and batch background work. Per-feature cost dashboards show overspending as it happens, not at the end of the month.
The best starting point is one feature and a handful of real examples. Share a short description of the feature and the stack behind it, and we’ll suggest how to scope and measure a first version.