AI App Development: What Models Do Well and How to Ship Them
Which problems language models actually solve well, how an AI feature is built, evaluated and budgeted, and how we take it from prototype to production.
You have a product, or an idea for one, where a language model could do real work: answer questions from your documentation, pull the important fields out of contracts, route incoming tickets, draft a first reply. Good AI app development starts from that job rather than from the model. It should end with a measured, affordable feature that the people relying on it actually trust.
From the outside, a feature like that looks unremarkable. It’s right often enough to be worth using, and it says so when it isn’t sure. Each task costs a predictable amount. When it gets something wrong, a person can notice the mistake and correct it. What follows is how we build AI products at Inferne, including the cases where we advise against using a model at all.
What language models do well, and where rules win
Models earn their place on messy, unstructured input, in jobs where a person or an automated check can catch the occasional imperfect answer. In practice, that covers work like this:
- Turning free text into structure. An email becomes a ticket with a category and priority; a contract becomes a table of parties, dates and obligations.
- Finding and summarizing information across a large body of documents, with references back to the sources.
- Sorting items into many fuzzy categories, where you’d never finish writing a rule for every case.
- Drafting text that a person reviews before it goes anywhere, such as replies, descriptions, summaries or translations.
- Understanding a request written in plain language and mapping it onto actions your software already supports.
Where the answer has to be exact and repeatable, a model is the wrong tool. Calculating a fee, checking eligibility against a policy, validating an account number and matching a SKU all belong in ordinary code. A rule runs faster and costs almost nothing. It also behaves the same way every time, and you can explain it to an auditor.
So if the rule can be written down, we write the rule. The model gets the part of the problem that resists rules, and the rest stays deterministic. Sometimes a regular expression, a search index or a decision table beats any model, and if that’s the case for you, we’ll say so plainly in discovery.
Product patterns we build
Most AI features turn out to be one of a handful of patterns. It pays to name the pattern early, since each one needs a different architecture, fails in different ways and has its own measure of quality.
- Assistants live inside your product, answer questions about the user’s own data and can trigger a few well-defined actions. We keep their scope narrow on purpose.
- Search and question answering is usually built as retrieval-augmented generation (RAG): the system finds the relevant passages in your content, then answers from them with citations. It’s common in help centers, policy libraries, legal research and internal knowledge bases.
- Extraction turns documents such as invoices, claims, medical forms and contracts into validated records. Every field carries a confidence signal, and uncertain records go to a review queue.
- Classification and routing sorts tickets, leads, transactions or user reports into categories and sends them to the right queue. When a person overrides a decision, that override feeds back into evaluation.
- Generation produces drafts of text, and more and more often images, audio and video. Media generation brings its own concerns around GPU cost, safety and provenance, which we cover in generative AI development.
Once a feature decides its own next steps and acts in your systems, it has become an agent, and permissions and control turn into the main questions. Our article on AI agent development covers when that’s warranted and when a fixed workflow is the better design.
The architecture of an AI feature
The model call is usually the smallest part of the system. In production, an AI feature is ordinary software with one probabilistic component, and most of the engineering goes into the code around that component. A typical request goes through seven stages:

- Validate and normalize the input, and establish what this user is allowed to see.
- Retrieve the relevant documents or records, filtered by the user’s permissions, and trim them down to what the task needs.
- Fill in the prompt template. Templates are versioned and live in the code repository, not in someone’s notes.
- Call the model through a thin internal interface, so the provider or model can change without anyone touching product code.
- Check the response against a schema and your business rules. If it fails, retry, fall back or hand the case to a person.
- Present the result with its sources, let the user edit before anything is sent, and show uncertainty instead of hiding it.
- Log the inputs, outputs, versions, latency and cost, along with whether the user accepted, edited or rejected the result.
For a feature added to an existing product, our article on LLM integration goes through the practical details: provider abstraction, retrieval, structured outputs, caching and fallbacks.
Evaluation: knowing whether it works
Until you can measure it, an AI feature is still a demo. So before anyone tunes a prompt, we build an evaluation set together with your team. It’s made of real inputs, anonymized where needed, and each one is paired with what a correct result looks like.

What counts as correct depends on the pattern. We score extraction field by field. For classification, we look at how often each category is right and how often it’s missed. Question answering is judged on whether each answer is supported by the sources it cites, while drafted text is checked against a written rubric.
The scoring itself mixes methods. Deterministic checks catch schema errors and exact-match fields. A second model can grade open-ended answers against the rubric, but we only rely on it once its grades agree with human judgments. People still review a sample, because some failures only a domain expert will spot.
Every change to a prompt, a model or the retrieval pipeline reruns the evaluation, which means regressions show up in review and not in front of customers. After launch, sampled production traces and user feedback keep adding hard cases to the set.
Cost, latency and privacy budgets
Quality isn’t the only thing shaping the design. Cost, latency and privacy matter as much, and we agree on a budget for each of them during discovery.
Cost
We estimate cost per completed task rather than per call, since a single task can involve retrieval, several model calls and retries. To bring it down, we use smaller models for the easy steps, keep prompts and context tight, cache repeated work and move non-urgent jobs into background batches.
Latency
An inline suggestion and an overnight document review tolerate very different delays. Interactive features stream their output so users see progress right away. Slow work goes to a queue, and the user is notified when it’s done.
Privacy
You need to decide what data may leave your infrastructure, and on what terms. For us that means reviewing the provider’s retention and training policies, using regional processing where it’s offered and redacting personal data the task doesn’t need. Access control is enforced during retrieval, so the model never sees a document the user couldn’t open. In MedTech, FinTech and LegalTech, sensitive data is the norm, and self-hosted open-weight models are an option there. The price is running GPUs and owning their operation.
How we run AI app development, from prototype to production
The stages are the ones any software project goes through. The work inside them is specific to AI.
Discovery comes first. Our business analysts map how the job is done today, what data exists, what a wrong answer costs and which metric will show that the feature is worth having. If rules would serve you better than a model, this is where we recommend them.
A prototype then tests the core behavior on your real data against the evaluation set, in weeks rather than months. Some ideas stop here, cheaply, and that’s a good outcome.
In design, our designers work out how the product presents suggestions, sources, confidence and corrections, so users know when to trust the output and how to fix it. Engineers then integrate the feature with validation, fallbacks and monitoring. QA testers work from the evaluation set and from adversarial inputs, including attempts to make the feature ignore its instructions.
The feature launches behind a flag, first for internal users and then for a growing share of customers, while we watch quality, cost and latency. After that, the work is iteration. We review failures on a regular cadence, add them to the evaluation set and fix them. New models get evaluated on your data before anyone switches.
You get regular demos and written reports throughout, and everything stays yours: the code in your repository, prompts under version control, the evaluation set and its results, the dashboards, and a runbook for the people who will operate the feature.
Frequently asked questions
Do we need to train our own model?
Rarely. A capable existing model with good retrieval and careful prompts is enough for most products. Fine-tuning helps when you have many examples of a consistent format or style and prompting has plateaued on your evaluation set. Training from scratch is almost never justified for a product feature.
Which model or provider should we use?
Whichever performs best on your evaluation set within your cost, latency and data-handling constraints. Public benchmarks are fine for a shortlist but shouldn’t make the decision. We keep the provider behind an internal interface, so the choice can be revisited later.
How do you stop the model from making things up?
We can reduce it, not eliminate it. Grounding answers in retrieved sources helps, and so do required citations, output validation, a narrow scope and letting the model say it doesn’t know. The product then has to make a wrong answer visible and cheap to correct.
Can our data stay private?
Yes, with the right setup. That means commercial terms under which the provider doesn’t train on your data, regional endpoints and redaction before anything is sent. Where data must not leave your infrastructure at all, we use self-hosted models.
Can you add AI to a product you did not build?
Yes. We review the existing code, data and infrastructure first, then add the feature as a well-bounded module instead of scattering model calls through the codebase.
If you have a job in mind for a model, the quickest test is to run it against your own data. Describe the task to us, and we’ll help you work out whether it calls for a model, a rule or a combination of the two.