Skip to content

ServiceAI systems

AI agent and LLM development, built for production

We design and build LLM applications and AI agents that work from your own data, call your systems and complete multi-step tasks, then keep them accurate and affordable once real users arrive. It’s for founders, scale-ups and enterprise teams who need an AI feature in production, and a Build Project starts from £12,000.

Software Stories is a UK software consultancy, and AI is our flagship service. We’ve shipped in-product AI assistants, retrieval over technical documents, LLM content pipelines and multi-model orchestration inside products we built end to end, so the AI and the software around it come from one partner.

What we build: LLM apps, RAG and AI agents

Most of the AI work we’re asked about falls into a few shapes, and many products need more than one. We design them to share one data layer, one set of logs and one way of measuring quality.

The model is rarely the hard part. The hard part is the data around it: documents in awkward formats, records spread across several systems and answers that need to point back to a source. We treat ingestion and retrieval as production software, with background jobs, retries and monitoring.

  • AI assistants and copilots inside your product, on OpenAI, Claude or Gemini
  • Retrieval-augmented generation (RAG) that grounds answers in your documents, manuals and records
  • AI agents that call your tools and APIs and pass work between steps, with human approval where a mistake would be costly
  • LLM pipelines that turn inputs into structured output, such as JSON, scores or ready-to-send content
  • Document pipelines that read invoices and files with OCR before anything reasons over them

From prototype to production: evaluation, guardrails and cost

A prototype answers the questions it was built for. Production gets the ones nobody planned, from users who paste in odd documents, ask the same thing five ways and sometimes try to misuse it. Evaluation, guardrails and cost control make the difference, so we build them in from the start.

We begin with a test set drawn from real examples of your users’ tasks, with expected answers agreed with your domain experts. It runs on every change to prompts, models or retrieval, so you can see whether a change made the system better or worse before your users do.

Cost matters as much as accuracy, because model spend grows with every user. We route simple steps to cheaper models, cache repeated work, limit how many steps an agent can take and log the cost of each request, so you can see what each customer costs to serve.

  • Checks on inputs and outputs for prompt injection, personal data and off-topic requests
  • Least-privilege tools, so an agent can only do what the user in front of it is allowed to do
  • Human confirmation before actions that send, delete, pay for or change records
  • Tracing of every prompt, retrieval result and tool call, so failures can be found and reproduced
  • A model layer that lets you switch providers, or use different models for different steps

Case studies: Kaivo, Cosmic IDE, Senderly and Biotech News Monitoring

Kaivo is an AI operating system for advertising agencies, which were juggling client ad accounts across Google, Meta, TikTok, LinkedIn, Microsoft, Reddit, Spotify and Shopify with no single view of performance. We designed and built a multi-service platform with an in-product AI assistant, “Kai”, and a signals engine that flags ROAS decline, budget pacing and creative fatigue, on Next.js, FastAPI microservices and LLM agents.

Cosmic IDE is a desktop IDE for embedded engineers. Its AI copilot runs on a local RAG package that grounds answers in datasheets and repository context, so the help it gives is based on the hardware and code the engineer is working with.

Senderly turns product and brand inputs into ready-to-send marketing emails for eCommerce brands. GPT-4 writes structured email copy as JSON, a second stage produces image layout blueprints, and webhook-driven async jobs render the final creatives, with job state and retries tracked at each step.

Biotech News Monitoring runs LLM sentiment and materiality scoring over real-time news for investors and researchers. A custom orchestration layer calls OpenAI, Anthropic and OpenRouter behind a short-lived cache, so responses stay fast and repeated questions don’t pay for the same answer twice.

When an AI feature is worth building

Not every problem needs a model. AI earns its place where the work involves reading, sorting, scoring, summarising or drafting at a volume people can’t keep up with, and where a wrong answer can be caught before it does harm.

In the intro call we’ll ask what a wrong answer would cost you, where the data lives and who needs to approve what the system does. If a rules engine or a plain search would do the job better, we’ll say so.

The stack we use

We choose tools to fit your existing systems and team. These are the ones we use most for AI work, all of them in products we’ve shipped.

The system runs in your own cloud account, under your access controls.

  • Models: OpenAI, Claude (Anthropic), Gemini on Vertex AI and OpenRouter
  • Retrieval: RAG over documents, with answers grounded in your own sources
  • Agents: tool-calling workflows with human approval steps
  • Backend: Python (FastAPI, Flask, Django) and Node.js
  • Data and jobs: PostgreSQL, MongoDB, SQLite, Redis and Celery
  • Documents: OCR with AWS Textract

What you get

  1. A written architecture and scope covering models, data, tools and how quality will be measured
  2. A production LLM application, RAG pipeline or AI agent, deployed in your cloud
  3. An evaluation test set that runs on every change
  4. Guardrails, tool permissions and human-approval steps matched to the risk of each action
  5. Tracing and cost tracking for model spend, latency and failures
  6. Regular demos throughout the build
  7. Documentation and a handover session, so your team can run and extend the system

How it works

  1. 1.Discovery and scope

    Before the build

    An NDA if you need one, then a working session on the problem, your data, your users and what a wrong answer would cost. You get a written architecture, scope and price before any work starts.

  2. 2.Test set and prototype

    Early in the build

    We build the evaluation set with your domain experts, then a working prototype on real data, so quality is measured from the start.

  3. 3.Build

    Agreed milestones

    Retrieval, tools, guardrails, integrations and the interface your users see, with regular demos. The evaluation set runs on every change.

  4. 4.Harden and launch

    Before go-live

    Cost and load testing, checks for prompt injection, data leakage and unsafe tool use, then a staged rollout to real users.

  5. 5.Hand over

    At launch

    Documentation, runbooks and a handover session, so your team can run, change and extend the system. If you’d like us to stay on, an ongoing retainer is priced on request.

Proof

“Steady progress, predictable delivery, and code that’s easy to review and integrate.”

Sean MetcalfFounder, Kaivo

Sean Metcalf

Founder, Kaivo · Client

He’s careful with contracts and edge cases, stays disciplined on scope, and communicates clearly when something needs a decision. What I value most is his consistency: steady progress, predictable delivery, and code that’s easy to review and integrate.

Price

From£12,000

For a scoped AI build from architecture to launch, with regular demos and a full handover. Every engagement starts with a free one-hour call, and you get a written scope with a fixed or capped price before any work starts.

FAQ

Our Build Projects start from £12,000 for a scoped AI feature, from architecture to launch. The price depends on the data sources, tools and integrations involved, and how much hardening the risk calls for. Model and cloud usage is billed to your own accounts, so you see those costs directly.

Yes, and it’s a common starting point. We measure where the prototype fails using test questions drawn from real use, then add the retrieval fixes, guardrails, logging and cost controls it needs. If parts are worth rebuilding, we’ll say which and why before the scope is agreed.

Most products should start with RAG, because answers stay grounded in your current documents and the data can change without retraining. Fine-tuning can help with a fixed format or a narrow, repetitive task, once testing shows prompting and retrieval have reached their limit. Many production systems never need it.

It depends on the task, your data residency needs and your budget, so we test candidate models against your own examples rather than public leaderboards. A model layer between your product and the provider lets you switch, or use different models for different steps, without a rewrite.

It can be, with the right setup. We check the provider’s data-use and retention terms for the API you use, keep personal data out of prompts unless the task needs it, enforce document permissions in retrieval and log access.

By limiting what it can do as well as what it’s told. Each tool gets the narrowest permissions it needs, risky actions need human confirmation and every step is traced. Before launch, we test the agent for prompt injection and unsafe tool use.

Tell us what you’re building.

Book a free one-hour call, or send a short brief and we’ll reply within two working days.