AI agent development

An agent with access to your systems is useful exactly to the degree that what it may do is bounded. We build agents with explicit permissions, a traceable chain of decisions and a human in the loop wherever the stakes are high, and we say upfront when your process needs ordinary automation instead.

What’s included

  • Bounding the use case and assessing feasibility
  • Scoped tools and permissions, step by step
  • A human in the loop wherever a step can’t be undone
  • Decision-trace logging and auditability
  • Evaluation suites and regression tests
  • Cost and latency monitoring, with ceilings

A chatbot answers. An agent acts

The difference isn’t the model, it’s the consequences. When a chatbot is wrong, the user gets a bad answer and is annoyed. When an agent is wrong, the invoice is already paid, the email already sent, the record already deleted. The same technology, an entirely different risk profile — which is why the most important part of an agent is what it is not allowed to do.

So building one starts from permissions and boundaries rather than from the model. Which tools it can reach, what value or volume it may decide on by itself, when a human has to approve, and what happens when it gets something wrong — whether the step can be undone. With those answered, choosing a model is a fairly ordinary technical decision. Without them, what you have isn’t an agent but an automated risk.

What we do

AI agent development services

Six services. Most projects begin with the first and some stop there — often the right outcome, because not every process needs an agent.

  1. 01

    Strategy and feasibility

    We work through your processes and assess which suit an agent at all: whether the step is reversible, whether success is measurable, whether the data exists, and which risk class the use case falls into under the EU AI Act. Then we choose the agent type and the model and give you an architecture and a cost. Most good candidates are boring — repetitive, rule-shaped and well documented — and some processes need no agent at all, which we will also say.

  2. 02

    Single-agent development and customisation

    One agent, one clearly bounded task: customer support, knowledge management, billing, data processing. For conversational use cases the same agent can run as a chatbot or virtual assistant wired into your systems. Each is tuned on your domain’s data and connected to the systems the task requires. A narrow agent you can test beats a broad one you can’t — and reaches production in months rather than a year.

  3. 03

    Multi-agent system design

    Multi-step processes with several decisions, several data sources and several integrations. Agents divide responsibility, hand work on, supervise one another’s output, and can use different models and tools depending on the task. Useful when the steps are genuinely separate — not as a way to make a simple thing more complicated, which is this pattern’s most common misuse.

  4. 04

    Orchestration and integration

    An agent is useful exactly to the extent of what it can see and do. We connect it to your CRM, ERP, databases and APIs and tie the steps into multi-stage workflows where each tool has its own permissions and its own log. The orchestration layer decides which agent acts when, what happens on failure and where the approval point sits — and it is where agent systems most often break.

  5. 05

    Model selection and fine-tuning

    The model is chosen for the task, the latency and the price rather than for a leaderboard position — often the right answer is a smaller, cheaper model for a narrow job. Fine-tuning comes up only once better context has stopped helping: we prepare the training set, tune the model on your domain and test the result for bias as well as accuracy. Most projects never reach this step, which is good news, because a fine-tuned model has to be maintained afterwards too.

  6. 06

    Maintenance and evaluation

    An evaluation set built from real cases, and regression tests that run on every change. Models shift even when your code doesn’t — a provider updates a version and the behaviour moves, and without tests you find out in production. We also track cost, latency and which cases the agent gives up on most often, because that list is what shows you where to improve next.

The agent types we build

Most projects are one of these seven, or a combination of two. Which one depends on how many steps the process has and how expensive a mistake is — not on which sounds the most capable.

What our agents can do

Eight capabilities that separate an agent from scripted automation. None of them is an end in itself — they are means, and we use only the ones your task actually calls for.

Process

How an agent gets built

Five stages. The first three decide whether the project is viable at all — and not one of them involves building the agent.

01

Framing and feasibility

We analyse the workflows, the decision points and the automation opportunities, and establish where an agent delivers measurable value. We agree the use cases, user roles, compliance constraints and success metrics — and decide whether this calls for a conversational agent, an agentic workflow, multi-agent orchestration, or ordinary automation instead.

Stack

Tech stack

Agent frameworks are young and turn over fast. We build so that the framework and the model are replaceable parts — your business logic and permission model must not depend on which library is popular this year.

What you get

The guardrails an agent doesn’t ship without

An agent doing things in your name needs the same constraints as an employee — only written down and enforceable. These four are agreed before the first line of code.

Permissions and scope

The agent gets exactly the permissions the task requires — and no more, even where more would be simpler.

  • Each tool separately scoped rather than one master key
  • Ceilings on amounts, volumes and number of calls
  • Write access only where the task requires it
  • Access to personal data decided separately and logged

A human in the loop

The question isn’t whether a person approves, it’s when. That is agreed beforehand rather than after the first incident.

  • Approval points wherever a step cannot be undone
  • On low confidence it stops rather than guesses
  • Fallback logic when a tool or model doesn’t respond
  • A person can stop the agent at any moment

Traceability

For every action it has to be possible to answer “why did it do that” afterwards — a month later included.

  • The decision trace logged: input, tools used, output
  • Every action tied to a specific trigger
  • Changes to data traceable and reversible
  • Logs retained according to an agreed policy

Quality and cost

The two things that most often surprise people about agents: how well it actually works, and what it costs per month.

  • An evaluation set of real cases rather than a demo
  • Regression tests after every change and model swap
  • Cost tracked per call and per completed task
  • Ceilings set so the invoice can’t creep up quietly
Working together

How an agent project runs

Five stages. First, bounding it: which process, what the measure of success is, and whether the step can be undone. Second, data and context — an agent is only as good as what it can see, and this is usually where it emerges that some of the necessary knowledge isn’t written down anywhere. Third, architecture and model selection. Fourth, the guardrails: permissions, approval points, fallback logic and an evaluation set. And only fifth, integration and deployment.

An agent never goes into production deciding things straight away. First it runs in shadow mode: it proposes, a person decides, and we compare the two. Only once the agreement rate is high enough and the failure modes are known do we give it authority to act — and even then within narrow bounds at first, on small amounts or low-risk cases.

When choosing a partner, ask what happens when the agent is wrong. If the answer doesn’t contain the words “log”, “approval” and “rollback”, no guardrails have been designed. And ask how it will be measured — without an evaluation set, every claim about an agent’s quality is an opinion.

Why choose Techbaltics

  • We’ll say when you don’t need one

    Most processes people want an agent for are really rule-based automation. That is cheaper, faster and predictable — and we’ll recommend it.

  • Shadow mode before authority

    The agent proposes and a person decides until the agreement rate has been measured. Only then does it get to act, and at first within narrow bounds.

  • Permissions before the model

    Every tool gets its own permissions and its own log. An agent must not reach anything its task doesn’t require, even where that would be convenient.

  • An evaluation set, not a demo

    Quality is measured against a set of real cases that runs after every change and every model swap. A demo shows the best case; a test shows the average.

  • The framework is replaceable

    The business logic and the permission model don’t live inside an agent framework. This field moves fast and today’s library may not exist in a year.

  • The EU AI Act and your data

    We work out which risk class your use case falls into under the EU AI Act, and what that requires in documentation and human oversight. Hosting can stay entirely inside the EU.

Frequently asked questions

How much data do we need to train a model?

Less than people usually assume — for many tasks a few thousand well-labelled examples is enough. Data quality and labelling consistency affect the result more than volume does.

Do we need a large language model?

Often not. For classification, forecasting and anomaly detection a smaller, purpose-built model is cheaper, faster and more accurate. We recommend a language model when the task genuinely involves free text.

How do you stop the model giving misleading answers?

Evaluation suites before release, confidence thresholds, routing uncertain cases to a person, and continuous monitoring in production. None of these is sufficient alone; only the combination works.

How does an agent differ from ordinary automation?

Ordinary automation follows a fixed sequence of steps. An agent decides the sequence itself, based on the situation. That’s useful when cases vary too much to encode as rules — and unnecessary when they don’t.

How do you stop an agent doing something harmful?

Tools are explicitly scoped, write permissions are assessed separately, and irreversible actions — payments, deletions, customer communication — are confirmed by a person. Every decision is logged so you can see afterwards why something happened.

Are agents reliable enough yet?

For narrowly scoped tasks, yes; for broadly defined ones, no. Our recommendation is to start with one clearly bounded workflow, measure the result, and widen only when the numbers support it.

What is an AI agent?

Software that uses a language model to carry out a task rather than to talk about it. It can break a task into steps, use tools such as search, APIs or a database, and decide the next step from the result of the last.

The practical difference from a chatbot is in the consequences. A chatbot gives you an answer and you decide. An agent decides itself and acts: sends the email, changes the record, initiates the payment. The same technology, an entirely different risk profile — which is why the most important part of an agent is what it is not allowed to do.

Which processes suit an agent?

Boring ones. Repetitive, rule-shaped, well-documented processes where success is measurable and the step can be undone — processing documents, classifying requests, reconciling data, a first pass before a person looks.

Poor fits are where an error is expensive and irreversible, where the rules aren’t written down anywhere, or where the right answer depends on context the agent can’t see. And a good share of what people want an agent for is really ordinary rule-based automation — cheaper, faster and predictable. We say so upfront, including when it means a smaller project for us.

How do you validate an agent before production?

In two stages. First an evaluation set: we collect real cases with known correct answers and measure how often the agent gets them right. That set runs after every change and every model swap — models shift even when your code doesn’t, and without tests you find out in production.

Then shadow mode. The agent runs on real data and makes its proposal, but a person makes the decision — and we compare the two. That shows the real agreement rate and surfaces the failure modes before they have consequences. Only then does the agent get authority to act, and at first within narrow bounds: small amounts, low-risk cases.

What does the EU AI Act require of us?

It depends what the agent is used for. The Act sorts use cases into risk classes, and most internal business agents — document processing, internal support, data reconciliation — sit in the lighter part, where the main obligation is transparency: a person has to know they are dealing with AI.

Stricter requirements arise where an agent affects decisions about people — hiring, credit, access to services. There you need a documented risk assessment, human oversight and traceability. Our job is to establish that class at the start and build so the requirements are met, rather than discovering later that the system needs reworking. We don’t give legal advice, but we’ll tell you when it’s worth asking for.

Related services

Let’s talk about it.

Describe your situation in a couple of sentences. We reply within one working day.