A guide for companies · Updated August 2026

AI Agents in Your Company

Almost every vendor is selling you an “AI employee” right now. Hardly any of them tells you what happens when it acts without supervision, which task is actually worth starting with — and why your first agent should not sit in sales.

This is the short version of what I work through with small and mid-sized companies, from around 20 people upwards: which task pays off first, who keeps control, what it costs and where the limits really are.

Book a free consultation

Straight talk

What an agent can do — and what it can’t

Recurring workflows with a clear outcome
Searching and pulling together from your own documents
Drafts that a human approves
Not: consequential decisions without oversight
Not: sending, promising or deleting on its own
An AI agent's working loop: read the task, use a tool, check the result, carry on. A human makes the decision.1 · Read the taskWhat needs doing?2 · Use a toolCRM, files, email3 · Check the resultCorrect? Complete?4 · Carry onor hand overThis is where the agent stopsA human makes the decision

An AI agent is not a digital employee. It is a program that runs a language model in a loop: read the task, use a tool, check the result, carry on. It never clocks off — but it has no judgement either.

A reality check against the general excitement: in 2025 the research institute METR had experienced developers work with and without AI tools in a controlled trial — on tasks in their own projects, so under the most favourable conditions. With AI they took 19 percent longer. Afterwards they estimated they had been 20 percent faster. Adoption is not the same as impact, and a feeling is not a measurement. That is why a pilot measures instead of believing.

Rule of thumb
An agent may prepare. A human decides.
The map

AI agents in your company: ten use cases, honestly ranked

The traffic light means: 🟢 workable today · 🟡 only with supervision and clear limits · 🔴 I would advise against it today.

FieldWhat the agent takes onMaturityThe catch
Knowledge & internal search“Where is that written? Have we done this before?”🟢Getting your files in order is the real work
Software developmentWriting code, testing, preparing reviews, documentation🟢Without review discipline you get code nobody understands
Marketing & contentDrafts, variants, research, structure🟢Image rights, and consent rules for outbound email
Customer serviceStandard enquiries, pre-qualification, draft replies🟡Must be recognisable as AI; always keep a path to a human
IT operations & helpdeskTriaging tickets, standard cases, reading logs🟡An agent with admin rights is a security incident with a calendar entry
DesignAsset variants, image prep, maintaining the design system🟡Purely AI-generated graphics may not be protectable — check before you rely on them
SalesQualifying enquiries, preparing quotes, keeping the CRM clean🟡No automated cold email — consent rules apply in most markets, B2B included
Accounting & financeReading incoming invoices, preparing records🟡Payments and approvals stay with a human
Purchasing, planning, productionStandard orders, stock levels, delivery dates🟡Needs clean master data — but a classic place to start
HR & recruitingStructuring applications, drafting follow-up questions🔴Never automate the selection itself; involve employee representation where you have it

The first candidate is almost always the knowledge agent. No customer contact, no personnel decision, no outward effect. If it talks nonsense, a colleague notices and nobody loses money. Your team learns how to handle a tool like this on the lowest-risk case.

The front-runners see it the same way: in a University of St. Gallen study of agentic AI at Otto Group, Deutsche Telekom, Siemens and Bosch, stage one is to start in areas with high volume and low ambiguity — support requests, standard orders, basic planning. Under close human supervision.

In detail

The three agents people ask about most

The knowledge agent. It answers questions from your quotes, minutes, manuals and past projects. Its architecture is also the reason it is the easiest one to govern: because the knowledge sits in reference documents rather than inside the model, you can update, delete and disclose it. Access, correction and deletion requests stay practically answerable — so the lowest-risk agent is also the cleanest one to run.

The sales agent. It qualifies enquiries, prepares quotes, keeps the CRM clean and summarises the history before a customer call. What it does not do: send cold email. In most markets that needs prior consent — business-to-business included. An agent that blasts out cold outreach does not scale sales, it scales your exposure. So it works behind the first contact, not in front of it.

The HR agent. It structures applications, spots missing documents, drafts follow-up questions and answers internal questions about leave and processes. The selection stays with a human. A ranking that gets adopted unchecked is, in substance, an automated decision — and rejecting a candidate is exactly the kind of decision people are entitled to have a human make. Recruiting is also the area regulators everywhere are watching most closely, so treat it as the last place to hand over control, not the first.

Approach

Rolling out AI agents: five phases, so it still runs after eight weeks

1 — Find the task, don’t shop for a tool
2 — Draw the limits before anything is built
3 — Its own identity and an audit trail
4 — Pilot: one team, four weeks, measurable
5 — Scale up, or switch off honestly
The five phases of rollout, from finding the task to scaling up or switching off. Phase four is the four-week pilot.1Find the task2Set the limits3Identity and logging4Pilot: 4 weeks, measurablemeasured here, not believed5Scale up or switch off

Phase 1. Write down three candidates: frequent, clear outcome, mistakes survivable. Anyone who starts with “which tool should we use?” builds a solution to a problem they never described.

Phase 2. Four questions, in writing: What may the agent read? What may it write or send? Where does it stop and a human decides? What happens when it is unsure? Answer those four and you have done most of the governance work before anyone writes a line of code. This is not paperwork, this is the architecture.

Phase 3. The agent gets its own account, never an employee’s. Every action stays traceable. And you need a clear answer to who switches it off when it misbehaves. Where several agents work together, the St. Gallen research recommends a guardian agent that monitors the rules, plus defined escalation paths.

Phase 4. Decide up front how you will measure success. Bring the people whose work changes into it before the pilot, not after — including any employee representation you have. And train the people who will use it; unskilled users are the most common reason a good agent gets abandoned.

Phase 5. Do the maths with all the costs: model usage, operation, maintenance, onboarding. An agent that does not pay for itself gets switched off. That is not failure, that is the point of a pilot. It only gets expensive if you still do not know after twelve months.

Orchestrating

From the second agent on, you are leading a team

Leading beats programming
One orchestrator, eight below it at most
Whoever creates something doesn’t review it
Oversight stays with a human
More agents means more management work

With the first agent the question is technical. From the second on it is a leadership question — and that is perhaps the best news in this guide for you.

In 2025 researchers at the Harvard Kennedy School ran a pre-registered experiment into who is good at leading AI agents to a result. The finding: skill at leading agents predicts skill at leading human groups — and the other way round. People who are good at both ask more questions and allow more back-and-forth, instead of issuing instructions. And, notably: there were no differences by gender, age, background or level of education.

For you that means the competence that counts here does not sit in your IT department. It sits with the people who already work well with people — the team lead on the shop floor, the sales manager who can give feedback. Anyone who can explain what they want, what a good result looks like and how they would recognise it can work with agents. Anyone who cannot will not learn it from the best tool either.

The limit at eight. Management theory has long held that more than about eight direct reports costs you quality. Reports from practice suggest the same applies to agents: give an orchestrating agent too many subordinates and the work becomes unreliable — it starts picking the wrong ones for the job. That is experience, not a study — but it matches what I see myself.

Separate creating from reviewing. An agent that built something is a poor reviewer of its own work — exactly like people. So you need separate roles: someone researches, someone creates, someone objects, someone coordinates. Oversight of the whole is the one role that does not go to a machine.

And the uncomfortable truth: if you rebuild your existing organisation one-to-one in agents, you get its flaws too. A marketing agent team that runs beautifully internally still knows nothing about sales — the same silo you have between human departments, only faster. And every additional agent costs management time. Anyone who believes agents save time without creating new management work has miscalculated.

The benchmark
An agent nobody leads is not an employee. It is an unsupervised program with access to your data.
Tools

Which tool — and why that question comes last

There is no best platform, only a fitting one. Three routes that are realistic for a mid-sized company:

The most common mistake happens earlier. A company rolls out a tool, people don’t take to it, so the tool gets swapped. Then swapped again. On the third attempt someone asks whether the next tool might finally be the right one — and in the meantime nobody has written down which specific workflow it was supposed to improve. That order of operations is the reason for most failed rollouts I see.

That is why the tool question comes sixth in this guide. Start with the workflow that annoys your people, not with the platform everyone is writing about.

n8n / DifyMicrosoft Copilot StudioBuzz (Block)
A fit when …you need to connect many systems (CRM, mail, ERP)you already run entirely on Microsoft 365people and several agents should work in the same space long-term
Self-hostingyesnoyes
Data controlhigh when self-hostedin the Microsoft cloudhigh, your own server
Getting startedmoderate, visualeasy for M365 teamstechnical, Docker + command line
Maturityestablishedestablishedyoung: since March 2026, “very rough” per the vendor
Limitationyou build the governance yourselfvendor lock-ina developer tool; CRM/HR would need integrating

Why Buzz is worth a paragraph. Block — the company behind Square and Cash App — started an open-source project in March 2026 that solves one thing better than any commercial vendor: the identity question. Every agent there has its own cryptographic key. It does not act under an employee’s account, it acts as itself, authorised by a human. Every action lands in a tamper-evident log. If an agent loses its key, you revoke the agent — not the person.

That is the benchmark to measure any tool against, even if you end up somewhere else: its own identity per agent, clear permissions, a complete log.

I work in it myself — self-hosted on my own server, with Claude Code as the agent. Which is also why I can tell you where it hurts: Buzz is a tool for developer teams, a few months old, with close to a thousand open issues on GitHub. The approval gates for workflows are, per the vendor, still “in progress”, and sign-in runs on cryptographic keys rather than your company login — for an IT department that has to on- and off-board staff cleanly, that is a real catch today. Block itself writes in the documentation that you should not build your compliance planning on it yet.

For me as a one-person business running my own agents, that fits. For the HR department of a machine builder I would not pick it as the foundation today — and that distinction is the whole point: the most exciting tool does not win, the one whose weaknesses you know and can live with does.

If you run SAP, look there first: SAP now ships more than 40 ready-made agents inside its business software, plus a kit for building your own. The topic is already in the building.

Governance

What to settle before you start

A defined basis for each purpose you process data for
A processor agreement with your model provider
An impact assessment where personal data is involved
An updated privacy notice
A written rule about private AI accounts
Logging: who triggered what
A decision on where data lives (region or your own hosting)
A kill switch that has actually been tested
Three properties that make an agent governable anywhere: attributable, revocable, logged.01AttributableWho authorised this action?02RevocableStep in or switch it off at any time03LoggedDelete and disclose on request

The most common misconception is: “there are no rules for AI agents yet.” That is wrong — and it is the expensive kind of wrong, because the rules that do apply were mostly written for processing personal data, which is what an agent does all day.

None of the eight points above is paperwork for its own sake. Each one is a question you will be asked the first time something goes wrong. Who authorised this? What did it have access to? Where is the record? Who switched it off? If you cannot answer those in the calm, you will not answer them in the incident either.

What the rules actually demand, almost everywhere, reduces to three properties. An action must be attributable — traceable to one agent acting under a named human’s authority, which is why the agent gets its own account and never an employee’s. It must be revocable — someone can step in or shut it down, and that path has been tested rather than assumed. And it must be logged — a record complete enough that you can delete, correct or disclose what the agent touched.

Build for those three and the legal work gets small. Skip them and it gets expensive to retrofit, because you are then reconstructing history that was never recorded. That is the whole reason the knowledge agent is the recommended first project: its data sits in reference documents rather than inside a model, so updating, deleting and disclosing stay possible.

The specifics — which basis you rely on, which assessment you owe, what your notice has to say — depend on where you operate and are still moving. I am not your lawyer and this page is not legal advice: have your own obligations checked in your jurisdiction. What I can tell you is the engineering, and the engineering is the part that decides whether the legal answer is cheap or painful.

Security

Security: the rule that explains almost every major incident

1 — The agent can reach internal data
2 — It reads content from outside
3 — It can send to the outside
All three together: exploitable
Close one of them: the attack goes nowhere

There is one attack that hits agents specifically and that hardly any guide mentions: indirect prompt injection. The instruction is not in your input, it is in the material the agent reads — an email, a PDF, a calendar entry, a ticket. The agent cannot tell the difference between your instruction and text it was merely supposed to process.

Security research has distilled a pattern from this that works as a test rule: an agent is vulnerable when it simultaneously has access to private data, processes third-party content, and can communicate outward. Remove any one of the three and the attack goes nowhere.

That this is not theory: in the documented “EchoLeak” case, a single email was enough to make Microsoft 365 Copilot read internal files and send content to an external server — without the recipient clicking anything. And in February 2026 an agent deleted its user’s emails after ignoring stop commands.

How everyday the attack surface is shows in a figure from Check Point’s AI Security Report 2026: roughly one in seventeen prompts entered into AI tools carries a serious risk of leaking data unintentionally. The share of especially risky prompts doubled within a year, to four percent.

Note what was affected: finished products from major vendors — Microsoft 365 Copilot, Slack AI, developer tools. So this is not a problem of home-built solutions, it is a property of the design. Buy it in and you buy this in too.

That is why the knowledge agent comes first in this guide: it reads internally and sends nothing out — factor three is missing. The sales agent, which reads mail and sends it, would have all three. This is not caution on principle, it is the difference between an agent that is a tool and one that is an open window.

Your test question
Can this agent read internal data, process third-party text and send outward — all at the same time? If yes: close one of the three routes.
Adoption

Is anyone else actually doing this?

University of St. Gallen study: Otto Group, Telekom, Siemens, Bosch
Deutsche Telekom runs customer communication on an agent platform
SAP ships more than 40 ready-made agents inside its software
Almost every public case study is a large enterprise
And many of the numbers come from vendors with something to sell

A bit of both. What is documented: a University of St. Gallen study analysed agentic AI at Otto Group, Deutsche Telekom, Siemens and Bosch and derived a three-stage model from it. Telekom runs customer communication on an agent platform. And SAP now ships more than 40 ready-made agents, plus a kit for building your own.

And now the caveat nobody else will give you: almost all public case studies are large corporations, and many of the numbers come from vendors who want to sell you something. For companies your size there are barely any solid case studies — not because nothing is happening there, but because nobody writes about it.

So what an agent will do for you is not something a study can tell you. Only a measurement at your place can. That is exactly what the pilot is for.

The offer

The 4-week pilot

One agent, one task, four weeks — and then a decision you can rely on.

What happens
Week 1Review three candidate processes, pick one. Limits and success measure in writing. Bring in the people whose work changes, check whether an impact assessment is needed, and settle the processor-agreement situation.
Week 2I build the agent with access to exactly the data it needs — no more. Its own account, a complete log.
Week 3One team works with it. I sit alongside, correct, and document what goes wrong.
Week 4Evaluation with numbers: what it took over, what it didn’t, what it costs. Recommendation: scale up, rebuild or switch off.

What you have afterwards: a running agent, or a well-founded no. Either way with documentation you can hand to whoever asks. And a team that knows how this works.

What I don’t do: sell you an “AI strategy” as a slide deck.

What is happening to running costs right now. The list prices of the major vendors are more stable than the headlines suggest — a developer tool like GitHub Copilot still costs $10 for individuals and $19 to $39 per person on a team. Something else has shifted, and it matters more for your planning: the billing logic. Since 1 June 2026 Copilot bills by consumption. The monthly price includes an allowance; anything beyond it costs extra. A predictable licence becomes a variable invoice.

At the same time new tiers have appeared at the top — $100 a month for Copilot, €229 for the highest ChatGPT tier. The reason is not arbitrary: on a Bridgewater estimate the four largest providers invested around $410 billion in 2025, with roughly $650 billion expected for 2026. That has to pay for itself eventually. The trade press calls it “The Free Lunch Is Over”.

For you that means two things: budget for consumption, not for a licence per head — and make sure your agents are not using the most expensive model for routine work. Reading invoices needs no flagship model. Which is exactly why “which model for which task” belongs in the pilot, not in the renewal.

Cost. A pilot for an internal process starts at €4,900 net. Scaling it into a permanently running agent connected to your systems starts at €6,000. Ongoing care — watching, adjusting, switching models — is €249 a month, plus model usage by consumption and, if you self-host, roughly €15–30 of infrastructure. An agent nobody watches drifts.

Book a free consultation

Why me

Why work with me and not an “AI agency”?

Over 12 years of web design and development, plus seven apps of my own
My agent workspace runs self-hosted, with Claude Code
My open-source project Bricks MCP is used by developers worldwide
I work with agents myself every day — not as a demo
You call me, not a hotline

AI agencies have been springing up since 2024. Many of them were selling social-media packages two years ago. How to tell the difference: ask to see something real.

I build the integration myself instead of reselling an off-the-shelf widget with a markup — the same way I work on software and web applications. Website and AI come from one place; you don’t need a second supplier who has to get up to speed first. How I build websites is on the services page, and you can see delivered projects under work.

If your question is more “a bot answers customer questions on our website” than “an agent works inside our internal processes”, have a look at AI chatbot & automation — that is where the matching offer is.

Tips

Seven things that make the difference

Start with the task you understand best yourself
One agent, one task — no do-it-all
Demand evidence, not assertions
Its own account per agent, never an employee’s
Don’t make the most expensive model the default
Version your instructions like code
Thirty minutes of review, once a month

Start where you can judge the result. The first agent should take on a task where you can see immediately whether the output is right. Automate the area you understand least and you will be the last to notice mistakes.

One agent, one task. The do-it-all sounds tempting and becomes unreliable. Three narrow agents with clear limits are easier to check, correct and — if it comes to it — switch off than one that does everything a little.

Demand evidence. An agent that answers “it’s in quote 2024-118, section 3” can be checked. One that just states a result is asking for trust it has not earned. That is not a convenience feature, it is your only practical form of control.

Don’t make the most expensive model the default. Reading invoices and sorting appointments need no flagship model. Use the strongest one for every routine task and you pay ten times over for the same work — and only notice on the monthly bill.

Version your instructions like code. The instruction you give an agent is the actual value of your work. Without a history, nobody knows three months later why it behaves the way it does — and one small change quietly breaks something else.

Thirty minutes once a month. Agents drift, not because they get worse, but because your data, prices and processes change. A quick look at ten cases catches it early. Without that appointment in the calendar, a customer will find it for you.

The most common mistake
Starting too big. A small agent that runs is worth more than a big project nobody signs off on.
Conclusion

What I take from my own practice

I have been designing and building websites for over twelve years, and I work with agents every day — not as a demo for clients, but because that is how my own tooling came about. Bricks MCP, my open-source project, lets agents build websites; developers worldwide use it. And the workspace where my agents and I work together is Buzz — the same tool in the comparison table above, self-hosted on my own server, with Claude Code as the agent. So I am not writing here about something I read.

What convinced me: the gain is not where the advertising promises it. It is not the spectacular automation that makes the difference, it is the disappearance of small frictions — looking things up, sorting notes, preparing drafts. Added together, that is more time than any single large automation I have seen.

What I underestimated: how much work sits in the describing. The technology is the easy part. The hard part is writing a process down precisely enough for a program to execute it — and realising in the process that half of it only ever existed in one person’s head. That is exactly why many projects fail: not because of the model, but because nobody knew the process.

Where I stay sceptical: agents are not a tool for cutting headcount. Introduce them that way and you get a team that sabotages them — rightly so. Introduce them so that people spend less time searching and retyping, and you get allies. That is not a moral point, it is the practical difference between an agent still running after six months and one that gets quietly switched off.

And quite soberly: you lose nothing by not being among the first. The tools will be better and cheaper in twelve months. What you can lose is trust — with your people and your customers — if you start without limits. Hence the pilot: small, measured, with a real decision at the end.

Frequently asked questions

What is the difference between a chatbot and an AI agent?

A chatbot answers. An agent acts: it uses tools, fetches data, writes into systems and works through several steps until a task is done. In short: the chatbot talks to your customers, the agent works inside your processes.

Which rules apply to us?

That depends on where you operate, what data the agent touches and what it is allowed to do — and the rules are still moving. This page deliberately gives you no legal advice. What it gives you is the engineering: build the agent so every action is attributable, revocable and logged, and you can bring it into line with almost any rule set. Then have your own obligations checked locally — including any employee-representation duties, which exist in more countries than people expect.

Do we have to tell people they are talking to an AI?

Assume yes. Several jurisdictions now require it, and where none does, people expect it anyway. Disclosure is the cheapest trust you will ever buy — and a support agent that pretends to be human is one screenshot away from a problem. Always keep a visible route to a human.

Do we have to label every text written with AI?

No — and labelling everything devalues the label. The cases that matter are synthetic media that could be mistaken for a real person or event, and content presented as independent information. For an internal draft that a human edits and signs off, nobody needs a badge.

Can an AI agent take over our cold outreach by email?

No. Unsolicited commercial email needs prior consent in most markets, business-to-business included — and an agent does not make that go away, it just makes you faster at breaking it. In sales, an agent should work behind the first contact: preparing, summarising, keeping records clean.

Can an AI agent screen out job applications?

Pre-sorting and preparing, yes. Deciding, no. Rejecting a candidate is exactly the kind of consequential decision people are entitled to have a human make — and a ranking adopted without review is, in substance, an automated decision. Recruiting is also the area regulators watch most closely, so treat it as the last place to hand over control, not the first.

Do we need a formal assessment before we start?

If the agent processes personal data, assume yes and check what form it has to take where you are. Either way the exercise is worth doing for its own sake: it forces you to write down what the agent may read, what it may send, and who switches it off — which is the same work the architecture needs anyway.

Can employees use their private AI accounts for company data?

No — and that is exactly what you should put in writing. Without a rule it happens anyway, just uncontrolled: company documents in a personal account you cannot audit, revoke or delete.

Does our data have to go into a US cloud?

No. There are self-hosted routes: workspace and data on your own server, with the language model interchangeable. What has to leave your systems and what does not is something we settle before building, not after.

We handle confidential client data — is this even possible?

With extra care. In law, medicine, accounting and similar fields, professional confidentiality sits on top of ordinary data protection, and a processor agreement alone is usually not enough. In practice the route runs through systems that stay in your own house, with the model swapped for one you host or one contractually barred from training on your data.

How long until the first result?

A simple agent runs within days. Getting it reliable enough for everyday use realistically takes four weeks — most of that clarifying the process, not the technology.

Let’s talk for 30 minutes

Tell me which three workflows cost you the most time. I’ll tell you honestly which of them suits an agent — and which doesn’t. You leave the call with a clear plan, not a sales pitch.

Book a free consultation