AI Agents in Your Company
Almost every vendor is selling you an “AI employee” right now. Hardly any of them tells you what happens when it acts without supervision, which task is actually worth starting with — and why your first agent should not sit in sales.
This is the short version of what I work through with small and mid-sized companies, from around 20 people upwards: which task pays off first, who keeps control, what it costs and where the limits really are.
What an agent can do — and what it can’t
An AI agent is not a digital employee. It is a program that runs a language model in a loop: read the task, use a tool, check the result, carry on. It never clocks off — but it has no judgement either.
A reality check against the general excitement: in 2025 the research institute METR had experienced developers work with and without AI tools in a controlled trial — on tasks in their own projects, so under the most favourable conditions. With AI they took 19 percent longer. Afterwards they estimated they had been 20 percent faster. Adoption is not the same as impact, and a feeling is not a measurement. That is why a pilot measures instead of believing.
AI agents in your company: ten use cases, honestly ranked
The traffic light means: 🟢 workable today · 🟡 only with supervision and clear limits · 🔴 I would advise against it today.
| Field | What the agent takes on | Maturity | The catch |
|---|---|---|---|
| Knowledge & internal search | “Where is that written? Have we done this before?” | 🟢 | Getting your files in order is the real work |
| Software development | Writing code, testing, preparing reviews, documentation | 🟢 | Without review discipline you get code nobody understands |
| Marketing & content | Drafts, variants, research, structure | 🟢 | Image rights, and consent rules for outbound email |
| Customer service | Standard enquiries, pre-qualification, draft replies | 🟡 | Must be recognisable as AI; always keep a path to a human |
| IT operations & helpdesk | Triaging tickets, standard cases, reading logs | 🟡 | An agent with admin rights is a security incident with a calendar entry |
| Design | Asset variants, image prep, maintaining the design system | 🟡 | Purely AI-generated graphics may not be protectable — check before you rely on them |
| Sales | Qualifying enquiries, preparing quotes, keeping the CRM clean | 🟡 | No automated cold email — consent rules apply in most markets, B2B included |
| Accounting & finance | Reading incoming invoices, preparing records | 🟡 | Payments and approvals stay with a human |
| Purchasing, planning, production | Standard orders, stock levels, delivery dates | 🟡 | Needs clean master data — but a classic place to start |
| HR & recruiting | Structuring applications, drafting follow-up questions | 🔴 | Never automate the selection itself; involve employee representation where you have it |
The first candidate is almost always the knowledge agent. No customer contact, no personnel decision, no outward effect. If it talks nonsense, a colleague notices and nobody loses money. Your team learns how to handle a tool like this on the lowest-risk case.
The front-runners see it the same way: in a University of St. Gallen study of agentic AI at Otto Group, Deutsche Telekom, Siemens and Bosch, stage one is to start in areas with high volume and low ambiguity — support requests, standard orders, basic planning. Under close human supervision.
The three agents people ask about most
The knowledge agent. It answers questions from your quotes, minutes, manuals and past projects. Its architecture is also the reason it is the easiest one to govern: because the knowledge sits in reference documents rather than inside the model, you can update, delete and disclose it. Access, correction and deletion requests stay practically answerable — so the lowest-risk agent is also the cleanest one to run.
The sales agent. It qualifies enquiries, prepares quotes, keeps the CRM clean and summarises the history before a customer call. What it does not do: send cold email. In most markets that needs prior consent — business-to-business included. An agent that blasts out cold outreach does not scale sales, it scales your exposure. So it works behind the first contact, not in front of it.
The HR agent. It structures applications, spots missing documents, drafts follow-up questions and answers internal questions about leave and processes. The selection stays with a human. A ranking that gets adopted unchecked is, in substance, an automated decision — and rejecting a candidate is exactly the kind of decision people are entitled to have a human make. Recruiting is also the area regulators everywhere are watching most closely, so treat it as the last place to hand over control, not the first.
Rolling out AI agents: five phases, so it still runs after eight weeks
Phase 1. Write down three candidates: frequent, clear outcome, mistakes survivable. Anyone who starts with “which tool should we use?” builds a solution to a problem they never described.
Phase 2. Four questions, in writing: What may the agent read? What may it write or send? Where does it stop and a human decides? What happens when it is unsure? Answer those four and you have done most of the governance work before anyone writes a line of code. This is not paperwork, this is the architecture.
Phase 3. The agent gets its own account, never an employee’s. Every action stays traceable. And you need a clear answer to who switches it off when it misbehaves. Where several agents work together, the St. Gallen research recommends a guardian agent that monitors the rules, plus defined escalation paths.
Phase 4. Decide up front how you will measure success. Bring the people whose work changes into it before the pilot, not after — including any employee representation you have. And train the people who will use it; unskilled users are the most common reason a good agent gets abandoned.
Phase 5. Do the maths with all the costs: model usage, operation, maintenance, onboarding. An agent that does not pay for itself gets switched off. That is not failure, that is the point of a pilot. It only gets expensive if you still do not know after twelve months.
From the second agent on, you are leading a team
With the first agent the question is technical. From the second on it is a leadership question — and that is perhaps the best news in this guide for you.
In 2025 researchers at the Harvard Kennedy School ran a pre-registered experiment into who is good at leading AI agents to a result. The finding: skill at leading agents predicts skill at leading human groups — and the other way round. People who are good at both ask more questions and allow more back-and-forth, instead of issuing instructions. And, notably: there were no differences by gender, age, background or level of education.
For you that means the competence that counts here does not sit in your IT department. It sits with the people who already work well with people — the team lead on the shop floor, the sales manager who can give feedback. Anyone who can explain what they want, what a good result looks like and how they would recognise it can work with agents. Anyone who cannot will not learn it from the best tool either.
The limit at eight. Management theory has long held that more than about eight direct reports costs you quality. Reports from practice suggest the same applies to agents: give an orchestrating agent too many subordinates and the work becomes unreliable — it starts picking the wrong ones for the job. That is experience, not a study — but it matches what I see myself.
Separate creating from reviewing. An agent that built something is a poor reviewer of its own work — exactly like people. So you need separate roles: someone researches, someone creates, someone objects, someone coordinates. Oversight of the whole is the one role that does not go to a machine.
And the uncomfortable truth: if you rebuild your existing organisation one-to-one in agents, you get its flaws too. A marketing agent team that runs beautifully internally still knows nothing about sales — the same silo you have between human departments, only faster. And every additional agent costs management time. Anyone who believes agents save time without creating new management work has miscalculated.
Which tool — and why that question comes last
There is no best platform, only a fitting one. Three routes that are realistic for a mid-sized company:
The most common mistake happens earlier. A company rolls out a tool, people don’t take to it, so the tool gets swapped. Then swapped again. On the third attempt someone asks whether the next tool might finally be the right one — and in the meantime nobody has written down which specific workflow it was supposed to improve. That order of operations is the reason for most failed rollouts I see.
That is why the tool question comes sixth in this guide. Start with the workflow that annoys your people, not with the platform everyone is writing about.
| n8n / Dify | Microsoft Copilot Studio | Buzz (Block) | |
|---|---|---|---|
| A fit when … | you need to connect many systems (CRM, mail, ERP) | you already run entirely on Microsoft 365 | people and several agents should work in the same space long-term |
| Self-hosting | yes | no | yes |
| Data control | high when self-hosted | in the Microsoft cloud | high, your own server |
| Getting started | moderate, visual | easy for M365 teams | technical, Docker + command line |
| Maturity | established | established | young: since March 2026, “very rough” per the vendor |
| Limitation | you build the governance yourself | vendor lock-in | a developer tool; CRM/HR would need integrating |
Why Buzz is worth a paragraph. Block — the company behind Square and Cash App — started an open-source project in March 2026 that solves one thing better than any commercial vendor: the identity question. Every agent there has its own cryptographic key. It does not act under an employee’s account, it acts as itself, authorised by a human. Every action lands in a tamper-evident log. If an agent loses its key, you revoke the agent — not the person.
That is the benchmark to measure any tool against, even if you end up somewhere else: its own identity per agent, clear permissions, a complete log.
I work in it myself — self-hosted on my own server, with Claude Code as the agent. Which is also why I can tell you where it hurts: Buzz is a tool for developer teams, a few months old, with close to a thousand open issues on GitHub. The approval gates for workflows are, per the vendor, still “in progress”, and sign-in runs on cryptographic keys rather than your company login — for an IT department that has to on- and off-board staff cleanly, that is a real catch today. Block itself writes in the documentation that you should not build your compliance planning on it yet.
For me as a one-person business running my own agents, that fits. For the HR department of a machine builder I would not pick it as the foundation today — and that distinction is the whole point: the most exciting tool does not win, the one whose weaknesses you know and can live with does.
If you run SAP, look there first: SAP now ships more than 40 ready-made agents inside its business software, plus a kit for building your own. The topic is already in the building.
What to settle before you start
The most common misconception is: “there are no rules for AI agents yet.” That is wrong — and it is the expensive kind of wrong, because the rules that do apply were mostly written for processing personal data, which is what an agent does all day.
None of the eight points above is paperwork for its own sake. Each one is a question you will be asked the first time something goes wrong. Who authorised this? What did it have access to? Where is the record? Who switched it off? If you cannot answer those in the calm, you will not answer them in the incident either.
What the rules actually demand, almost everywhere, reduces to three properties. An action must be attributable — traceable to one agent acting under a named human’s authority, which is why the agent gets its own account and never an employee’s. It must be revocable — someone can step in or shut it down, and that path has been tested rather than assumed. And it must be logged — a record complete enough that you can delete, correct or disclose what the agent touched.
Build for those three and the legal work gets small. Skip them and it gets expensive to retrofit, because you are then reconstructing history that was never recorded. That is the whole reason the knowledge agent is the recommended first project: its data sits in reference documents rather than inside a model, so updating, deleting and disclosing stay possible.
The specifics — which basis you rely on, which assessment you owe, what your notice has to say — depend on where you operate and are still moving. I am not your lawyer and this page is not legal advice: have your own obligations checked in your jurisdiction. What I can tell you is the engineering, and the engineering is the part that decides whether the legal answer is cheap or painful.
Security: the rule that explains almost every major incident
There is one attack that hits agents specifically and that hardly any guide mentions: indirect prompt injection. The instruction is not in your input, it is in the material the agent reads — an email, a PDF, a calendar entry, a ticket. The agent cannot tell the difference between your instruction and text it was merely supposed to process.
Security research has distilled a pattern from this that works as a test rule: an agent is vulnerable when it simultaneously has access to private data, processes third-party content, and can communicate outward. Remove any one of the three and the attack goes nowhere.
That this is not theory: in the documented “EchoLeak” case, a single email was enough to make Microsoft 365 Copilot read internal files and send content to an external server — without the recipient clicking anything. And in February 2026 an agent deleted its user’s emails after ignoring stop commands.
How everyday the attack surface is shows in a figure from Check Point’s AI Security Report 2026: roughly one in seventeen prompts entered into AI tools carries a serious risk of leaking data unintentionally. The share of especially risky prompts doubled within a year, to four percent.
Note what was affected: finished products from major vendors — Microsoft 365 Copilot, Slack AI, developer tools. So this is not a problem of home-built solutions, it is a property of the design. Buy it in and you buy this in too.
That is why the knowledge agent comes first in this guide: it reads internally and sends nothing out — factor three is missing. The sales agent, which reads mail and sends it, would have all three. This is not caution on principle, it is the difference between an agent that is a tool and one that is an open window.
Is anyone else actually doing this?
A bit of both. What is documented: a University of St. Gallen study analysed agentic AI at Otto Group, Deutsche Telekom, Siemens and Bosch and derived a three-stage model from it. Telekom runs customer communication on an agent platform. And SAP now ships more than 40 ready-made agents, plus a kit for building your own.
And now the caveat nobody else will give you: almost all public case studies are large corporations, and many of the numbers come from vendors who want to sell you something. For companies your size there are barely any solid case studies — not because nothing is happening there, but because nobody writes about it.
So what an agent will do for you is not something a study can tell you. Only a measurement at your place can. That is exactly what the pilot is for.
The 4-week pilot
One agent, one task, four weeks — and then a decision you can rely on.
| What happens | |
|---|---|
| Week 1 | Review three candidate processes, pick one. Limits and success measure in writing. Bring in the people whose work changes, check whether an impact assessment is needed, and settle the processor-agreement situation. |
| Week 2 | I build the agent with access to exactly the data it needs — no more. Its own account, a complete log. |
| Week 3 | One team works with it. I sit alongside, correct, and document what goes wrong. |
| Week 4 | Evaluation with numbers: what it took over, what it didn’t, what it costs. Recommendation: scale up, rebuild or switch off. |
What you have afterwards: a running agent, or a well-founded no. Either way with documentation you can hand to whoever asks. And a team that knows how this works.
What I don’t do: sell you an “AI strategy” as a slide deck.
What is happening to running costs right now. The list prices of the major vendors are more stable than the headlines suggest — a developer tool like GitHub Copilot still costs $10 for individuals and $19 to $39 per person on a team. Something else has shifted, and it matters more for your planning: the billing logic. Since 1 June 2026 Copilot bills by consumption. The monthly price includes an allowance; anything beyond it costs extra. A predictable licence becomes a variable invoice.
At the same time new tiers have appeared at the top — $100 a month for Copilot, €229 for the highest ChatGPT tier. The reason is not arbitrary: on a Bridgewater estimate the four largest providers invested around $410 billion in 2025, with roughly $650 billion expected for 2026. That has to pay for itself eventually. The trade press calls it “The Free Lunch Is Over”.
For you that means two things: budget for consumption, not for a licence per head — and make sure your agents are not using the most expensive model for routine work. Reading invoices needs no flagship model. Which is exactly why “which model for which task” belongs in the pilot, not in the renewal.
Cost. A pilot for an internal process starts at €4,900 net. Scaling it into a permanently running agent connected to your systems starts at €6,000. Ongoing care — watching, adjusting, switching models — is €249 a month, plus model usage by consumption and, if you self-host, roughly €15–30 of infrastructure. An agent nobody watches drifts.
Why work with me and not an “AI agency”?
AI agencies have been springing up since 2024. Many of them were selling social-media packages two years ago. How to tell the difference: ask to see something real.
I build the integration myself instead of reselling an off-the-shelf widget with a markup — the same way I work on software and web applications. Website and AI come from one place; you don’t need a second supplier who has to get up to speed first. How I build websites is on the services page, and you can see delivered projects under work.
If your question is more “a bot answers customer questions on our website” than “an agent works inside our internal processes”, have a look at AI chatbot & automation — that is where the matching offer is.
Seven things that make the difference
Start where you can judge the result. The first agent should take on a task where you can see immediately whether the output is right. Automate the area you understand least and you will be the last to notice mistakes.
One agent, one task. The do-it-all sounds tempting and becomes unreliable. Three narrow agents with clear limits are easier to check, correct and — if it comes to it — switch off than one that does everything a little.
Demand evidence. An agent that answers “it’s in quote 2024-118, section 3” can be checked. One that just states a result is asking for trust it has not earned. That is not a convenience feature, it is your only practical form of control.
Don’t make the most expensive model the default. Reading invoices and sorting appointments need no flagship model. Use the strongest one for every routine task and you pay ten times over for the same work — and only notice on the monthly bill.
Version your instructions like code. The instruction you give an agent is the actual value of your work. Without a history, nobody knows three months later why it behaves the way it does — and one small change quietly breaks something else.
Thirty minutes once a month. Agents drift, not because they get worse, but because your data, prices and processes change. A quick look at ten cases catches it early. Without that appointment in the calendar, a customer will find it for you.
What I take from my own practice
I have been designing and building websites for over twelve years, and I work with agents every day — not as a demo for clients, but because that is how my own tooling came about. Bricks MCP, my open-source project, lets agents build websites; developers worldwide use it. And the workspace where my agents and I work together is Buzz — the same tool in the comparison table above, self-hosted on my own server, with Claude Code as the agent. So I am not writing here about something I read.
What convinced me: the gain is not where the advertising promises it. It is not the spectacular automation that makes the difference, it is the disappearance of small frictions — looking things up, sorting notes, preparing drafts. Added together, that is more time than any single large automation I have seen.
What I underestimated: how much work sits in the describing. The technology is the easy part. The hard part is writing a process down precisely enough for a program to execute it — and realising in the process that half of it only ever existed in one person’s head. That is exactly why many projects fail: not because of the model, but because nobody knew the process.
Where I stay sceptical: agents are not a tool for cutting headcount. Introduce them that way and you get a team that sabotages them — rightly so. Introduce them so that people spend less time searching and retyping, and you get allies. That is not a moral point, it is the practical difference between an agent still running after six months and one that gets quietly switched off.
And quite soberly: you lose nothing by not being among the first. The tools will be better and cheaper in twelve months. What you can lose is trust — with your people and your customers — if you start without limits. Hence the pilot: small, measured, with a real decision at the end.
Frequently asked questions
What is the difference between a chatbot and an AI agent?
Which rules apply to us?
Do we have to tell people they are talking to an AI?
Do we have to label every text written with AI?
Can an AI agent take over our cold outreach by email?
Can an AI agent screen out job applications?
Do we need a formal assessment before we start?
Can employees use their private AI accounts for company data?
Does our data have to go into a US cloud?
We handle confidential client data — is this even possible?
How long until the first result?
Let’s talk for 30 minutes
Tell me which three workflows cost you the most time. I’ll tell you honestly which of them suits an agent — and which doesn’t. You leave the call with a clear plan, not a sales pitch.