Chatbot, workflow or agent: which fits the task?
| One answer is enough | ChatExample: rephrase a paragraph to make it easier to understand. If no further actions are needed, a single model response is enough. |
|---|---|
| The steps are fixed in advance | Fixed workflowExample: read in and total every approved CSV export according to the same rules. A script or an existing spreadsheet function is often the right solution for this. |
| The next steps depend on the content | AgentExample: read several differently structured quotes, flag missing information, ask a follow-up question and then create a comparison file. |
The distinction follows Anthropic's post “Building effective agents” of 19 December 2024: in workflows, the model and tools run along predefined paths; in agents, the model directs the process and its use of tools itself. In everyday product use, the terms overlap; a chat interface can also include tools and agent functions. What matters for your choice is the specific task, not the label on the website.
The Federal Statistical Office's survey for 2025 shows how widespread this already is: 26% of companies in Germany used AI. Of these, 35% used AI to generate text or program code and 27% to automate workflows or decisions; among small companies with 10 to 49 employees, the latter figure was 26%. The statistics do not show agents in the narrow sense separately.
Examples: three office tasks with acceptance criteria
| Comparing documents | Preparing quotes as a basis for a decisionInput: approved copies of the quotes. Result: a table of scope of services, exclusions, date and price with document and page references. Check: verify every item against the original; missing information must appear as missing. |
|---|---|
| Analysing Excel or CSV | Creating a traceable monthly overviewInput: a copy of the approved export. Result: totals broken down by month and category plus a list of unclear rows. Check: reconcile the grand total, number of rows, date format, currency and excluded values with the source. |
| Preparing a presentation | Turning an existing report into slidesInput: approved report and formatting guidelines. Result: draft slides with sources for figures. Check: every key figure, the message of each chart and legibility in the slide deck as actually opened. Whether a tool produces a finished presentation file or only the slide text has to be checked for each product. |
These are example tasks, not promised success rates. The four cases on the Biwak home page are re-enacted, without a model call; they show the way of working, not measured performance. Check whether your PDF files, Excel formulas or templates work using representative sample files that are approved for processing. Scanned PDFs without a text layer, embedded images, macros and external spreadsheet links need a separate check.
How do you write a usable task?
A good task names the input, result, limits and check, for example like this:
Compare the three approved quotes in the working folder. Create a table with service, exclusions, total price and delivery date. For each value, state the source file and page. Do not fill in missing information; write “missing” instead. Save the result as a new CSV file and do not change the originals.
- Input
- Three quotes as PDFs with a text layer, approved for the trial
- Result
- A new CSV file with one row per quote and the file and page for each value. Anything a quote does not state appears there as “missing”.
- Check
- Look up every total price and every date on the page given, search the quote for each “missing” item, and check from the modification date that the originals are unchanged.
- Input: Which files and which period are included? Which version is authoritative?
- Result: Specify the file format, columns, target audience and storage location.
- Limits: Specify which sources are permitted and when a follow-up question is needed.
- Acceptance: Define which totals, source references and formatting requirements are checked before use.
The article on the first task for an AI agent shows how to start small and refine a result.
Can you build an AI agent yourself?
Yes, with very different levels of effort. Every agent consists of three parts: a language model, tools it is allowed to call, and instructions with limits. That is how OpenAI's practical guide to building agents describes it. If you build your own, you decide on all three parts and are also responsible for operation, testing and security.
Assembling an agent on a platform
Many office and automation platforms offer building blocks: write instructions, attach knowledge sources, enable actions. This is quick to set up; the limits and data routes are those of the platform.
- Check permissions and data access for each action
- Create test cases with a known result
Your own agent via a model API
Via a model provider's API, with an agent framework or a few lines of your own code. Anthropic recommends starting by working directly with the API and only using a framework if you understand its code.
- Limits on steps and retries
- Human approval before risky or irreversible actions
- Tests in a sandboxed environment
Using an existing agent
Desktop and web agents for office tasks come ready to use; you describe the task rather than the process. In return, the vendor determines the tools, model access and data route, and you check whether that suits your data.
Both guides advise starting with the simplest set-up. Anthropic recommends the simplest possible solution, adding complexity only when needed, plus extensive testing in sandboxed environments. OpenAI advises making full use of a single agent's capabilities before several agents work together, and involving a human when there are repeated failures or risky actions.
What works locally in Biwak and what is transmitted
Biwak is a desktop app for macOS and Windows with an AI agent for office tasks. You choose a working folder and describe the task; the agent reads files there, for example PDFs with a text layer, Word, Excel and PowerPoint, and creates new files, such as a CSV table or an HTML page. The files stay on your computer. For each model task, your question, the conversation context needed and the file contents used go via your Biwak account to the language model. Processing happens in the EU: at Microsoft Azure in the EU Data Zone, or as a fallback via Prem AI (contracting party PREM SA, Switzerland). The full chain with providers and locations is set out in section 14 of the privacy policy.
Before each task, Biwak backs up the working folder; one click restores the previous state. The checkpoint only covers this folder and does not bring back anything that has already gone out. In the app from version 0.2.27, you choose at the input field from four approval levels how independently the agent works: from a confirmation before every step to full access. From version 0.2.30, they are called “Always ask”, “Read”, “Write” and “Full access”, and one rule applies: whatever the level permits runs without asking; anything beyond that, Biwak puts to you first. On “Read”, the agent reads and searches freely and asks before every change; on “Write”, it can also change files in the working folder. Below full access, the file tools only write in the working folder; reading and a permitted command, however, reach as far as your user account on the computer. In addition, two switches in the settings turn off terminal commands and internet access completely. In OWASP terms, this means: switch off what a task does not need, and approve consequential steps yourself. What each level guarantees technically is set out in the security evidence. A checklist for your IT is on the security page. A separate article explains why a local program is not yet a locally run language model.
When is a task not yet suitable?
The risk of an agent lies less in a wrong answer than in a wrong action. In its list of the biggest risks of language model applications, the security initiative OWASP calls this “Excessive Agency”: too many functions, too broad permissions or too much autonomy. The damage is triggered by information the model has made up or by hidden instructions in files or web pages that the agent reads. OWASP recommends giving an agent only the tools and permissions it needs and having consequential actions approved by a human.
- There is no one who can judge the result professionally.
- The source data is incomplete or contradictory, but the result is still supposed to be treated as reliable.
- A wrong action would have an effect that is hard to reverse, such as a payment or a binding declaration.
- The data and the intended processing route have not yet been approved for this use.
In these cases, start with a narrower task, for example a draft instead of automatic execution. The guide Introducing AI in SMEs shows how to set up a pilot with a baseline, error checks and a clear decision. Product differences are covered in the comparisons with ChatGPT, Claude and local model tools.
Frequently asked questions
Is every chatbot an AI agent?
No. A plain text answer does not need an agent. When a system itself chooses tools and further steps depending on intermediate results, it is working as an agent. Many products combine both ways of working in one interface.
Does an AI agent replace Excel, Word or PowerPoint?
No. It can help you work on files, but the programs remain important for checking, further professional work and verifying the file format. Test formulas, links and layout in the program in which the result will later be used.
Are an AI agent's results automatically correct?
No. A generated document only proves that a file was created. Sources, figures, completeness and effects must be checked against criteria set in advance. If data is missing, the result should name the gap instead of filling it.
Does a company need several AI agents straight away?
Usually not. Start with the simplest process that solves the task reliably; OpenAI's practical guide also advises developing a single agent first. Additional agents make coordination and control more complex. Whether they help must be shown by the quality and total effort for the same task.
Sources
- OWASP Top 10 for LLM Applications 2025: LLM06 Excessive Agency
- OWASP Top 10 for LLM Applications 2025: LLM01 Prompt Injection
- Federal Statistical Office: Companies using artificial intelligence technologies, 2025 (as of 24 November 2025)
- Anthropic: Building effective agents, 19 December 2024
- OpenAI: A practical guide to building agents (PDF)
- Biwak: Privacy policy, section 14 (requests to a language model)
