Skip to content
biwak

Guides

Comparisons

Chat or agent: what changes when AI runs commands itself?

A chat delivers text, an agent works in your folder: what terminal commands mean, explained with real commands, with safeguards and limits.

Responsible:

Published

Updated 12 min read

Two ways of working: answers or manual steps

Most people know AI as a chat window: you type a question, the machine answers. For explanations, drafts and quick research, that is great. But everything that is supposed to come out of the answer, you then do yourself: the table in the right folder, the amended contract, the recalculated total. You copy, save, name and check.

  • Chat in the browser

    1. kontoauszug.pdfPlease put all transactions in a table.
    2. Here are the 16 transactions:08-001;Dauerauftrag;-1250,0008-002;SEPA-Lastschrift;-186,40… 14 more rows

    QuestionAnswerYou take it from there

  • Biwak in the working folder

    • Accounting
    • kontoauszug.pdf
    • buchungen-2026-08.csvnew
    1. ReadFolder and files
    2. ActRun a command
    3. CheckRun a cross-check
    4. Fixor report the gap
On the left, the chat window: the file is uploaded, the answer comes back as text, and the rest is up to you. On the right, Biwak in the working folder: the agent reads, runs commands, checks and fixes; whatever cannot be resolved, it reports. A diagram, not a screenshot.

An AI agent works differently. It is given a goal and a place and chooses the steps itself. In Biwak, this place is a working folder that you specify. The agent opens the PDFs, spreadsheets, Word files and earlier versions in it itself, and its results end up as files in the same folder. How agents work in general and which office tasks are suitable is covered in the article What is an AI agent?.

What a terminal command is, without jargon

Every Mac and every Windows computer has a text window in which you can write instructions to it: the Terminal, or PowerShell on Windows. Instead of clicking, you type a short command, the computer runs it and replies with text. Software developers have worked like this for decades, because a command is precise, repeatable and can be chained with others. Four examples from the log further down:

Four commands and what they mean by hand
CommandWhat it doesBy hand, this would be
lsShows what is in the folderOpen the folder in Finder or Explorer and look
biwak text auszug.pdfOutputs the text of a PDF, Word, Excel or PowerPoint file (the app's read command, abbreviated here)Open the file, select all, copy
awk …Adds up a column or finds rows that stand outCalculator or a formula in Excel
Write fileCreates a new file in the folderSave as, type a name, choose a folder

A chat can suggest such commands to you, but you would have to type them in yourself. An agent runs them, reads the computer's response and decides on the next step accordingly. That is the step from advice to action.

In software development, this way of working is already widespread. Anthropic describes Claude Code as a tool that reads code, edits files and runs commands; according to OpenAI, its Codex CLI runs “locally on your computer”. Even there, it is not yet everyday practice: in the Stack Overflow 2025 developer survey, 84% of respondents use or plan to use AI tools, but only 14.1% work with AI agents daily. Biwak transfers this way of working from source code to the office folder: instead of program files, it contains invoices, contracts and spreadsheets.

Same task, two routes: who does the manual steps?

Take an everyday task: a bank statement is available as a PDF, the transactions need to go into a table, and the total has to match the account balance. In ChatGPT, data analysis handles this. For many analyses, according to OpenAI, it writes Python code and runs it in an environment that works with the files made available to the session, i.e. with your uploaded copy. OpenAI states limits for uploads: 512 MB per file, around 50 MB for spreadsheets, up to 80 files in three hours, and three uploads a day on the free plan. You get the result as an answer in the chat, for example as a table or chart.

Same task, two routes

TaskTransfer all transactions from kontoauszug.pdf into a table and check them against the account balance.

6manual steps for you
out of 7 steps

  1. YouFind the bank statement in the folder and upload it
  2. YouType in the task
  3. ChatGPTReads the uploaded copy and answers with a table
  4. YouTake over the table, for example copy it and paste it into Excel
  5. YouPut the file in the right folder and name it
  6. YouCheck the total against the account balance, or ask for it to be checked
  7. YouIf there is a discrepancy: describe the error, get a new version, file it again

2manual steps for you
out of 7 steps

  1. YouType in the task. The working folder is already chosen
  2. BiwakBacks up the folder before anything changes
  3. BiwakLooks at what is in the folder
  4. BiwakReads both pages of the statement
  5. BiwakWrites buchungen-2026-08.csv to the folder
  6. BiwakRuns the check and, if there is a discrepancy, looks for the cause
  7. YouCheck the result, keep it or reset it with one click
Switch between the two routes. In the chat window, six of seven steps stay with you; in the working folder, two: giving the task and checking the result. A diagram based on the steps of this task, not a measurement.

To be fair: ChatGPT can also recalculate the total if you ask. The difference lies not in how clever the model is, but in where the work happens. In the chat, the AI works with a copy in an environment at the provider. In the working folder, it works on the spot, next to all the other files it may still need: the price list, last month's spreadsheet, last week's contract.

A log, translated line by line

What does it look like when an agent works? Below are the commands of such a task. We ran them in the Terminal on a MacBook Pro on 24 September 2026, with the invented bank statement from our sample pack. On the left is what the computer saw, on the right what it means.

AccountingLog · 6 steps

  1. ls
    kontoauszug.pdf

    Looks at what is in the folder. Nobody had to upload the file.

  2. biwak text kontoauszug.pdf
    --- Seite 1 von 2 ---
    Alter Kontostand am 31.07.2026  14.850,32 EUR
    03.08.2026  Dauerauftrag  Beleg 08-001  -1.250,00
    …
    16 Buchungen | Summe Gutschriften: +7.953,80 EUR | Summe Belastungen: -6.124,75 EUR
    Neuer Kontostand am 31.08.2026  16.679,37 EUR

    Reads both pages as text, in about a tenth of a second on our computer. The new balance at the bottom is the check for later.

  3. (schreibt buchungen-2026-08.csv)
    Buchungsdatum;Wertstellung;Belegnummer;Buchungsart;Empfaenger_Auftraggeber;Verwendungszweck;Betrag_EUR
    03.08.2026;03.08.2026;08-001;Dauerauftrag;Hausverwaltung Fichtenhof GmbH;Gewerbemiete August 2026 / Objekt Lindenweg 12;-1250,00
    … 15 weitere Zeilen

    Saves the table as a new file in the folder. The statement stays as it was.

  4. awk -F';' 'NR>1{b=$7; gsub(/[+,]/,"",b); s+=b; n++} END{print n" Buchungen, neuer Stand "1485032+s" Cent"}' buchungen-2026-08.csv
    16 Buchungen, neuer Stand 1489437 Cent

    First check, in whole cents: €14,894.37 instead of €16,679.37. That doesn't add up.

  5. awk -F';' 'NF!=7{print NR": "NF" Felder"}' buchungen-2026-08.csv
    8: 8 Felder

    Looks for the row that stands out: row 8 has one field too many. Its payment reference itself contains a semicolon.

  6. awk -F';' 'NR>1{b=$NF; gsub(/[+,]/,"",b); s+=b; n++} END{print n" Buchungen, neuer Stand "1485032+s" Cent"}' buchungen-2026-08.csv
    16 Buchungen, neuer Stand 1667937 Cent

    Recalculates using the last column: €16,679.37, the statement's balance to the cent. The table was right; the first check had counted wrong.

Commands and outputs are real: run in the Terminal of a Mac on 24 September 2026, with synthetic sample data. The sequence recreates a Biwak run. No language model was involved; the table in step 3 comes from the sample pack. We abbreviate the app's read command as biwak text; the agent calls it via the app's program path, and the output is the same.

The interesting step is the fourth. The first check gave €14,894.37 instead of €16,679.37. It was not the table that was wrong, but the check: one payment reference itself contains a semicolon, exactly the character that separates the columns, and in row 8 the calculation picked the wrong column. We had not planned this error; it happened while we were putting the log together. It shows nicely what matters. A chat answer with a wrong total sounds just as convincing as one with the right total. An agent that recalculates in the folder notices the gap, because the number does not match the statement.

Why the folder matters more than the answer

In the chat, you decide what the model gets to see: whatever you paste in, upload or attach from a connected service such as Google Drive or OneDrive. Whatever you forget is missing, and the answer still sounds fluent. In the working folder, the agent looks for itself. It finds the price list next to the order and the template next to the draft. The folder is the context.

Volume

Many files in one go

Instead of fifty uploads: when there are many files, Biwak instructs the agent to write a loop that runs over all of them and outputs only the result. It reads PDFs with a text layer, Word, Excel, PowerPoint and OpenDocument. It does not read scanned receipts without a text layer.

Storage

The result is where it belongs

The agent puts tables as CSV, pages as HTML and texts directly in the folder. Every new or changed file appears in the answer as a card with a preview.

Evidence

What happened can be checked

Logging happens on your computer: the tasks, which files the agent read and changed, the checkpoints. If you don't like a result, you reset the folder.

What this looks like with a larger volume is shown in the demo on our home page: a clause is replaced in 30 contracts, 28 are changed, two were already up to date, and a list of changes is created. The demo is a recreation without a model call; it shows the process, not a measurement.

Commands on your own computer: what protects you?

Anything allowed to run commands can also do damage. That is the price of real work, and we say so openly: the agent's commands reach as far as your user account on the computer. Biwak works in the folder you choose and saves the results there, but that is not a technical barrier for the rest of the computer. The terms and conditions expressly point this out in § 9. That is why there are these guard rails:

What Biwak does before and during a task
  • Before every task, Biwak backs up the entire working folder. One click restores the previous state, and the reset itself can be undone too.
  • If the backup fails, the task does not start.
  • Terminal commands and internet access can be switched off in the settings.
  • The checkpoint only covers the working folder and is not a substitute for a backup. So start with a separate folder full of copies.
  1. Your computer

    • Working folder with your files and results
    • Commands run here, with this computer's programs
    • Conversations and log are stored here
    • Checkpoint before every task
  2. Task, conversation context, required file contents Answer and the next commands
  3. Biwak relay

    • Frankfurt am Main
    • Account and quota
    • Does not store task contents
  4. Language model

    • Processing happens in the EU
    • On Microsoft Azure in the EU Data Zone
    • Alternatively via Prem AI (PREM SA, Switzerland)
    • No training on your content
The hands work on your computer; the thinking happens in the EU. What is transferred is the task, the necessary conversation context and the file contents the agent needs for the answer. The graphic shows the desktop app.

Your files stay in the folder on your computer. Biwak also stores the app's conversations on your computer, outside the working folder; they do not reach us, except in an error report or a message that you send yourself. The app does not report anything to us of its own accord: no telemetry, no usage statistics. For an answer, your question, the necessary context and the file contents used go via the Biwak relay server in Frankfurt am Main to the language model. Processing happens in the EU: on Microsoft Azure in the EU Data Zone, alternatively via Prem AI (contracting party PREM SA, Switzerland). Biwak does not store this content in the process and does not use it for training. What the model provider promises, and where this promise ends, is set out in the privacy policy, section 14 and on the security page.

What ChatGPT does better

  • Questions, drafts, research: that is what a chat is built for, in the browser, on the desktop and on the phone. The Biwak desktop app practically does not search the web; web search is only available in the browser workspace.
  • Teamwork: ChatGPT Business and Enterprise offer shared projects and administration for many seats. With Biwak, any number of colleagues work in the Rope Team plan with their own access from one quota; we do not promise a jointly edited workspace.
  • OpenAI has agents too: Codex CLI works on the developer's own computer. In Business and Enterprise, ChatGPT Work can use approved files, programs and browser sessions on the computer; OpenAI itself stresses that the model still does not run on the device. According to OpenAI, the earlier “ChatGPT agent”, which worked on a virtual computer at OpenAI, no longer exists.
  • Images: ChatGPT generates images, according to the pricing page also in Business and Enterprise. Biwak does not generate images and does not read scanned receipts without a text layer.

So the question is less “chat or Biwak?” than “what work needs doing?”. For questions, drafts and research, a chat is enough. As soon as a task consists of files that have to be read, changed, filed and checked, the working folder pays off. A comparison of the business ChatGPT plans with Biwak is available under Biwak or ChatGPT.

Try it yourself: a first task with a check

  1. Download

    Biwak is available for macOS and Windows on the Download page. The app is free; model requests run via a Biwak account.

  2. Create account

    After email confirmation, you get a one-off 100 trial credits, with no payment details. Create account.

  3. Create a folder with copies

    For example three PDF invoices or a spreadsheet export, and choose this folder as the working folder in Biwak.

  4. Write the task with a check

    Say how the result can be checked: “Transfer all invoices into a table and check whether the sum of the net amounts matches the individual invoices.”

  5. Check the result

    The new file appears as a card with a preview. If something is wrong, reset the folder with “Undo”.

More ideas for getting started are in the article A good first task. After that there are two plans: Base Camp for €19 and Rope Team for €99 a month, both cancellable monthly (prices at a glance). Prefer to watch first? In a twenty-minute demo, we work on a real task from your day-to-day work.

How this comparison was made

This text was written by Biwak, i.e. a provider writing about its own product; it does not replace an independent test. The information on ChatGPT, Codex and Claude Code comes from the vendors' help and documentation pages and was checked against the wording on 24 September 2026. OpenAI's help pages cannot be retrieved automatically; for those, we read the Internet Archive copies of 2, 15, 17 and 18 September 2026. The information on Biwak comes from the privacy policy, the terms and conditions and the app's source code. The commands in the log ran on 24 September 2026 on a MacBook Pro with M1 Pro, with invented sample data and without a language model. The graphic of the manual steps is a diagram. We have not carried out a performance comparison between ChatGPT and Biwak. Please report errors to kontakt@biwak.ai.

Frequently asked questions

Does Biwak run commands without asking first?

Yes. Within a task, the agent runs its commands itself; that is the point of an agent. Protection comes from the checkpoint before every task and the two switches in the settings with which you turn off terminal commands and internet access. Without terminal commands, the agent no longer runs commands or scripts.

Do I need to be able to use the terminal myself?

No. You write the task in normal sentences, just as you would give it to a colleague. The agent chooses the commands. Afterwards, you see the new and changed files with a preview and can reset the folder.

So does the AI run on my computer?

No. The commands run on your computer; the language model runs at the provider. Processing happens in the EU. Biwak therefore needs an internet connection for new model requests; your files and the checkpoints stay on your device.

Can't ChatGPT also edit files on my computer?

Partly. OpenAI's Codex CLI works on the developer's own computer, and ChatGPT Work can use approved files and programs in the Business and Enterprise plans. In the browser chat, ChatGPT works with uploaded copies or with files from connected services such as Google Drive, OneDrive and SharePoint. A more detailed comparison of the business plans is available under Biwak or ChatGPT.

What does it cost to try this out?

The app is free. After email confirmation, you get a one-off 100 trial credits, with no payment details. After that, Base Camp costs €19 and Rope Team €99 a month, each cancellable monthly.

Sources

  1. OpenAI: Data analysis with ChatGPT (help page, read as an Internet Archive copy of 15 September 2026)
  2. OpenAI: File Uploads FAQ, upload limits (read as an Internet Archive copy of 18 September 2026)
  3. OpenAI: ChatGPT agent, notice of discontinuation (read as a copy of 2 September 2026)
  4. OpenAI: Introducing ChatGPT agent, 17 July 2025 (read as an Internet Archive copy)
  5. OpenAI: ChatGPT Business and Enterprise, plans and pricing (read as an Internet Archive copy of 17 September 2026)
  6. OpenAI: ChatGPT Work local security, local tasks and what is transferred (undated, checked on 24 September 2026)
  7. OpenAI: Codex CLI on GitHub (checked on 24 September 2026)
  8. Anthropic: Claude Code overview (undated, checked on 24 September 2026)
  9. Stack Overflow: Developer Survey 2025, AI section (survey May to June 2025)
  10. Biwak: Privacy policy, section 14 (requests to a language model)
  11. Biwak: Terms and conditions, § 2 and § 9

Further reading

All articles