Skip to content
biwak

Guides

Getting started

How do you get the first task for an AI agent right?

How to write your first task for an AI agent: goal, files, result format and a check that exposes errors, with examples you can copy.

Responsible:

Published

Updated 6 min read

Start with a task whose result you know

The first task does not have to map an entire business process. Choose a task you have already done by hand: the framework agreement you know, the monthly list whose total is fixed, the quote that has already been sent. Then you see immediately whether the result is right, and do not first have to work out what would be correct.

Work with copies. In its article “Building effective agents”, Anthropic advises starting with the simplest solution and warns that an agent's errors can compound over several steps; agents should therefore first be tested in sandboxed environments. In the office, this means: a separate folder with copies, no originals, no customer data until the data path has been checked.

What belongs in the task: goal, context, expectations, source

In its guide to Copilot prompts, Microsoft describes four parts: goal, context, expectations and source; at minimum, a clear goal is needed. For an agent that works with files, the four parts translate like this:

The four parts of a task, applied to file work
PartQuestionExample
GoalWhat should be there at the end?“Summarise the framework agreement on one page.”
ContextWhat for, and for whom?“For the renewal decision in November, read by the management.”
ExpectationsIn what form, and what must not happen?“As a new Markdown file in the same folder; for every statement, give the place in the contract; do not change the contract.”
SourceWhich file is authoritative?“The basis is rahmenvertrag-mueller-2024.pdf, no other file.”

A good cross-check is in Anthropic's guidance on writing prompts: show the task to a colleague who doesn't know it. If they would be confused, so will the model. Put together, the task then looks like this:

Task

Summarise the framework agreement rahmenvertrag-mueller-2024.pdf in the folder “Supplier Müller” on one page, for the management's renewal decision. State the term, notice period and price adjustment, list open questions at the end and refer to the place in the contract for every statement. Save the summary as a new Markdown file in the same folder and do not change the contract.

Input
A framework agreement as a PDF with a text layer that you know well yourself
Result
A new Markdown file with a one-page summary, references and a list of open questions.
Check
Look up every deadline and amount at the stated place in the contract. If a detail is missing from the contract, it must appear as missing, not as a guess.

Write a check into the task that exposes errors

The most effective sentence in a task is often the one that asks for a check. With numbers, that is a total; with lists, a count; with texts, the reference. Anthropic describes the same thing from the other side: at each step, an agent needs feedback from its environment, such as the result of a tool or a calculation, to assess its progress. A check in the task provides exactly this feedback, and you see it in the result.

How much this helps is shown by our measurement of 24 September 2026 on an invented test bank statement with 16 transactions: two widely used table-recognition tools did not find a single complete transaction row in it, and one text output did contain all the values, but spread over three lines per transaction. Anyone building a table from it, by hand or with a model, can easily lose a row. The balance check, opening balance plus the sum of the transactions equals the closing balance, caught the error as soon as we left out a single transaction.

Task

Read kontoauszug-2026-08.pdf and create buchungen-2026-08.csv: one row per transaction in the order of the statement, with booking date, value date, payee or payer, payment reference and amount with sign. The opening balance, carry-overs and totals are not transactions. At the end, check whether the opening balance plus the sum of the transactions equals the closing balance, calculate in whole cents and do not adjust any amount to make it add up. Write the result of the check in the last row.

Input
A bank statement as a PDF with a text layer, names redacted or replaced with test data
Result
A CSV file with one row per transaction and a control row for the balance
Check
Does the last row say “adds up”? Then count the rows against the statement and spot-check three transactions. If the check fails, a transaction is missing or a sign is wrong; look at the page break first.

Narrow down the material

Only put the files needed for this task in the working folder, and give them clear names. State explicitly which file is authoritative, where the result should be created and that originals remain unchanged. A small, clearly defined scope makes misunderstandings visible earlier.

Check the result, refine it and repeat it once

Open the result and compare it with your original question. Are the figures, names and references correct? Is anything important missing? Is the text understandable for the intended readers? Then give feedback that says exactly what should be different, and where.

Feedback an agent can act on
Instead ofBetter
“Make it better.”“Shorten the first section to five sentences and keep the three dates.”
“That's wrong.”“The notice period is in § 12(2), not § 11. Correct the line and give the reference.”
“Turn it into a table.”“Create a CSV file with the columns deadline, date and reference.”
“Is this complete?”“List all sections of the contract that do not appear in the summary.”

If the result is correct, run the same task with the same files a second time. Microsoft itself points out that the same prompt can produce slightly different answers each time. If the two results differ in figures or references, the task is not yet unambiguous enough, or it needs one more check. Only when two runs deliver the same result is the task suitable as a template for next month.

Frequently asked questions

May I use real customer data for the first attempt?

Better not. Whatever an agent uses from your files goes to the provider's language model for the answer, and whether this is permissible for customer data is something to clarify beforehand: legal basis, DPA and, for professionals bound by confidentiality, also Section 203 of the German Criminal Code (StGB). For the first attempt, documents without personal data, or copies in which names have been replaced with numbers, are enough.

Why does the same task give a different result the second time?

Language models choose their answer with an element of randomness; Microsoft writes that the same prompt can produce slightly different results each time. Differences in wording are harmless; differences in figures or references are not. In that case, a more precise task, a fixed list of columns or a check that makes an error visible will help.

Should the agent be allowed to change the original files?

Not on the first attempt. Have the result created as a new file and write in the task that originals remain unchanged. In any case, only put copies in the working folder: a checkpoint like Biwak's only covers this folder and does not replace a backup.

Sources

  1. Microsoft Support: get started writing prompts in Microsoft Copilot (as of February 2026)
  2. Anthropic: Prompting best practices, section “Be clear and direct”
  3. Anthropic: Building effective agents, 19 December 2024

Further reading

All articles