From Assistants to Agents: Multi-Step Work You Can Delegate
An agent is a model in a loop with tools and a goal. Learn to delegate multi-step work safely as a user, and to build the loop — then graduate to the full Agent Workflows track.
The idea
An assistant answers; an agent pursues a goal. Give it "book me a table for four near the office on Thursday, then put it in my calendar and message the team" and it will search, choose, book, create the event and send the message — calling tools in whatever order the situation needs, reading each result, and deciding what to do next until it is done. The difference from stage 4 is that the model, not you, decides the sequence of steps.
As a user in 2026 you meet agents everywhere: Claude Code, Gemini CLI and Codex write and test software autonomously; browser agents (Claude in Chrome, ChatGPT agent, Gemini in Chrome) operate websites for you; deep-research modes in every major assistant run dozens of searches and produce a cited report; workplace copilots file tickets and update records. The skills that matter are delegation skills: a clear goal, explicit constraints ("don’t spend more than €50", "ask before sending anything"), a way to check the result, and starting with low-stakes tasks.
As a builder, an agent is a short loop — send messages, execute tool calls, append results, repeat until the model stops asking — wrapped in guardrails: a step budget, permission checks on side-effecting tools, logging, and a way to stop. Frameworks (Google ADK, the Claude Agent SDK, the OpenAI Agents SDK) package that loop with sessions and tracing. The Agent Workflows track covers all of it in depth; this stage gets you to your first working agent.
Vocabulary
- Agent
- A model that plans and executes multiple tool calls toward a goal without step-by-step instructions.
- Agent loop
- Send → model asks for tool → run it → return result → repeat until done.
- Guardrail
- A limit on what the agent can do: budgets, allowed tools, approval gates, forbidden actions.
- Computer / browser use
- Agents that operate a screen or a browser like a person would.
- Deep research
- An agent mode that runs many searches and synthesises a cited report.
No-code walkthrough · non-developers
Delegate a research task, then a browser task — and check the work
For: Analysts, marketers, founders, students, anyone who does multi-step desk work
- 1Research agent: in ChatGPT, Claude or Gemini, switch on deep research mode and ask: "Compare the top 5 [project management tools] for a 20-person non-profit: pricing, integrations, limits. Cite sources. Finish with a recommendation and what you are unsure about." Walk away for 10 minutes.
- 2Check the work: open three citations at random. Check one price on the vendor site. Note what the agent said it was unsure about — a good agent tells you.
- 3Browser agent (Claude in Chrome, ChatGPT agent, or Gemini): start with something reversible — "Find three flights Dublin→Lisbon on 12 Oct under €150 and put them in a table; do not book." Watch it work.
- 4Coding agent (even if you don’t code): open Gemini CLI, Claude Code or Codex on an empty folder and ask for "a single-page website for my bakery with a menu and a contact form". Then ask for changes. This is the most concrete way to feel what an agent does.
- 5Write down your delegation template: goal, constraints, what "done" looks like, what it must ask before doing. Reuse it.
Goal: [what you want, with the finished output described]. Constraints: [budget, time, sources to use or avoid, tone]. Before doing anything irreversible (sending, buying, deleting, posting), stop and ask me. When finished, list what you did, what you could not verify, and anything you decided on my behalf.
Check your work: Judge an agent like a new hire: did it follow the constraints, did it flag uncertainty, and can you audit what it did? If it hid a decision, tighten the constraints next time.
Developer walkthrough · Gemini · Claude · OpenAI
Your first agent with a real framework: a goal, two tools, a step budget, and an approval gate on the side-effecting tool. Each framework below is the provider’s own; the shape is identical. Once this runs, the Agent Workflows track takes you through workflows, skills, multi-agent topologies, context engineering and evals.
# pip install claude-agent-sdk
import asyncio
from claude_agent_sdk import query, ClaudeAgentOptions, tool, create_sdk_mcp_server
@tool("create_ticket", "Create a support ticket.", {"title": str, "body": str})
async def create_ticket(args):
return {"content": [{"type": "text",
"text": str(TICKETS.create(args["title"], args["body"]))}]}
options = ClaudeAgentOptions(
system_prompt="You triage customer reports. Research with WebSearch, then "
"create ONE ticket with a clear title and reproduction steps.",
mcp_servers={"tickets": create_sdk_mcp_server("tickets", tools=[create_ticket])},
allowed_tools=["WebSearch", "mcp__tickets__create_ticket"],
permission_mode="default", # side-effecting tools prompt for approval
max_turns=12, # step budget
)
async def main():
async for msg in query(prompt="Customer says checkout fails on Safari 18 with "
"a blank page after 'Pay'. Triage it.",
options=options):
print(msg)
asyncio.run(main())- 1Run it on five real reports. Read the trace for each: how many steps, which tools, where it wasted calls.
- 2Remove the approval gate and watch what happens with an ambiguous report. Put it back. This is why gates exist.
- 3Move on to Agent Workflows lesson 1 — you now have the loop; the track teaches you to structure it.
Practice
Everyone
Delegate one real task per day for a week to an agent mode — research, comparison, a browser task, a document. Keep a log: time saved, errors caught, what you had to specify better.
Developers
Build an agent that closes the loop on something you own: it reads an inbox/queue, researches, drafts, and asks you before acting. Ship it with logging and a daily cost cap.
Pitfalls at this stage
- •Vague goals. "Sort out my inbox" is not delegable; "archive newsletters older than 30 days and draft replies to anything from a customer, don’t send" is.
- •No budget. Agents that can loop will loop. Always cap steps, time and spend.
- •Trusting the summary instead of the log. Read what the agent did, not just what it says it did.