harnsy

Task management for AI agents: board and worktrees

When an agent does the task, a line in a list is not enough. The agent needs criteria it can check off, an owner that stays when the agent is replaced, an order that other tasks respect, and its own place to work. Here is how the harnsy board does it, on one real bug we fixed on 6 October 2026.

7 min readPavel Buchnev

Key points

86 MB/s
what the live node read from its database before the fixes (our measurement, one node, 6 October 2026)
about 7 min
from the bug to the first cause, found by a developer agent working in a worktree
about 5 MB/s
what the same node read after the last landing, about 17 times less, in four fixes, each reviewed before it landed (our measurement, 2 to 7 MB/s)
In this article
  1. What a task needs when an agent does it
  2. One task, from “new” to “done”
  3. The task belongs to a role
  4. Order of work: dependencies and chains
  5. Where the work happens: a worktree on the card
  6. Documents and screenshots on the card
  7. One real case: a node that read its database at 86 MB a second
  8. What you do in all this
  9. What it is not
  10. Install it with one sentence

What a task needs when an agent does it

A line in a list says what to do. It does not say when the work is done, who owns it if the agent is replaced, what must happen first, or where on disk the change is made. A person keeps those answers in their head. An agent does not carry them from one session to the next, so the board has to.

The harnsy board gives each task four things: criteria the agent checks off with evidence, an owner that stays, a place in an order, and a git worktree.

One task, from “new” to “done”

A task is an item on the project’s board. It has a type (bug, feature, research or task) and a priority (high, medium or low). It moves through states: new, taken, in progress, blocked, in review and done. It can also be dropped, or parked in the backlog until its time comes. A lead can hand a part of a task to a role below it, as a sub-item.

Each criterion is a check. The agent that does the task closes a check with evidence, and a reviewer adds what it finds as new checks. When the work is ready, it waits in review. On the card you accept it, send it back for rework with a comment, or defer it.

#452Add a CSV export to the reports pageIn review
  • Typefeature
  • PriorityHigh
  1. New
  2. Taken
  3. In progress
  4. In review
  5. Done

Checks2/4

  • Export button on the reports pageEvidence · commit
  • The columns match the tableEvidence · test run
  • An empty report gives an empty file
  • Reviewer: the header row is missing in the fileFinding
Schematic. A task card: type, priority, criteria as checks, evidence, a reviewer’s finding and your three actions. Sample data.

The task belongs to a role

The owner of a task is a role, not a session. When an agent runs out of context, it writes a handover note and a fresh agent takes the role. The task, its checks and its history stay on the board.

Order of work: dependencies and chains

Some tasks have to wait for others. In harnsy you say which task blocks which, also across projects. A card names the tasks that hold it, and when they are done it says “Ready to start”. When a blocker is done or dropped, the blocked task’s agent and its lead are told once.

A dependency is a plan, not a lock: nothing stops you from sending a task to work early. Tasks › Chains draws linked tasks as chains, step by step from left to right, and the Dependencies of a task show its whole chain.

Chains

Five tasks: Cart API is done and holds up Checkout form and Empty-cart screen; Checkout form blocks Payment step, which blocks Release 1.4. Empty-cart screen is ready to start.

  • #31Cart APIDone→ blocks 2
  • #34Checkout formIn progress→ blocks 1
  • #35Empty-cart screenReady to start
  • #36Payment stepBlockedafter 1 open item→ blocks 1
  • #39Release 1.4Blockedafter 1 open item
Schematic. Five tasks as a chain: what is done, what is in progress, what waits and what is ready to start. Sample data.

Where the work happens: a worktree on the card

Two agents in one checkout step on each other’s changes. The usual answer in git is a worktree: a second working folder on its own branch. A worktree works with Claude Code and with any other coding agent, because it is plain git. In harnsy a task can name its worktree, and the card shows the branch and the folder, with a button to copy the path.

The Changes tab collects every worktree of a project in one place. For each one it shows what it changed against the base branch, the task it belongs to and the agent that works in it. It points out worktrees left behind, and when two live worktrees change the same file it tells the lead. When a change has landed, you remove its worktree with one button. The commands an agent runs are linked to the task whose worktree they ran in.

Changes · kestrelbillBase branch: main
  • feat/csv-exportnot merged

    #452 ·developer ·3 files changed

  • fix/report-headernot merged

    #455 ·developer ·2 files changed

  • feat/old-filterleft behind

    #431 ·developer ·landed

api/export.go is changed in two live worktrees: the lead is told.

Schematic. The Changes tab: each worktree with its branch, task and agent, a landed one to remove, and a file that two live worktrees both change. Sample data.

Documents and screenshots on the card

A task keeps more than notes. Whoever works on it can attach a document: a report, a plan, a review, as Markdown or as a self-contained HTML page, or a screenshot. The files sit on the card, grouped by kind (reports, plans, mockups, screenshots, logs, data), and the lead and the person read them there, next to the task’s history, instead of in a chat or in a file left on someone’s disk. The same file name attached again becomes a new version.

On our own board the editor’s review of a text sits on the card as a report, and the designer’s mockups as screenshots. The real case below kept its record as notes on the task and had no attachments.

#452Add a CSV export to the reports page

Artifacts6

Reports1

  • audit.mdv2editor

Plans1

  • plan.mddeveloper

Mockups2

  • mock-390.pngdesigner
  • mock-1440.pngdesigner

Screenshots1

  • live-check.pngreviewer

Data1

  • numbers.mdlead
Schematic. A task card’s artifacts, grouped as the dashboard groups them: reports, plans, mockups, screenshots and data, each with the role that attached it. Sample data.

One real case: a node that read its database at 86 MB a second

On 6 October 2026 the lead agent of our own harnsy project measured the node we run every day. harnsy’s process was using one to four cores and reading its database at about 86 MB a second. The lead filed it as a bug, with the measurements in the task’s criteria.

A developer agent took the task, found the first cause about seven minutes after the bug was filed and fixed it in a worktree. A reviewer read the change, ran the tests and accepted it. The lead landed it.

That was the first of four fixes. After it the node read 30 to 50 MB a second: better, not enough, so the card stayed open. The second fix came from the notification bell, which read the whole task table for every row.

Then a temporary diagnostic ran for an hour on the live node and ranked who reads most. It pointed at two things: the Links and traces view, which re-read the whole message log on every call and became a task of its own, and the dashboard’s flood of requests. Each got a fix, and every fix was reviewed before it landed.

After the last landing the node reads 2 to 7 MB a second, about 5 on average, and uses 7 to 15 percent of one core. The target was under 5, so it sits right at the edge. What is left is a bigger change, and the lead decides on it.

The task holds every step as notes: the measurements, the cause, each fix with its review and what landed. We did not rebuild it afterwards. It is the record the agents wrote as they went.

Reads from the database: 86 MB/s about 5 MB/s

6 October 2026
  1. 17:31

    Bug filedlead86 MB/s

    The live node reads the database at 86 MB/s and uses one to four cores. The bug carries the measurements.

  2. 17:38

    Cause founddeveloper

    A background pass re-read the last week of messages every second or two. Found on a copy of the database.

  3. 17:40

    Fixeddeveloper

    One index, so the query stops reading whole rows.

  4. 17:43

    Reviewed and acceptedreviewer

    Read the change and ran the checks that touch it.

  5. 17:47

    Landedlead30–50 MB/s

    Landed, then measured again on the live node: better, not yet enough. The card stays open.

  6. 17:56

    Fixeddeveloper

    The notification bell read the whole task table for every row, and the Links view refreshed constantly. Now an index, and at most every 30 seconds.

  7. 17:58

    Reviewed and acceptedreviewer

    Accepted; one comment fixed before landing.

  8. 18:05

    Diagnosticdeveloper

    A temporary diagnostic ran for an hour on the live node and ranked who reads most: the background pass, then the dashboard’s flood of requests.

  9. 18:50

    Fixeddeveloper

    Found in the same search, a task of its own: the Links and traces view re-read the whole message log on every call, up to 12 GB. Now one index, and the same answers byte for byte.

  10. 19:13

    Reviewed and acceptedreviewer

    Accepted: full test suite, new queries compared with the old on every kind of duplicate.

  11. 19:26

    Fixeddeveloper

    One request instead of a fan-out per project, counts from indexes, cached transcript tails. The diagnostic is removed.

  12. 19:41

    Reviewed and acceptedreviewer

    Accepted: full test suite, and 455 board requests return the same answers as before.

  13. 19:56

    Landedleadabout 5 MB/s

    Landed and live. Measured afterwards: 2 to 7 MB/s and 7 to 15 percent of one core.

Schematic from the real record of the task, roles only: from the measurement to the landed fixes. Times are from the record.

What this shows is a bug with criteria, an owner role, fixes made in worktrees and a review before every landing. It does not show that harnsy is faster than another way of working; we did not measure that.

What you do in all this

You decide and accept. You answer an agent’s question when it asks, accept finished work, send it back with a comment or defer it. The agents keep the board current, and Home shows what waits for you in one place.

Several coding agents are hard to manage, and hard to join into a process you control. The board is one answer: you can see the process, and you step in when something needs your attention.

What it is not

  • It does not replace your company’s tracker. harnsy does not read or write Jira or Linear.
  • Checklists are read-only in the dashboard: the agents close the checks, you read them.
  • A dependency orders the work, it does not lock it.
  • Milestones, a release of tasks with progress and statistics, are part of harnsy Max.

Install it with one sentence

Tell your agent: “Install harnsy following https://harnsy.dev/llms.txt”. Your agent shows the plan and then installs it.

Task management for your agents

The board is part of harnsy. Install it with one sentence to your agent.

Install harnsy

harnsy is not affiliated with Anthropic, OpenAI or OpenCode; their names are theirs.

site-3flead · Claude Codeliveview only
❯ Plan #42 with the team.

Waiting for the breakdown from analyst-7a…

from analyst-7a through harnsy❯ #42 broken down: three acceptance criteria, including a retry after 24 h.

I’ll hand #42 to site-9a.

❯