Back to Blog
product Sofiya Brenner

How the Agent Run Inspector Works

A deep-dive into the Agent Run Inspector: how it captures every decision an agent makes, why step-level tracing matters for ops teams, and what an audit trail looks like in practice.

Abstract layered trace visualization showing nested agent decision steps in teal on dark navy background

When we were building the first version of MicroAGI, we spent a lot of time debating what "inspectable" actually means in practice. It was easy to say: "every step should be visible to the manager responsible for that workflow." It was harder to figure out what that meant at the level of a running agent that calls five different tools in sequence, where one of those calls makes a write to a production ERP system.

The Agent Run Inspector is the answer we built. This post explains how it works, what it captures, and where we made deliberate tradeoffs in how it presents information to ops teams rather than engineers.

What a Run Trace Actually Records

Every time a MicroAGI agent executes a workflow, it produces a run trace: a sequential, timestamped log of every step the agent took. Not just "the agent ran successfully" and a final output. Every tool call, every intermediate value, every branching decision, every piece of data that moved between steps.

The trace has four data points per step:

  • Tool called: the name of the integration or internal function the agent invoked (e.g., sap_post_invoice, gmail_fetch_inbox, jira_create_ticket)
  • Input payload: what data was passed into the call, sanitized for any credentials but showing the business-level values (invoice number, vendor ID, amount, etc.)
  • Output payload: what came back from the tool, including any error or partial result
  • Status: complete, pending, flagged for review, or failed, with a reason string when not complete

The status field is where most of the interesting work happens. "Complete" is straightforward. "Flagged for review" is what we spent the most time designing.

How Flagging Works

An agent step gets flagged when it crosses a threshold that the workflow designer set, or when the agent encounters an ambiguous match it was not configured to resolve autonomously. Consider a concrete example from our invoice processing workflow: when the agent matches an invoice to a purchase order, it looks for an exact vendor ID match and a total amount within a configurable tolerance window (say, 2%). If the amount is outside that window, the step is flagged, not failed and not silently approved.

The flagged state means the step enters a review queue. The agent pauses on that step, and anyone with reviewer access on the workspace sees it in their inspection panel. The reviewer sees the exact values that caused the flag: the invoice amount, the PO amount, the variance percentage, and the raw response from the matching function. They can approve with a note, reject with a note, or escalate.

We are not saying flagging is the same as human approval gates on every transaction. For most steps in most workflows, the agent should run to completion without pausing. The flagging threshold exists so that the cases where a human judgment call is genuinely needed get routed there, while the routine cases go through without interruption. The ratio in a well-tuned workflow should be well under 5% flagged, closer to 1-2%.

The Inspection Panel Layout

The inspection panel is built for ops managers, not for engineers reading API logs. We made specific choices about what to show and what to hide.

Each run trace opens as a vertical timeline. Steps are shown in order, numbered from 1. Completed steps show a brief summary line: tool name, key output value, timestamp. The full input and output payloads are collapsed by default and expand on click. This keeps the main view readable even for a workflow with 12 steps.

Flagged steps are visually distinct: amber background, a flag icon, and the flag reason shown inline without needing to expand. The idea is that an ops manager can scan the timeline in a few seconds to see where the run paused, without having to read every detail of every completed step.

Failed steps show in red with the full error message and the last known state of the step's inputs. This matters for debugging: if the SAP API returns a 403 because a credential rotated, the reviewer can see that the call was made correctly with the right data, and the failure was environmental, not a logic error in the workflow definition.

Rollback and What It Can and Cannot Do

The run trace powers the rollback feature, but rollback has clear limits. For read operations, there is nothing to undo. For write operations, MicroAGI records the full payload of what was written, so a reviewer can see exactly what was posted. Whether that write can be reversed depends entirely on whether the target system supports it.

For systems like Slack or Gmail, "rollback" in practice means flagging the run as a mistake and queuing a corrective action (delete the message, send a follow-up). For an ERP ledger post, it depends on whether the ledger entry is still in a reversible period state. We do not make guarantees that a rollback will succeed end-to-end, and we are explicit about this in the product. The run trace gives you the information you need to make the reversal; it cannot always execute the reversal for you.

This is an honest limit. Any tool that claims it can fully roll back an arbitrary sequence of distributed writes across multiple production systems is overstating what is technically possible.

Retention and Access Control

Run traces are stored at the workspace level. On the Starter plan, traces are retained for 7 days. On the Team plan, 90 days. Enterprise retains indefinitely, subject to the customer's own data retention configuration.

Access to traces is governed by the workspace role system. Viewers can read traces. Reviewers can approve or reject flagged steps. Admins can delete traces within the retention window (subject to an audit log of the deletion). The trace itself cannot be edited: the record of what the agent did is append-only.

For teams operating under audit requirements, whether that is SOX controls for finance workflows or internal change management processes, the append-only trace is the key property. It means the trace is a reliable record, not a mutable log that someone could retroactively alter.

How We Built the Trace Storage

We kept the trace storage simple by design. Each step's record is an immutable document written at the time the step executes. We do not reconstruct the trace from distributed event logs after the fact. The step record is written synchronously before the agent proceeds to the next step, which means if the agent crashes mid-run, the trace up to the last completed step is intact.

The tradeoff is that writing step records synchronously adds a small latency to each step. For back-office workflows that are not real-time (invoice processing, vendor onboarding, expense reconciliation), this is not a meaningful cost. For workflows that need sub-second response times at scale, it would be a different conversation.

We made this tradeoff deliberately: back-office ops automation is about correctness and traceability, not raw throughput. An agent that processes 50 invoices per batch with full per-step traces is more useful to an ops team than one that processes 500 per batch with only a pass/fail outcome.

What the Inspector Does Not Show

A few things are explicitly excluded from the trace by design. LLM reasoning chains are not stored. If a step uses a language model to parse an unstructured field (like extracting a vendor name from a free-text email), the input and output of that call are logged (the raw email text, the extracted vendor name), but not the intermediate token-by-token reasoning. The output is what the workflow acted on, and that is what the trace records.

We also do not log credential values. Tool call inputs are sanitized to show the business-level parameters while masking any authentication tokens. This is standard, but worth stating explicitly.

The Inspector is a tool for ops managers to understand what an agent did with their data in their business systems. It is not a debugging console for the underlying model. Those are different audiences with different needs, and mixing them in one view would make both views worse.

See the Inspector in action

Request early access and we will walk you through a live run trace from your first workflow.

More from the blog