Most back-office automation pilots do not fail because the agent does not work. They stall because the path from "it works in a demo" to "we have authorization to run it in production" involves a series of decisions, approvals, and tests that no one explicitly planned for, and that become blockers one by one until the project loses momentum.
This post is the checklist we use with teams going through that path. It is not a technical deployment guide; it is a decision and approval checklist, based on the blockers we have seen teams hit, in roughly the order they tend to appear.
Before You Start: Define What "Production" Means for Your Workflow
The most common source of misaligned expectations on back-office automation projects is that different stakeholders have different definitions of "production-ready." For the IT team, production means the system is stable, the credentials are provisioned, and the error handling is in place. For the ops manager, production means they trust the output enough to not check every result manually. For the finance director or compliance team, production means the audit trail is in place and the approval controls meet their requirements.
Before any approval conversations, write down a single shared definition of what production means for your specific workflow. It should cover: what the agent does and does not do, what the approval gate configuration is and who is in the reviewer role, what the audit trail looks like and how long records are retained, what the escalation path is when a step fails, and what the rollback procedure is if a run needs to be reversed. This definition becomes the reference for every subsequent conversation about whether the system is ready.
The IT/Security Sign-Off
Before going live in production, IT and information security need to evaluate the integration. The key questions they will typically ask: How are credentials managed? (They should be stored in the MicroAGI workspace secrets vault, not hardcoded in workflow definitions.) What permissions does the service account have? (Least privilege: the agent should have only the permissions it needs for the specific workflow, not broad ERP admin access.) Where does data go? (MicroAGI runs in EU infrastructure; if your security policy requires data residency documentation, provide the DPA and infrastructure details.) How is access logged? (Every authentication event and API call is logged in the audit trail.)
Prepare these answers before the IT review meeting, not in response to questions at the meeting. IT and security sign-off is a process, not a conversation. The review goes faster when the reviewer has documentation to review rather than having to extract information from the project team in real time.
The Data Privacy Review
For any workflow that processes personal data (employee records, vendor contact information, customer data), a data privacy review is required before production deployment. This is not optional and it is not fast to do at the last minute. In EU organizations, this typically involves reviewing and documenting the processing basis under GDPR, updating the Record of Processing Activities, and confirming the DPA with MicroAGI GmbH if not already in place.
For teams in Germany or the broader EU, allocate 2-3 weeks for this process if it has not been started. The privacy review is the most common cause of late-stage delays in production authorization. It does not block the technical work, but it blocks the go-live authorization, and it is hard to compress.
The Ops Manager Walkthrough
The ops manager who will own the running workflow in production needs to see the actual approval and review interface before the workflow goes live. Not a demo of what the interface looks like in general: a walkthrough using real data from the pilot, showing a run trace from a workflow instance that was flagged, the decision the reviewer made, and the outcome.
This walkthrough has two purposes. The first is practical: the ops manager needs to know how to use the review queue, how to read a run trace, how to approve or reject a flagged step, and what happens when they reject one. The second is about authorization: an ops manager who does not understand the review interface will not sign off on a production deployment, and they should not. Automating a process that the responsible manager cannot inspect or intervene in is exactly the scenario that erodes trust in automation programs.
The Test Run Protocol
Before going live in production mode, run the workflow against production-grade data in a review-only mode. Review-only mode means every step is presented for human approval before executing, regardless of whether it would normally pass the automatic threshold. This is not a staging environment test; it is a live-data test where a human reviews every result.
Run at least 30-50 instances in review-only mode. This is not a magic number, but it is enough to see the full distribution of inputs including several edge cases, and enough for the reviewers to build intuition about what the agent's decisions look like. After the review-only run, analyze the results: what percentage of decisions were correct without modification, what percentage required corrections, and what were the correction patterns? This analysis informs the final threshold calibration.
Define pass/fail criteria for the test run before you start, not after. Common criteria: automated decision accuracy above a target threshold (typically 92-96% for invoice matching, higher for access control workflows), no instances where an error would have posted to production without the review-only gate catching it, reviewer feedback that the decision context presented in each flagged step is sufficient to make an informed judgment. If any criterion fails, identify the root cause and fix before requesting production authorization.
The Production Authorization Conversation
With the IT sign-off, data privacy review, ops manager walkthrough, and test run results in hand, you have the evidence package for a production authorization conversation. This should be a meeting, not a document: bring the ops manager, the IT lead, and whoever owns the process from a compliance perspective into the same room (physical or virtual).
Present the test run results, the workflow definition with the approval gate configuration, the audit trail design, the rollback procedure, and the monitoring plan for the first two weeks of production. Answer the question "what do we do if something goes wrong?" with a specific, written procedure rather than "we'll handle it."
Authorization at this stage is typically straightforward if the groundwork has been done. The delays happen when any of the above elements are missing and have to be prepared reactively. The approval conversation becomes a discovery session rather than a confirmation meeting, and discovery sessions do not end with authorization.
The First Two Weeks of Production
Going live is not the end of the project. The first two weeks of production are the highest-value monitoring period. Flag everything that looks unexpected, even if the agent handled it correctly. Monitor the flag rate daily and compare it to the expected rate from the test run. If the production flag rate is significantly higher than the test run rate, the production data distribution is different from what you tested, and the thresholds need adjustment.
Schedule a review meeting at the end of week two: go through the flag log, identify recurring patterns in the flagged cases, and make threshold adjustments based on what you found. This meeting is also the time to identify any edge cases that the test run did not surface and to decide whether they need additional handling in the workflow definition.
After week two, the workflow should be running stably and the flag rate should be close to the expected steady-state level. At that point, the production go-live is complete. The project transitions from "deployment mode" to "operational maintenance," and the team's attention can shift to the next workflow on the automation roadmap.