Skip to main content

Evaluation Rollout Guide

The Evaluation tier exists for the moment when Community is no longer enough to answer the real question. The real question is not "does AxonFlow compile and run?" It is "can this become part of a production-grade operating model?"

This guide is for the engineer or platform owner running that next step.

Use the paid production path when the decision is already sponsored

Evaluation is deliberately self-serve. If you have a dated audit, security approval, incident, board commitment, or production deadline; written control requirements; and an executive sponsor who needs a production decision, use the 60- or 75-day Paid Production Program instead. It establishes scope, success criteria, conversion pricing, and the procurement path before production access begins.

What Evaluation Is For​

Evaluation is best used to validate the controls that usually matter just before production:

  • policy simulation before rollout
  • evidence export for governance review
  • larger policy, execution, and provider limits
  • more realistic workload and reviewer behavior

The current evaluation tier gives you:

  • 50 tenant policies
  • 5 organization policies
  • 5 custom policy connectors
  • 3 LLM providers
  • 14-day audit retention
  • evidence export up to 5000 records and 3 exports per day
  • 300/day policy simulation cap
  • 50 inputs/run impact report limit

Not included: approval queues. Creating HITL approval entries requires Professional or above. A require_approval policy on Evaluation still holds the step and reports approval_enqueue: "tier_disabled", and if you are upgrading from a release in which Evaluation did create entries, the ones you already hold stay approvable and still expire - see HITL Approval Gates.

See Community vs Evaluation vs Enterprise for the complete limit profile.

That is enough to run a meaningful internal pilot rather than only a developer proof of concept.

Ready to start? Request an Evaluation License

Keep The Evaluation Bounded​

The best evaluations feel serious without becoming risky. In practice that usually means:

  • one or two real workflows, not a platform-wide rollout
  • self-hosted deployment first
  • explicit telemetry and provider choices
  • clear success criteria for engineering, security, and reviewers
  • a rollback path if the evaluation is stopped

If your team is doing this in healthcare, banking, government, or another high-scrutiny environment, use Assessing AxonFlow in Regulated Environments as the companion guide.

Pick The Right Evaluation Scope​

A strong evaluation scope usually has all three of these:

  1. one real application or workflow that matters
  2. one workflow path with meaningful governance or approval risk
  3. one stakeholder beyond the core engineering team

Weak evaluations are usually too small. They prove that the platform starts, but they do not prove that the organization can operate it.

Good examples:

  • a customer-support assistant with governed connector access
  • an internal research assistant with redaction and evidence requirements
  • a multi-step workflow that requires review before execution of risky actions

What To Capture During The Evaluation​

StakeholderEvidence to collectUseful docs
Application engineersRequest path, SDK or framework changes, latency, failure behaviorRuntime Request Paths, Failure Modes
Security reviewersData boundary, provider path, telemetry setting, policy hierarchyRegulated Environment Assessment, Policy Hierarchy
Governance or complianceApproval decisions, audit records, evidence exports, retention expectationsHITL Approval Gates, Evidence Export
Platform ownersDeployment shape, capacity signals, rollback path, ownership modelDeployment Mode Matrix, Capacity Planning

Phase 1: Prove Technical Fit​

Use the first phase to answer:

  • does the SDK integration fit the app architecture?
  • do the right connectors and providers exist?
  • do policies catch the right classes of risky behavior?

This is where Community To Enterprise Migration and Deployment Mode Matrix are most useful.

Phase 2: Prove Operational Fit​

This is the real evaluation phase. Validate:

  • simulation and evidence-export behaviour on real traffic
  • policy simulation before policy rollouts
  • execution visibility and incident handling
  • evidence export for internal governance review
  • whether the current limits are enough for the intended pilot

This is where pages like Human-in-the-Loop and Execution Viewer matter.

Phase 3: Prove Organizational Fit​

Use the final phase to answer:

  • would security sign off on this rollout model?
  • can reviewers use it without engineering babysitting everything?
  • does the pilot already point toward identity, portal workflows, or enterprise connectors?

That is the phase where the enterprise decision usually becomes obvious.

Exit Criteria For A Good Evaluation​

Before you call the evaluation successful, you should have answers to:

  1. which workflows deserve approval gates?
  2. which policies need simulation before rollout?
  3. how will reviewers inspect, replay, and export executions?
  4. what are the first scale or governance limits you are likely to hit?
  5. is Evaluation enough for the intended production pilot, or is Enterprise the realistic landing zone?

If those answers are still fuzzy, the evaluation probably measured developer excitement more than platform fit.

Signals That Evaluation Should Turn Into Enterprise​

The strongest signals are:

  • several teams want to share the platform
  • non-engineers need approval or portal workflows
  • SSO or SCIM becomes mandatory
  • security, procurement, or compliance wants stronger operational evidence
  • enterprise connectors or provider management become part of the plan

That is when Enterprise Overview and Enterprise Rollout Checklist become more relevant than one more pilot iteration.

If the workflow needs founder-led rollout support and a fixed sponsor decision, review the Paid Production Program rather than renewing an unbounded evaluation.

Operational Readiness Checklist​

Before relying on this page in a production rollout, pair it with the core operations docs: