Skip to main content

Advisory Risk Scoring for Agentic Payments

The Fraud & Risk Add-on for Agentic Payments includes an ML risk score that augments the deterministic controls of the FinCrime Policy Pack. It is advisory by design: a high score holds a transaction for a person to review, and nothing the model says can block a transaction. This page states exactly what scoring does, what each decision records about it, and what its published numbers do and do not mean.

Enterprise add-on

Advisory risk scoring ships with the Fraud & Risk Add-on for Agentic Payments, a separately priced add-on to AxonFlow Enterprise. It is not included in the base Enterprise license, the Evaluation tier, or the source-available Community edition. This page describes platform v11.1.0, the release in which the score reaches the v11 decision engine. The add-on is currently available through the early-access program.

The score is a fact, and the policy decides​

The scoring service returns a number. It does not return a decision that AxonFlow acts on: the platform reads only the score, and ignores the threshold and above-threshold flag the scoring service reports about itself.

That number reaches the v11 decision engine as a fact about the request, and one control in the FinCrime pack reads it:

Controlpack:fincrime:fincrime__ml__risk__stepup
KindA mandatory requirement carrying the approval_challenge obligation. It has no constraint, so it can never deny.
Applies whenThe score is at or above 0.011591929942369461
EffectThe request is held for approval (see Where the score is read)
When there is no scoreThe control does not apply. The request continues, and the decision records why there was no score

There is no path where the model alone denies or approves a transaction. A score can hold a transaction for review; only the pack's deterministic constraints deny, and they apply regardless of the score.

Where the score is read​

The score is requested on the two scopes where the pack applies, and only after the identity plane has admitted the caller. The scoring service is sent the admitted principal:

  • POST /api/v1/decide, from context.fincrime_transaction.
  • The MCP request pass: POST /api/v1/mcp/check-input, POST /mcp/resources/query, POST /mcp/tools/execute, and the MCP server's check_policy tool, from parameters.fincrime_transaction.

A request that declares no fincrime_transaction is not scored, and the scoring service is not called for it. A deployment without the pack installed never calls the scoring service.

At or above the threshold, on an Enterprise deployment licensed for human approval, the call is held as a pending approval. Decide answers verdict: "needs_approval", the MCP routes answer HTTP 403, and both carry a pending_approval object. A person approves or rejects it, and the caller retries the same call naming the approval. The full contract, including the retry and its conditions, is in Pending approval on MCP and decide.

Your enforcement point must be able to carry out the hold

An enforcement point that sends a PEP capability handshake declares the obligations it can carry out. If its handshake does not advertise approval_challenge, a score above the threshold cannot be held for it. The call is refused with unsupported_obligation, which Enterprise explains as pep_capability_unsupported, rather than held. To receive holds, advertise approval_challenge in the handshake. An Enterprise request with no handshake header is held, because decide and the MCP planes carry the capability by default. See the PEP capability handshake for the X-Axonflow-PEP-Handshake header.

Degradation is visible, never silent​

A missing score never blocks a request. If the scoring service is not configured, unreachable, slower than its budget (100 milliseconds by default, configurable), refuses the agent's credentials, or answers with something the agent cannot read, the control does not apply and the decision proceeds on the deterministic controls alone. The record states why there was no score, so a reviewer can always distinguish "scored and below threshold" from "not scored at all".

What a decision records​

On every decision where the score control is active, the audit record carries policy_details.fincrime_risk_score. With the pack installed, that is every decision on decide and on the MCP request pass. A reviewer reads it to answer one question from the row alone: was this transaction scored, and if not, why?

status is always one of seven values:

statusMeaning
scoredThe scoring service answered with a valid score
unconfiguredThis deployment configures no scoring service
no_transactionThe request declares no fincrime_transaction, so there was nothing to score and the service was not called
timeoutNo answer within the budget
auth_rejectedThe scoring service refused the agent's credentials (HTTP 401 or 403): the internal service secret differs between them
unavailableThe service was unreachable, or answered with any other non-200 status, including a rejection of the request itself
malformed_responseThe answer could not be used: a missing model identity, a score that is missing or outside 0 to 1, an out-of-range feature coverage, a feature without a numeric contribution, or a body larger than 1 MiB

A scored record also carries:

  • score: the overall score, present only when status is scored. A failure is never recorded as a number;
  • signal: the attribute the score was stated at (signal.scorer.fincrime__fraud);
  • model_id and model_version;
  • feature_coverage: the fraction of the fields this model version can use that the request supplied;
  • top_features: the top contributing features, each with feature, value, contribution and direction (increases_risk or decreases_risk);
  • controls: every control that read the score on this decision, with the threshold it compared against. When your organization has moved the threshold, this shows the value that actually applied.

So a reviewer sees why a transaction was held, not just that it was, and the record lands in the same audit surface as every other decision; see audit logging and decision explainability.

Two reading notes. Feature contributions are SHAP values in the model's margin (log-odds) space, not percentages of risk; treat a near-zero magnitude as neutral regardless of its direction label. A feature's value can be the empty string when the field was not supplied, which is itself an explanation: a score can rise because something is missing.

Feature coverage is the fraction of the fields this model version can actually consume that your request supplied with usable values. It is declared by the model bundle, so it cannot drift from the model. Use it to weigh score confidence: a low-coverage score was computed from less evidence.

Threshold and operating point​

The pack's threshold, 0.011591929942369461, was chosen to flag at most 1% of transactions on the Sparkov validation partition (model 0.1.0, chronological split). That is a property of that simulated corpus, not of your traffic. The model has no agent-traffic training data, so the 1% review rate will not hold on your transactions, and we publish no accuracy claims for them. Treat the threshold as a starting point, and choose your own from your review capacity using the review-rate table below.

To move the threshold, up or down, publish a copy of the control under its own id (pack:fincrime:fincrime__ml__risk__stepup) in your organization's typed policy document, with your threshold as its comparison literal. A policy under the pack's id replaces the pack's copy for your organization. A requirement under any other id cannot move the pack's threshold: it can only add a second hold beside it. See tuning a control for your organization.

The scoring service's own AXONFLOW_FINCRIME_SCORER_THRESHOLD setting does not move this threshold. It changes only what the scoring service reports about itself, which AxonFlow does not read.

The model, honestly​

The v0 scoring model (fincrime-fraud, version 0.1.0) is a gradient-boosted tree ensemble trained in house from a published method on a public dataset: the Sparkov Credit Card Transactions Fraud Detection corpus (CC0 1.0 Public Domain; 1,852,394 simulated consumer card transactions, 0.52 percent fraudulent). Every number below is measured on that corpus.

No number on this page is an agent-traffic claim

Every metric below was measured on the Sparkov corpus, which is simulated consumer card transactions. None of it was measured on agent traffic, and none of it predicts what this model will do on agent traffic. No public labeled corpus of agent-initiated transaction fraud exists yet, and we have no agent-traffic training data either. That gap is why scoring is advisory, why feature coverage is recorded per score, and why calibration against real traffic precedes any accuracy claim. We publish no detection-rate or accuracy claims for agent traffic.

Headline metrics, with the label that keeps them honest​

Measured on the Sparkov corpus, on both a chronological split (train on the past, test on the future; the setting production resembles) and a random split (reported for comparison only, because it lets the model interpolate across transactions it has effectively seen):

Split (Sparkov corpus)PartitionRowsFraud rowsPR-AUCROC-AUCFraud precisionFraud recallReview rateAccuracy
chronologicaltest370,4791,3490.46780.968022.06%61.16%1.01%99.07%
chronologicalvalidation185,2397960.41430.960423.27%54.15%1.00%99.04%
randomtest370,4791,9300.92310.998548.15%94.87%1.03%99.44%
randomvalidation185,2399650.90670.998448.81%93.68%1.00%99.46%

The shipped model is the chronological one. A random split roughly doubles this model's apparent PR-AUC (0.9231 against 0.4678 on the Sparkov test partitions), which is exactly why the chronological numbers are the ones to plan around.

Why the accuracy column is nearly meaningless here​

On a corpus that is 0.52 percent fraud, accuracy measures the majority class and almost nothing else:

Sparkov test partitionsModel accuracyBaseline accuracyModel fraud recallBaseline fraud recall
chronological split99.07%99.64%61.16%0.00%
random split99.44%99.48%94.87%0.00%

The baseline is the model that predicts "not fraud" always. On the chronological split it beats the shipped model on accuracy while catching nothing. That is the entire argument for reading the minority-class columns (fraud precision and recall) first, and for distrusting any fraud product's unlabeled accuracy headline.

The review-rate trade-off​

The threshold is a capacity dial. On the Sparkov chronological test partition (370,479 rows, 1,349 fraud), different target review rates buy:

Target review rateFlaggedFraud caughtFraud precisionFraud recall
0.10%37031384.59%23.20%
0.50%1,85269337.42%51.37%
1.00%3,70482522.27%61.16%
2.00%7,40996212.98%71.31%
5.00%18,5231,1266.08%83.47%

The trade-off is steep and worth stating to anyone who will operate a review queue: at a 0.1 percent review rate the queue is 85 percent fraud but three quarters of fraud is missed; at 5 percent almost everything is caught and nineteen of every twenty reviews is a false positive.

What the v0 model actually consumes​

The v0 model's consumable-field list, declared in the model bundle and used as the feature-coverage denominator, is four fields: amount, timestamp, and the two cohort aggregates (historical_mean_amount, txn_frequency_1h). On traffic conforming to the documented context formats, the v0 model's discriminative signal comes from those four fields only: its geography and merchant features were fitted on the training corpus's own location and merchant vocabulary, which does not overlap the ISO 3166 and ISO 18245 formats the context guide specifies, so conforming values in those fields are accepted but carry no learned signal. The metrics above were measured on corpus-format inputs where those features are live; no number on this page quantifies the reduced conforming-traffic setting. A retrain aligning the geography and merchant vocabularies with the context formats is planned.

Training integrity​

The training pipeline deliberately deviates from its published source method where the published construction inflates results: aggregate features and categorical encodings are fitted on the training partition only (fitting them on the whole corpus, as published, roughly doubles apparent PR-AUC on this corpus: 0.9403 against 0.4678 on the Sparkov chronological test partition), model selection never touches the test partition, and the shipped chronological headline is reported next to the flattering random split rather than instead of it. Columns carrying personal attributes in the training CSVs (names, addresses, birth dates, and similar) are never loaded; every model feature is a pure function of the documented wire contract, which has no field that could carry them.

Data handling and performance​

  • The scoring service runs inside your deployment and is called by the agent over an internal HTTP contract authenticated with the internal service secret. Transaction context is not sent to any third-party service; see the enablement page for the wiring.
  • The agent sends the scoring service the decision id, the scope (decide or mcp), the fincrime_transaction and fincrime_cohort objects as the caller sent them, and the admitted principal's id and type. The request text, statement and prompt are never sent.
  • Server-side inference latency for the shipped v0 model measured 4.49 milliseconds at p99 (3,000-request benchmark of the scoring service's inference path, Apple M4 Pro), against the 100 millisecond default budget on the decision path. Network transport between the agent and the service is the deployment's share of that budget.