01 One Loop · Team WGG, National University of Singapore · AMEX AI Hackathon 2026 · Round 1 · A 29-page proposal, pages 1 to 22 and 84 to 90; pages 23 to 83 are evidence you may skip

One loop and one model for three questions: who to sign, what to offer, and where growth comes next.

Executive summary

One Loop is one transaction backbone serving three growth heads, one for each question above, entered under Amex's Growth theme, with a measurement layer that puts a behind every partner campaign so the lift can be certified rather than asserted. The operator is the Singapore Decision Science CoE; the customer is a GNS partner bank in APAC; what changes for them is that a signing list, an offer campaign and a corridor forecast each arrive with a number someone can check.

The scarce resource in payments AI is a number that still stands after an audit, rather than another model. This report presents such a number. 27 comparisons in this document carry an interval, one row per metric, depth, seed and arm the pre-registrations named: 7 cleared zero, 7 landed on the wrong side of it and 13 straddle it, all drawn at one size in Section 5, the wrong side including our own offers exhibit's primary endpoint. Most evaluations stop before that picture exists; the apparatus that produces it is the part of this we built. The revenue it builds toward is modelled, not measured: US$10.5M a year by year three of an APAC rollout, across eight partner markets, counting only discount revenue on , the base case of the three scenarios Section 2 prints. Section 2 builds it bottom-up and then shows what is left of it when the lanes that lost are struck out.

Watch the film: One Loop in ten minutes The system, the evidence and the losses, animated · 10 min 31 s Play

Opens the YouTube player in place. Also at youtu.be/g0ptMP5NRPo.

Listen to this project as a podcast Two voices discuss the whole entry, including the evidence · 22 min Play
Real data 5.9x the average visit effect, concentrated into the top tenth of a randomized real holdout of 4,193,878 rows

In the unit a budget owner uses: targeting a tenth of customers this way captures 59.3 percent of every incremental visit the campaign produced, against 48.6 percent for response ranking at the same reach. That is this tile's own figure at one tenth of the reach, not a second measurement. The version that carries a : against response ranking at that depth, this ranking adds 0.011 per customer targeted, interval [0.009, 0.014], entirely above zero.

What travels with it The ratio is point estimates over point estimates and carries no interval of its own. Most of that concentration is what response ranking already reaches alone, and the is what the model adds. The corpus is an advertising experiment. Visit was the robustness endpoint, and on conversion, the pre-registered primary, this same ranking lost at every depth, which Section 6 prints at the same size.

See the measurement, Section 6
Synthetic corpus 0.235 to 0.257 of next spending categories called correctly, about one in four, before and then with the backbone, under the strictest of our four protocols

This is roughly two more correct answers in every hundred, on an interval above zero. This tile is included to show the number we could have printed instead. Scored the permissive way, the same frozen model on the same task shows a gain of 0.0590 rather than 0.0218, more than two and a half times as much. The guards, not the model, account for the difference, and the smaller number is the one we ship.

See the ladder, Section 5
Real-data facts GenAI, gated Three layers on every generated sentence: a numeric match, a cross-examination, and a direction check that fails 5 of the 10 bundles it can read and refuses to pass the other 20

Every ranked list and lift report leaves with a narrative written by a self-hosted model over the committed numbers, never over free text. The five the direction check fails say the model beat the while their own facts say the naive was ahead: every numeral was correct and only the direction was wrong, which is the one thing the first two layers never checked. This layer explains decisions and never makes one, by choice; the model that does feed a decision is the pretrained transformer whose per-field likelihoods score fraud in Section 5.

What travels with it The cause was our fact bundle, not the model. It stated a false comparison rule, and layer two could not catch that because it is the same model reading the same bundle. The rule is fixed; the text is not regenerated, which needs a cluster GPU. Our red team also beats this gate on two of seven attack classes, fixed and re-measured in Section 6.

Step through the replay console, Section 6

Self-assessment of the evidence

  • The question, and the answer The question is whether adding the model to the cheap control improves that control, rather than whether the model beats the control on its own. On protection, on synthetic TabFormer, adding it more than doubled the area, and that combination was built after the first run rather than pre-registered, so it is counted in none of the 27. On the same synthetic corpus it added again on fraud, but only at the pretraining scales where the counting baseline still has room and not at the full corpus we ship, and on . On real corpora it added on the store head and on the corridor forecast; each of those gains that carries an interval clears zero, and the corridor gain is a point comparison that says so.
  • What did not hold, reported in full Alone, the model still loses to the counting control on protection, and on corridors it is worse than the it is added to, so the line above is the blend winning and not the model winning. That blend win is scale-free too: totalled in arrivals the naive is ahead, and Section 7 prints both. On the real corpus the embeddings subtracted from the baseline instead of adding to it, the ladder's own argument on a second axis: a method that only travels when the corpus flatters it is not a method. The added nothing over plain density on its forward check, though a post hoc re-read, counted in nothing, finds its other channels order formation inside bands of comparable density. The offers head is mixed, reported as mixed and not as the better seed. The two-sided leg itself, the merchant-axis runs and the pair-retention task built to test it, is a null so far, so the closed-loop advantage is argued, not measured. And our own red team still beats our generated-text gate on two of seven attack classes, found by us, fixed, and replayed.
  • What the evaluation protocol contributes One frozen model read under four evaluation protocols, guards turned on one step at a time. The permissive number is far larger than the one we ship, and the guards, not the model, explain the gap. Section 5 draws every interval behind these lines against one zero line.

The full scoreboard, every interval drawn against one zero line, is in Section 5.

If you read nothing else

The product is summarized in two rows of a randomized experiment, both on the visit endpoint. The segment a response model ranks 26 of 28, so it never gets mailed, is the one the randomized arms measure moving most: visit uplift 5.82 points, ninety-five percent interval [3.96, 7.67], clear of zero. The segment it ranks 5 of 28, and mails, measures 2.31 points, interval [-1.25, 5.87], which cannot be told from zero, and its spend uplift spans zero too. The budget therefore goes to the row whose measurement cannot be distinguished from zero, and skips the row whose measurement can. Both intervals are the normal approximation on the stored standard error. Nothing in a response score sees either fact; only a randomized holdout does, and the first tile above is the same mechanism at 4,193,878 rows. It was built for about US$28 of compute at cited public rates, cash US$0, and stop three below describes how to check any number on this page.

  • Stop one, the proposal The idea, the value and the architecture are Sections 1 to 4. The business actions, the plan and the plain list of what we have not shown are Section 8. Sections 5 to 7 are the evidence behind them, kept because it can be checked, and none of it is required reading; each section can be consulted independently when a reader wants to verify a claim. In the supplied NUS_WGG_OneLoop.pdf the proposal is 29 pages: pages 1 to 22 and 84 to 90. Pages 23 to 83 are the evidence, and the last 6 are the numbered sources. A build check asserts every one of those boundaries against this PDF each time it is rendered, so the promise is measured rather than asserted.
  • Stop two, the scoreboard The three lines directly above: what won, what lost, and what the turned out to be worth.
  • Stop three, the audit Press Show sources at the top of the screen, a screen feature, then click any measured figure, amber: the file it was read from opens at the line that number sits on, inside this page. Blue marks are declared rather than measured, and they open a panel that says so.You are reading the printed copy, so the audit is not in your hands here. Open NUS_WGG_OneLoop.html in any browser, press Show sources, and click any measured figure, amber: the committed file it was read from opens at the line that number sits on, inside the page, with nothing fetched and nothing installed. Blue marks are declared rather than measured, and open a panel that says so. Uplift measurement in that same bar is the result named above, printed beside the outcome where the same ranking loses.
This table reproduces the audit in printed form. On screen these figures open the committed file at the line shown; the rows below are the same rows a click would produce, so a printed copy carries the mechanism rather than only a description of it. Every column is generated from the same index the interactive drawer reads, so this table cannot drift from what a click shows. The checker behind it is attacked too: six committed tamper cases must each turn the build red while an untouched page stays green. The same table at width, one headline figure from every exhibit, opens the evidence block in Section 5.
FigureCommitted filePath inside itValue as storedLineFile checksum, leading characters
Offer targeting, paired difference at the top decileresults/​uplift.json/​criteo/​targeting_​at_​k/​outcomes/​visit/​k/​10/​difference_​cate_​minus_​response/​value0.01104852282e6322028c7aaabf
Fraud ranking, model added to the counting control, PR-AUCresults/​protection.json/​comparisons_​by_​key/​pll_​plus_​rarity_​vs_​rarity_​all_​fields/​pr_​auc/​difference0.201423498c9735e2ceace5bb6
Corridor forecast, the pre-registered equal-weight combination, macro MASEresults/​corridor_combination.json/​macro/​mase_​combination0.50608190cce3eac9ab789275

The unfavourable results are reported deliberately, because the value of a decision product depends on the reliability of its numbers, including in the cases where a result shows none.

Scroll down for the full report

02Problem and ValueApplicability at American ExpressGrowth, primary

Only the smallest lane has a measurement under it

Amex's CEO named coverage as the constraint. Three lanes model US$10.5M a year by year three at base case; strike the lanes that lost their tests and the measured offers lane, the smallest, remains.

US$10.5Ma year by year three, three lanes, base casemodelled from declared assumptions, not measured; only the offers lane's mechanism is measured, randomized

Coverage is the CEO's named constraint; the CoE operates it for a GNS partner bank; the Phase 1 bar is priced on the one lane a randomized measurement stands under.

Results that did not hold and are still reported

  • Whitespace, the largest lane at base case, lost its forward check to plain venue density; venue formation is not merchant signing, so this settles the bar must beat in the pilot.
  • 0.623 model against 0.53 for a per-corridor , pooled; it wins on five of the twelve corridors, so this head sells coherence and per-corridor explanation, not pooled accuracy.
  • The corridor lane applies only on the five of twelve corridors where the forecast beats the naive; its dollar figure states that restriction rather than being scaled by it.
  • 0.3% QRIS against 1.5% to 3% card rates: merchants decline Amex on fee, so only the fee-tolerant stratum of the non-accepting universe is signable at a card rate.
  • US$10.5M at base case falls to under half when the whitespace lane is struck, and to the offers lane alone, the smallest of the three, when only the measured mechanism is kept.
  • The struck whitespace lane is not rescued: a post hoc re-read in Section 7 is not and in no tally, a reason to run the Phase 1 gate, not to restore the lane.
  • US$0.4244 incremental spend per customer treated, randomized Hillstrom arms: a measured dollar figure from a different setting, retail e-commerce rather than billed business on a card; it establishes the counting rule but not the size.
  • The signing-conversion uplift, the input the model is least able to anchor, drives the largest lane at base case; no public source fixes it, so a reader should move that number first.
  • A one million dollar Phase 1 clears in months of offers-lane revenue at base case and in years at the conservative scenario; the lane's size input is still an assumption.
  • The retention payoff of holding GNS partners on the network is counted in no scenario and sized nowhere; the page gives only the threshold of billed business held that would match all three lanes.

The problem Amex names itself

0.3%QRIS charges merchants 0.3% against card rates of 1.5% to 3%; APAC is Amex's fastest-growing market family and smallest developed region, at $5.22 billion (7.2% of the company total) in FY2025 revenue.

On the Q4 2025 earnings call, Steve Squeri said Amex must "continue to build coverage, obviously, in international as that continues to grow and continues to be the fastest-growing overall part of our business"6. The gap is visible in Amex's own filing. APAC revenue was $5.22 billion (7.2% of the company total) in FY2025, the smallest developed region despite being the fastest-growing market family, and the card network business runs through partner relationships in approximately 110 countries and territories. Every figure in this paragraph is the 10-K's, not a third party's reading of it.

The value proposition

US$10.5M a year by year three, base case, three lanes, modelled; strike the lanes that lost and the measured offers lane remains.

Three scenarios, or one input at a time
Move the inputs
  • Conservative, all three lanesUS$0.52M
  • Base, all three lanesUS$10.5M
  • Stretch, all three lanesUS$108.5M

Modelled, US dollars a year, declared assumptions not measurement; each scenario moves every input together, which obscures the input that most deserves scrutiny. Whitespace, the largest lane, lost its forward check.

  • Base total, nothing moved10,480,000
  • Markets or partners live, swing15,720,000
  • Signing-conversion uplift, swing9,600,000
  • Corridor yield gain, swing800,000

One input moved across its own range, the rest held at base; the swing in US dollars a year. Partners live feeds all three lanes; a pilot measures the second.

Modelled, US dollars a year, declared assumptions not measurement; each scenario moves every input together, which obscures the input that most deserves scrutiny. Whitespace, the largest lane, lost its forward check.

Bound figures, read from the committed results file and checked by the numeral gate; nothing is computed at view time.

Offers alone clears Phase 1 in months, or years
  • One million, stretch0.5
  • One million, base4.2
  • Four million, base16.7
  • One million, conservative66.7

Months, modelled on the offers lane alone: size input assumed, mechanism measured on randomized data; a longer bar indicates a slower payback; the figures are the threshold to clear rather than an assumed budget, with the conservative scenario included.

Bound figures, read from the committed results file and checked by the numeral gate; nothing is computed at view time.

Per partner, only the measured lane survives the subtraction
  • All three lanes, per partner1,310,000
  • Offers lane alone, per partner360,000

Modelled, base case, US dollars per partner a year; whitespace struck after losing its pre-registered forward check, corridors lose pooled to a seasonal naive; only offers carries a randomized measurement.

These figures are bound to the committed results file and checked by the numeral gate; nothing is computed at page-view time.

US$2.88Moffers lane · base case · mechanism measured, dollars assumed

03Why AmexApplicability at American ExpressGrowth, primary

The recipe is public, and only Amex holds both sides of the transaction

Visa, Stripe, Nubank and Revolut have run one backbone with many heads in production; each sees one side of a transaction. Amex's closed loop joins cardmember, network and merchant in one event stream.

+111%improvement in abnormal-behavior detection, Visa TREASUREa production figure Visa published, cited here, not measured by us

Under Applicability, the data to run a proven recipe on both sides sits inside Amex's perimeter; Phase 1 needs only what the CoE holds. The advantage is argued, not measured.

Results that did not hold, reported here regardless

  • Null on the pair-retention task: both intervals span zero on synthetic data, so the merchant-axis leg is unmeasured and the closed-loop advantage is a structural argument, not a measured one.
  • Two earlier merchant-axis tasks were too degenerate to answer, so the merchant side has no measured result yet; Phase 1 is where it is tested on data that could settle the question.
  • American Express is not mentioned in SAFR's contributor list, where both open-loop networks appear. A contributor list is not a participation list, so we read this as a visible opening rather than an exclusion; it is a Round 2 direction and has not been built.
  • The row both sides joined in one event stream is the closed loop's structural argument, not a measured result; Phase 1 runs inside the perimeter on what the CoE already holds.

In 2025 and 2026 the payments industry reached a consensus on one question. One pretrained transaction backbone feeding many task heads beats per-task feature engineering, and the proof came with production numbers: Visa's TREASURE reports +111% on abnormal-behavior detection, Stripe reports card-testing detection moving from 59% to 97%, Nubank runs nuFormer in production for 100+ million customers, and Revolut's PRAGMA scaled the same idea to a billion parameters.

Why only Amex (capability matrix)

One row reads Yes, No, No: both sides joined in one event stream, and that row is about the proprietary markets.

Only the closed loop joins both sides

Public positions by row, cited where a source exists; the marked row is argued, not measured: the pair-retention task came back null on synthetic data, both intervals spanning zero.

These figures are bound to the committed results file and checked by the numeral gate; nothing is computed at page-view time.

CapabilityAmex closed loop Open-loop networksSingle banks and fintechs
Issuer-side view of the cardmemberYesPartialYes
Acquirer-side view of the merchantYesPartialPartial (acquiring banks hold it, one side only, never joined to an issuer view)
Both sides joined in one event streamYesNoNo
Premium cross-border cardmember baseYesMixedNo
Partner-bank distribution across APAC (GNS)17YesYesNo
Incrementality culture (a 2024 uplift contest run by this CoE)22YesUnpublishedRare
Decision-science home under MAS in Singapore30YesPartial (data-science hubs in Singapore, without a closed-loop decision-science charter)Some

The gap the regulator has already named

SAFR page 9 calls for agent-declared carts to be authenticated against their origin, the merchant side; a closed loop holds that side, and Amex is not a printed contributor.

04ArchitectureTechnical depth of the all-around system designProtection, demonstrated

One backbone with three heads, GenAI that explains but does not decide, and measured guards

One pretrained transaction backbone serves three growth heads and an always-on randomized measurement service; merchant-level wiring is the Phase 1 gate; GenAI explains on top of classical scores; data-rights guards are measured, not asserted.

0.9395Matthews coefficient, a linear probe reads the amount decile out of one embeddinglinear probe, public synthetic TabFormer, one-transaction merchant, point shown, interval in file

Amex's red line is designed in: models score, GenAI explains, a measures, and the embeddings that leak stay inside the perimeter; the control plane is designed, not built.

Results that did not hold, reported here regardless

  • 47.1% AML drop when cross-user graph structure is missing, published by Revolut's PRAGMA: the reason the graph layer is Phase 2 and drawn dashed, and not priced as free.
  • Matthews 0.9395: a linear probe on one merchant's single-transaction embedding recovers the amount decile on public synthetic TabFormer, so raw embeddings are an internal artifact, never a partner deliverable.
  • Cells of at least 10 venues is the shipped whitespace guard, priced by the ladder with the two stronger guards above it, and the PDPC's own caveat about transactional data attached.
  • The household offers head on real retail data is mixed, null in one seed and positive in the other; merchant-level wiring is demonstrated at small scale, not proven.
  • We have not checked, and cannot check here, whether the GNS partner agreements as written permit any of it. A data perimeter does not constitute a permission, and Phase 1 should not open before Amex legal answers.
  • Agentic volumes are small today: the ACE hook is a roadmap item, not a demo claim, with two of five ACE services still under development.
  • Nothing in the safety card measures American Express exposure; each result measures a mechanism on a public corpus.
  • Every control bounds the loss, none prevents it: a verified agent with a valid mandate can still act against a manipulated cardmember, and no public dataset carries authorized-scam labels.
  • Caps are evaded by decomposition, novelty is a model output that drifts and can be poisoned, and anything absent from the mandate schema, data scope above all, is unbounded.
  • Four threats assessed and dropped with the reason printed: model theft and single-stream training poisoning unmeasured, unbounded consumption covered by rate limits, prompt leakage moot since no decision depends on a prompt.
  • The control plane is designed and not built; the friction-only red line is built for generated text in Section 6 only, and origin authentication is not closed either.
  • No obligation row is a compliance claim: PDPC guidelines are not legally binding, the Veritas Toolkit's latest release is dated 2023 and not current tooling, MAS FEAT is principles, not a rulebook.
  • 0.9689 of member accounts recovered by counting alone against the model's 0.2724 at the same floor: a no-model attack wins, so the score adds no attack power over the population difference.
One Loop system architecture CLOSED-LOOP EVENT STREAM Issuer signals cardmember spend · engagement Network events authorizations · settlement Acquirer signals merchant acceptance · terms TRANSACTION BACKBONE dual-axis objectives · cardmember axis + merchant axis field-tokenized masked-LM pretraining · as-of entity embeddings Graph layer, dashed Phase 2 · PRAGMA-justified Whitespace head who to sign · merchant-embedding similarity + real public signals Offers / uplift head what to offer · incremental, not correlational, targeting Corridor head where growth comes next · cross-border corridor foresight ALWAYS-ON MEASUREMENT SERVICE ghost-ads holdouts productized · per-partner auditable lift reports GOVERNANCE & DATA RIGHTS MAS Veritas 2.0 · AIRG · SAFR · mapped per section 08 region sharding: PDPA · RBI localization · PIPL · partners consume scores and aggregates · embeddings never leave · GenAI explains, never decides credit
The architecture consists of one backbone, three heads, a measurement service, and built-in governance.
  1. Closed-loop event stream: issuer, network, and acquirer events in one sequence per entity.
  2. Pretrained transaction backbone: field-tokenized masked-LM over transaction sequences, with dual-axis objectives (a cardmember axis and a merchant axis). At Amex scale this becomes a transaction foundation model; our prototype is the same recipe at prototype scale.
  3. Graph layer, drawn dashed: Phase 2. Revolut's PRAGMA published a 47.1% drop on AML when cross-user graph structure is missing. That published failure is the reason to add a graph layer, and also the reason to account for its cost rather than treating it as free.
  4. Three task heads: merchant-signing whitespace, uplift-measured offers, corridor forecasting. The three heads consume the same embeddings. Wired on this page: the backbone reaches the whitespace head at category level, the offers and corridor heads are proven on their own public benchmarks, and merchant-level wiring is demonstrated at small scale on real retail data in Section 5, where the store head is positive in both seeds and the household offers head is mixed, null in one seed and positive in the other. Wired at Amex: all three heads consume merchant-level backbone embeddings on real closed-loop data, which is the Phase 1 gate.
  5. Always-on measurement service: every partner campaign ships with a by design; lift reports are per-partner and auditable38.
  6. Governance and data rights: sharded by region, mapped to MAS instruments, GenAI as explanation only. The operating half, retrain cadence, serving paths and the control plane, is in Section 8's lifecycle paragraph and Section 4.

Where GenAI sits

GenAI writes reason codes and partner narratives on top of classical scores; it never makes credit or approval decisions, Amex's own red line.

GenAI writes reason codes and partner-facing narratives on top of classical model scores. It never makes credit or approval decisions, which is Amex's own stated red line39, and it runs the way Amex already runs GenAI: a ConnectChain-style layer40 behind the AI firewall41, inside the GenAI Council's use-case funnel42.

Data rights, by design

10 venuesEmbeddings never leave Amex: a linear probe recovers amount decile at Matthews 0.9395 on synthetic TabFormer; whitespace cells carry at least 10 venues.

We have not checked, and cannot check, whether the GNS partner agreements as written permit any of it. A technical perimeter does not constitute contractual permission. Whether a partner's agreement covers training on Amex-owned inbound spend into their market, and scores flowing back, is a contracts question for Amex legal and the partnership owner, and Phase 1 should not open before they answer. A partner-facing product that has not asked is a demo, not a plan.

The agentic hook, stated honestly

Two of five ACE services are still in development and intent-to-settlement traceability is unsolved; our forecasts and embeddings are a substrate and a roadmap hook rather than a demonstration.

Amex's ACE kit defines an intent-intelligence service with two of five services under development43, and Amex's engineering blog names intent-to-settlement traceability as unsolved44. The corridor forecasts and merchant embeddings here are a natural substrate for it. Agentic volumes are small today, so this is a roadmap hook rather than a demonstration claim.

A defensible number has no value if the system that produced it leaks or can be steered, so the privacy work below applies the 's method to the machinery itself: one frozen artifact, with guards turned on one at a time and the cost of each recorded beside it. The shape is the R-U confidentiality map of statistical disclosure control45, priced one rung at a time on the product metric a partner acts on. Where a guard cost something, the cost is here; where a check did not run, we say so. Nothing here measures American Express exposure: each result measures a mechanism on a public corpus.

What these controls do not protect against, and what each obligation actually got

Every control bounds the loss, none prevents it; four threats dropped with reasons; obligation rows are built or designed, none a compliance claim.

Interlude · Safety, measured

Protection theme. Privacy and safety priced one guard at a time, on the output a partner actually receives, with the failures printed beside the numbers.

The security section is measured as well

A defensible number has no value if the system that produced it leaks or can be steered, so the privacy work below applies the 's method to the machinery itself: one frozen artifact, with guards turned on one at a time and the cost of each recorded beside it. The shape is the R-U confidentiality map of statistical disclosure control45, priced one rung at a time on the product metric a partner acts on. Where a guard cost something, the cost is here; where a check did not run, we say so. Nothing here measures American Express exposure: each result measures a mechanism on a public corpus.

The disclosure ladder, on what a partner actually receives

0% churnEach guard is priced one rung at a time on the metric a partner acts on; at the shipped rung, top-twenty churn is 0.

The shipped guard has no cost at the head of the ranking; stronger guards reduce utility
Utility metric
  • Raw output, no protection0
  • Small-cell suppression, shipped0
  • Contribution bounding, density channel10
  • Calibrated noise, mean over seeds17.70

Rows entering or leaving the top twenty relative to raw output, on a scale of zero to forty, where lower is better; the data are real public Foursquare venues on one frozen ranking, measuring a mechanism rather than exposure; the privacy unit is a venue.

  • Raw output, no protection100
  • Small-cell suppression, shipped97
  • Contribution bounding, density channel83
  • Calibrated noise, mean over seeds77.35

Share of the raw top hundred still present after the guard; higher is better. Deeper in the ranking, suppression is no longer free. Noise spreads over seeds from 73.95 to 80.00.

Real public Foursquare venues, one frozen ranking; the figure measures the mechanism, not exposure. Suppression is free at the head but not deeper; bounding costs 10 churn, noise 7.70 more, seeds 12.95 to 23.05.

These figures are bound to the committed results file and checked by the numeral gate; nothing is computed at page-view time.

The free guard, with the mechanism that makes it free

The reason is a property of this scorer, not a general result, so it is printed with the number: the top twenty raw cells carry at least 113 contributing venues, so a threshold of 10 cannot reach them, and neither can 50. Density is one of the four scoring signals, so dense cells score high by construction. Deeper down it is not free: 338 of P0's 400 released rows are still in the released list at the shipped threshold, and 225 at the stricter one. The control makes that null informative rather than flattering: removing the same number of cells at random moves a mean churn of 22.20 and 33.70, against 0 for the rule.

The two guards the product lacks are the ones that cost. Bounding one venue to at most 20 recipients replaces 5 of the top twenty rows, a churn price of 10, paired, because both rungs score the same buckets. Noise on top adds 7.70 more. Turning suppression on changes which buckets are published, so that first step is unpaired: no interval exists for it and none is implied.

The membership attack, at the operating point that matters

0.2724 of member accounts caught at the smallest measurable false-positive floor, 0.0024; interval 0.1576 to 0.5629. An attacker learns something, and the amount is measured.

Share of member accounts caught at the floor, by attack arm
  • Model score, behavioural arm0.2724
  • Counts-only control0.0019
  • Random-label negative control0.0026
  • No model: account tenure0.9689

Synthetic public TabFormer, account unit, share of member accounts caught at the false-positive floor 0.0024; model interval 0.1576 to 0.5629; the no-model attack scores higher, so the figure measures a confound rather than memorization.

These figures are bound to the committed results file and checked by the numeral gate; nothing is computed at page-view time.

At the account false-positive floor of 0.0024 the model arm recovers 0.2724 of member accounts, interval [0.1576, 0.5629], against 0.0019 for the counts-only control and 0.0026 for a random-label negative control. At the account operating point this attack DOES separate: the model arm recovers far more member accounts than the false positive rate it pays, while the counts-only control and the random negative control both sit at that floor, which is what tells you the pipeline is measuring something rather than nothing. What it separates is not established to be membership. By the membership rule itself a non-member account is an account that arrived after the corpus cut, so the two arms differ in tenure and in calendar position as well as in membership, and a no-model attack that uses only account size is reported beside the model for exactly that reason. Read this as a measurement of an attack against a confound, not as a measurement of what the model memorized, and not as evidence that no record is identifiable. A population average says nothing about the most exposed record: Aerni, Zhang and Tramer 2024 built a defence that fully leaks one training sample and still passes existing evaluations including TPR at low FPR, and measured their most vulnerable CIFAR-10 sample at 99.9 percent TPR at 0.1 percent FPR against a population figure of 4 percent. A population-average result says nothing about the most exposed record. This caveat belongs in the SAME paragraph as the number, never in a footnote.

The more informative result is the negative one. Counting an account's transactions needs no model, no checkpoint and no access to us, and at the same floor it recovers 0.9689 of member accounts against the model's 0.2724. The in account AUC is -0.0504, interval [-0.0617, -0.0395], a direction the file records as b_wins. A no-model attack winning means the model score buys an attacker nothing over the difference between the two account populations.

The injection red team, per class, against our own gate

28 of 30 narratives pass unattacked, against 30 of 30 committed; two of seven attack classes beat the gate.

Residual pass rate of each attack class through both gate layers
  • Instruction aimed at the auditor1.000
  • Ignore the supplied facts1.000
  • Figure inflated past the facts0.000
  • Digits written into a data field0.267
  • Rider aimed at the parser0.933
  • Unsupported claim, no numeral0.033
  • Reveal the system prompt0.000

Static attacks, one attempt each, public-data bundles, measured; residual share, point shown, intervals in table; marked hole closed by replay, others unchanged; a zero is a null result and does not demonstrate a defence.

These figures are bound to the committed results file and checked by the numeral gate; nothing is computed at page-view time.

An adverse finding against the gate we ship

Layer one's allowed pool is built by collect_fact_numbers, which recurses the bundle and also harvests every numeral out of every fact string (faithcheck.py:42-43). A string written into a data field therefore enlarges the pool the defence checks against. No model is involved in this measurement. So a planted figure the matcher flags against a bundle's own facts is accepted once one attacker-written string sits in one field: 30 of 30 bundles, the median allowed pool growing from 14 numerals to 15. The flip is deterministic. A defence that checks a narrative against a document the attacker partly wrote does not have a trustworthy reference.

The agent control plane, designed and not built

0 violationsDesigned and checked, not built: 8,294,400 mandate triples enumerated, 0 monotonicity violations, no model-derived read on the deterministic path.

The disposition order DETERMINISTIC BLOCK identity, mandate binding, caps, velocity, category, geography, revocation no model value is read here GRADED BLOCK value, novelty, reversibility the model score enters only here, and only to move an action down the list on the right DISPOSITION AUTO_EXECUTE OBSERVE ESCALATE DENY more restrictive
DESIGNED AND NOT BUILT. The house rule: a model output may add friction and never remove it. Status DESIGNED-AND-CHECKED. The agent control plane is designed and not built. No line of it runs in any product. This exhibit checks one property of the written specification against a reference implementation of it. DESIGNED-AND-CHECKED is not deployed, and nothing here may be presented as a measurement of a running system. This exhibit opens no dataset. The sibling safety exhibits run on the public IBM TabFormer benchmark, which is synthetic, or on public Foursquare venue records. This one runs on records constructed here from declared plan constants, so it carries no privacy question and no sampling uncertainty. The counts are exact over the enumerated space.

, and not run

SAFE-B1, unicity on the pretraining corpus did not run: scripts/safety/unicity.py was never written, so no result exists. The consequence for the page, stated as a rule: the de Montjoye reference stays a citation and never becomes a comparison. No claim about how many spatiotemporal points single out an account in this corpus may appear anywhere, in either direction.

Inside the injection red team, a narrower list of probes did not run; the reasons are recorded in the results file: adaptive attack, meaning search or gradient guided payloads; injection into the merchant name field of a live partner bundle; a second served model as an independent auditor.

What these controls do not protect against, and what each obligation actually got

Every control bounds the loss, none prevents it; four threats dropped with reasons; obligation rows are built or designed, none a compliance claim.

05BackboneFeasibility to develop and implementProtection, demonstratedProductivity, supporting

Built end to end, with every loss printed in full

The backbone pretrained on 24 million synthetic rows, and the we ship is the one that survived every guard, 0.0218; the same protocol went negative on a real corpus.

0.0218next-category gain, strictest rungsynthetic corpus, task, entity-level clear of zero, measured rather than modelled

Phase 1 can start inside the perimeter with a , because the recipe, the guards and the controls already run end to end at prototype scale on synthetic data.

What lost, and is printed anyway

  • 0.0590 to 0.0218: the same frozen checkpoint under a permissive protocol and under ours; the guards, not the model, account for the whole difference.
  • -0.0160 and -0.0147 on the real Complete Journey corpus under the identical L3 protocol, both household-clustered intervals below zero, the with-embedding arm under the 0.5101 majority floor: the corpus is a variable too.
  • The gain itself falls to -0.0240 once the split is , interval below zero; our reading: most of the permissive lift was account identity, not .
  • Eight fraud intervals across the ladder's four rungs all span zero: that task is saturated on this synthetic corpus before the protocol can matter.
  • -0.0345 on , interval [-0.0523, -0.0159] below zero, and a tie on : the label-free alone loses to a counts-only rarity control; the model does not carry the task on its own.
  • 0.985 baseline accuracy and 0.972 PR-AUC on the merchant axis: both tasks degenerate there, both deltas null, so two-sided beats single-sided stays an open question.
  • -0.0329 on top-five when the baseline gets per-entity aggregates, interval clear of zero: the last guard costs lift, and the two metrics disagree on how much.
  • 0.99389 to 0.996 fraud PR-AUC at the full corpus, a null on a saturated baseline; the fraud AUC delta spans zero at every corpus size.
  • A null at three million rows on next-category, interval spanning zero: consistent with a corpus effect or a smaller downstream sample, and the table does not separate the two.
  • +0.0141 with interval [-0.0094, 0.0382] spanning zero in one seed and +0.0269 clear of zero in the other: the offers head is mixed, quoted as mixed, and not causal.
  • +0.1957 store-head accuracy in one seed with interval [-0.0217, 0.4130] spanning zero: a null on the reported metric, while the decision metric cleared zero in both seeds, intervals wide, on a 46-store test set.
  • 0.53 for a per-corridor beats the corridor model pooled, which wins five of twelve corridors; the fixed average with the naive wins, a point comparison with no interval claimed.
  • Conversion, the rare outcome on the randomized Criteo holdout: loses to response ranking at every depth measured, all paired intervals below zero; only a randomized layer can tell a partner that.
  • 0.000589 ROC-AUC and 0.002829 PR-AUC, intervals [-0.000475, 0.001567] and [-0.000422, 0.006311] spanning zero: the merchant-native task is a null on this corpus at this scale, with the question properly put.

One question organizes what follows: whether adding the model to the cheap control improves what a partner would actually run, rather than whether the model beats that control on its own. The rows are read that way; the model-alone results remain printed.

Next-category top-1, embeddings added to the baseline 0.0218 0.0103 to 0.0343 Next-category top-5, embeddings added to the baseline 0.0218 0.0106 to 0.0336 Offer targeting, site visits, top decile 0.0110 0.0085 to 0.0145 Offer targeting, site visits, top 20 percent 0.0012 -0.0002 to 0.0025 Offer targeting, site visits, top 30 percent 0.0005 -0.0005 to 0.0012 Offer targeting, conversions, top decile (the primary) -0.0010 -0.0017 to -0.0003 Offer targeting, conversions, top 20 percent (the primary) -0.0005 -0.0008 to -0.0002 Offer targeting, conversions, top 30 percent (the primary) -0.0005 -0.0007 to -0.0002 Fraud, model alone against the counting control, PR-AUC 0.0056 -0.0391 to 0.0506 Fraud, model alone against the counting control, ROC-AUC -0.0345 -0.0523 to -0.0159 Whitespace composite against venue density, forward check -0.1392 -0.1785 to -0.1020 Real-corpus transfer, households, top-1, seed 7 -0.0160 -0.0207 to -0.0113 Real-corpus transfer, households, top-1, seed 8 -0.0147 -0.0197 to -0.0101 Real-corpus transfer, households, top-5, seed 7 0.0024 -0.0026 to 0.0075 Real-corpus transfer, households, top-5, seed 8 0.0047 -0.0006 to 0.0108 Real-corpus store head, macro-F1, seed 7 0.1912 0.0064 to 0.3865 Real-corpus store head, macro-F1, seed 8 0.2242 0.0449 to 0.4343 Real-corpus store head, accuracy, seed 7 0.1957 -0.0217 to 0.4130 Real-corpus store head, accuracy, seed 8 0.2174 0.0217 to 0.4136 Real-corpus offers head, AUC, seed 7 0.0141 -0.0094 to 0.0382 Real-corpus offers head, AUC, seed 8 0.0269 0.0023 to 0.0551 Real-corpus offers head, PR-AUC, seed 7 0.0121 -0.0697 to 0.1100 Real-corpus offers head, PR-AUC, seed 8 0.0661 -0.0300 to 0.1633 Pair retention, merchant-axis embedding, ROC-AUC 0.0006 -0.0005 to 0.0016 Pair retention, merchant-axis embedding, PR-AUC 0.0028 -0.0004 to 0.0063 Pair retention, cardholder-axis embedding, ROC-AUC -0.0008 -0.0034 to 0.0014 Pair retention, cardholder-axis embedding, PR-AUC -0.0013 -0.0067 to 0.0036 Not pre-registered: the combination, built after the first run. Drawn, and counted in no tally. Fraud, model added to the counting control, PR-AUC 0.2014 0.1480 to 0.2450 Fraud, model added to the counting control, ROC-AUC 0.0129 0.0041 to 0.0226 zero
The whole pre-registered record, one row per metric the pre-registration named for that comparison, per registered depth, seed and arm, against a single zero line: 27 rows, 7 clear it upward, 7 clear it downward and 13 straddle it. A comparison registered at three depths is three rows, and a registered metric that was not the decision metric is a row, so nothing registered is summarised away. The two rows under the dashed rule are the protection combination, built after the first run and drawn here because the page leans on them; they are in no count. Each row is scaled to its own largest bound, since a , an AUC difference and a rate difference share no scale, so lengths are not comparable between rows and the only claim here is which side of zero an interval sits on. The first page prints these counts, and scripts/protocol_value.py refuses to write them if any row fails to resolve.

What we built

20.54MThe model has about 20.54M parameters and was pretrained for 8 epochs on 24 million synthetic public transactions; the loss curve shown is from the actual run.

A field-tokenized masked language model over transaction sequences, in the TabBERT lineage48, about 20.54M parameters, pretrained for 8 epochs on an A100 on the public IBM TabFormer benchmark (24 million synthetic transactions, labeled as such). The loss curve below is the actual run. This direction extends work Amex itself has published on sequence models for credit monitoring49.

The closest public prior art, named

A public blueprint of about 29M parameters uses the same corpus, so we are not the first to work on it; our contribution is the protocol rather than the architecture.

The multi-task transfer table

0.2570.235 to 0.257 next-category accuracy with embeddings, interval [0.0103, 0.0343] clear of zero, both tasks pre-registered; full-corpus fraud is a null on a saturated baseline.

The leakage design, named

400 of 2,000 accounts held out ; labels are excluded from the vocabulary, embeddings are computed as-of, identifiers are hashed, and four checks are asserted in the results file.

TabFormer · synthetic, labeled 20.54M params · 8 epochs · corpus cut 2017-08-25T09:37:00Z

Backbone pretraining loss curve over training steps. 0 2 4 0 20,000 40,000 60,000 Step Masked-LM loss Pretraining loss · x=10 · y=4.108 Pretraining loss · x=7,190 · y=0.6928 Pretraining loss · x=14,050 · y=0.6359 Pretraining loss · x=21,230 · y=0.6039 Pretraining loss · x=28,410 · y=0.5858 Pretraining loss · x=35,600 · y=0.5654 Pretraining loss · x=42,450 · y=0.5534 Pretraining loss · x=49,640 · y=0.5439 Pretraining loss · x=56,820 · y=0.5362 Pretraining loss · x=64,000 · y=0.5291 Pretraining loss · x=70,860 · y=0.5253 Pretraining loss · x=78,040 · y=0.526 Pretraining loss
Pretraining loss on the TabFormer corpus (single series; the title names it). The curve is downsampled for rendering; the results file holds every step.
Multi-task transfer · LightGBM with per-entity temporal aggregates (house baseline) vs the same + backbone embeddings. Entity-level paired-bootstrap 95% CIs. Deltas reported as obtained.
TaskMetricBaseline +EmbeddingsΔ (95% CI)
Fraud (Protection head v0)AUC 0.999940.99961 [-0.0012, 0.0001]
PR-AUC 0.993890.99600 [-0.0041, 0.0086]
Next merchant categoryTop-1 accuracy 0.234930.25678 [0.0103, 0.0343]
Top-5 accuracy 0.530840.55260 [0.0106, 0.0336]
  • ✓ label column excluded from pretraining vocab
  • ✓ corpus time-truncated before the test window
  • ✓ as-of / prefix-only entity embeddings
  • ✓ user & card IDs hashed / excluded

seed 7 · data labels: synthetic · generated by scripts/fm/transfer_eval.py --check-able · lightgbm 4.7.0 · numpy 2.5.2 · polars 1.43.2 · pyarrow 25.0.1 · python 3.12.3 · sklearn 1.9.0 · torch 2.13.0+cu130

What the protocol costs: the leakage ladder Synthetic

0.02180.0590 under a permissive protocol, 0.0218 under ours, same frozen checkpoint: the difference is attributable entirely to the guards rather than to the model.

The same checkpoint under four protocols and the gain that remains
Turn the guards on
  • L0 permissive, shared accounts0.05900.0547 to 0.0634

L0: temporal split only, same accounts on both sides, one pooled embedding that sees the scored row. This is the largest figure in the ladder; it is scored on different rows and is unpaired.

  • L0 permissive, shared accounts0.05900.0547 to 0.0634
  • L1 entity-disjoint split-0.0240-0.0373 to -0.0112

L1 holds out whole accounts. The lift does not merely decrease; it falls below zero. Our reading is that most of the L0 gain reflected account identity rather than transfer. This step changes three things at once.

  • L0 permissive, shared accounts0.05900.0547 to 0.0634
  • L1 entity-disjoint split-0.0240-0.0373 to -0.0112
  • L2 as-of embeddings0.01590.0077 to 0.0246

L2 truncates the embedding before the scored transaction, and the transfer signal appears. It uses the same rows as L1, so the guard is priced directly: top-one accuracy moves 0.0399, clear of zero.

  • L0 permissive, shared accounts0.05900.0547 to 0.0634
  • L1 entity-disjoint split-0.0240-0.0373 to -0.0112
  • L2 as-of embeddings0.01590.0077 to 0.0246
  • L3 baseline gets aggregates, shipped0.02180.0103 to 0.0343

L3 is the shipped protocol behind every other number on this page. Top-one accuracy moves 0.0060, with an interval spanning zero; top-five accuracy changes by -0.0329. Fraud cannot discriminate between protocols because it is saturated at every rung.

Synthetic corpus, next-category top-one delta, rungs pre-registered, entity-level paired intervals, measured. L0 scores different rows and is unpaired; L1 sits below zero; the last guard costs top-five lift, -0.0329.

These figures are bound to the committed results file and checked by the numeral gate; nothing is computed at page-view time.

One of those steps is not a single switch, and a reader should have that before the numbers rather than after. The first step, L0 to L1, moves three things together in our own code, because they all hang off the same entity-disjoint flag. It replaces the row selection: L0 trains on the earlier part of the post-cut pool and tests on the later part, while L1 onward uses the shipped split, which holds out whole accounts. It therefore also changes what the test set is a sample of, from a later slice of time to a set of accounts drawn across the whole post-cut window, which is why L0 and L1 do not score the same rows or the same number of them. And it restricts the rows the category, city and state frequency encodings are fitted on, from every pre-cut row to the pre-cut rows of accounts that are not held out, so the baseline's own features tighten at that step as well. The other two steps each move exactly one thing, which is the other reason only they carry a on the change. We would rather name the compound step than let the phrase "one guard at a time" stand for all three.

This ladder is not an audit of any other team's work. It measures what our own pipeline reports when its evaluation guards are removed one step at a time, on a synthetic public corpus. We are not claiming that any particular published result used any particular rung, and we put no other team's numbers next to ours anywhere on this page. The use we do claim is narrower and more useful to a reviewer: when two transfer results disagree, the protocol is a measurable variable, and this is what measuring it looks like. The idea underneath is not ours. Constraining and measuring the rather than assuming it is what TESSERACT did for malware classifiers across time and distribution54, and a study of recommender offline evaluation found that a split ignoring the global timeline moves reported accuracy unpredictably55. What we add is a price rather than a warning: one frozen checkpoint, four rungs, and a paired interval on each step whose rungs score identical rows.

Next merchant category, Δ top-1 accuracy: the transfer delta with its 95% confidence interval at each of the 4 protocol rungs, with a zero reference line -0.05 0 0.05 Next merchant category, Δ top-1 accuracy L0 · delta 0.0590 · 95% CI [0.0547, 0.0634] L0 0.0590 L1 · delta -0.0240 · 95% CI [-0.0373, -0.0112] L1 -0.0240 L2 · delta 0.0159 · 95% CI [0.0077, 0.0246] L2 0.0159 L3 · delta 0.0218 · 95% CI [0.0103, 0.0343] L3 0.0218
Next merchant category, Δ top-1 accuracy at each protocol rung, with entity-clustered paired-bootstrap 95% CIs and a zero reference line. One guard is turned on per step; the value under each rung is the delta itself. The first step also changes which rows are scored and the frequency-encoding fit set, which the prose above names in full.
What each rung changes · one frozen checkpoint, the same two downstream tasks, one guard turned on per row. The last column is the number of baseline features the next-category task and the fraud task see, which is what the third guard moves. The first guard carries two further changes with it, the row selection and the frequency-encoding fit set, named in full in the prose above. L3 is the protocol every other number on this page already uses.
RungEntity-disjoint splitAs-of embeddingsAggregate baselineBaseline features
L0 permissiveoffoffoff4 / 11
L1 entity-disjoint splitonoffoff4 / 11
L2 as-of embeddingsononoff4 / 11
L3 aggregate-equipped baseline (shipped)ononon10 / 18
The ladder itself · with-embeddings minus baseline at each rung, with entity-clustered paired-bootstrap 95% CIs, reported as obtained. The fraud columns carry one extra caveat: the split guard changes which rows are scored, so the fraud positive rate is not identical across rungs and the results file records it per rung.
RungNext-MCC Δtop-1 (95% CI)Next-MCC Δtop-5 (95% CI)Fraud ΔAUC (95% CI)Fraud ΔPR-AUC (95% CI)
L00.0590 [0.0547, 0.0634]0.1064 [0.1010, 0.1122]-0.000017 [-0.000044, 0.000005]-0.0002 [-0.0162, 0.0160]
L1-0.0240 [-0.0373, -0.0112]0.0000 [-0.0158, 0.0140]-0.000789 [-0.002698, 0.000149]0.0052 [-0.0098, 0.0218]
L20.0159 [0.0077, 0.0246]0.0547 [0.0421, 0.0674]-0.000006 [-0.000221, 0.000158]0.0076 [-0.0022, 0.0212]
L30.0218 [0.0103, 0.0343]0.0218 [0.0106, 0.0336]-0.000335 [-0.001213, 0.000093]0.0021 [-0.0041, 0.0086]
What each guard changes · how the transfer delta moves when that guard is turned on, on all four metrics the ladder measures. A negative entry means the guard removes apparent lift and a positive one means it uncovers lift the looser protocol was hiding. The L1 to L2 and L2 to L3 steps are scored on identical test rows, so their changes carry a paired entity-clustered bootstrap interval. The L0 to L1 step changes which rows are scored, so its change is a difference of point estimates with no interval, and this table marks it as not paired rather than reporting an interval. The fraud AUC column prints at seven decimals because the metric it sits on is already saturated at this corpus size, so every movement there is immaterial to a queue whichever way its interval falls.
Step and guard turned onNext-MCC Δtop-1 change (95% CI)Next-MCC Δtop-5 change (95% CI)Fraud ΔAUC change (95% CI)Fraud ΔPR-AUC change (95% CI)
L0 to L1 · entity-disjoint split by account-0.0831 not paired-0.1064 not paired-0.0007720 not paired0.0054 not paired
L1 to L2 · as-of prefix-only embeddings0.0399 [0.0249, 0.0561]0.0547 [0.0385, 0.0742]0.0007821 [-0.0000067, 0.0025274]0.0024 [-0.0046, 0.0125]
L2 to L3 · baseline equipped with per-entity temporal aggregates0.0060 [-0.0020, 0.0152]-0.0329 [-0.0436, -0.0219]-0.0003282 [-0.0010353, -0.0000002]-0.0055 [-0.0169, 0.0055]

Merchant-side embedding: pre-cut pooled at every rung; only the account-side embedding varies, so L0 is a lower bound on how much a fully permissive protocol inflates, not an upper bound

seed 7 · data labels: synthetic · generated by scripts/fm/ladder_eval.py --check-able · lightgbm 4.7.0 · numpy 2.5.2 · polars 1.43.2 · pyarrow 25.0.1 · python 3.12.3 · sklearn 1.9.0 · torch 2.13.0+cu130

The ladder's second axis: the same protocol on a real corpus Real

-0.0160 and -0.0147 household transfer on real retail data, both intervals below zero; store head positive both seeds, offers mixed. The result depends on the corpus.

The same protocol on a different corpus gives a different result
  • Transfer top-one, seed seven-0.0160-0.0207 to -0.0113
  • Transfer top-one, seed eight-0.0147-0.0197 to -0.0101
  • Store growth F-score, seed seven0.19120.0064 to 0.3865
  • Store growth F-score, seed eight0.22420.0449 to 0.4343
  • Offers AUC, seed seven0.0141-0.0094 to 0.0382
  • Offers AUC, seed eight0.02690.0023 to 0.0551

Real retail basket lines, not card data; pre-registered, entity-clustered paired intervals, measured, predictive not causal. Three metrics share one zero line: transfer negative, store positive but wide, offers mixed.

These figures are bound to the committed results file and checked by the numeral gate; nothing is computed at page-view time.

Measuring protection: a label-free surprise score and the control that runs against it Synthetic

0.32960.1282 to 0.3296 PR-AUC with the added to counting, interval clear of zero; on its own it loses on ROC-AUC. This analysis is post hoc and was not pre-registered.

What the fraud queue holds under each scoring rule
Scoring rule
  • PR-AUC0.12820.0950 to 0.1717
  • ROC-AUC0.93670.9008 to 0.9653
  • Recall, top one percent0.32280.2850 to 0.3687
  • Recall, top tenth of a percent0.10480.0719 to 0.1307

The control with no model: how globally rare each field value is, fitted on pre-cut rows. Synthetic TabFormer, 945 fraud positives, entity-level intervals, measured, no labels in the score.

  • PR-AUC0.13380.0983 to 0.1751
  • ROC-AUC0.90220.8742 to 0.9265
  • Recall, top one percent0.37880.3354 to 0.4233
  • Recall, top tenth of a percent0.11640.0912 to 0.1430

The backbone's surprise score is read off the frozen checkpoint without retraining and measures how unlikely each value is given the account's history. On its own it scores below counting on ROC-AUC and ties with counting on PR-AUC.

  • PR-AUC0.32960.2705 to 0.3881
  • ROC-AUC0.94970.9170 to 0.9742
  • Recall, top one percent0.59890.5481 to 0.6454
  • Recall, top tenth of a percent0.19260.1643 to 0.2278

Both scores are standardized and added without weights; nothing is fitted and no labels are used. The top one percent holds 0.5989 of the fraud against 0.3228 for counting alone. This comparison is post hoc and was not pre-registered.

  • PR-AUC0.41330.3532 to 0.4691
  • ROC-AUC0.98420.9796 to 0.9882
  • Recall, top one percent0.67830.6370 to 0.7182
  • Recall, top tenth of a percent0.21800.1863 to 0.2494

The same combination is repeated with the calendar fields and the error flag dropped, keeping the six fields that describe what the transaction did. The result is stronger, which indicates that the surprise signal comes from behaviour rather than from an unseen year token.

The control uses no model and scores how globally rare each field value is, fitted on pre-cut rows. The data are synthetic TabFormer with 945 fraud positives; the intervals are entity-level, the figures are measured, and no labels enter the score.

These figures are bound to the committed results file and checked by the numeral gate; nothing is computed at page-view time.

Surprise alone loses to counting; both together win
  • Surprise minus counting, PR-AUC0.0056-0.0391 to 0.0506
  • Surprise minus counting, ROC-AUC-0.0345-0.0523 to -0.0159
  • Both minus counting, PR-AUC0.20140.1480 to 0.2450
  • Both minus counting, ROC-AUC0.01290.0041 to 0.0226
  • Both minus surprise, PR-AUC0.19580.1654 to 0.2227

Synthetic corpus, card-fraud labels, paired on identical rows, entity-level intervals, measured; the combination is post hoc, not pre-registered. Alone the model is behind on ROC-AUC and not separated on PR-AUC.

These figures are bound to the committed results file and checked by the numeral gate; nothing is computed at page-view time.

The limitation is real and we state it here explicitly. No public dataset carries authorized-scam labels, so we validate on card-fraud labels. A ranking that separates card fraud is not thereby shown to separate scam-induced authorized payments, and this exhibit does not claim that it is. What it does establish is narrower and still useful: a frozen masked-field backbone yields a transaction-level ranking with no labels at all, the ranking can be measured with entity-level intervals, and it has to be measured against a counts-only control before the model can be credited with the improvement. Testing the same score against authorized-scam outcomes needs labels only a network holds, which is what the Section 8 pilot is for. The corpus here is the same synthetic one as the rest of this section, and it is marked as such.

What this proves, and what it does not

The recipe runs end to end at prototype scale on synthetic data; production performance at Amex is not proven here.

Capacity: the scaling check Synthetic

0.9040.638 to 0.904 fraud PR-AUC at ten million rows, with the PR-AUC interval above zero and the AUC interval spanning zero; the next-category interval clears zero from ten million rows onward, and the merchant axis result is null.

The merchant embedding atlas Synthetic

Per-merchant embeddings projected to two dimensions, coloured by category group; this structure is what the whitespace head's similarity signal runs on.

A two-dimensional projection of the backbone's per-merchant embeddings, colored by merchant category group. Structure in this map is what the whitespace head's embedding-similarity signal runs on.

Two-dimensional projection of per-merchant backbone embeddings, colored by merchant category group.

Misc Retail & DiningRetail GoodsBusiness & EntertainmentApparelUtilities & TelecomPersonal ServicesProfessional & MembershipTransportationCar RentalLodgingFinancial ServicesGovernmentContracted ServicesAirlines

UMAP projection of 238,615 pseudonymous merchant embeddings, drawn as a 3,000 point stratified sample; color = category group, no merchant names.

seed 42 · data labels: synthetic · generated by scripts/scale/atlas_build.py --check-able · lightgbm 4.7.0 · numpy 2.5.2 · polars 1.43.2 · pyarrow 25.0.1 · python 3.12.3 · sklearn 1.9.0 · torch 2.13.0+cu130 · umap 0.5.12

The merchant-native task, built and run Synthetic

838,8630.002829 PR-AUC, interval [-0.000422, 0.006311] spanning zero, on 838,863 pre-cut pairs: the merchant-native task is a null, with the question properly put.

What the null is worth is the question, not the answer. The two earlier merchant-axis nulls came from tasks that could not have shown a gain; this one could have. The baseline here scores 0.940356 ROC-AUC and 0.863579 PR-AUC against a test positive rate of 0.177578, which leaves real headroom, where the next-category baseline on the same axis already sits at 0.985 top-1. A null with the question properly put is a different and more useful thing than a null from a task that was never going to answer it.

06OffersTechnical depth of the all-around system designFeasibility to develop and implementGrowth, primary

The measurement layer, rather than the model, is the product

On 4,193,878 randomized advertising holdout rows, uplift ranking beats response ranking at the top decile on visit, interval above zero, and loses on conversion, the primary, at every depth.

0.011049incremental visit rate, uplift over response ranking, top decilereal randomized advertising holdout, not payments; entirely above zero; measured

This CoE ran the same metric as its 2024 contest; One Loop ships it as a standing per-partner, per-campaign lift report, interval attached, whichever way it lands.

What lost, and is printed anyway

  • -0.001022 at ten percent, -0.000522 at twenty, -0.000479 at thirty: on the rare conversion outcome the uplift ranking loses to response ranking at every depth, all three paired intervals entirely below zero.
  • 68.6 percent of incremental conversions captured by response ranking against 59.2 by uplift ranking: on the primary endpoint, response ranking wins. The choice of endpoint therefore determines which ranking wins.
  • 0.000199 against 0.000345: on conversion, X-learner against response ranking, the fully below zero. We committed before the full run to report it whichever way it landed.
  • 0.002993 against 0.003015: on visit, the endpoint we lead with, paired difference spanning zero. The win is a depth result at the top decile, not a whole-curve one.
  • 0.001199 at twenty percent and 0.000454 at thirty on visit, both intervals spanning zero: past the top decile the two rankings are not separated here.
  • 5.9 times the average visit effect at the top decile is a ratio of point estimates with no interval; response ranking reaches most of that alone, the model adds the paired difference.
  • Commit 2fd8c23 fixed the endpoint hierarchy before the 13,979,592-row run, but a one-million-row development pass of the same script and seed ran first: fixed before the full run, not before every result.
  • 5 of the 10 corridor narratives the third gate layer can read are reversed: our fact bundle stated a false rule, now fixed; text not regenerated; 20 bundles reported unchecked.
  • 28 of 30 narratives pass the same gate unattacked on a live re-run, against all thirty on the committed run: the false-reject budget a pilot has to beat.
  • Layer two reads with the same served model as the writer, which is the shipped design, so it is not an independent check.

The product is the measurement layer

Every campaign carries a by design; each partner gets an auditable incremental-lift report. This is the ghost-ads pattern, productized for partner banks.

The judges are familiar with this problem. The 2024 edition of this hackathon was an uplift contest scored on incremental activation across 12.6 million customer x merchant pairs. One Loop turns that contest's metric into a standing service.

The case for uplift targeting, in one finding

Up to 6.8 percentage points more churn reduction from treatment-effect targeting than from response targeting aimed at the wrong customers, in Ascarza's randomized field experiments.

The exhibit: the same measurement gives two verdicts on two outcomes

13,979,592 randomized Criteo rows: the winning ranking changes with the endpoint. Response ranking wins on conversion, the pre-registered primary; uplift ranking wins on visit at the top decile.

The choice of endpoint determines which ranking wins
  • Uplift ranking, visit59.3
  • Response ranking, visit48.6
  • Uplift ranking, conversion59.2
  • Response ranking, conversion68.6

The data are from Criteo, a real randomized advertising holdout rather than payments data. Share of incremental outcome captured: point estimates, modelled rankings scored against randomized arms. Conversion is the pre-registered primary; uplift ranking loses it.

These figures are bound to the committed results file and checked by the numeral gate; nothing is computed at page-view time.

Where uplift targeting wins, on the same randomized data Real

0.011049 paired difference over response ranking at the top decile on visit, interval entirely above zero, on 4,193,878 holdout rows; the advantage fades beyond the decile.

How far the top-of-ranking advantage reaches
Targeting depth
  • Uplift ranking, visit rate0.0612410.057974 to 0.064709
  • Response ranking, visit rate0.0501930.046261 to 0.053774
  • Whole holdout, average effect0.0103260.009832 to 0.010820

Incremental visit rate, treated minus control, among rows each ranking would target at this depth; the data are the randomized Criteo dataset, not payments data. The decile carries 5.9 times the average effect, point estimates only.

  • Uplift ranking, visit rate0.0394060.037444 to 0.041248
  • Response ranking, visit rate0.0382070.035907 to 0.040377
  • Whole holdout, average effect0.0103260.009832 to 0.010820

When the budget is widened to a fifth of the holdout, the two rankings converge; the paired difference at this depth, 0.001199, has an interval that spans zero.

  • Uplift ranking, visit rate0.0294160.028094 to 0.030610
  • Response ranking, visit rate0.0289620.027386 to 0.030600
  • Whole holdout, average effect0.0103260.009832 to 0.010820

At the top third the advantage has faded, difference 0.000454 with an interval spanning zero. A budget is spent at a particular depth, and the decile is the depth at which a partner would act.

Incremental visit rate, treated minus control, among rows each ranking would target at this depth; the data are the randomized Criteo dataset, not payments data. The decile carries 5.9 times the average effect, point estimates only.

These figures are bound to the committed results file and checked by the numeral gate; nothing is computed at page-view time.

Uplift ranking leads on visit at the top decile and trails on conversion
Endpoint
  • Top ten percent0.0110490.008545 to 0.014453
  • Top twenty percent0.001199-0.000175 to 0.002534
  • Top thirty percent0.000454-0.000474 to 0.001231

Visit is the denser outcome. The figure shows the paired difference of uplift minus response with bootstrap intervals on the Criteo randomized advertising holdout, which is not a payments dataset. The difference clears zero at the top decile only; at the deeper cutoffs the two rankings are not separated.

  • Top ten percent-0.001022-0.001670 to -0.000277
  • Top twenty percent-0.000522-0.000842 to -0.000184
  • Top thirty percent-0.000479-0.000675 to -0.000249

Conversion is the pre-registered primary and the rare outcome. On the same holdout, arms and depths, every interval sits entirely below zero, so plain response ranking performs better on this one dataset.

Criteo, real randomized advertising holdout, not payments. Paired difference, uplift minus response, bootstrap intervals, measured. Visit clears zero only at the top decile; conversion, the pre-registered primary, stays below zero.

These figures are bound to the committed results file and checked by the numeral gate; nothing is computed at page-view time.

What that means for an offers programme

Dense outcome, uplift won here; rare outcome, response won. Nothing in the model says which regime you are in; the does.

The interpretable half

28.4 percent of an in-sample response model's top-third budget lands where Hillstrom's randomized arms measure the least uplift. The dataset is retail e-commerce, so this demonstrates the mechanism rather than the size of the effect.

Where a response model's top-third budget lands
  • Best segment the experiment measures11.68
  • Response pick, bottom-third segment2.31
  • Response pick, another bottom-third segment1.26

Hillstrom, real randomized email arms, not payments; mechanism, not size. Measured visit uplift, no model. 28.4 percent of an in-sample response model's top-third budget lands here; halves, 63.8 percent.

These figures are bound to the committed results file and checked by the numeral gate; nothing is computed at page-view time.

The response model therefore concentrates on the group the experiment says responds least, and measuring incrementality instead of response corrects it. That is one proxy on one retail corpus and the models behind the selection columns are in-sample, exactly as the segment table above is flagged, so this describes the mechanism rather than certifying an outcome. Phase 1 is required to deliver the real version: subgroup lift and selection rates on the partner's own protected attributes, pre-registered before the first campaign opens and reported whichever way they move.

Every lift report ships with a plain-language narrative written on top of the exhibit numbers. Facts go in from the committed results files, the text comes out of a self-hosted model, and a gate then decides whether that text is allowed to stand. The gate is the part we would defend in a review. The layer explains scores and reports. It never makes a decision.

Interlude · GenAI explanation layer

This section covers technical depth. The contribution is the gate rather than the prose that passes through it.

The two-layer gate on generated text

Every lift report ships with a plain-language narrative written on top of the exhibit numbers. Facts go in from the committed results files, the text comes out of a self-hosted model, and a gate then decides whether that text is allowed to stand. The gate is the part we would defend in a review. The layer explains scores and reports. It never makes a decision.

n = 30 narratives · strict 30 of 30 pass both layers · model Qwen/Qwen3-32B · self-hosted vLLM on NUS GPU

checker: numeric-claim extraction vs source JSON + LLM cross-check failures ship by rule

Replay console

A step through the committed narratives file, one example at a time: the facts that went in, the text that came out, what layer one matched, and what layer two reported. It runs entirely inside this page. Nothing here calls a model and nothing here reaches the network.

seed 42 · data labels: genai-explanation-layer facts-from-committed-results explains-never-decides · generated by scripts/narratives/generate.py (check: scripts/narratives/check.py --check) · python 3.11.15 · torch 2.11.0+cu130 · transformers 5.14.1 · vllm 0.26.0

corridor corridor-00 strict verdict: pass direction layer: pass

Picked because the model loses to the seasonal-naive baseline on this corridor, and the generated text says so rather than talking around it.

Facts in, read from the committed results file

what
held-out monthly backtest of Singapore inbound visitor arrivals by origin market (public proxy for cross-border spend corridors); MASE below 1 beats the seasonal-naive baseline
corridor
Taiwan
region
Greater China
holdout_months
13
mase_model
1.137818
mase_seasonal_naive
0.840431
mase_model_reconciled
1.016972
top3_model_attributions[0].feature
lag1
top3_model_attributions[0].mean_abs_shap_log1p_units
0.244893
top3_model_attributions[1].feature
corridor_id
top3_model_attributions[1].mean_abs_shap_log1p_units
0.132313
top3_model_attributions[2].feature
lag3
top3_model_attributions[2].mean_abs_shap_log1p_units
0.12867
attribution_label
model attributions (mean |SHAP| on holdout rows, log1p-arrivals units) — associations learned by the model, not causal drivers

Narrative out, generated from those facts only

We conducted a 13-month holdout backtest of Singapore inbound visitor arrivals from the Taiwan corridor, part of the Greater China region, using a model with a MASE of 1.137818, which underperformed the seasonal-naive baseline of 0.840431. The reconciled version of the model improved to a MASE of 1.016972. The top three model attributions were "lag1" (0.244893), "corridor_id" (0.132313), and "lag3" (0.12867), measured in mean absolute SHAP values on log1p-arrivals units. These attributions reflect associations learned by the model, not causal drivers.

Layer one, deterministic numeric match

  • 13 matches holdout_months
  • 1.137818 matches mase_model
  • 0.840431 matches mase_seasonal_naive
  • 1.016972 matches mase_model_reconciled
  • 1 matches what
  • 0.244893 matches top3_model_attributions[0].mean_abs_shap_log1p_units
  • 0.132313 matches top3_model_attributions[1].mean_abs_shap_log1p_units
  • 3 matches top3_model_attributions[2].feature
  • 0.12867 matches top3_model_attributions[2].mean_abs_shap_log1p_units
  • 1 matches what

Layer one: every numeral in this text is carried by a fact above. Every numeral the extractor finds is marked, including a digit sitting inside a name, because this layer matches numerals and does not read sentences. The match is also percent-aware, so a fact and the same quantity printed as a percentage both count as support, which widens what layer one accepts. The field named beside each numeral is the closest supporting fact, not the first one found.

Layer two, model cross-examination

Layer two: the cross-examination reported no claim the facts do not carry.

whitespace whitespace-10 strict verdict: pass direction layer: unchecked

This example is included deliberately. Layer one passes this text, although a careful reader should not: the count of places inside the bucket is read out as a count of buckets. The gate checks the numerals and does not check the nouns to which those numerals attach.

Facts in, read from the committed results file

what
pseudonymized Singapore merchant-signing whitespace bucket, ranked by real public signals only (scores in 0-1); sensitivity = this bucket's rank if a simulated demand-weighted acceptance-gap signal were added at the stated strength
bucket_label
Personal Services cluster — Clarke Quay area
area
Clarke Quay
category
Personal Services
rank_real_signals
28
n_pois
151
score_real_signals
0.785897
reasons[0]
Premium-demand corridor (0.4 km to Clarke Quay; zone score 0.79)
reasons[1]
MDR-sensitive category mix (Personal Services, prior 0.70: 1.5-3% card MDR vs 0.3% QRIS anchor)
reasons[2]
Dense commercial cluster (median 493 POIs within 250 m)
sensitivity[0].leakage_strength
0.1
sensitivity[0].rank
29
sensitivity[1].leakage_strength
0.25
sensitivity[1].rank
24
sensitivity[2].leakage_strength
0.5
sensitivity[2].rank
23

Narrative out, generated from those facts only

The Personal Services cluster — Clarke Quay area is ranked 28th out of 151 pseudonymized merchant-signing whitespace buckets in Singapore, with a score of 0.785897 based on real public signals. Key contributing factors include its proximity to a premium-demand corridor (within 0.4 km of Clarke Quay, with a zone score of 0.79), a mix of MDR-sensitive categories (Personal Services, with card MDRs at 1.5-3% versus 0.3% for QRIS anchors), and a dense commercial environment (median of 493 POIs within 250 m). Sensitivity analysis shows that if a simulated demand-weighted acceptance-gap signal were added, the bucket’s rank would shift to 29 at 0.1 strength, 24 at 0.25 strength, and 23 at 0.5 strength.

Layer one, deterministic numeric match

  • 28 matches rank_real_signals
  • 151 matches n_pois
  • 0.785897 matches score_real_signals
  • 0.4 matches reasons[0]
  • 0.79 matches reasons[0]
  • 1.5 matches reasons[1]
  • -3 matches reasons[1]
  • 0.3 matches reasons[1]
  • 493 matches reasons[2]
  • 250 matches reasons[2]
  • 29 matches sensitivity[0].rank
  • 0.1 matches sensitivity[0].leakage_strength
  • 24 matches sensitivity[1].rank
  • 0.25 matches sensitivity[1].leakage_strength
  • 23 matches sensitivity[2].rank
  • 0.5 matches sensitivity[2].leakage_strength

Layer one: every numeral in this text is carried by a fact above. Every numeral the extractor finds is marked, including a digit sitting inside a name, because this layer matches numerals and does not read sentences. The match is also percent-aware, so a fact and the same quantity printed as a percentage both count as support, which widens what layer one accepts. The field named beside each numeral is the closest supporting fact, not the first one found.

Layer two, model cross-examination

Layer two: the cross-examination reported no claim the facts do not carry.

uplift uplift-21 strict verdict: pass direction layer: unchecked

This panel presents one Hillstrom segment in the form a partner would see it in a lift report.

Facts in, read from the committed results file

what
customer segment from the randomized Hillstrom e-mail experiment (Womens E-Mail vs No E-Mail); ranks are 1=highest among 28 segments; response_rank orders segments by a response model, uplift_rank by an uplift model, measured_uplift_rank by uplift measured directly from the randomized arms; the verdict compares response-model targeting against measured uplift
segment
3) $200 - $350 x recency 10-12m
n_treated
815
n_control
788
response_rank
23
uplift_rank
18
measured_uplift_rank
18
measured_visit_uplift_pp
3.9581
measured_visit_uplift_se_pp
1.6262
verdict
aligned
verdict_rule
wasted-budget: response rank in top tercile AND measured uplift <=0 or in bottom tercile; hidden-gem: response rank bottom tercile AND measured uplift top tercile; else aligned

Narrative out, generated from those facts only

The customer segment "3) $200 - $350 x recency 10-12m" includes 815 treated and 788 control customers from the Hillstrom e-mail experiment. This segment ranks 23rd in response rank and 18th in both uplift rank and measured uplift rank. The measured visit uplift per person is 3.9581 with a standard error of 1.6262, and the verdict is "aligned," indicating that the response model and measured uplift are consistent.

Layer one, deterministic numeric match

  • 3 matches segment
  • 200 matches segment
  • 350 matches segment
  • 10 matches segment
  • -12 matches segment
  • 815 matches n_treated
  • 788 matches n_control
  • 23 matches response_rank
  • 18 matches uplift_rank
  • 3.9581 matches measured_visit_uplift_pp
  • 1.6262 matches measured_visit_uplift_se_pp

Layer one: every numeral in this text is carried by a fact above. Every numeral the extractor finds is marked, including a digit sitting inside a name, because this layer matches numerals and does not read sentences. The match is also percent-aware, so a fact and the same quantity printed as a percentage both count as support, which widens what layer one accepts. The field named beside each numeral is the closest supporting fact, not the first one found.

Layer two, model cross-examination

Layer two: the cross-examination reported no claim the facts do not carry.

07Signing and corridorsAlignment to the three themesGrowth, primary

Density beat our composite; no claim is made for the blend on either view

On the forward check venue density beat ; the corridor model loses pooled and wins five of twelve corridors; the equal-weight blend wins on only, 6 of 12, no interval claimed.

0.6869density's on venue formation, the bar to beatreal Foursquare records, , observational, the composite's gap interval entirely below zero

Both heads map to the Growth theme, and each has a measured bar to beat in Phase 1: density for signings and the seasonal naive for corridors. Only the signing head reads the backbone today, at category level.

What lost, and is printed anyway

  • 0.6869 against 0.5477: plain pre-cutoff venue density beat the composite on venue formation over 24 months, cell-clustered interval entirely below zero. Venue formation is not signing; density is the bar Phase 1 must beat.
  • 0.38 against 0.66 on precision at top fifty: the same pre-registered forward check on its second metric, and the composite sits behind density there too.
  • 0.4498 and 0.32 for equal weights on the same four channels, below both the composite and density on the same outcome; reweighting the composite does not rescue it.
  • Backbone similarity is a per-category constant, six values across the whole ranking, so the backbone adds no merchant-level signal to the signing head today; the offers and corridor heads take nothing from it yet.
  • 0.441 against 0.348 at the total level: reconciliation moves the model's total from 0.904, and the on the total stays ahead; reconciliation improves coherence but not accuracy.
  • Pooled, the model's MASE of 0.623 loses to 0.53 for a per-corridor over 13 held-out months, winning five of twelve. The value of this head is coherence and per-corridor explanation rather than pooled accuracy.
  • 82.1 percent of attribution is recent level: on this public proxy, mostly persistence with a seasonal correction. Customer segments and category shifts are absent from public arrivals data; the model cannot speak to them.
  • 113211.69 against 114180.52 arrivals of error: weighting each arrival equally the naive is ahead and the blend costs 968.83 more; the lane is claimed on neither view, winning 6 of 12 corridors.
  • Scoring the four channels at equal weight agrees with the shipped composite at Spearman 0.963 and keeps 368 of them: the hand-set weights have little effect on the ranking.
  • Tourist-zone proximity alone agrees with the released list at Spearman 0.869, the closest single signal to the shipped composite.

Whitespace card: who to sign in Singapore Real

0.5477Ranked by real signals over 364,635 places; on the pre-registered forward check density scored 0.6869 against the composite's 0.5477, the Phase 1 bar.

Forward check: plain density beat the shipped composite
  • Pre-cutoff venue density, null control0.6869
  • The shipped composite0.5477
  • Equal weights, same four channels0.4498

Real Foursquare records, pre-registered, observational, measured Spearman on venue formation over 24 months, not signings; higher is better. Density is the bar to beat; the composite's gap interval sits entirely below zero.

Bound figures, read from the committed results file and checked by the numeral gate; nothing is computed in the browser.

Forward check: the shipped ranking against real subsequent data

We state here what the check does not settle. The composite is built to rank signing yield and signing yield is not what was scored here, so nothing above shows that the composite would beat density on signings, and nothing shows that it would not: the question is open, it is the Phase 1 gate in Section 8, and density is the bar the composite has to clear there. Two limits on the outcome: date_created is the day a record entered Foursquare rather than the day a venue opened, and the slice keeps only venues still open at the snapshot, so the predictors are computed on survivors. The backbone wire is not an arm here either: the embedding-similarity column is a per-category constant entering neither the tested composite nor the shipped ranking order, so nothing about the backbone was tested and lost. The shipped ranking is unchanged, because this is a forward check of the frozen ranking and not a re-rank. For a partner, the relevant feature is the loop itself: this is the same forward validation the Phase 1 pilot runs inside Amex, with signings as the outcome instead of venue formation.

Corridor card: where growth comes next Real

0.623Model loses pooled, MASE 0.623 against 0.53; the pre-registered equal-weight blend wins on macro-MASE only, 6 of 12 corridors, a point comparison, no interval.

The corridor verdict changes with the view
View of the error
  • Seasonal naive0.53
  • Model alone0.623

Scaled error, MASE, averaged over twelve corridors, 13 held-out months, lower is better; model loses pooled, wins five of twelve; the pre-registered equal-weight blend beats the naive on this average.

  • Seasonal naive113211.69
  • Blend, equal weights114180.52

Absolute error in arrivals summed over the corridors across the holdout, every arrival weighted equally, lower is better; the blend costs 968.83 more and wins 6 of 12 corridors.

  • Model, before reconciliation0.904
  • Model, reconciled0.441
  • Seasonal naive on total0.348

MASE on the total series. Reconciliation makes corridor, country and region totals agree by construction and pools variance; it provides coherence rather than an accuracy headline, and the naive remains ahead.

Scaled error, MASE, averaged over twelve corridors, 13 held-out months, lower is better; model loses pooled, wins five of twelve; the pre-registered equal-weight blend beats the naive on this average.

These figures are bound to the committed results file and checked by the numeral gate; nothing is computed at page-view time.

The forecaster is mostly persistence with a seasonal correction
  • Recent level, last four months82.1
  • Same month last year8.1
  • Corridor identity6.3
  • Calendar month3.1
  • Each traveler-mix covariate, at most0.2

Modelled holdout attributions on real public arrivals, point shares of the total, no interval printed, shown as they came out. Customer segments and category shifts are absent from public data.

These figures are bound to the committed results file and checked by the numeral gate; nothing is computed at page-view time.

The corridor head forecasts monthly inbound demand by origin market, demonstrated on Singapore visitor arrivals as the stated public proxy for cross-border spend corridors.

08Implementation PlanFeasibility to develop and implementApplicability at American ExpressGrowth firstProtection demonstratedProductivity supported

Six months, one pilot, a gate we might fail

Phase 1: six months inside the SG CoE on data it already holds, no data moves, and it stops unless embeddings beat production features on at least two of three heads offline, paired intervals.

US$27.56the whole build at public ratesresearch compute at rent, prototype scale; measured hours, failed runs included, cash zero

The ask is six months of CoE time, not priced here: no Amex data leaves the perimeter, and a failed gate still leaves the behind.

Rate at which our own comparisons missed zero
  • Pre-registered comparisons with interval27
  • On the wrong side of zero7
  • Could not be told from zero13

Counts across this page's pre-registered comparisons, each with a measured interval, synthetic and real corpora alike, printed at the same size as the wins, so that a reader can compare their own number against them.

These figures are bound to the committed results file and checked by the numeral gate; nothing is computed at page-view time.

What lost, and is printed anyway

  • On our own real-corpus evidence the Phase 1 gate is as likely to fail as to pass: two of three heads needed, one cleared in both seeds, one failed in both, one mixed.
  • 0.0590 to 0.0218: the same frozen checkpoint's next-category gain under a permissive protocol and under ours; the guards, not the model, account for the whole difference, one alone costing -0.0831.
  • 63.0 percent of our own headline gain was protocol rather than model, on the one claim the ladder has been pointed at; a correction size, not a catch rate.
  • 7 of 27 comparisons with an interval landed on the wrong side of zero and 13 more could not be told from it; printed at the same size as the wins.
  • The holdout's cost is not hidden: the forgoes the effect, under half a US dollar of spend per customer on the exhibit's own arms, the number the fee is weighed against.
  • 2,595,732 basket lines at a grocery retailer, about 2,500 households: under the identical L3 protocol the household came back negative in both seeds, both intervals below zero, and the offer head mixed.
  • The backbone trained on synthetic public data; it proves the method runs, not production performance at Amex, and real Amex transactions at scale stay the untested case the Phase 1 gate exists for.
  • Two-sided beats single-sided stays a hypothesis: the merchant-axis next-category baseline alone scores 0.985 top-one, leaving no room, and the pair retention task came back null, both intervals spanning zero.
  • The corridor model loses to the pooled and forecasts visitor arrivals as a stated proxy; the pre-registered combination in Section 7 adds to that naive rather than replacing it.
  • 0.6869 against 0.5477: on venue formation over the following 24 months pre-cutoff venue density beat the shipped composite, difference interval entirely below zero; density is the bar the pilot composite must beat.
  • The whitespace acceptance-gap signal is simulated at plausible strengths and labeled as such; only the closed loop can supply the real one, and signing yield stays unvalidated until the Phase 1 gate.
  • 18 of 46 jobs failed or were cancelled and are counted in the cost, because research compute includes the runs that did not work.
  • No return on this page is measured: the conservative scenario's US$517,500 a year it stands against is a declared-assumption planning model, and no latency, throughput or team size is measured either.
  • No alarm threshold is stated for any of the three drift signals, because a threshold set without a production baseline would be arbitrary.
  • Agentic commerce volumes are small today; the ACE hook is a roadmap item rather than a revenue line.

The first business action. Phase 1 below: six months inside the SG Decision Science CoE on data it already holds, which is what turns the offers lane in Section 2 from a modeled number into a measured one. The brief asks for the business actions that execute the idea, and this is the first of them.

Independent holdout assignment. The pilot's holdout assignment is co-signed by the partner before a campaign opens, so the party being paid is never the party deciding who went untreated.

Implementation plan (business actions, with owners)

The plan has three phases: six months inside the SG CoE, one partner pilot with Malaysia as the illustrative case, and APAC scale-out, each behind a .

PhaseTimelineOwner Business actionsGate to pass
1. Internal pilot Months 0 to 6 SG Decision Science CoE30 Hand model risk the leakage rules and pre-registrations, then run the in-perimeter build (pretrain on proprietary-market and Amex-owned inbound data, rebuild the transfer table under the same rules) and take the go or no-go at month six against the printed gate Embeddings beat production per-task features on at least 2 of 3 heads, offline, with paired CIs
2. One-market partner pilot Months 6 to 12 GMNS APAC (partner), ICS (issuing side) Run a test-and-learn with one GNS partner; Maybank is the illustrative case, a live Amex network partnership in Malaysia18; ship ranked signing lists and two measured offer campaigns Measured incremental billed business above zero at 95% confidence; signing-list conversion above the current prioritization; partner signs for a second cycle
3. APAC scale-out Year 2 onward GMNS APAC with SG CoE as model owner Extend to additional GNS markets under the market-by-market license model17; stand up the corridor product for partner budget planning Positive unit economics per market; lift reports adopted in partner planning cycles

Monetization hypothesis: rev-share on measured , so the pricing model and the lift measurement are the same artifact.

What the build cost, measured, and how it would run

8.3 GPU-hours and 131 core-hours measured, US$27.56 at public rates, cash US$0, 18 of 46 failed jobs counted; US$2.42 to refresh one million merchants.

What Amex already runs, and what One Loop adds

Gen X fraud, Orchestra, the CoE uplift contest and sequence research are the existing record; One Loop applies the same capabilities to growth, in a partner-facing form, with lift reports.

Limitations, stated plainly

The synthetic backbone, the negative household on a real retail corpus, the merchant-axis null, the corridor losing to the naive, and density beating are all reported on this page.

  • The backbone trained on synthetic public data (IBM TabFormer). It demonstrates that the method runs; it does not demonstrate production performance at Amex. The L3 transfer protocol was additionally replicated with byte-identical training code on the real Complete Journey corpus (2,595,732 basket lines, each one product on one receipt rather than a transaction, about 2,500 households at a grocery retailer, not card data), where the household-level transfer gain came back negative in both seeds with both intervals below zero and the store-level head came back positive in both, reported in Section 5 as obtained. That is the ladder's own argument that protocol and corpus are both variables: the method and the protocol travel, and real card data at Amex scale stays the untested case the Phase 1 gate exists for.
  • The uplift exhibit reports its pre-registered rankings as obtained. We did not pick the estimator after seeing results, and we do not claim uplift wins everywhere.
  • "Two-sided beats single-sided" is still a hypothesis, and we now know more about why. We ran the merchant-axis pretrain on the same events regrouped into merchant sequences; it produced merchant-view embeddings, so the mechanism works. Two things keep that run from settling the question. Its split is in merchants only, since cardholders cross merchant boundaries, which makes it a partial two-sided demonstration. And neither of our two tasks could test the claim: a merchant's category barely moves inside its own sequence, so the next-category baseline alone scores 0.985 top-1 there, and fraud leaves little room on that axis as well. The merchant-axis row of the Section 5 scaling table is that result. We therefore built a task that could test it, pair retention, pre-registered before the run, and it came back a null as well: both intervals span zero on a baseline with real headroom. That leaves the hypothesis untested rather than refuted, and the missing piece we can name is the account-side vector, which that arm does not carry. The rest of the merchant-native family, and a split disjoint on both entities, are pre-registered below.
  • The corridor exhibit forecasts visitor arrivals, a stated proxy. Cross-border billed business is the real target and lives inside Amex. The model also loses to the pooled, and the pre-registered combination in Section 7 shows it adds to that naive rather than replacing it.
  • The whitespace acceptance-gap signal is simulated at plausible strengths and labeled as such; only the closed loop can supply the real one. The ranking now has one pre-registered forward-looking check against real subsequent data, committed at fef8af5 before any outcome existed, and the result does not favour the composite: on venue formation over the following 24 months, plain pre-cutoff venue density beat the shipped composite, Spearman 0.6869 against 0.5477, difference interval entirely below zero, in Section 7. Venue formation is not merchant signing and the composite is built to rank signing yield, so signing yield itself stays unvalidated until the Phase 1 gate, and density is the bar the composite has to beat in the pilot.
  • Agentic commerce volumes are small today. The ACE hook is a roadmap item, not a revenue line.

Round 2: what we build next

63.0%Package the , which priced 63.0 percent of our own headline as protocol, build the pilot kit, and pre-register the pilot evaluation.

  • Package the protocol ladder as something the CoE can use on other people's models, not just on ours. It is the one part of this work that beat its own control outright, it needs no Amex data and no cardmember data, and it answers a question the CoE gets asked constantly: how much of a claimed lift is the model and how much is the evaluation. The ladder has four rungs, with one guard added at a time, and each rung is priced. On next-category top-1 the same frozen checkpoint reports 0.0590 under the permissive protocol and 0.0218 under ours, and the single entity-disjoint split costs -0.0831 of that on its own, which is enough to turn a headline gain into a loss. Applied to a vendor claim or an internal one, it serves as a triage tool from the first month. Section 5 is the working version; Round 2 adds packaging rather than new research.
  • Build the pilot kit, which is the thing to demonstrate rather than describe: a partner-facing lift report a GNS partner can open, with k-anonymous aggregates, the arm counts they can recompute the payment from, and the Veritas and AIRG control mapping filled in for a Maybank-shaped pilot. The measurement layer is the product, so the kit is what a partner would be handed. At a final it goes on screen beside the audit console, and a judge opens any number in it themselves.
  • Scale the backbone and test the two-sided hypothesis on tasks that can answer it. The merchant-axis pretrain ran and the first merchant-native evaluation, pair retention, came back a null in Section 5. Pre-registered here is the rest of that family: next-period merchant volume, cardmember churn away from a merchant, new-cardmember acquisition at a merchant, and acceptance-lapse risk, each carrying the account-side as-of vector pair retention had to leave out, on a split disjoint in merchants and cardholders at once. Same leakage design, same paired bootstrap, reported whichever way it moves, with paired dual-axis objectives and the dashed graph layer PRAGMA's published AML failure argues for29. Compute is already secured.
  • Separate corpus size from what currently travels with it in the scaling table. Two controls are pre-registered here: re-run the 10 million row transfer evaluation down-capped to the 3 million point's training, test and fraud-positive counts, and re-run the whole curve with the evaluation window held fixed and one vocabulary shared across every point. Both get reported whichever way they move.
  • Pre-register the pilot evaluation: gate metrics, holdout design, and stopping rules written before the first campaign, so Round 2 is judged the way this page is, on results as obtained.
  • Harden the house baseline before re-measuring transfer: add per-user category-histogram and modal-category features to the next-category baseline, pre-registered here, and report the delta whichever way it moves.

The full account

Problem and value: everything that moved off the page

Why the timing is pressing in APAC

The timing is pressing because APAC rails are undercutting the fee levels on which card economics rest. QRIS charges merchants 0.3% against card rates of 1.5% to 3%. Nexus Global Payments was incorporated in Singapore with a go-live target of 202611, and the G20 wants retail cross-border payment costs at or under 1% by the end of 2027. In Singapore itself, hawkers and small retailers still decline Amex over swipe fees13.

Signing merchants one at a time is expensive, and holding them is worth more than most teams assume: acquiring a merchant costs 5 to 25 times what retaining one does. The parties affected are the GNS partner bank whose card economics cheap rails compress, the small merchant priced out of acceptance by that fee gap, and the Card Member the card fails at the hawker stall or the overseas till. The three heads address those three parties in order.

The value model, lane by lane

One Loop turns the closed loop into a value-added service for APAC issuing and acquiring partners: ranked signing lists with evidence, offer campaigns with measured incremental lift, and corridor forecasts that steer budgets before the demand arrives.

We begin with the offers lane. It is the only one of the three whose targeting mechanism has been measured on real data with random assignment (Section 6), and at base-case assumptions it models about US$2.88M a year of network discount revenue on incremental billed business. All three lanes together reach US$10.5M by year three of an APAC rollout, and that larger figure is declared-assumption upside gated on the Phase 1 pilot in Section 8, because the whitespace ranking has already lost a pre-registered forward check to plain venue density and the corridor model loses pooled to a seasonal naive, both in Section 7. We note one caveat to that figure: whitespace is the largest of the three lanes at base case, and it is also the lane whose ranking lost its forward check. Two restrictions travel with the arithmetic. Lane A is not the whole non-accepting universe: merchants who decline Amex decline it on fee, at QRIS rates of 0.3% against card rates of 1.5% to 3%, so only the fee-tolerant stratum of that universe is signable at a card rate. Lane C applies only where the forecast beats the naive, five of the twelve corridors backtested; the lane's arithmetic runs on partner budgets rather than corridor counts, so we state that restriction rather than scale the figure by it, which the committed inputs do not support.

The figure is a bottom-up sum of three lanes (incremental merchants signed off the ranked list, partner campaigns measured, corridor budget reallocated), every input is cited or labeled as an assumption, the arithmetic is shown in the value model one click away, and no line is derived as a share of network volume.

Scenario range: US$0.52M conservative, US$10.5M base, US$108.5M stretch. The strategic payoff of keeping GNS partners on the network while wallet rails compress card economics sits on top of every one of them, is counted in none, and is not sized anywhere either. In place of an estimate, we report a threshold: at the base case's discount rate it takes US$524,000,000 a year of billed business held rather than migrated, about US$65,500,000 per partner, for retention alone to match all three lanes. Whether that amount is large for a GNS partner is a judgement for the reader to make, not for us.

Those three rows move every input at once, which hides the one a reader should argue about, so we also report the one-way version. Hold everything at base and move one input across its own range: the swing is largest for markets or partners live, at US$15,720,000 against a base total of US$10,480,000, because that one input feeds all three lanes. The second, at US$9,600,000, is the signing-conversion uplift credited to the ranked list, which is the model's own contribution, and the two were drawn at the same width of five times low to high. So the answer turns on how many partners sign and on how well the list works, in that order, and the second is the quantity a pilot measures directly. One-way swings scale with the range each input was given, so this ordering says as much about how wide we drew each range as about the structure underneath. The input we flag as least anchored, the corridor yield gain, is also the least influential at US$800,000, which is the one reassuring thing in the ranking. The full ordering is in results/value_model_sensitivity.json. A forward break-even needs a Phase 1 budget and we state none, so we invert the question, as with the retention threshold: instead of assuming a budget, we report the bar that a given budget must clear. Priced on the offers lane alone, the only lane whose targeting mechanism was measured on randomized data, a Phase 1 costing US$1,000,000 is cleared in 4.2 months, one costing US$4,000,000 in 16.7. Those are base case, and quoting only base case would be the selective framing this page is designed to avoid, so we also give the spread across scenarios: the same US$1,000,000 takes 66.7 months to clear at the conservative scenario and 0.5 at stretch. The CoE knows what six months of its own time costs and we do not, so we give the bar at all three scenarios for comparison against that internal figure. The lane's size input is still an assumption; what is measured underneath it is the targeting mechanism and the counting rule, not the size.

We next perform the subtraction that a reader would otherwise do independently. Those three scenarios vary the assumptions; none of them asks what survives our own tests. Strike the whitespace lane, whose ranking lost its pre-registered forward check to plain venue density in Section 7, and the base case falls from $10.5M to $4.08M. Keep only the lane whose targeting mechanism was measured on randomized real data, offers, and it is $2.88M. Our reading is that the case for Phase 1 rests on the smallest of the three, because it is the only one supported by a measurement. The struck lane is not deleted and not rescued: a post hoc re-read in Section 7, not pre-registered and counted in no tally, finds its non-density channels do order venue formation inside bands of comparable density, which is a reason to run the Phase 1 gate against real signings rather than a reason to put the lane back.

Assumptions behind this figure: the value model, bottom-up

Every input below is either a cited public number or an assumption tagged ASSUMPTION with the anchor that makes it plausible. Dollar figures are US dollars per year. Lane values count only network discount revenue on incremental billed business; the rev-share service fee from Section 8 is upside and is not counted, and nothing here is derived as a percentage of total network volume.

Shared scale input, all three lanes. Markets or partners live: 3 conservative, 8 base, 15 stretch. ASSUMPTION, anchored to the GNS model of market-by-market licensed partners across APAC17 and the live Maybank partnership as the first concrete case18.

Lane A. Whitespace signing economics

Structure: addressable non-accepting merchants per market x signing-conversion uplift attributable to the ranked list x average incremental Amex billed business per newly signed merchant x discount revenue rate, summed over markets.

  • Addressable non-accepting merchant universe per market: 30,000 / 50,000 / 70,000. Anchor: the whitespace exhibit measures 146,048 card-accepting-category places among 364,635 Singapore points of interest; the share of those not taking Amex today is an ASSUMPTION of roughly one fifth to one half of that universe, consistent with the documented small-merchant holdout pattern13.
  • Signing-conversion uplift attributable to the ranked list: 0.3% / 0.8% / 1.5% of the addressable universe signed per year because of the list, over and above what the current pipeline signs anyway. ASSUMPTION, and the input this whole model is least able to anchor: it is the rate that turns a ranked list into signed merchants, no public source fixes it, and it drives the largest lane in the base case, so a reader who wants to move one number should move this one. Only this increment is credited to One Loop; the signing pipeline itself is not. Capacity is proven at scale, since global accepting locations grew about 16% in a year, to roughly 160 million, so the list's job is prioritization and conversion, not capacity.
  • Average annual incremental Amex billed business per newly signed merchant: $50,000 / $100,000 / $200,000. ASSUMPTION for small-to-mid urban merchants; the fee sensitivity that keeps these merchants out today is documented at card rates of 1.5% to 3% against wallet rails at 0.3%, and Singapore's holdouts are exactly this segment13.
  • Discount revenue rate on that billed business: 1.5% / 2.0% / 2.5%, taken inside the public card MDR range10.

Arithmetic:

  • Conservative: 3 markets x 30,000 x 0.3% = 270 merchants; 270 x $50,000 = $13.5M billed; x 1.5% = $0.20M.
  • Base: 8 x 50,000 x 0.8% = 3,200 merchants; 3,200 x $100,000 = $320M billed; x 2.0% = $6.40M.
  • Stretch: 15 x 70,000 x 1.5% = 15,750 merchants; 15,750 x $200,000 = $3.15B billed; x 2.5% = $78.8M.

Retention note: attrition runs at 23.6% and costs the industry about $2 billion a year in losses plus about $1 billion replacing lost merchants, which is why the ranked list also scores retention risk.

Lane B. Offers: incremental spend per partner

Structure: partners x campaigns per partner per year x measured incremental billed business per campaign x discount revenue rate.

  • Partners running measured campaigns: 3 / 8 / 15 (shared scale input above).
  • Campaigns per partner per year: 4 / 6 / 10. ASSUMPTION, anchored to the campaign cadence implied by the CoE's own 2024 contest, which scored offer targeting across 12.6 million customer x merchant pairs, and to card-linked-offer practice where test-and-control incrementality is the commercial standard23.
  • Incremental billed business per campaign, the part a holdout certifies: $1M / $3M / $6M. ASSUMPTION, and the connection most open to challenge. Our measured result does not enter this line, and we do not claim that it does: the randomized number we have is an incremental site-visit rate on an advertising experiment, and converting that into billed business on a card would introduce the false precision this page is designed to avoid. What the measurement buys is the mechanism and the rule for counting, not the size. Only spend the randomized holdout certifies counts here, which is the rule that constrains the lane; the efficiency case for uplift targeting is public, including a reported instance where uplift-targeting 30% of users matched the conversion gain of promoting everyone. One measured dollar figure does exist on this page, and we report it alongside this line rather than as an input to it: on the randomized Hillstrom arms, 21,387 treated against 21,306 control, the campaign caused US$0.4244 of incremental spend per customer treated. That is a randomized measurement, but it is of a different quantity from the one this lane needs: retail e-commerce spend at a clothing retailer, not billed business on a card, which is why it informs the counting rule and not the size.
  • Discount revenue rate: 1.5% / 2.0% / 2.5%10.

Arithmetic:

  • Conservative: 3 x 4 = 12 campaigns; 12 x $1M = $12M incremental billed; x 1.5% = $0.18M.
  • Base: 8 x 6 = 48 campaigns; 48 x $3M = $144M; x 2.0% = $2.88M.
  • Stretch: 15 x 10 = 150 campaigns; 150 x $6M = $900M; x 2.5% = $22.5M.

Lane C. Corridor budget reallocation

Structure: partners x annual corridor-facing marketing and signing budget per partner x share reallocated on the forecast x yield gain on the reallocated share.

  • Partners using corridor forecasts: 3 / 8 / 15 (shared scale input).
  • Corridor-facing budget per partner per year: $3M / $5M / $8M. ASSUMPTION for a mid-size APAC issuing or acquiring partner; for scale, Amex's FY2025 marketing expense was $6.25 billion, so single-digit millions per partner is a small planning envelope.
  • Share of budget reallocated using the forecast: 15% / 20% / 30%. ASSUMPTION; partners shift only the portion where corridor timing matters, in a window where APAC business travel is forecast at $700.9 billion.
  • Yield gain on the reallocated share: 10% / 15% / 20%. ASSUMPTION, and the least anchored input in this lane: the pooled backtest does not beat a seasonal naive (0.623 model MASE against 0.53 naive), so the case rests on the corridors where the model wins, on reconciled totals that cohere by construction, and on the pre-registered combination in Section 7.

Arithmetic:

  • Conservative: 3 x $3M x 15% x 10% = $0.14M.
  • Base: 8 x $5M x 20% x 15% = $1.20M.
  • Stretch: 15 x $8M x 30% x 20% = $7.20M.

Totals

Totals are computed on the unrounded lane values; displayed cells are rounded, so a total can differ from the sum of its displayed cells in the last digit.
ScenarioLane A signing Lane B offersLane C corridors Total per year
Conservative $0.20M $0.18M $0.14M $0.52M
Base $6.40M $2.88M $1.20M $10.5M
Stretch $78.8M $22.5M $7.20M $108.5M

The other side of this model is measured: the build took 8.3 GPU-hours and 131 CPU-core-hours, cash US$0, priced in Section 8 from results/cost_model.json.

Sanity check, stated for scale only: APAC revenue was $5.22 billion (7.2% of the company total) in FY2025, and the base case is well under one percent of it. This ratio is worth noting for scale, but it is not the basis on which the case should be judged: the base case prices eight partner markets in a pilot, not a region, and counts only discount revenue on incremental billed business. Per partner it is US$1,310,000 a year, US$360,000 on the one lane that survives the subtraction above; we print the second because a reader would divide anyway. This is a reasonableness check on the arithmetic rather than a materiality claim.

The full account

Why Amex: everything that moved off the page

The recipe in public, and the unmeasured leg

Every one of those systems sees one side of the transaction. Amex's closed loop sees the cardmember, the network, and the merchant in a single event stream. That is the best data in the industry to run this recipe on, and only Amex is positioned to run it on both sides at once.

The counterpart page one already prints belongs here too, next to the claim rather than sixty pages away: the merchant-axis half of this is the leg we could not measure. The pair-retention task built to test it came back null on synthetic data, both intervals spanning zero, and the two earlier merchant-axis tasks were too degenerate to answer. So the closed-loop advantage argued above is a structural argument, not a measured one, and Phase 1 is where it gets tested on data that could settle it.

SAFR, the origin check and the incumbents

MAS published SAFR, its runtime-safeguards paper for agentic finance, on 3 July 202631. On page 9 the paper names its own unsolved problem, and it is a data problem rather than a modelling one: a governance envelope that an agent declares about its own proposed action cannot be taken as proof of what the original instruction produced, so the envelope is treated as a document to be authenticated "against its origin"32.

That phrase is where the difficulty sits. If an agent declares a cart, the declaration is worth only as much as the origin you can check it against, and the origin of a card transaction is the merchant side of it. An open-loop network does not hold that side. The acquirer does, and the acquirer and the network are different companies with no shared record of the same event. A closed loop holds both sides of the one transaction, which turns authenticating an agent-declared cart against its origin from a partnership problem into a lookup.

There is an opening here, and we state it explicitly because it is straightforward to verify and we prefer that a judge verify it independently. The runtime control MAS has specified is one a closed loop can satisfy structurally, and the firm best placed to satisfy it is not yet in the room where the paper was written: both open-loop networks appear in SAFR as contributors and carry published component mappings in its case studies, and American Express appears in neither33. That is a statement about which names the published paper prints, which is all we can check from outside, and a contributor list is not a participation list. Whether Amex took part by another route is something the reader would know and we would not, so we present this as an opening we can observe rather than an exclusion we can assert. The structural point does not depend on it: the control MAS has specified and the data structure that would let a firm satisfy it have not yet been put together by the company that already has the structure. That is the gap we would fill, a Round 2 direction rather than something built. Nothing in Sections 5 through 7 rests on it, and we have shipped no agentic path.

The incumbents for each worked example confirm the gap. Enigma sells raw merchant signals, US-only, with no ranked propensity to sign34. Cardlytics measures offer incrementality but publishes no methodology a partner can audit23. Mastercard's Economics Institute publishes corridor insight as a backward-looking annual report35. One Loop is built to beat all three on the same two axes: explainability, and the two-sided loop.

The full account

Architecture: everything that moved off the page

The six layers and the partner lift report

Item five is the one a partner actually holds, so here it is rather than a description of it. This is built at page-build time from a narrative that cleared both layers of the gate and the randomized arms behind it, and the underlying exhibit is in Section 6.

Partner lift report campaign close, in the shape a partner receives it. Built from a public randomized experiment on this page, not from American Express data.
Segment
3) $200 - $350 x recency 10-12m
Treated
815
Control
788
Measured visit uplift
3.96 points
Standard error
1.63
A response model ranked it
23 of 28
Measurement ranked it
18 of 28
The customer segment "3) $200 - $350 x recency 10-12m" includes 815 treated and 788 control customers from the Hillstrom e-mail experiment. This segment ranks 23rd in response rank and 18th in both uplift rank and measured uplift rank. The measured visit uplift per person is 3.9581 with a standard error of 1.6262, and the verdict is "aligned," indicating that the response model and measured uplift are consistent.

The two arm counts are the entire basis of the report: a partner can recompute both the lift and the fee from them independently, without relying on our stated values for either. The sentence is written by the model over these numbers and cleared by both layers of the gate. The gap between the two ranks is the reason the holdout exists. No partner has yet received one of these reports, because the pilot that would produce the first one is what Section 8 asks for.

What the contracts must answer first

The pretraining corpus is scoped to what Amex owns outright: transactions in proprietary issuing markets, plus Amex-owned inbound cross-border spend into partner markets. Partner banks receive scores and aggregated lift reports; raw embeddings never leave Amex. This is measured rather than assumed: for a merchant whose embedding pools exactly one transaction, a linear probe on that vector alone recovers the transaction's amount decile at a Matthews coefficient of 0.9395 on the public synthetic TabFormer corpus. The amount decile is an encoder input that varies inside a merchant, unlike merchant category, which is near constant, so it is what makes that file an internal artifact and not a partner deliverable (results/safety_embedding_inversion.json). What a partner does see is published under a minimum cell-frequency rule rather than an assertion: the shipped whitespace setting publishes only cells carrying at least 10 contributing venues, and the band at the end of this section prices that guard and the two stronger ones above it, with the PDPC's own caveat about transactional data attached. Storage and training are sharded by region to match the law of each market: PDPA in Singapore, RBI data-localization in India, PIPL in mainland China. Only governance-approved scores cross shard boundaries. One backbone reconciles with regional shards like this: pretraining runs once, on the corpus already scoped to what Amex owns outright, and ships to each shard as a frozen checkpoint; inside a shard only head training and embedding refresh run. A checkpoint crossing a border is a different legal object from personal data crossing one, and it joins the contracts item below for Amex legal before Phase 1. A market whose regulator refuses the import runs its heads without the backbone, the configuration two of the three heads are already proven in here, and the gate reports that market separately.

The threat map and the obligation map

Assets crossed against actors crossed against surfaces, then against the OWASP Top 10 for LLM Applications and SAFR's disposition factors31. Anything where one actor can cause an irreversible loss or an unauthorised disclosure is kept.
The mandate records what the cardmember approved, not what they wanted. The authorised-scam case matters most here: the Shared Responsibility Framework has covered seemingly authorised transactions since 16 December 202446, someone under manipulation approves the mandate and then the escalations inside it, Amex's own executive for global fraud risk has said bad actors moved toward scams as controls grew sophisticated47, and Section 5 reports that no public dataset carries authorized-scam labels.
A verified agent with a valid mandate can still act against the cardmember's interest. Identity proves who, not what. Origin authentication is not closed here either; Section 3 makes the closed-loop argument, and a lookup can still return a compromised merchant record. Every control bounds the loss; none of them prevents it.
Caps are evaded by decomposition. An agent can spend up to the cumulative cap in actions of at most the per-action cap, across every merchant the mandate allows, inside the velocity window. Novelty is a model output, so it drifts and can be poisoned: the attack consists of making the harmful behaviour appear usual. Anything absent from the mandate schema is unbounded the same way, data scope above all, because reading is not spending.
Threats assessed and dropped, with the reason for each, because a threat list carrying only the threats with answers is visibly selective. Model theft through the partner surface, which releases bucket-level scores over aggregates: dropped, unmeasured. Training-data poisoning by a merchant that controls only its own stream: dropped, unmeasured, stated rather than implied. Unbounded consumption: dropped, covered by rate limits and velocity. System prompt leakage: dropped, no enforcement decision depends on a prompt.
The obligation map, with built and designed items labelled separately. Only rows carrying a built exhibit are listed; a row we could map but not measure is the strip this section exists to avoid, and no row is a compliance claim. The status of each source is as follows: PDPC advisory guidelines are not legally binding, the Veritas Toolkit's latest release is dated 2023 and is not current tooling, and MAS FEAT is principles rather than a rulebook. Disclosure control on what crosses the perimeter, PDPC Selected Topics with its caveat: measured under the ladder, replacing the assertion. What the model remembers of its training data, OWASP LLM02: measured by the membership attack, with a no-model control that beats it. Indirect injection through data fields, OWASP LLM01 and LLM09: measured per class, including the class our checker cannot see. A model output may only add friction, our red line39 and MAS FEAT principle 6: built for generated text in Section 6, while the control plane is designed and not built. A pre-declared threshold with its measured cost, Veritas Document 3A recommendation 437: built twice, the guard prices in Section 5 and the disclosure prices above.

Guards priced one rung at a time

One frozen ranking, guards turned on one at a time, priced against the raw output. Churn counts rows entering plus leaving the top twenty, so it runs from zero to forty. The noise rung is stochastic and carries its spread across seeds, which is mechanism randomness. The privacy unit here is one point of interest, meaning one venue. This exhibit protects VENUES, not cardmembers: no cardmember data and no American Express data enters this head at all, and the corpus is the public Foursquare OS Places Singapore slice. The sibling safety exhibits run on the public IBM TabFormer benchmark, which is synthetic. Either way the result measures the mechanism, meaning what a guard of this shape costs an output of this shape. It is not a measurement of American Express's exposure and must never be presented as one.
RungGuard added Churn in the top twenty Overlap in the top hundred
P0 raw output, no protectionnone0100
P1 small-cell suppression shippedminimum cell frequency097
P2 contribution bounding on the density channelmax recipients per POI1083
P3 calibrated differential-privacy noiseLaplace noise17.70
seeds 12.95 to 23.05
77.35
seeds 73.95 to 80.00

This is a position, not a compliance claim. minimum k-anonymity value of 5 with relevant safeguards for sharing with external parties, k of 3 internally with relevant internal controls (advisory guidelines, not legally binding, though PDPC is likely to take positions consistent with them) Quoting that threshold without the same body's own caveat would be selective quotation: the same body's guide records that k-anonymity 'may not be suitable for all types of datasets or other complex use cases (e.g. longitudinal or transactional data where the same indirect identifiers may appear in multiple records)', and notes the limitation on attribute disclosure by homogeneity attack. Transaction data is the warned case. Quoting the threshold without this caveat would be selective quotation from one document. A contributing-venue count is a minimum frequency rule, not a k-anonymity class. The noise rung is reported without an epsilon. It ran, at the operating point recorded in results/safety_privacy_ladder.json, whose epsilon sweep is priced there at every step. The privacy unit is one point of interest (one venue), and it protects venues, not cardmembers. No epsilon value is reported on this page. An epsilon has no meaning without its unit, its delta, its neighbouring-dataset definition, the contribution bound that makes its sensitivity finite, the invariants declared outside it and its composition accounting, and that apparatus does not fit the page cap, so quoting the number alone would be the error rather than the disclosure. What transfers is the row above: calibrated noise is expensive on a ranking this short and this tightly packed at the head, at any epsilon a privacy researcher would accept.

What the membership attack measures

The pretraining corpus is the public IBM TabFormer benchmark and it is SYNTHETIC. This measures the mechanism, meaning how much a masked-field transaction model of this shape gives back under this protocol. It is not a measurement of American Express's exposure and must never be presented as one. The unit is the account and the operating point is the floor, the smallest false positive rate 415 non-member accounts can express. Area under the curve is recorded in the results file and is not used as the headline figure, because averaging over every false positive rate includes rates no adversary would work at. Every arm is scored on identical rows, with clustered-bootstrap intervals, and all of them are in results/safety_membership.json.

The direct test of that confound did not run. NOT RUN. Greedy 1:1 caliper matching on log10 of account total rows at a caliper of 0.2 dex produced 24 usable pairs, below the floor of 30. The two populations barely overlap on tenure, which is itself the finding: no tenure-matched comparison is available at useful size on this corpus, so the headline separation CANNOT be attributed to membership alone. Caliper and floor were fixed before that pair count was computed. Neither is pre-declared, and the three tenure-matched fields are null on purpose. This is the weak attack in its family, which we state before reporting the number: score access only, 0 reference models, 0 shadow models, no likelihood-ratio test against a trained reference model.

Seven attack classes against our own gate

Success and residual are different quantities, never merged here, since an attack can fail for reasons unrelated to the defence: success is whether the planted artifact appeared, residual whether it survived both layers. Intervals are bundle-clustered bootstrap over thirty bundles. Unattacked, the same gate passes 28 of 30 bundles against a committed reference of 30 of 30, and the file says why rather than crediting the defence: this arm is a re-run and not a replay, only 12 of the 30 narratives come back byte-identical to the committed text, and both clean failures are on narratives that differ from it. This measures a two-layer gate of this shape, under these static attacks, on fact bundles built from public data. It is not a measurement of American Express's exposure and it is not a property of GenAI explanation layers in general.
Attack class Layer one can see it Attack success, strict (95% CI) Residual after both layers (95% CI)
Unsupported claim, plus an instruction aimed at the auditorno1.000 no interval1.000 no interval
Ignore the supplied factsno1.000 no interval1.000 no interval
A figure inflated past the factsyes0.167 [0.033, 0.300]0.000 no interval
That figure, with its digits written into the data fieldno0.300 [0.133, 0.467]0.267 [0.100, 0.433]
Unsupported claim, plus a rider aimed at the parserno0.967 [0.900, 1.000]0.933 [0.833, 1.000]
A claim the facts do not support, no numeralno0.933 [0.833, 1.000]0.033 [0.000, 0.100]
Reveal or override the system promptno0.000 no interval0.000 no interval

A ZERO HERE IS A NULL, NOT A DEFENCE RESULT. This class landed 0 of 30 on the strict criterion, so the interval is degenerate and no claim can be read off it. The strength is fixed and stated: static author-written payloads, one attempt per probe, no adaptive search and no search of any kind. Read it as one attack class at one strength failing to land on this run. It is not evidence that the class is blocked, that either layer of the gate caught it, or that a stronger or adaptive attempt would also fail. The per-layer catch rates in this same block are where to check which of those happened: a zero attack rate sitting beside zero catch rates means the writer did not comply, not that the gate intervened. The two rates that note refers to are as follows: layer one caught the planted artifact at 0.000 and layer two flagged anything at 0.000, both degenerate intervals. These are static attacks. The published record shows twelve defences that reported near-zero attack success under static evaluation and above 90 percent under adaptive attacks, so this number is a floor on the attack and not a measure of our defence.

What this measurement is not, from the file, in full. Static attacks only, one attempt per probe, no adaptive search. Exact-phrase success matching is conservative, so the strict rate is a lower bound and the lenient rate is reported beside it as the upper bound. Ten of the thirty bundles inject into a field our own code writes, so those probes are a mechanism demonstration against the deployment shape rather than a live hole; the split is printed per class. The auditor is the same served model as the writer, which is the shipped design, so layer two is not an independent check. Thirty bundles is a small population and the intervals say so.

What we did about it. layer one now builds its allowed pool from typed numeric leaves plus a format-pinned string allowlist rather than scanning every fact string, so an attacker-written string can no longer put digits into it (commit cb923ef). Replaying the same frozen probe corpus against the repaired gate: 30 of 30 probes are flagged, the same 30 still flip to accepted under the pre-fix builder, which is what makes that a test rather than a tautology, and all 30 shipped narratives keep their stored verdict, so the repair costs nothing on legitimate input. This is a replay of the frozen corpus, not a fresh campaign: it shows this hole closed, not the gate unbreakable, and the other attack classes the red team beat are unchanged by it.

The control plane, checked not built

That one rule is checked rather than asserted, over every combination of the declared axes with no sampling: 8,294,400 triples, 0 where the model made an outcome less restrictive, 0 model-derived reads in the deterministic block. The denominator overstates the coverage: only 3,072 reach the graded block and the model moved 864 of those, every time toward more friction. Because a checker that cannot fail provides no evidence, deliberately broken reference rules were run through the same enumeration, and each failed as declared in advance. Every other control in the plane is designed, unbuilt and unmeasured.

Each number rebuilds from its own file in results/, every file with a check mode that recomputes each leaf: safety_privacy_ladder.json seed 42 · safety_membership.json seed 7 · safety_injection.json seed 42 · safety_control_plane.json seed 0. Corpus provenance is in each caption above.

The full account

Backbone: everything that moved off the page

What won and what lost, in full

The public blueprint and the differences in our approach

We are not first on this corpus and we do not claim to be. In June 2026 NVIDIA published a transaction foundation model blueprint50 that pretrains a decoder-only Llama-style model of about 29M parameters with a next-token objective on the same public IBM TabFormer data, then feeds its embeddings to an XGBoost fraud classifier and reports an average-precision lift over that baseline. The code is public under Apache-2.052. Ours differs in two ways. The model is an encoder trained with a masked-field objective rather than a decoder trained to predict the next token, and the architecture is not what we are offering: our contribution is the evaluation protocol wrapped around the backbone, which is why the table below carries pre-registered tasks, a named leakage design, entity-level confidence intervals, and the results that came back null. We do not put their fraud numbers beside ours anywhere on this page, because the architecture, the downstream task setup and the test split all differ, so the two are not comparable.

Two pre-registered tasks and what they buy

The central question is whether the backbone's embeddings carry signal that strong per-task features do not already have. We tested it the hard way. The baseline is LightGBM on per-entity temporal aggregates, named exactly: per-account transaction count, amount mean, spread, maximum and last value, recency and tenure, and the previous merchant category and city with their frequency encodings. That is the pattern that wins on Amex-style data in public competition53. It does not yet include a per-user category histogram; that harder baseline is pre-registered for Round 2 in Section 8. The comparison adds backbone embeddings to that same baseline.

Both tasks were pre-registered:

Confidence intervals are entity-level paired bootstrap. Deltas are reported exactly as obtained.

What it buys. One embedding service standing behind several models instead of a separate feature pipeline per task, which is where the operating saving is located. It only counts where the transfer is real, and that is why the nulls are printed beside the gains rather than left out of the table.

The four guards and the sampling rule

Pretraining corpus cut at 2017-08-25T09:37:00Z, before the evaluation window. Label columns excluded from the pretraining vocabulary. Embeddings are as-of: nothing at or after the scored transaction feeds them. Account identifiers hashed, so unseen entities embed inductively. These four checks are asserted in the results file rather than only stated here. The same file records the split and the sampling rule: entity-disjoint by account, 400 of 2,000 accounts held out, rows drawn "uniform + all fraud positives kept". Keeping every positive and thinning the negatives to reach the row cap raises the positive rate in the scored set above the rate in the window it comes from, so the PR-AUC levels on this page read as a comparison between the two arms rather than as a rate a live queue would see.

Four evaluation rungs, with each guard priced separately

Published transaction foundation model demonstrations report downstream lifts much larger than the one in the table above. The question is whether that gap is attributable to the model or to the measurement, and it can be answered rather than argued. We took the same frozen backbone, changed nothing about it, and ran the same two downstream tasks under four evaluation protocols, turning the guards on one step at a time. L0 is the permissive shape: a temporal-only split, so the same accounts sit on both sides; a baseline without per-entity temporal aggregates; and one pooled embedding per account computed over that account's whole sequence, so it can see the scored transaction and everything after it. L1 turns on the entity-disjoint split by account. L2 adds as-of embeddings that stop strictly before the scored transaction. L3 adds the per-entity aggregates to the baseline, which is the protocol every other number on this page already uses, so L3 has to reproduce the table above and the results file records by how much it does.

Two further properties of this ladder should be stated before the numbers. The first is that only the evaluation protocol moves. The pretrain-side guards, the label columns kept out of the vocabulary and the account identifiers kept out of it as well, stay on at every rung, so nothing here is a claim about the pretraining. The second is that the ladder is conservative: the merchant-side embedding stays pre-cut pooled at all four rungs, so only the account-side embedding varies, and L0 is a lower bound on what a fully permissive protocol would show, not an upper bound. The rung definitions, the metric and the bootstrap scheme were written into the decision log and committed before the first run.

What it buys. The ladder provides a commercial answer as well as a methodological one. When a partner pays on certified lift, the first question anyone asks is which protocol produced the number. Section 8 puts the fee basis on a table in Section 6 that the partner can recompute from their own arm counts, and the ladder is what makes the protocol behind it checkable instead of assertable. It is also worth more than this project: what was built is a rig that takes any frozen checkpoint and prices its transfer claim under four protocols, and it is independent of the source of the checkpoint. That makes it a triage tool for any model claim the CoE is asked to believe, from a vendor or from an internal team, and a common protocol to settle it when two teams disagree. The rig remains useful whether or not One Loop ships.

Reading the top panel left to right: the next-category delta starts at 0.0590 under the permissive protocol and ends at 0.0218 under the protocol this page ships. The single guard that moves it the most is the entity-disjoint split, which takes it from 0.0590 to -0.0240. That step also changes which rows are scored, so its change is a difference of point estimates and carries no interval.

L0's temporal split puts 1,530 accounts on both sides of the split, which is the guard being removed, asserted in code rather than assumed; L3 reproduces the eight shipped point metrics to a largest absolute difference of 0; the pretrain-side guards stay on at every rung; only the evaluation protocol moves.

The result is not a discount on our number; it is that the permissive protocol and the shipped protocol are not measuring the same thing. Under L0 the next-category delta is 0.0590 on top-1 and 0.1064 on top-5, the largest figures anywhere in this ladder. When the split is switched to entity-disjoint, that lift does not settle toward the shipped value; it disappears: top-1 lands at -0.0240, with an interval of [-0.0373, -0.0112] that stays below zero, and top-5 lands on a null at [-0.0158, 0.0140]. Our reading of that reversal is that most of the L0 lift was account identity rather than transfer: with the same accounts on both sides, one pooled vector per account works as a key to that account's own category mix, and it stops working the moment the accounts are ones the model has not seen. The with-embeddings arm at L1 stops almost immediately under its own entity-disjoint validation fold, which is the model detecting the same thing. That step also changes which rows are scored, so we report no interval on the change itself and let the L1 interval carry the point.

The transfer signal appears once the embedding is made as-of. At L2 top-1 is 0.0159 and top-5 is 0.0547, and because L1 and L2 are scored on identical rows that guard can be priced directly: it moves top-1 by 0.0399, interval [0.0249, 0.0561], and top-5 by 0.0547, interval [0.0385, 0.0742], both clear of zero. The last guard is the one that reduces our lift, and the two metrics disagree about the size of the reduction. Equipping the baseline with per-entity aggregates moves top-1 by 0.0060 with an interval of [-0.0020, 0.0152] that spans zero, so on that metric L2 and L3 are not distinguishable from each other; on top-5 it takes lift away, -0.0329, interval [-0.0436, -0.0219], clear of zero. The rung this page ships is therefore the strictest one rather than the most favourable one, and it still reports a gain whose interval clears zero on both metrics.

On the fraud task the ladder cannot discriminate, and we say that rather than reading point estimates. The baseline already scores 0.985 PR-AUC at its weakest rung and 0.99389 at the shipped one, and all eight per-rung fraud intervals, four on AUC and four on PR-AUC, span zero. That task is saturated on this corpus before the protocol gets a chance to matter, which is the same reading the transfer table above already gives. Two further details the exhibit does not hide. The L0 rung scores a different set of rows, 1,943 test accounts against 391 at the guarded rungs and 809 fraud positives against 945, so nothing in the L0 column is paired with the rest. L3 also reproduces the transfer table above exactly, to a largest absolute difference of 0 across the eight point metrics, which is the check that keeps this ladder attached to the rest of the page.

We also re-ran this ladder from scratch rather than trusting the first run. On the machine that produced the file the comparison returns CHECK OK, exit 0 across 424 numeric leaves, at a largest difference of 1.110e-16, which is double-precision rounding. Run on a different CPU vendor the same comparison reproduces 398 of 424 leaves bit-identically and differs on the rest at a largest absolute difference of 1.349e-05, all of it inside L0's next-category arm, where a many-class model breaks argmax ties differently under a different floating-point reduction order. Rendered numerals on this page that move as a result: none, since every affected quantity is printed here at four decimals and both values round to the same figure. We report that because a reviewer who reruns this on their own hardware should know what to expect, and because a reproducibility claim is more useful when its tolerance is stated alongside it.

A real corpus, three heads, two seeds

The ladder above changes the protocol and holds the corpus fixed. Corpus is a variable in the same way, so the next question is what happens when the corpus changes and the protocol does not. We ran the shipped L3 protocol again on real retail transactions from dunnhumby's Complete Journey, with byte-identical training code, the same guards, the same entity-clustered paired bootstrap and the same decision rule: two pretraining seeds, three heads, every split entity-disjoint and pre-registered (household-disjoint for the transfer and offers evaluations, store-disjoint for the store head). Everything was pre-registered at 6c9c7a9 with a pre-run amendment at 51c69e5, both committed before the run, and the rule that decides what counts was fixed in the same place: "pre-registered in CJ-REPLICATION-PREREG.md: positive only if both seed intervals sit above zero; anything else ships as a null or mixed result; no third seed, no protocol changes after seeing numbers".

The same L3 protocol, a different corpus. Two backbone seeds against one pre-registered split, entity-clustered paired bootstrap, 95% CI, every result as obtained.
Head and outcomeBackbone seed 7 Backbone seed 8Call
Household transfer, next-commodity top-1, delta on the L3 baseline -0.0160 [-0.0207, -0.0113] -0.0147 [-0.0197, -0.0101] Negative, both seeds
Store head, post-cut sales growth, macro-F1 delta on a counts-only control +0.1912 [0.0064, 0.3865] +0.2242 [0.0449, 0.4343] Positive, both seeds, intervals wide
Household offers head, post-cut coupon redemption, AUC delta on counts plus demographics +0.0141 [-0.0094, 0.0382], null +0.0269 [0.0023, 0.0551], positive Mixed, and quoted as mixed

Under the L3 protocol on the real Complete Journey corpus the backbone's transfer gain on next-commodity top-1 is negative in both pretraining seeds (-0.0160 and -0.0147), with both 95% household-clustered intervals below zero. Per the pre-registration this ships as obtained. The with-embedding arm lands at 0.5041 and 0.5053 on top-1, both under the 0.5101 majority-class floor, so here the embeddings do not only fail to add accuracy; they reduce it below the level the model would have reached by predicting the most common class. Our reading is the ladder's second axis behaving as designed. The ladder showed the protocol is a variable; this shows the corpus is one too, and on this corpus the answer moved.

On this corpus the backbone's store embeddings add measurable signal over the counts-only control for post-cut sales growth (macro-F1 delta +0.1912, 95% interval [0.0064, 0.3865] above zero). The second pretraining seed repeats it: On this corpus the backbone's store embeddings add measurable signal over the counts-only control for post-cut sales growth (macro-F1 delta +0.2242, 95% interval [0.0449, 0.4343] above zero). The control is counts only, three pre-cut features and nothing more, and the outcome is growth rather than volume by design: raw post-cut volume is mostly pre-cut size restated; the counts control would win by construction without answering the embedding question. The universe is 231 eligible stores split into a 46-store test set, which is why the intervals are wide and why they are printed wide. Accuracy was pre-registered alongside macro-F1 as a reported metric, with the decision rule keyed to macro-F1, and it is the weaker of the two: the seed 7 accuracy delta is +0.1957 with an interval of [-0.0217, 0.4130] that spans zero, so on that endpoint that seed is a null, while seed 8 clears it at +0.2174 [0.0217, 0.4136]. This is predictive and not causal: nothing here estimates what a backbone embedding causes. The offers head came back mixed and is reported as mixed, null in the seed 7 run and positive in the seed 8 run, and on causality the file is explicit: campaigns in this corpus were targeted by the retailer, not randomized; nothing here estimates what an offer causes. The causal offers exhibit in this project remains the Criteo randomized-uplift result.

The scale caveat travels with every number in this band. The corpus is 2,595,732 basket lines from about 2,500 households at a grocery retailer, against TabFormer's 24 million rows. These are basket lines rather than transactions: each row is one product on one receipt, so the row count is larger than the number of shopping trips behind it by roughly the basket size, and the event unit itself differs from TabFormer's card transactions. That is a second variable sitting alongside protocol and corpus, so the negative household transfer here should not be read as a pure corpus effect. It is a protocol replication on a small real corpus, not a scale replication, and it is not card data and not card-network data. It settles that the method and the protocol travel to a real corpus, and that the result they produce on that corpus differs from the result they produce here. It leaves open real card data at Amex scale, which is what the Phase 1 gate in Section 8 exists for.

Surprise, rarity and the scam gap

American Express has reported the lowest US credit card fraud rate among major networks for 19 straight years58, on a stack that runs 8 billion-plus automated risk decisions on over $1 trillion of volume59. Nothing in this subsection improves on that, and nothing in it is a criticism of it. The gap we are pointing at is one Amex's own fraud leadership has described in public: once the transaction controls became good enough that bad actors struggled to make money against them, the attacks came back around to scams and social engineering47. Singapore quantifies that shift. The police report that 81.8% of reported scam cases in 2025 involved self-effected transfers, where the account holder moves the money themselves after being deceived4. A model trained to answer "was this really the cardholder" cannot see those cases, because the answer is yes.

Singapore's Shared Responsibility Framework has been in force since 16 December 202446, and its scope should be stated precisely rather than invoked loosely. The framework covers seemingly authorised transactions, meaning a scammer obtained the account credentials and transacted, and it places a duty on financial institutions to run real-time fraud surveillance when an account is drained of a material sum quickly. Self-effected transfers, the larger class, sit outside that defined scope. As a result, the largest category of scam loss in this market is the one that neither unauthorized-fraud scoring nor the payout framework is built around, and the practical question is what signal a firm could run against it without waiting for a labeled scam dataset that does not exist.

One such signal can be read directly from the backbone we already have. The model was pretrained with a masked-field objective at a mask probability of 0.15, so it can be asked, with no retraining and no fine-tuning, how surprised it is by a transaction. Mask one field of the scored transaction at a time, take the negative log-likelihood of the value that was actually there, and sum across fields. That total is the transaction's behavioural surprise given the account's own history. Pseudo-log-likelihood used this way is Salazar and colleagues' recipe for scoring masked language models, ACL 2020. Masking every field at once would be out of distribution for a model pretrained at 15 percent, so we do not do it, and the results file records that decision rather than leaving it implied.

No fraud label enters that score, and the guarantee is structural rather than stated in prose. The scoring code is handed the token array and nothing else, the label file is opened only after every score exists and the hash of the score matrix has been recorded, and that hash is checked again when the numbers are written. That is the purpose of the exercise: the score is a detector that could be deployed on a population with no labels at all, which is the situation authorized-scam detection is currently in.

Label-free score · frozen checkpoint 300,000 scored transactions · 945 fraud positives · 391 held-out accounts · an uninformed ranking scores 0.00315 PR-AUC.

  • ✓ no fraud label reaches the score
  • ✓ labels opened only after every score exists
  • ✓ label column excluded from the pretraining vocabulary
  • ✓ account and card identifiers never in the vocabulary
  • ✓ pretraining corpus truncated before the test window
  • ✓ every scored row is post-cut, so none was pretrained on
How well each label-free ranking separates real fraud · no fraud label enters any of these scores, and the labels are opened only afterwards to measure the ranking. Intervals are entity-clustered bootstrap over held-out accounts. The last two columns are the share of the fraud positives that land in the top slice of the ranking, which is what an operations queue receives.
ScoreFieldsPR-AUC (95% CI)ROC-AUC (95% CI)Caught in the top tenth of a percent (95% CI)Caught in the top one percent (95% CI)
Contextual surprise, from the frozen backboneall ten fields0.1338 [0.0983, 0.1751]0.9022 [0.8742, 0.9265]0.116 [0.091, 0.143]0.379 [0.335, 0.423]
Global rarity, counts only, no modelall ten fields0.1282 [0.0950, 0.1717]0.9367 [0.9008, 0.9653]0.105 [0.072, 0.131]0.323 [0.285, 0.369]
Both, unweightedall ten fields0.3296 [0.2705, 0.3881]0.9497 [0.9170, 0.9742]0.193 [0.164, 0.228]0.599 [0.548, 0.645]
Contextual surprise, from the frozen backboneerrors field dropped0.1323 [0.0977, 0.1729]0.9026 [0.8747, 0.9268]0.114 [0.089, 0.141]0.379 [0.334, 0.421]
Global rarity, counts only, no modelerrors field dropped0.1676 [0.1263, 0.2233]0.9382 [0.9013, 0.9674]0.134 [0.110, 0.171]0.330 [0.289, 0.407]
Both, unweightederrors field dropped0.3496 [0.2862, 0.4078]0.9510 [0.9180, 0.9757]0.200 [0.169, 0.233]0.628 [0.576, 0.670]
Contextual surprise, from the frozen backbonebehavioural fields only0.1881 [0.1410, 0.2352]0.9381 [0.9252, 0.9488]0.138 [0.110, 0.167]0.421 [0.378, 0.463]
Global rarity, counts only, no modelbehavioural fields only0.1706 [0.1300, 0.2268]0.9721 [0.9626, 0.9806]0.134 [0.107, 0.169]0.340 [0.299, 0.388]
Both, unweightedbehavioural fields only0.4133 [0.3532, 0.4691]0.9842 [0.9796, 0.9882]0.218 [0.186, 0.249]0.678 [0.637, 0.718]
The control, run against the model · each row is the first score minus the second, on identical rows inside every bootstrap replicate, so the differences are paired. Each metric carries its own verdict, read off its own interval, because the two metrics do not agree here and reading only one of them would hide the rows where the counts-only control wins. The result recorded for each row is the one that is deployed.
Comparison Δ PR-AUC (95% CI) Verdict on PR-AUC Δ ROC-AUC (95% CI) Verdict on ROC-AUC
contextual surprise minus global rarity0.0056 [-0.0391, 0.0506]not separated-0.0345 [-0.0523, -0.0159]second wins
the combined rule minus global rarity0.2014 [0.1480, 0.2450]first wins0.0129 [0.0041, 0.0226]first wins
the combined rule minus contextual surprise0.1958 [0.1654, 0.2227]first wins0.0475 [0.0366, 0.0580]first wins
contextual surprise minus global rarity, errors field dropped from both-0.0353 [-0.0898, 0.0181]not separated-0.0356 [-0.0537, -0.0163]second wins
contextual surprise minus global rarity, behavioural fields only0.0175 [-0.0381, 0.0737]not separated-0.0340 [-0.0490, -0.0206]second wins
the combined rule minus global rarity, behavioural fields only0.2427 [0.1817, 0.2955]first wins0.0121 [0.0045, 0.0201]first wins

Reading the comparison the way it landed: on PR-AUC the paired difference is 0.0056 with an interval of [-0.0391, 0.0506], spanning zero, so the two are not separated; on ROC-AUC the paired difference is -0.0345 with an interval of [-0.0523, -0.0159], below zero, so the control beats the model. The reading used for deployment is taken from the interval rather than from the point estimate.

the error flag on its own ranks fraud at 0.5085 ROC-AUC, which is why every score above is also reported with that field dropped; the share of scored rows with no prior transaction in the window is 0.00013; the share carrying at least one value the pre-cut vocabulary never saw is 0.8871, and both the model score and the counts control see those the same way; the unseen-calendar-year flag on its own ranks fraud at 0.5432 ROC-AUC, which is what the behavioural-only rows are there to price.

Full-mask variant: NOT SHIPPED. Masking every field of a row at once is out of distribution for a 15% mask-probability pretrain. Any such variant would have to be labelled separately.

seed 7 · data labels: synthetic · generated by scripts/fm/protection_pll.py --check-able · lightgbm 4.7.0 · numpy 2.5.2 · polars 1.43.2 · pyarrow 25.0.1 · python 3.12.3 · sklearn 1.9.0 · torch 2.13.0+cu130

This section addresses the question that determines whether any of this is worth deploying. A surprise score can be nothing more than "this value is globally rare", so we built exactly that with no model in it: marginal token frequencies fitted on pre-cut rows of accounts outside the test set, scored the same way and measured the same way. On the same 300,000 scored transactions carrying 945 real fraud positives across 391 held-out accounts, the model's surprise score reaches 0.1338 PR-AUC and 0.9022 ROC-AUC against an uninformed baseline of 0.00315. The counts-only control reaches 0.1282 and 0.9367. Paired on identical rows, contextual surprise minus global rarity is 0.0056 on PR-AUC with an interval of [-0.0391, 0.0506] that spans zero, and -0.0345 on ROC-AUC with an interval of [-0.0523, -0.0159] that sits entirely below it. The model therefore does not beat counting on this task, and on the metric where the intervals do separate it is behind. We report this finding directly rather than leading with a number that the control beats.

The combination of the two scores does outperform both, and it also requires no labels: standardize the two scores over the scored rows and add them, unweighted, with nothing fitted. That combination reaches 0.3296 PR-AUC and 0.9497 ROC-AUC, which is 0.2014 PR-AUC above global rarity alone, interval [0.1480, 0.2450], and 0.1958 above contextual surprise alone, interval [0.1654, 0.2227], both clear of zero on both metrics. In queue terms, the top one percent of that combined ranking holds 0.5989 of the fraud positives against 0.3228 for rarity alone. The deployable rule here is therefore rarity plus novelty rather than either one on its own, and the model contributes as the second term rather than the first.

What it buys. The result is a rule that costs almost nothing to run and requires no labels, reported alongside the control that beats it. For a model risk committee, that pair is the reviewable object: a candidate signal, the inexpensive baseline it has to be better than, and an interval indicating whether it is. In our experience, work that reports only its favourable results takes longer to approve, not less.

Two controls are reported in the same table so that these results can be verified rather than assumed. TabFormer carries an Errors? column that could plausibly co-occur with fraud, so every score is reported a second time with that field dropped. On its own the error flag ranks fraud at 0.5085 ROC-AUC, which is close to chance level, and dropping it moves the model score from 0.1338 to 0.1323 PR-AUC. The second control answers a question our own diagnostics raised rather than one a reviewer had to ask. The vocabulary was fitted on pre-cut rows, so a post-cut calendar year is an unknown token on 0.8631 of the scored rows, and it is fair to ask how much of the surprise is that rather than behaviour. The behavioural-only rows drop the calendar fields and the error flag and keep the six that describe what the transaction did. The scores on those rows are higher rather than lower: the combined score there reaches 0.4133 PR-AUC with the top one percent holding 0.6783 of the positives, and the unseen-year flag on its own ranks fraud at only 0.5432 ROC-AUC. The verdict does not move with the field set either: on all three, the control beats the model on ROC-AUC with an interval clear of zero, the two are not separated on PR-AUC, and the unweighted combination beats the control on both.

What prototype scale settles

It demonstrates that the recipe runs end to end and that the transfer question can be answered rigorously at prototype scale. It does not demonstrate production performance at Amex: the data is synthetic, and we state that limitation explicitly on this page.

Each row is one full pretrain plus the same leakage-hardened transfer evaluation, on the axis shown, with the places the runs are not identical named in the card above. Points: a 3M row earliest-window subset (seed 7); a 10M row earliest-window subset (seed 7); the full 24.4M row corpus (seed 7, the main run above); a seed 1337 repeat of the full 24.4M row corpus as a seed stability check; the same events regrouped into merchant sequences (merchant axis, seed 7). Deltas are reported as obtained.

Pair retention, the null and its limits

The scaling card says settling the two-sided question needs a task that is not degenerate on the merchant axis. We built one and ran it, and the design was committed before the first run.

The task is pair retention. Take every account and merchant pair that transacted at least once before the corpus cut, and ask whether that same pair transacts again after it. That is the merchant-side question a merchant embedding could plausibly carry, and unlike next category it is not settled by the merchant alone: 838,863 pre-cut pairs with a return rate of 0.1691. The split is entity-disjoint by merchant, 167,245 training merchants against 71,370 test merchants sharing 0. The baseline is deliberately strong: it gets pair frequency and recency, the account's pre-cut activity, and the merchant's own pre-cut transaction count, distinct-account count and recency, so the embedding has to beat a model that already knows how busy and how recent the merchant is. Every feature reads pre-cut rows only and every label reads post-cut rows only.

The result is a null, and we report it as such. Adding the merchant-axis embedding moves ROC-AUC by 0.000589, interval [-0.000475, 0.001567], and PR-AUC by 0.002829, interval [-0.000422, 0.006311]. Both point estimates are positive and the PR-AUC interval comes close to clearing zero without doing it. We put the cardholder-axis merchant embedding through the identical pipeline as a control, and it reads the other way on both metrics, -0.000847 and -0.001342, with intervals that also span zero. On this task, at this scale, on this corpus, the merchant-axis backbone does not add measurable signal over a pair-history baseline.

Two limitations would need to be addressed before the two-sided question could be considered settled. The merchant vector is pooled over that merchant's whole pre-cut history and carries nothing specific to the account, so it can only help through merchant-level structure the merchant counts miss; the account-side as-of vector, which is the half most likely to carry a return signal, is not in this arm, and it is absent because this run was built to stay on a laptop rather than because we decided it did not belong. And the split is merchant-disjoint only, since cardholders cross merchant boundaries, which is the same caveat the merchant-axis pretraining run carries.

Pair retention on the merchant axis · one LightGBM setting, one test set, only the feature block moves. Deltas are the embedding arm minus the baseline arm, with merchant-clustered paired-bootstrap 95% CIs.
ArmROC-AUC PR-AUCΔ ROC-AUC (95% CI) Δ PR-AUC (95% CI) Features
Baseline only (pair frequency and recency)0.9403560.863579referencereference13
Baseline + merchant-axis embedding0.9409450.8664080.000589
[-0.000475, 0.001567]
0.002829
[-0.000422, 0.006311]
526
Baseline + cardholder-axis embedding0.9395090.862237-0.000847
[-0.003390, 0.001358]
-0.001342
[-0.006704, 0.003589]
526

838,863 pre-cut (account, merchant) pairs built from 24,386,900 transactions, of which 0.1691 transact again after the cut. Split entity-disjoint by merchant: 167,245 train merchants against 71,370 test merchants, sharing 0 merchants. Scored on 257,239 test pairs across 71,370 merchants, positive rate 0.1776. Bootstrap 1,000 resamples, clustered on the test merchant. Every feature reads pre-cut rows only and every label reads post-cut rows only; the embedding join left 0 test pairs without a merchant vector.

What it buys. Read commercially this is a sizing answer rather than a disappointment. It says roughly how much corpus a two-sided backbone needs before it pays for the compute, it keeps the two-sided claim out of the pitch until a merchant-native task supports it, and it leaves that task already built and pre-registered for the next round rather than promised.

Five pretrains, headroom and the cap

The same pretrain recipe and the same leakage-hardened transfer evaluation, run at three corpus sizes, on a second seed, and on a second grouping axis, with the places the runs are not identical named below. All five points are in the table, reported as obtained. The next-category gain is not significant at 3 million rows, where the interval spans zero. It clears zero at 10 million and stays there at the full corpus on both seeds, so the gain emerges with corpus size and then flattens rather than keeping on rising. The 3 million row point is also the only one where the evaluation itself is smaller, at 430,084 training rows and 109,918 test rows against the 800,000 and 300,000 cap that binds at every larger point, so its null is consistent with a corpus effect or with a smaller downstream sample, and this table does not separate the two. Corpus size is not varied on its own here either: the 3 million and 10 million row points are earliest-window subsets, and each one recomputes its own time cut and builds its own vocabulary from the rows it holds, so the smaller runs also see an earlier time window and a different vocabulary. A clean single-variable version holds the evaluation window fixed and shares one vocabulary across every point. Both controls are pre-registered in Section 8. The baseline moves with the data as well, from 0.22 top-1 accuracy at 3 million rows to 0.235 and 0.241 at the full corpus, so the embeddings are adding to a target that is itself getting better.

On the fraud task the embeddings add PR-AUC exactly where the baseline has headroom, and that holds at both of the corpus sizes where it has any. Those two points are not independent replications of each other: the 3 million row corpus is an earliest-window subset of the 10 million row one, so the smaller run's rows and its users are contained in the larger run's. Two separate pretrains scored on disjoint test rows carries some evidential weight; it does not constitute two independent samples, and we state this explicitly rather than leave the term unqualified. The baseline here never sees the backbone, so its own movement across these rows is a property of the downstream evaluation rather than of the pretraining corpus. The fraud test set carries 42 positives at 3 million rows, 322 at 10 million and 945 at the full corpus, so headroom in this table means headroom on that evaluation. At 3 million rows PR-AUC moves from 0.485 to 0.551, and at 10 million from 0.638 to 0.904, delta CI [0.2001, 0.3436] entirely above zero, on 322 positive test cases in 300,000 scored transactions. Both intervals sit above zero and both are in the table. AUC on the same task does not confirm it and does not contradict it either: the fraud AUC delta spans zero at all five points in the table, including that 10 million row point, where the point estimate is positive but the interval runs [-0.0034, 0.0115]. We lead with PR-AUC on fraud because positives are that rare, so AUC is carried by the negatives. On the record of which decision came first: both metrics have been reported side by side since the first run and neither was named the lead before results existed, so the reason for leading with PR-AUC is the one just given rather than a pre-registration. One caveat on the level: our evaluation keeps every fraud positive and thins the negatives to reach the row cap, so wherever that cap binds, which is every point from 10 million rows up, the positive rate in the scored set sits above the rate in the window it is drawn from and the PR-AUC levels here sit above what a live queue would score. The 3 million row point is the one place the cap does not bind, so its rows are unthinned. The comparison between the two arms is unaffected either way, since both arms are scored on the same rows. The full-corpus null in the table above is saturation rather than a failure of the embeddings: that baseline already scores 0.99389 PR-AUC, which leaves close to nothing to add. Before reading the next-category curve above 10 million rows as flat, we note that the 10 million and full-corpus deltas are not distinguishable from each other on their own intervals, so what follows reads point estimates rather than a measured difference. Our reading of that curve starts inside our own pipeline rather than with the corpus: the evaluation cap named above binds at every point from 10 million rows up, so the downstream training and test samples stop growing while the pretraining corpus keeps growing. A ceiling in the synthetic corpus is the other candidate, and this table does not separate the two. The production-scale evidence in Section 3 (Visa's TREASURE27, Nubank's nuFormer28) is the scale argument this table cannot supply, and the Section 8 pilot re-measures it on real closed-loop data.

The merchant-axis row is the two-sided leg, and it is an honest negative. We pretrained on the same events regrouped into merchant sequences, and it produced merchant-view embeddings, so the dual-axis mechanism is implementable. On that row the split is entity-disjoint in merchants only, since cardholders cross merchant boundaries, so it is a partial two-sided demonstration. Both of that row's deltas are nulls, and leakage of that kind would flatter the embeddings rather than hide a gain, so the negative below is the conservative reading. What it does not show is a two-sided gain, because both evaluation tasks are degenerate on that axis. A merchant's category is close to constant across its own transaction sequence, so next category is trivial there by construction: the baseline alone scores 0.985 top-1, leaving no room for a transfer signal to show up. Fraud leaves little room on that axis as well: the baseline there already scores 0.972 PR-AUC, and at that level this run cannot separate a ceiling from an absence of merchant-side signal for a label defined on the cardholder. So the run turns two-sided beats single-sided into a measured open question instead of an assertion, and settling it needs merchant-native tasks. We built one and put the question to it, below.

Scaling entries · prepared dataset rows and parameter budget vs the next-MCC transfer delta and the fraud PR-AUC transfer delta (entity-level paired-bootstrap 95% CIs). The row count is the whole prepared dataset for that run, not the pretraining corpus: each pretrain reads only the rows before that run's corpus cut, and the rest is what the evaluation is drawn from. Reported as obtained: on the cardholder axis the fraud delta clears zero at the two corpus sizes where the baseline still has headroom and the full-corpus runs are nulls on a saturated baseline, while the merchant-axis row is a null the card reads separately; the next-MCC delta clears zero from 10M rows upward on the cardholder axis.
Dataset rowsParams (M) AxisSeedNext-MCC Δ (95% CI) Fraud ΔPR-AUC (95% CI)
3,000,000 20.45 cardholder 7 [-0.0067, 0.0160] [0.0024, 0.1662]
10,000,000 20.51 cardholder 7 [0.0165, 0.0383] [0.2001, 0.3436]
24,386,900 20.54 cardholder 7 [0.0103, 0.0343] [-0.0041, 0.0086]
24,386,900 20.54 cardholder 1337 [0.0110, 0.0334] [-0.0084, 0.0085]
24,386,900 20.54 merchant 7 [-0.0078, 0.0217] [-0.0162, 0.0143]

The full account

Offers: everything that moved off the page

The measurement service and the judges' own field

The offers head does not ask partners to trust a targeting model. It ships an always-on measurement service: every campaign carries a randomized holdout by design, and each partner receives an auditable incremental-lift report for every campaign. The design is the ghost-ads pattern from the marketing-science literature, already run in production at platform scale38. Productizing it for partner banks is the new part.

Response models against uplift models

Response models find people who will buy; uplift models find people who will buy because of the intervention. Ascarza's randomized field experiments showed treatment-effect targeting recovering up to 6.8 percentage points more churn reduction than response-based targeting aimed at the wrong customers. Spending offer budget on people who would have converted anyway is the failure mode a measurement layer exists to catch.

The two endpoints and their two verdicts in full

Read in the unit a budget owner uses, the two verdicts are the argument for the whole measurement layer rather than a blemish on it. The comparison uses the same customers, the same randomized arms, and the same depth. Ranking by predicted incremental effect captures 59.3 percent of the incremental site visits and 59.2 percent of the incremental conversions. Ranking by predicted response captures 48.6 percent and 68.6 percent. So the winner flips with the endpoint: response ranking wins on the outcome we pre-registered as primary and loses on the other, and its share moves by 20.0 points between the two while the uplift ranking's moves by 0.1 points. Two endpoints is not evidence of a general stability property and we claim none. What we claim is narrower and harder to dispute: the choice of endpoint determines which model wins, the size of that effect here is larger than the gap between the two models, and only a holdout reveals to a team that this is happening. That is the case for the measurement layer stated in its own evidence rather than in ours.

We ran the comparison on the largest public randomized uplift dataset, Criteo v2.1: 13,979,592 rows with randomized treatment assignment. On a pre-registered split, an X-learner61 ranks customers against a response-propensity ranking. The corpus should be described accurately: Criteo v2.1 is an advertising experiment and not a payments one, where treatment is being targeted by a campaign and the outcomes are a site visit and a conversion, not a card offer and a spend at a merchant. What carries across to a partner campaign is the targeting question and the randomized design that answers it, not the domain.

The endpoint hierarchy was fixed before the full run and it does not flatter us: the producing script names conversion the primary outcome and visit the robustness outcome in its own docstring, first committed at 2fd8c23 before the 13,979,592-row run, and the results file still stores them that way. We are exact about that word because this page depends on it: unlike the two standalone pre-registrations elsewhere here, that commit is page plumbing that carried the script, and a one-million-row development pass of the same script and seed ran before it. The hierarchy was fixed before the full run, not before every result, and we would rather correct our own wording than have it read as more than it earned. Conversion is the endpoint the uplift ranking lost, at every targeting depth. We lead with visit because its base rate is the one a per-customer effect estimate can resolve, and we print the primary endpoint at the same size below.

Qini coefficient on conversion, X-learner: 0.000199. Qini coefficient, response ranking: 0.000345. Paired difference CI: [-0.000194, -0.000089]. Uplift at the top slices of the ranking is in the table below, with stratified bootstrap CIs.

The same metric on the endpoint we lead with does not separate the two rankings either. On visit the Qini coefficients are 0.002993 against 0.003015, with the paired difference at [-0.000229, 0.000159], spanning zero. So our win is a depth result and not a whole-curve one: at the top decile the paired difference clears zero, and across the whole ranking the two are indistinguishable. A budget is spent at a depth, which is why the decile is what a partner acts on, but the curve metric is a null result and we report it here rather than leave it to be discovered.

We committed before the full run to reporting this ranking whichever way it landed. As obtained, the response ranking wins on conversion Qini here, with the paired CI fully below zero. We do not read that as an argument against the product. It is the intended function of the product: only a randomized measurement layer can tell a partner which targeting rule buys incremental spend on their book, and here it did so, on the largest public test available. At Amex the same layer answers that question per partner, per campaign, on closed-loop data, and the answer is worth paying for either way.

The loss has an identifiable mechanism. Our reading is that conversion on this dataset is rare (holdout rate 0.00304 treated, 0.001947 control), so per-customer treatment-effect estimates are noise-dominated at exactly the resolution the X-learner needs, and the treatment effect here correlates with baseline propensity, so ranking by response propensity proxies the uplift ranking.

Results by targeting depth for both outcomes, with intervals

The same holdout carries a second outcome, visit, at a much higher base rate (0.048391 in the treated arm against 0.003040 for conversion), and there the uplift ranking wins. The measurement is the one in the table below: 4,193,878 randomized holdout rows, 0.85 of them treated, and for each targeting depth the treated-minus-control outcome rate among the rows that ranking would have targeted.

Targeting the top ten percent by predicted CATE returns an incremental visit rate of 0.061241 (CI [0.057974, 0.064709]) against an average treatment effect of 0.010326 (CI [0.009832, 0.010820]) across the whole holdout. On point estimates that is 5.9 times the average effect, concentrated into a tenth of the population. Against response ranking at the same depth the paired difference is 0.011049, CI [0.008545, 0.014453], entirely above zero. The advantage is a top-of-ranking effect and it fades as the budget widens: at twenty percent the difference is 0.001199, CI [-0.000175, 0.002534], and at thirty percent 0.000454, CI [-0.000474, 0.001231]. Both of those intervals span zero, so past the top decile the two rankings are not separated here.

On the rare conversion outcome the same table reads the other way at every depth, and we print it at the same size. The paired difference is -0.001022 at ten percent, CI [-0.001670, -0.000277]; -0.000522 at twenty percent, CI [-0.000842, -0.000184]; and -0.000479 at thirty percent, CI [-0.000675, -0.000249]. All three sit entirely below zero. Uplift targeting concentrates incremental visits on the denser outcome and loses to plain response ranking on the rare one, on one dataset, one holdout and one pair of models.

Which regime a campaign is in

Both facts are true at the same time, and a partner running offers has to know which regime their campaign is in before they pick a targeting rule. On a dense outcome with real heterogeneity, uplift ranking puts the budget where the treatment actually moves someone, and here it bought 5.9 times the average effect at a tenth of the reach, that figure again being a ratio of point estimates with no interval of its own. On a rare outcome the per-customer effect estimate is noisier than the ranking needs, and the simpler response model wins. Nothing in the model indicates which of those regimes a campaign is in. The randomized holdout does, and that is what we are proposing to ship. Rather than a targeting model that a partner is asked to trust, it is a standing measurement layer that reports, per campaign and per outcome, which rule bought the increment and by how much, with the interval attached. The failure mode it exists to catch is paying for people who were going to convert anyway, and the only way to catch it is to measure it.

Hillstrom: wasted budget, fairness and the tables

The Hillstrom email experiment, a randomized three-arm public dataset with human-readable features, supplies the segment table: side by side, the customers a response model prioritizes and the customers an uplift model prioritizes, with the disagreements called out in plain terms. This is the table a marketing lead reads before a campaign, and the one a partner sees in their lift report.

It also allows us to test the premise directly rather than cite it. Wasted offer budget is the reason this product exists, and elsewhere we argue it from other people's papers, whereas this experiment can answer it directly. Split the 28 segments into thirds. Of the 5,096 treated customers a response model puts in its top third, 1,448 sit in segments the randomized arms measure in the bottom third of actual uplift, at 2.31 and 1.26 points against the best segment the experiment measures, at 11.68. That is 28.4 percent of the response model's top-third budget spent where measurement finds the least. Where to cut is a choice and it moves the answer, so the file carries halves, thirds and quarters together and the third is what prints here. This is not the largest available figure: with a split into halves it is 63.8 percent, more than twice as large, and that would have been the more favourable figure to quote. The domain caveat stands with it, because Hillstrom is retail e-commerce and the outcome is a site visit rather than billed business on a card, so this measures the mechanism and not the size of the opportunity at Amex.

The same arms answer a question the obligation map in Section 4 owes and the rest of this document does not otherwise pay: MAS FEAT opens on fairness, and there is no fairness number anywhere else here. There is a reason for that gap, although we do not consider it sufficient on its own. No corpus in this project carries a protected attribute that can be read: TabFormer is synthetic, Criteo ships anonymized features, and Complete Journey's demographic file was anonymized into opaque classifications in its 2023 re-release. So we measured the one group variable in the project that sits behind a randomized assignment, Hillstrom's three-level zip class, and we note that geography is a proxy and not a protected attribute.

We report two quantities per group, because they answer different questions. The effect itself, as a direct arm contrast with no model in it, is 4.89 points suburban, 4.69 urban and 2.91 rural, all three intervals above zero. A campaign that works less well in one group is a property of the campaign rather than evidence of unfairness. The fairness question concerns which customers the targeting rule then selects. Ranking by predicted uplift selects the three groups into the top decile in the same order as their measured effect, at a parity ratio of 0.873 between the least and most selected. Ranking by predicted response does not: it puts 0.1560 of the rural group into the top decile, the highest rate of the three, while the randomized arms measure that group's effect as the smallest of the three, and its parity ratio is 0.517.

Criteo Uplift v2.1 · randomized 13,979,592 rows · split: pre-registered: stratified-by-treatment 70/30 holdout, seed 42, committed in scripts/uplift_exhibit.py before the full-data run · CI: stratified bootstrap

Qini coefficient (conversion, holdout): X-learner 0.000199 · response 0.000345 · paired Δ(X−response) 95% CI [-0.000194, -0.000089]

Qini curves on Criteo Uplift v2.1: X-learner uplift targeting versus response-propensity targeting, with a random-targeting reference diagonal. 0 0.00025 0.0005 0.00075 0 0.25 0.5 0.75 1 Fraction of population targeted Qini (incremental conversions) X-learner (uplift) · x=0 · y=0 X-learner (uplift) · x=0.0905 · y=0.0005 X-learner (uplift) · x=0.1809 · y=0.0007 X-learner (uplift) · x=0.2714 · y=0.0007 X-learner (uplift) · x=0.3618 · y=0.0007 X-learner (uplift) · x=0.4523 · y=0.0008 X-learner (uplift) · x=0.5477 · y=0.0007 X-learner (uplift) · x=0.6382 · y=0.0008 X-learner (uplift) · x=0.7286 · y=0.0006 X-learner (uplift) · x=0.8191 · y=0.0004 X-learner (uplift) · x=0.9095 · y=0.0008 X-learner (uplift) · x=1 · y=0.0009 Response targeting · x=0 · y=0 Response targeting · x=0.0905 · y=0.0006 Response targeting · x=0.1809 · y=0.0008 Response targeting · x=0.2714 · y=0.0008 Response targeting · x=0.3618 · y=0.0008 Response targeting · x=0.4523 · y=0.0009 Response targeting · x=0.5477 · y=0.0009 Response targeting · x=0.6382 · y=0.0009 Response targeting · x=0.7286 · y=0.0009 Response targeting · x=0.8191 · y=0.0009 Response targeting · x=0.9095 · y=0.0009 Response targeting · x=1 · y=0.0009

Qini curves. Hover for point values; the tables below are the no-JS carrier of every headline number.
Targeting concentration at k · both outcomes, both rankings, differences and the ATE reference
Visit (denser outcome) · incremental outcome rate inside the targeted top-k, CATE ranking vs response ranking, with the whole-holdout ATE as the reference line (stratified-bootstrap 95% CIs, paired differences).
Targeted fractionCATE rankingResponse rankingDifference (CATE − response)
whole holdout (ATE)0.010326 [0.009832, 0.010820]n/a
top 10%0.061241
[0.057974, 0.064709]
0.050193
[0.046261, 0.053774]
0.011049
[0.008545, 0.014453]
top 20%0.039406
[0.037444, 0.041248]
0.038207
[0.035907, 0.040377]
0.001199
[-0.000175, 0.002534]
top 30%0.029416
[0.028094, 0.030610]
0.028962
[0.027386, 0.030600]
0.000454
[-0.000474, 0.001231]
Conversion (rare outcome) · incremental outcome rate inside the targeted top-k, CATE ranking vs response ranking, with the whole-holdout ATE as the reference line (stratified-bootstrap 95% CIs, paired differences).
Targeted fractionCATE rankingResponse rankingDifference (CATE − response)
whole holdout (ATE)0.001093 [0.000976, 0.001213]n/a
top 10%0.006471
[0.005539, 0.007542]
0.007493
[0.006300, 0.008572]
-0.001022
[-0.001670, -0.000277]
top 20%0.003950
[0.003435, 0.004438]
0.004471
[0.003897, 0.005067]
-0.000522
[-0.000842, -0.000184]
top 30%0.002763
[0.002402, 0.003103]
0.003242
[0.002831, 0.003625]
-0.000479
[-0.000675, -0.000249]

What the measurement layer is for, in two rows

A response model ranks this Hillstrom segment 5 of 28, so it gets the mail. Measured on the randomized arms, its visit uplift is 2.31 points with a standard error of 1.82, and its measured spend uplift is US$-0.69 per customer across 937 treated. The segment ranked 26 of 28, which never gets mailed, measures 5.82 points plus or minus 0.95 on 1,891 treated, the tightest interval in the table below. The response score has no access to either fact; only the randomized holdout reveals them. This is the function the product provides, shown here on a public experiment before it is applied to any partner data.

Hillstrom segment table · segments on which a response model can misallocate budget
Hillstrom MineThatData e-mail experiment (randomized, three arms) · segment-level response rank vs uplift rank vs the uplift measured directly from the randomized arms.
SegmentTreatedControlResponse rankUplift rankMeasured uplift rankMeasured visit uplift (pp)SE (pp)Verdict
7) $1 000 + x recency 1-3m33432615105.5913.067aligned
6) $750 - $1 000 x recency 1-3m3813802957.0432.894aligned
4) $350 - $500 x recency 10-12m33131831111.6772.849aligned
7) $1 000 + x recency 4-6m58594427.167.204aligned
4) $350 - $500 x recency 1-3m937877522222.3131.817wasted-budget
5) $500 - $750 x recency 1-3m8638116376.3971.794aligned
3) $200 - $350 x recency 1-3m1,5631,499714193.851.363aligned
6) $750 - $1 000 x recency 4-6m1181448647.0744.141aligned
4) $350 - $500 x recency 4-6m511526923241.2632.332wasted-budget
7) $1 000 + x recency 7-9m26181024250.854712.818aligned
4) $350 - $500 x recency 7-9m4094031113115.2012.482aligned
7) $1 000 + x recency 10-12m1013122727-5.38513.789aligned
1) $0 - $100 x recency 1-3m2,0062,0391319203.7451.102aligned
5) $500 - $750 x recency 4-6m37939014237.1212.379aligned
5) $500 - $750 x recency 10-12m15917515766.5233.541aligned
6) $750 - $1 000 x recency 7-9m67561626260.34656.676aligned
2) $100 - $200 x recency 1-3m1,4021,5061716154.6691.263aligned
6) $750 - $1 000 x recency 10-12m2742182828-11.3769.632aligned
5) $500 - $750 x recency 7-9m26127619885.8852.803aligned
3) $200 - $350 x recency 7-9m8678692020144.7531.59aligned
3) $200 - $350 x recency 4-6m9108882125231.6731.608aligned
2) $100 - $200 x recency 4-6m1,0711,0822217174.3951.393aligned
3) $200 - $350 x recency 10-12m8157882318183.9581.626aligned
1) $0 - $100 x recency 4-6m1,7061,6772411134.9541.074aligned
2) $100 - $200 x recency 7-9m1,1061,1412512125.1971.246aligned
1) $0 - $100 x recency 7-9m1,8911,936261095.8160.9457hidden-gem
2) $100 - $200 x recency 10-12m1,1481,1072721212.8521.225aligned
1) $0 - $100 x recency 10-12m2,0311,9602815164.4160.8679aligned

seed 42 · data labels: proxy randomized-experiment · generated by scripts/uplift_exhibit.py --check-able · lightgbm 4.7.0 · numpy 2.5.2 · polars 1.43.2 · pyarrow 25.0.1 · python 3.13.1 · scikit-learn 1.9.0

The gate, its hole and its disposition

The gate also has a gap that we found and report here rather than conceal. Both shipped layers pass every narrative here. A third layer, written after a reviewer asked whether a numeric gate can catch a reversed comparison, fails 5 of the 10 corridor narratives it can read: they say the model beat the seasonal naive when their own facts say the naive was ahead. Every numeral in them matches its source, which is why layer one passed them. The cause is ours and it is not a hallucination: the fact bundle we handed the model stated that a MASE below one beats the seasonal-naive baseline, which is false for this comparison, and every reversed narrative has a model MASE below one. Layer two could not catch it either, because it is the same model reading the same bundle. The rule is corrected in the generator, the direction layer is committed at scripts/narratives/directioncheck.py, and the text is not regenerated because that needs a cluster GPU, so the narratives below are the ones written under the false rule and this paragraph is what stops them being read as clean. A gate that checks whether the numbers are right is not a gate that checks whether the sentence about them is right, and that is worth more than the pass rate it replaces. The layer reports the other 20 bundles as unchecked rather than passed, because a gate that silently passes what it cannot read is the same failure one layer down.

What happens when it blocks, because a gate without a disposition is not a control: it fails closed to human review, and that review is cheap by construction, since the narrative only explains numbers that were already certified, so a block degrades the report to its tables rather than stopping it. The reviewer is the model owner named in Section 8, and the clean-run rate below is the false-reject budget a pilot has to beat rather than one we put forward as acceptable. This disposition is a design and not a measured service.

Layer one, deterministic: every numeral in the generated text is matched against that example's own fact bundle, on CPU, with no model in the loop. Layer two, cross-examination: a model reads the same facts and the same text and reports claims the facts do not carry, and that auditor is the same served model as the writer, which is the shipped design and means layer two is not an independent check. A narrative passes strictly only when both layers come back clean.

These are the committed narratives. When the red team in Section 4 re-ran the same gate with nobody attacking, it passed 28 of 30, not because the defence changed but because greedy decoding is not bit identical across a different GPU and a larger prompt batch, so only 12 of them came back byte identical. That re-run is a different sample of the same model, and its lower rate is the more informative one.

The deterministic numeric match against the source JSON is the hard part, it reproduces on a laptop with scripts/narratives/check.py --check, and by rule any failing narrative ships on this page rather than being filtered out. The gate reads numbers, not grammar: a numeral can be carried by the facts and still be attached to the wrong noun, and the whitespace panel shows one, where the count of places inside the bucket is read out as a count of buckets. That is the next layer to build, and we report the gap rather than removing the example.

The match is permissive by design, and the console indicates where that permissiveness has a cost. It accepts a fact stated as a fraction or as a percentage, and it accepts the rounding a sentence applies, so a numeral can find a supporting fact by arithmetic coincidence rather than by meaning. Because every numeral is printed next to the fact it matched, that is visible rather than assumed.

The size of this set should be interpreted with care. This is a working check on a small set, not a benchmark, and no strict failure appears in it, so the failure the console reaches is the one a reader catches and the numeric layer does not. There is no semantic error rate here and we do not claim one: layer one matches numerals and does not read sentences. What we can price is the loose part of the match. Across the 30 committed narratives, 303 of 338 numerals are carried by a fact exactly as stored, 34 by that fact rounded, and 1 by reading a fact as a percentage. Those last two are the loose ones: rounding is a half-unit window and a percentage moves the decimal point, so either can carry a numeral by arithmetic rather than by meaning.

The full account

Signing and corridors: everything that moved off the page

The signing ranking, its forward check, its control

The signing head ranks Singapore merchants a signing team should call first. The universe is 364,635 real Singapore places from Foursquare OS Places63. Three signal families feed the ranking:

  1. Real observable signals: category, merchant density, tourist-zone proximity, chain versus independent, and category economics under card fees of 1.5% to 3%.
  2. Backbone embedding similarity, the wire from Section 5's pretrain, with its limit stated here rather than left for a reader to find. That column is a per-category constant, six values across the whole ranking, so the backbone contributes no per-bucket merchant-level signal to this head today: it is a diagnostic column and not a feature, and the ranking below is by real signals. The reason is the corpora, since TabFormer and Foursquare share no merchant key and category is the only granularity they have in common. For the whole system, this is the one head that reads the backbone today, at category level, while the offers and corridor heads are proven on their own public benchmarks without it. On the one real corpus with a shared merchant key, the Complete Journey replication in Section 5, a store-level head consuming backbone store embeddings beat the counts-only control on post-cut sales growth in both seeds (macro-F1 +0.1912 and +0.2242, both intervals above zero and both wide, across 231 eligible stores); Criteo and SingStat still share no key with any transaction corpus, so those two heads remain proven without the backbone.
  3. A demand-weighted acceptance-gap increment, which at Amex would come from closed-loop spend attempts. On this page it is simulated, labeled as simulated on the chart, and shown as a sensitivity analysis: how the ranking reorders as the signal is added at plausible strengths.

The ranking now carries one pre-registered forward-looking check, and it is unflattering. Cutoff, outcome window, arms and decision rule were committed at fef8af5 before any outcome was computed: rebuild the composite from venues that existed before 2024-08-10, then score it against where card-accepting-universe venues actually appeared over the following 24 months. Plain pre-cutoff venue density is the null control; the composite also carries a density channel of its own, a log-normalized nearby-venue count built differently from the raw venue count used here.

Three arms are scored against one outcome. Precision, top fifty, is the share of an arm's top fifty buckets that also sit in the outcome's own top fifty.
ArmSpearman Precision, top fifty
The shipped composite 0.5477 0.38
Pre-cutoff venue density, the null control 0.6869 0.66
Equal weights on the same four channels 0.4498 0.32

Over the 1510 buckets formed before 2024-08-10, plain pre-cutoff venue count predicts where card-accepting venues appeared in the following 24 months better than the shipped composite: Spearman 0.6869 against the composite's 0.5477, difference -0.1392 with a cell-clustered bootstrap 95 percent interval of [-0.1785, -0.1020], entirely below zero. The composite is not shown to add predictive value over density on this outcome, and that is the finding. The outcome is venue formation in Foursquare records, not merchant signing, and date_created is record creation, a proxy for opening, so this is a forward-looking check of the ranking against real subsequent data, not a validation against signings.

A second read of the same data, and it is not pre-registered

The loss above is the pre-registered result and it stands. This is a different question asked of the same numbers afterwards, so it is labelled post hoc everywhere it appears and it is counted in no tally on this page. The composite contains density at weight 0.25, so the comparison above asks whether the blend beats one of its own four channels. A signing team is never in that position: they work inside a band of comparable density and ask what orders formation there. Stratifying the same 1,510 buckets into 10 density bands and re-reading the same outcome, the composite's three non-density channels order venue formation at a size-weighted Spearman of 0.2383, interval [0.1807, 0.2916], while density's own ordering inside those same bands is 0.1086, a paired difference of 0.1297 with interval [0.0575, 0.2006]. Density's marginal win is concentrated in the densest band, where its within-band Spearman is 0.6790; across the seven least dense bands it sits at or below 0.0137. The decision rule that gates this claim was fixed in the producing script before the statistic was computed, and a result covering zero would have shipped saying so.

The scope of this result is limited: it does not reverse the pre-registered check, it is a secondary read of an outcome already observed, and it is excluded by construction from the pre-registered census on page one. What it does say is that the lane's ranking is not empty once density is held roughly fixed, which is the condition under which a signing list is used, and that the Phase 1 gate against real signings is still the test that settles it.

Merchants are pseudonymized into buckets, such as "F&B cluster, Maxwell Road". No fabricated score is attached to a real business. The top 100 ranked buckets render on the inline map.

pseudonymized buckets · no fabricated score is attached to a real business FSQ OS Places SG slice (gated HF) · 364,635 POIs

Top pseudonymized merchant buckets · ranked by real public signals; the embedding-similarity column is the category-level backbone wire (cosine to the best-accepting category groups). The simulated acceptance-gap increment appears only as a labeled sensitivity analysis. The stated reason is the first of each row's committed reasons, which the brief's explainability ask requires beside the score rather than in the file.
BucketAreaCategory Real signalsEmbedding similarity Stated reason
F&B cluster — Chinatown area ChinatownF&B 0.853-0.163 MDR-sensitive category mix (F&B, prior 0.85: 1.5-3% card MDR vs 0.3% QRIS anchor)
F&B cluster — Raffles Place area Raffles PlaceF&B 0.840-0.163 MDR-sensitive category mix (F&B, prior 0.85: 1.5-3% card MDR vs 0.3% QRIS anchor)
F&B cluster — Orchard Road area Orchard RoadF&B 0.838-0.163 MDR-sensitive category mix (F&B, prior 0.85: 1.5-3% card MDR vs 0.3% QRIS anchor)
F&B cluster — Little India area Little IndiaF&B 0.837-0.163 MDR-sensitive category mix (F&B, prior 0.85: 1.5-3% card MDR vs 0.3% QRIS anchor)
Personal Services cluster — Chinatown area ChinatownPersonal Services 0.8330.133 Premium-demand corridor (0.3 km to Chinatown; zone score 0.85)
F&B cluster — Bugis area BugisF&B 0.832-0.163 MDR-sensitive category mix (F&B, prior 0.85: 1.5-3% card MDR vs 0.3% QRIS anchor)
F&B cluster — Orchard Road area · sector 2 Orchard RoadF&B 0.831-0.163 MDR-sensitive category mix (F&B, prior 0.85: 1.5-3% card MDR vs 0.3% QRIS anchor)
Personal Services cluster — Orchard Road area Orchard RoadPersonal Services 0.8240.133 Premium-demand corridor (0.3 km to Orchard Road; zone score 0.81)
Personal Services cluster — Bugis area BugisPersonal Services 0.8220.133 Premium-demand corridor (0.3 km to Bugis; zone score 0.81)
F&B cluster — Clarke Quay area Clarke QuayF&B 0.820-0.163 MDR-sensitive category mix (F&B, prior 0.85: 1.5-3% card MDR vs 0.3% QRIS anchor)
F&B cluster — Kampong Glam area Kampong GlamF&B 0.814-0.163 MDR-sensitive category mix (F&B, prior 0.85: 1.5-3% card MDR vs 0.3% QRIS anchor)
Personal Services cluster — Raffles Place area Raffles PlacePersonal Services 0.8120.133 Premium-demand corridor (0.3 km to Raffles Place; zone score 0.85)
The backbone wire: which categories transact like the best accepting ones
Transaction-weighted MCC-group centroids from the backbone's synthetic-corpus merchant embeddings · cosine to an acceptance anchor built from the most card-accepting groups, weighted by one minus their MDR-sensitivity prior. Centered cosines correct for pooled-embedding anisotropy; anchor members score high partly by construction, so the readout is the ordering of the non-anchor groups. These six values are the whole embedding column, and the mapping is documented in the results file.
Category groupCosine to anchor (centered) Raw cosineRole Embedded merchants
Hotels & Lodging 0.8430.849 anchor member9,745
Health & Wellness 0.5440.722 anchor member8,062
Entertainment & Leisure 0.4430.662 anchor member6,931
Personal Services 0.1330.576 scored against anchor24,045
F&B -0.1630.535 scored against anchor42,193
Retail -0.2640.639 scored against anchor98,889

The control this ranking shipped without, registered before it ran

Four of the seven pre-registered arms, against the shipped composite. Churn counts rows entering plus leaving the top twenty, so it runs from zero to forty. Three sanity rungs pass beside them: the composite against itself agrees perfectly, negated it disagrees totally, and twenty seeds of uniform random score land at a mean Spearman of -0.0077. Agreement is not accuracy, there is no observed acceptance label in this corpus, and none was manufactured. All seven arms, the three rungs and the gate that checks this control against the shipped file are in results/whitespace_control.json.
Control ranking Spearman, all buckets Released rows kept, of 400 Top-20 churn
Raw venue count, the pre-registered density control0.573622024
The same four channels at equal weight0.96313686
Tourist-zone proximity alone0.652031128
The backbone embedding column alone-0.12769040
Map of Singapore whitespace buckets: the top ranked pseudonymized merchant buckets positioned by location over a simplified Singapore coastline outline, dot size tracking places per bucket. F&B cluster, Chinatown area · rank 1 · real-signals score 0.853F&B cluster, Raffles Place area · rank 2 · real-signals score 0.84F&B cluster, Orchard Road area · rank 3 · real-signals score 0.838F&B cluster, Little India area · rank 4 · real-signals score 0.837Personal Services cluster, Chinatown area · rank 5 · real-signals score 0.833F&B cluster, Bugis area · rank 6 · real-signals score 0.832F&B cluster, Orchard Road area · sector 2 · rank 7 · real-signals score 0.831Personal Services cluster, Orchard Road area · rank 8 · real-signals score 0.824Personal Services cluster, Bugis area · rank 9 · real-signals score 0.822F&B cluster, Clarke Quay area · rank 10 · real-signals score 0.82F&B cluster, Kampong Glam area · rank 11 · real-signals score 0.814Personal Services cluster, Raffles Place area · rank 12 · real-signals score 0.812Retail cluster, Chinatown area · rank 13 · real-signals score 0.809Personal Services cluster, Orchard Road area · sector 2 · rank 14 · real-signals score 0.806Personal Services cluster, Little India area · rank 15 · real-signals score 0.805Retail cluster, Orchard Road area · rank 16 · real-signals score 0.804F&B cluster, Marina Bay area · rank 17 · real-signals score 0.801F&B cluster, Katong area · rank 18 · real-signals score 0.801Retail cluster, Bugis area · rank 19 · real-signals score 0.796Retail cluster, Orchard Road area · sector 2 · rank 20 · real-signals score 0.795F&B cluster, Changi Airport area · rank 21 · real-signals score 0.794F&B cluster, Somerset area · rank 22 · real-signals score 0.792F&B cluster, Orchard Road area · sector 3 · rank 23 · real-signals score 0.791Personal Services cluster, Kampong Glam area · rank 24 · real-signals score 0.79Personal Services cluster, Marina Bay area · rank 25 · real-signals score 0.789Retail cluster, Little India area · rank 26 · real-signals score 0.788Retail cluster, Raffles Place area · rank 27 · real-signals score 0.787Personal Services cluster, Clarke Quay area · rank 28 · real-signals score 0.786F&B cluster, Tanjong Pagar area · rank 29 · real-signals score 0.784F&B cluster, City Hall area · rank 30 · real-signals score 0.783F&B cluster, Marine Parade area · rank 31 · real-signals score 0.778Retail cluster, Kampong Glam area · rank 32 · real-signals score 0.776Personal Services cluster, Somerset area · rank 33 · real-signals score 0.774Personal Services cluster, Katong area · rank 34 · real-signals score 0.773Personal Services cluster, Katong area · sector 2 · rank 35 · real-signals score 0.772F&B cluster, Katong area · sector 2 · rank 36 · real-signals score 0.771Personal Services cluster, Orchard Road area · sector 3 · rank 37 · real-signals score 0.771Personal Services cluster, Changi Airport area · rank 38 · real-signals score 0.77Personal Services cluster, City Hall area · rank 39 · real-signals score 0.77F&B cluster, Marine Parade area · sector 2 · rank 40 · real-signals score 0.766Personal Services cluster, Katong area · sector 3 · rank 41 · real-signals score 0.761Retail cluster, Clarke Quay area · rank 42 · real-signals score 0.759Health & Wellness cluster, Orchard Road area · rank 43 · real-signals score 0.755Entertainment & Leisure cluster, Chinatown area · rank 44 · real-signals score 0.751Health & Wellness cluster, Orchard Road area · sector 2 · rank 45 · real-signals score 0.749Retail cluster, Somerset area · rank 46 · real-signals score 0.749Entertainment & Leisure cluster, Raffles Place area · rank 47 · real-signals score 0.748F&B cluster, Tanjong Pagar area · sector 2 · rank 48 · real-signals score 0.747Retail cluster, Marina Bay area · rank 49 · real-signals score 0.747Personal Services cluster, Marine Parade area · rank 50 · real-signals score 0.746F&B cluster, Esplanade area · rank 51 · real-signals score 0.746Retail cluster, Katong area · rank 52 · real-signals score 0.746Personal Services cluster, Tanjong Pagar area · rank 53 · real-signals score 0.745F&B cluster, Lavender area · rank 54 · real-signals score 0.742Health & Wellness cluster, Chinatown area · rank 55 · real-signals score 0.741F&B cluster, Dhoby Ghaut area · rank 56 · real-signals score 0.741F&B cluster, Lavender area · sector 2 · rank 57 · real-signals score 0.739Retail cluster, Orchard Road area · sector 3 · rank 58 · real-signals score 0.738F&B cluster, Changi Airport area · sector 2 · rank 59 · real-signals score 0.735Retail cluster, City Hall area · rank 60 · real-signals score 0.735Entertainment & Leisure cluster, Orchard Road area · rank 61 · real-signals score 0.734Entertainment & Leisure cluster, Little India area · rank 62 · real-signals score 0.73Entertainment & Leisure cluster, Changi Airport area · rank 63 · real-signals score 0.728Retail cluster, Changi Airport area · rank 64 · real-signals score 0.726Entertainment & Leisure cluster, Clarke Quay area · rank 65 · real-signals score 0.726Health & Wellness cluster, Somerset area · rank 66 · real-signals score 0.725Health & Wellness cluster, Raffles Place area · rank 67 · real-signals score 0.725Retail cluster, Katong area · sector 2 · rank 68 · real-signals score 0.723Entertainment & Leisure cluster, Bugis area · rank 69 · real-signals score 0.722Personal Services cluster, Tanjong Pagar area · sector 2 · rank 70 · real-signals score 0.721Health & Wellness cluster, Bugis area · rank 71 · real-signals score 0.721Personal Services cluster, Changi Airport area · sector 2 · rank 72 · real-signals score 0.72Retail cluster, Katong area · sector 3 · rank 73 · real-signals score 0.72Entertainment & Leisure cluster, Orchard Road area · sector 2 · rank 74 · real-signals score 0.718Entertainment & Leisure cluster, Kampong Glam area · rank 75 · real-signals score 0.717Health & Wellness cluster, Clarke Quay area · rank 76 · real-signals score 0.716Personal Services cluster, Lavender area · rank 77 · real-signals score 0.715Health & Wellness cluster, Little India area · rank 78 · real-signals score 0.713Personal Services cluster, Dhoby Ghaut area · rank 79 · real-signals score 0.713Entertainment & Leisure cluster, City Hall area · rank 80 · real-signals score 0.713Personal Services cluster, Esplanade area · rank 81 · real-signals score 0.712Retail cluster, Tanjong Pagar area · rank 82 · real-signals score 0.711Health & Wellness cluster, City Hall area · rank 83 · real-signals score 0.71Retail cluster, Marine Parade area · rank 84 · real-signals score 0.709Entertainment & Leisure cluster, Katong area · rank 85 · real-signals score 0.708Entertainment & Leisure cluster, Marina Bay area · rank 86 · real-signals score 0.705F&B cluster, Sentosa area · rank 87 · real-signals score 0.7Entertainment & Leisure cluster, Marine Parade area · rank 88 · real-signals score 0.699F&B cluster, Joo Chiat area · rank 89 · real-signals score 0.699Personal Services cluster, Lavender area · sector 2 · rank 90 · real-signals score 0.699F&B cluster, River Valley area · rank 91 · real-signals score 0.698Health & Wellness cluster, Kampong Glam area · rank 92 · real-signals score 0.697F&B cluster, Tiong Bahru area · rank 93 · real-signals score 0.697Entertainment & Leisure cluster, Somerset area · rank 94 · real-signals score 0.696F&B cluster, Farrer Park area · rank 95 · real-signals score 0.695F&B cluster, Sentosa area · sector 2 · rank 96 · real-signals score 0.691Retail cluster, Esplanade area · rank 97 · real-signals score 0.69Personal Services cluster, Orchard Road area · sector 4 · rank 98 · real-signals score 0.69Retail cluster, Tanjong Pagar area · sector 2 · rank 99 · real-signals score 0.689Hotels & Lodging cluster, Chinatown area · rank 100 · real-signals score 0.688 ChinatownOrchard RoadKatongChangi AirportEsplanadeSentosa
Exactly the top 100 ranked buckets render here (sliced from results/whitespace_map_points.json); dot size tracks places per bucket, values on hover, the table carries the numbers. The map is inline SVG with no external tiles; the coastline is a simplified outline drawn for orientation only, not a data layer. real-signals base · pseudonymized
Sensitivity: how the ranking reorders as the simulated signal is added
Simulated demand-weighted acceptance-gap signal (labeled simulated) added at plausible strengths · Spearman rank correlation with the real-signals ranking. A lower correlation indicates a larger reordering.
Signal strengthSpearman, all buckets Spearman, map-ranked buckets
0.10.99840.9753
0.250.98950.8893
0.50.95180.7239

embedding-similarity column: category-level wire on synthetic-corpus embeddings; at Amex this runs at merchant level on real closed-loop data

The wire maps 55 merchant category codes into the exhibit's six category groups, covering 189,865 of 238,615 backbone-embedded merchants. Reported as obtained: the Spearman rank correlation between the real-signals ranking and the embedding-similarity ordering is 0.0451 across the 400 ranked buckets and -0.1276 across all 1,536 buckets.

The signing score is a hand-set weighted sum of four public channels: 0.30 on the category fee prior, 0.30 on tourist-zone proximity, 0.25 on local density and 0.15 on independent share. Those weights are a judgement call and not a fitted quantity: nothing here tunes or selects them, because there is no outcome to select them against. The question is therefore whether this is an expensive way to sort by density. The arms, the metrics and the decision rule with its numeric thresholds were written down before any of them were computed (WHITESPACE-CONTROL-PREREG.md, committed first and alone), then seven control rankings scored the same 1,536 buckets.

What the control run shows, including the two readings against us. The pre-registered call is material difference, so the null did not fire: the strongest density control keeps 220 of the released rows, against the 380 the null needed on that measure, so the four signals do something one column does not. Two results point the other way and are reported in the same block. Scoring the four channels at equal weight agrees at Spearman 0.9631 and keeps 368 of them, so the particular weights are close to not load bearing, which argues the weighting does little work rather than that it is right. On the released list the closest single signal is tourist-zone proximity at Spearman 0.8685, not density, so much of what a partner would receive is ordered by that one channel. The backbone column agrees least of the seven arms on both measures here, which is the bullet above measured a second way.

seed 42 · data labels: real-signals base simulated-increment sensitivity pseudonymized synthetic-corpus embeddings (category-level wire) · generated by scripts/whitespace_exhibit.py --check-able · numpy 2.5.2 · polars 1.43.2 · python 3.13.1 · scipy 1.18.1

The corridor forecaster, its baselines and drivers

What reconciliation buys is coherence, not an accuracy headline: grouped reconciliation65 makes corridor, country and region totals agree by construction, and pooling variance across the twelve series moves total-level backtest MASE from 0.904 to 0.441, against 0.348 for the same seasonal naive computed on the total, which stays ahead.

Against a per-corridor seasonal naive over 13 held-out months, the model's MASE of 0.623 sits against 0.53 and loses pooled, winning five of the twelve. The twelve were selected by 2025 volume; every metric is computed identically for model and baseline. Gravity covariates (migrant stocks66, GDP, distance, language) enter lagged, so nothing the model sees would be unknown at forecast time. Per-corridor drivers are labeled as attributions rather than causes.

The brief asks this head for "the key drivers of growth, such as seasonality, traveler mix, customer segments, and category shifts", so the following is what this model relies on, grouped into those families from the same holdout attributions and reported without adjustment. 82.1 percent of the attribution is recent level, meaning last month and the three months before it. Seasonality is 8.1 percent as the same month last year plus 3.1 percent as the calendar month. Corridor identity is 6.3 percent. The three traveler-mix covariates together, distance, migrant stock and shared language, are 0.2 percent or less each. Customer segments and merchant category shifts, which the brief also names, are absent from public arrivals data, so this model cannot address them. On this public proxy the forecaster is mostly persistence with a seasonal correction, and the drivers a partner would act on are the ones only closed-loop transaction data carries. For that reason the corridor head is the one whose Phase 1 value depends most on Amex's data and least on ours.

The deployment question, pre-registered at 418aef3 before any combination number existed: whether adding the model to the naive helps. Averaging the model with the seasonal naive at fixed equal weights beats the naive on the same 2025-01 to 2026-01 holdout: macro MASE 0.5061 against the naive's 0.5302, a margin of 0.0241 (4.5 percent lower error), with the model alone at 0.6230. The model does not replace the cheap baseline, but adding it to the baseline at equal weights improves the baseline, and the combination beats the naive in 6 of the 12 corridors. The weights were fixed at 0.5 and 0.5 before the run in CORRIDOR-COMBINATION-PREREG.md (commit 418aef3), no other weight was computed, and with 13 holdout months across 12 corridors this is a point comparison with no interval claimed.

That win is scale-free, and it reverses under one alternative view. Macro MASE weights each corridor equally; totalling the same twelve corridors' absolute errors weights each arrival equally, and on that view the naive is ahead, 113211.69 against 114180.52 arrivals of error, so the blend costs 968.83 more over the holdout. Those columns were in the file cited above before this paragraph was written, and quoting the average that favours us while leaving them unprinted is the practice this document states it avoids. The lane is therefore claimed on neither view: adding the model to the cheap baseline helps a corridor-weighted average, does nothing for total arrivals, and wins in 6 of 12 corridors.

SingStat visitor arrivals · monthly model MASE 0.623 · seasonal-naive MASE 0.530 · holdout 13 months

covariates: t-1+ · reconciliation: grouped MinT (coherent by construction); base-vs-reconciled accuracy compared

Per-corridor holdout MASE (✓ = model beats the seasonal naive on that corridor) · model attributions (mean |SHAP| on holdout rows, log1p-arrivals units): associations learned by the model, not causal drivers
OriginMASE model MASE seasonal naiveMASE reconciled Top model attributions
China 0.9050.662 0.880lag1 (2.099); lag3 (0.370); corridor_id (0.142)
Indonesia 0.5940.240 0.581lag1 (1.979); lag3 (0.329); roll12 (0.144)
Malaysia ✓ 0.4490.576 0.450lag1 (1.031); lag3 (0.254); corridor_id (0.194)
Australia ✓ 0.4960.558 0.516lag1 (1.067); lag3 (0.275); corridor_id (0.147)
India 0.4300.293 0.462lag1 (1.043); lag3 (0.259); corridor_id (0.161)
Philippines ✓ 0.6370.916 0.656lag1 (0.627); lag3 (0.208); lag2 (0.053)
USA 0.5170.376 0.508lag1 (0.579); lag3 (0.186); corridor_id (0.080)
Japan 0.6060.409 0.643lag1 (0.412); lag3 (0.154); lag12 (0.050)
United Kingdom ✓ 0.4290.553 0.460lag1 (0.368); lag3 (0.149); corridor_id (0.106)
South Korea 0.8290.436 0.733lag1 (0.300); lag3 (0.146); corridor_id (0.089)
Taiwan 1.1380.840 1.017lag1 (0.245); corridor_id (0.132); lag3 (0.129)
Thailand ✓ 0.4470.502 0.447lag1 (0.119); corridor_id (0.119); lag12 (0.088)
Forecast against actual over the holdout months 2025-01 to 2026-01. The two corridors were selected by a fixed rule rather than by inspection: the corridor with the widest MASE margin in the model's favour and the corridor with the widest margin against it. The model beats the seasonal naive on 5 of 12 corridors and loses on the pooled measure, as the table above shows; the two charts illustrate a winning corridor and a losing corridor as a partner would see them. Arrivals are a proxy for cross-border card spend and are labelled so.

Philippines, the model's best corridor by MASE margin: model MASE 0.637 against seasonal naive 0.916.

0 50,000 100,000 Jan Apr Jul Oct Jan visitor arrivals, monthly actual arrivals · 2025-01 · 49,088 actual arrivals · 2025-02 · 50,154 actual arrivals · 2025-03 · 51,623 actual arrivals · 2025-04 · 69,047 actual arrivals · 2025-05 · 74,560 actual arrivals · 2025-06 · 66,647 actual arrivals · 2025-07 · 62,381 actual arrivals · 2025-08 · 61,462 actual arrivals · 2025-09 · 51,355 actual arrivals · 2025-10 · 64,398 actual arrivals · 2025-11 · 59,859 actual arrivals · 2025-12 · 65,491 actual arrivals · 2026-01 · 42,526 model forecast · 2025-01 · 51,633 model forecast · 2025-02 · 52,476 model forecast · 2025-03 · 54,464 model forecast · 2025-04 · 53,716 model forecast · 2025-05 · 64,007 model forecast · 2025-06 · 75,252 model forecast · 2025-07 · 76,067 model forecast · 2025-08 · 67,707 model forecast · 2025-09 · 59,910 model forecast · 2025-10 · 65,473 model forecast · 2025-11 · 67,558 model forecast · 2025-12 · 72,948 model forecast · 2026-01 · 51,497 seasonal naive · 2025-01 · 46,287 seasonal naive · 2025-02 · 55,604 seasonal naive · 2025-03 · 100,271 seasonal naive · 2025-04 · 53,154 seasonal naive · 2025-05 · 58,311 seasonal naive · 2025-06 · 76,120 seasonal naive · 2025-07 · 72,589 seasonal naive · 2025-08 · 60,174 seasonal naive · 2025-09 · 55,976 seasonal naive · 2025-10 · 61,489 seasonal naive · 2025-11 · 66,942 seasonal naive · 2025-12 · 72,162 seasonal naive · 2026-01 · 49,088

South Korea, the model's worst corridor by MASE margin: model MASE 0.829 against seasonal naive 0.436.

0 25,000 50,000 75,000 Jan Apr Jul Oct Jan visitor arrivals, monthly actual arrivals · 2025-01 · 79,301 actual arrivals · 2025-02 · 72,230 actual arrivals · 2025-03 · 38,446 actual arrivals · 2025-04 · 32,026 actual arrivals · 2025-05 · 39,936 actual arrivals · 2025-06 · 36,356 actual arrivals · 2025-07 · 59,416 actual arrivals · 2025-08 · 64,983 actual arrivals · 2025-09 · 35,676 actual arrivals · 2025-10 · 50,782 actual arrivals · 2025-11 · 41,934 actual arrivals · 2025-12 · 35,924 actual arrivals · 2026-01 · 69,227 model forecast · 2025-01 · 55,643 model forecast · 2025-02 · 66,547 model forecast · 2025-03 · 46,501 model forecast · 2025-04 · 24,793 model forecast · 2025-05 · 36,467 model forecast · 2025-06 · 40,757 model forecast · 2025-07 · 56,002 model forecast · 2025-08 · 62,004 model forecast · 2025-09 · 45,994 model forecast · 2025-10 · 39,213 model forecast · 2025-11 · 44,607 model forecast · 2025-12 · 43,906 model forecast · 2026-01 · 49,264 seasonal naive · 2025-01 · 79,254 seasonal naive · 2025-02 · 69,112 seasonal naive · 2025-03 · 42,064 seasonal naive · 2025-04 · 37,346 seasonal naive · 2025-05 · 39,825 seasonal naive · 2025-06 · 39,496 seasonal naive · 2025-07 · 57,399 seasonal naive · 2025-08 · 60,986 seasonal naive · 2025-09 · 44,886 seasonal naive · 2025-10 · 42,088 seasonal naive · 2025-11 · 39,600 seasonal naive · 2025-12 · 42,842 seasonal naive · 2026-01 · 79,301

Backtest design and the COVID structural break

The training window contains the 2020-2022 COVID structural break (arrivals collapse to near zero, then a staggered reopening recovery). Pre-registered handling: keep the break in training (the model sees the regime through its lag features), hold out 2025-01..2026-01 (13 post-recovery months) and report accuracy only on the holdout. The MASE denominator also spans the break, shrinking MASE values for model and baseline alike; model-vs-naive comparison is scale-invariant.

seed 42 · data labels: proxy: tourism arrivals stand in for cross-border card spend · generated by scripts/corridor_exhibit.py --check-able · lightgbm 4.7.0 · numpy 2.5.2 · openpyxl 3.1.5 · python 3.13.1 · scikit-learn 1.9.0 · shap 0.52.0

The full account

Implementation plan: everything that moved off the page

Cost, payback shape and the lifecycle design

Building One Loop took a measured 8.3 GPU-hours and 131 CPU-core-hours on a university cluster: cash cost US$0, or US$27.56 at cited public on-demand rates with every failed run included, and refreshing embeddings for one million merchants projects to US$2.42 per run at the same rates. The annual value it stands against is a declared-assumption planning model, so the payback ratio is an assumption, not a measurement.

One-off build cost. Hours are measured from Slurm accounting across every job this project ran, failed and cancelled runs included; dollar figures are those hours at public on-demand rates read on one day, so the market-equivalent and refresh figures are declared assumptions rather than measurements; the cash column is the measured amount actually paid.
GPU-hoursCPU-core-hours Cash paidMarket equivalent Refresh, one million merchants
8.3 131 US$0 US$27.56 US$2.42

The following clarifies what that figure is and what it is not. It is the market equivalent of a run that was free to us on a university cluster, at prototype scale: most of those hours went to the public TabFormer corpus of 24 million synthetic transactions, and the rest to this page's other exhibits, including the real-corpus replication in Section 5. It is what those hours would have cost to rent, not what this would cost American Express in production. Every job this project ran is inside it, including the 18 of 46 that failed or were cancelled, because research compute includes the runs that did not work. The price mapping is conservative in one direction throughout: a MIG slice priced as a whole card, an H100 NVL at the top listed H100 tier, the older Titan and T4 cards at the A100 rate. Prices are public on-demand rates read on 2026-08-24, excluding tax68,69,70. The refresh figure is a ceiling and not a rate: the embedding phase's wall time is bracketed by two artifact timestamps, and everything inside that bracket is charged to the 238,615 merchant embeddings alone. Laptop-side development, operator time and data egress are not costed. Latency, throughput and headcount stay unmeasured, so no latency figure, no throughput figure and no team size appears here or anywhere else on this page.

The payback is stated as a shape rather than as a return. Against the conservative scenario's declared-assumption US$517,500 a year, a one-off cost of this size is negligible. The ratio is large because the denominator is small, not because the numerator is proven: the value side stays a declared-assumption planning model, and no return on this page is measured. Nothing in the lifecycle design below is measured either. No serving path has been benchmarked and no team has been sized, so what follows is the design a model owner signs off on before any of it runs, and every cadence in it is a planned constant a pilot moves, not a finding we are reporting.

The lifecycle design is summarized in this paragraph. Head retraining runs inside the perimeter on the regional shard that owns the data, a planned quarterly batch rather than a continuous one; backbone retraining is a single scheduled run where the pretraining corpus sits, shipping to the shards as a new checkpoint tag, because the leakage-hardened evaluation in Section 5 has to be read between checkpoints; the as-of embedding store refreshes nightly in between, so a merchant signed last week is servable without moving the model. Two of the three heads need no online call, since the signing list and the corridor forecast are scheduled files; the offers head is the only request-time path, and there the embedding is a lookup rather than a computation, which we have not measured, so we quote no latency for it. Every vector, list and lift report carries its checkpoint tag and a campaign is pinned to the tag in force when it opens, so a rollback is a tag switch and the rev-share basis stays recomputable. Drift has three signals with a producer already on this page: unseen-token share per field, the distribution of the label-free surprise score, and the randomized holdout every campaign carries. What we cannot state is the alarm threshold on any of them, because a threshold set without a production baseline would be an invented number, and this page contains no invented numbers. The owner is the SG CoE as model owner, with ICS and GMNS APAC on the issuing and partner sides.

The brief asks that every solution be sustainable, compliant, ethical, and designed to scale, and in this entry each of those four is supported by an exhibit rather than asserted as a description. The designed-to-scale requirement is addressed by the lifecycle paragraph above together with the scaling curve and its two pre-registered controls in Section 5. The compliance requirement is addressed by the obligation map in Section 4, which has a built exhibit on every row and makes no compliance claim on any row. The ethical requirement is measured where we could measure it, through the subgroup selection audit in Section 6, with the protected- attribute gap stated rather than omitted. The sustainability requirement is addressed by the operating footprint: the build consumed hours that we metered rather than capacity that we assert, and the retrain cadence above is a quarterly batch rather than a standing fleet.

The public record and what is added

Public record at AmexWhat One Loop adds
Gen X fraud stack: sequence deep learning plus GBMs making 8 billion-plus automated risk decisions on over $1 trillion of volume, and the lowest US fraud rate among major networks for 19 straight years Points the same sequence intelligence at growth tasks; the next-merchant-category task in Section 5 shows the embeddings carry over on the synthetic corpus, with the delta's confidence interval fully above zero
Orchestra, the internal engine personalizing consumer Amex Offers71 A partner-facing productization: offers for GNS partner campaigns, each with an auditable incremental-lift report
The 2024 SG CoE uplift contest, scored on incremental activation22 Turns the contest metric into an always-on measurement service partners subscribe to
Published sequence research for credit monitoring49 One pretrained backbone across issuer and acquirer views, serving three heads instead of one task
GenAI Council funnel, AI firewall, ConnectChain orchestration42,41,40 A MAS-mapped governance wrapper (Veritas 2.0, AIRG, SAFR) purpose-built for a partner-facing decision product

The gate, the ladder and who pays

The cost of being wrong is low. Phase 1 is a kill gate rather than a milestone. If embeddings do not beat production per-task features on at least 2 of 3 heads offline with paired intervals, it stops there and the partner phase never opens. That gate should be read against Section 5 before it is approved: on the one real corpus we could test, the store-level head cleared it in both seeds, the household transfer head failed it in both, and the offer head came back mixed. On that evidence the gate is close, and we would rather put up a gate we might fail than one we already know we pass.

And a failed gate still leaves something behind. Read the sentence above as a VP should: on our own real-corpus evidence this gate is as likely to fail as to pass, so the honest question is what Phase 1 is worth if it does fail. The answer is the protocol ladder, and it is not a promise. It is built, it is running in Section 5, and it needs no Amex data and no cardmember data. It is the rig that priced our own transfer claim from 0.0590 under a permissive protocol to 0.0218 under ours, with one entity-disjoint split alone costing -0.0831 of it, enough to turn a headline gain into a loss. On the one claim it has actually been pointed at, our own, 63.0 percent of the headline gain turned out to be protocol rather than model. It has never been pointed at anyone else's work, so that is a correction size and not a catch rate, and we are not implying one. Where it sits in your organisation is beside model risk and validation rather than instead of them: they own whether a model may ship, and this prices how much of a claimed lift the evaluation is holding up before that question is asked. You know what a wrongly-adopted model costs you and we do not, so we give the correction size and our own rate, 7 of 27 on the wrong side of zero and 13 more that could not be told from it, and you put your number against them. That is what a failed gate leaves behind, and we would rather say it now than have you assume it leaves nothing.

No data has to move. We are a university team with no access to American Express data, and Phase 1 is written so that stays true: it runs inside the perimeter on data the CoE already holds. What crosses the boundary is our side of it, the pretraining recipe, the leakage rules, the evaluation code and the pre-registration discipline this page is written in. The CoE supplies the data and the sign-off; we supply the build and the evidence standard it gets measured to. That is a smaller ask than access, and it is why six months is enough.

The allocation of cost and revenue. Discount revenue on incremental billed business accrues to the network either way; the partner pays a share only of the lift the measurement layer certifies, recomputable by that partner from their own arm counts. The partner's side is as follows: they keep their issuing economics on every certified incremental dollar, minus the share, both terms on the same recomputable base. The share itself is a pilot commercial term, not ours to set; the design fixes the base, certified lift only. The merchant is the brief's other half: offer spend steered by certified incremental effect instead of blanket discounting, on the same report. The holdout's cost is not hidden, the control arm forgoes the effect, US$0.4244 of spend per customer on this exhibit's own arms, and that is the number the fee is weighed against. A partner can run holdouts on its own book; what it cannot do alone is have the party it pays accept its own grading, the conflict the co-signed design below removes.

The full account

the first screen: the material that moved off the page

The brief scored, the themes and the record

The three questions in the title are the brief's own three worked examples, in its order, and the brief asks the second to carry "measurable lift for both merchants and issuing partners", which is why the measurement layer is the spine of this entry rather than a feature. Scored against those three questions, losses included:

The brief's three questions, and how this entry scores on each, from its own exhibits
The brief asksDelivered here The printed verdict
Identify merchant opportunities for the sales teams to sign for acceptance A ranked signing list of pseudonymized merchant clusters over 364,635 real Singapore places, with a stated reason per row The list, the reasons and the forward check are all delivered; that check LOST to plain venue density, printed in Section 7 and in page one's loss list. Clusters rather than named merchants because the prototype is on public data; merchant-level is a Phase 1 deliverable inside the perimeter, not something shown here
Partner-led personalized offers with measurable lift, packaged as a value-added service for issuing and acquiring partners across APAC The measurement layer, and the lift report a partner recomputes from their own arm counts Top decile captures 59.3 percent of incremental visits against 48.6 for response ranking; the pre-registered primary endpoint LOST and is printed at the same size
Predict and prioritize the highest-potential cross-border corridors, explainable, forward-looking, actionable A thirteen-month backtest over twelve corridors of Singapore inbound arrivals, a stated public proxy for cross-border card spend, drivers grouped into the brief's own families The pre-registered blend wins on macro-MASE, the scaled forecast error averaged over corridors, and not on total arrivals; the model alone loses pooled to the seasonal naive, so the lane is claimed on neither view. Half the question is unanswered: the public proxy carries no merchant category, so corridors are ranked and the categories inside them are not

We pretrained a transaction backbone before writing this proposal, aimed at the closed loop: cardmember, network, and merchant in one event stream. It trained on the public IBM TabFormer benchmark: 24 million synthetic card transactions, labeled as synthetic everywhere on this page. Each product head is prototyped on the best public dataset for its job, and every number traces to a committed script or a cited source.

  • What is built, and what is not The backbone is pretrained and frozen on a public synthetic corpus, never on Amex data. Each head is prototyped on the best public dataset for its job, the offers head on two randomized experiments, the rest observational. The wiring from backbone to head is measured on one real corpus at small scale, and Section 7 says what each head reads from the backbone today.
  • The record

    Sources

    1. Base-scenario modeled annual total of the bottom-up value model by year three of an APAC rollout, three lanes summed before rounding. Every lane is modelled from declared inputs, the offers lane included: its randomized result establishes the targeting mechanism and the counting rule, not the size. Derivation in copy/value-model.md. · value: US$10.5M · internal derivation, arithmetic in the value model · copy/value-model.md · accessed 2026-08-22
    2. The IBM TabFormer repository ships a synthetic credit-card transaction dataset of 24 million records under Apache-2.0. · value: 24 million · https://github.com/IBM/TabFormer · accessed 2026-08-22
    3. Productivity, protection, and growth is Amex's own GenAI framing, stated by EVP and CTO Hilary Packer. · value: productivity, protection, and growth · https://www.americanexpress.com/us/business/american-express-ventures/articles/gen-ai-dinner/ · accessed 2026-08-22
    4. 81.8% of Singapore scam cases in 2025 involved self-effected transfers, per the SPF Annual Scam and Cybercrime Brief 2025. · value: 81.8% · https://www.police.gov.sg/-/media/SPF/Media-Room/Statistics/Annual-Scams-and-Cybercrime-Brief-2025/Annual-Scam-and-Cybercrime-Brief-2025.pdf · accessed 2026-08-22
    5. Stripe's payments foundation model raised card-testing fraud detection for large users from 59% to 97% with no increase in false positives. · value: from 59% to 97% · https://www.ai-street.co/p/stripe-built-a-payments-llm-to-fight-fraud · accessed 2026-08-22
    6. On the Q4 2025 earnings call, CEO Steve Squeri said Amex will continue to build coverage in international as the fastest-growing part of the business. · value: continue to build coverage, obviously, in international as that continues to grow and continues to be the fastest-growing overall part of our business · https://www.fool.com/earnings/call-transcripts/2026/01/30/american-express-axp-q4-2025-earnings-call-transcript/ · accessed 2026-08-22
    7. American Express FY2025 total revenues net of interest expense in APAC (Asia Pacific, Australia and New Zealand) were $5,218 million of $72,229 million consolidated, which is 7.2 percent, per Table 23.2 Summary of Total Revenue and Pretax Income by Region in the FY2025 Form 10-K. · value: $5.22 billion (7.2% of the company total) · https://www.sec.gov/Archives/edgar/data/4962/000000496226000080/axp-20251231.htm · accessed 2026-08-25
    8. American Express maintains relationships with third-party banks and other institutions in approximately 110 countries and territories through its card network business, per the business description in the FY2025 Form 10-K. · value: approximately 110 countries and territories · https://www.sec.gov/Archives/edgar/data/4962/000000496226000080/axp-20251231.htm · accessed 2026-08-25
    9. Indonesia's QRIS standard merchant discount rate is 0.3%, waived entirely for transactions under Rp500,000 since December 2024. · value: 0.3% · https://digitalinasia.com/how-digital-payments-work-in-southeast-asia/ · accessed 2026-08-22
    10. Card merchant discount rates in Southeast Asia run 1.5% to 3% of transaction value. · value: 1.5% to 3% · https://digitalinasia.com/how-digital-payments-work-in-southeast-asia/ · accessed 2026-08-22
    11. Nexus Global Payments was incorporated in Singapore in March 2025 by five central banks, with go-live targeted for 2026. · value: go-live target 2026 · https://www.bis.org/about/bisih/topics/fmis/nexus.htm · accessed 2026-08-22
    12. The G20 target is retail cross-border payment cost at or under 1% by end-2027. · value: at or under 1% · https://www.fsb.org/2025/10/g20-roadmap-for-cross-border-payments-consolidated-progress-report-for-2025/ · accessed 2026-08-22
    13. Hawker stalls and small retailers in Singapore commonly do not accept Amex, citing swipe fees of about 1% to 3%. · value: swipe fees of about 1% to 3% · https://www.singsaver.com.sg/credit-card/blog/accept-american-express · accessed 2026-08-22
    14. Acquiring a new merchant costs 5 to 25 times as much as retaining an existing one. · value: 5 to 25 times · https://www.vahorizon.site/b2b/glossary/attrition-merchant/ · accessed 2026-08-22
    15. Conservative-scenario modeled annual total of the bottom-up value model, three lanes summed before rounding; derivation in copy/value-model.md. · value: US$0.52M · internal derivation, arithmetic in the value model · copy/value-model.md · accessed 2026-08-22
    16. Stretch-scenario modeled annual total of the bottom-up value model, three lanes summed before rounding; derivation in copy/value-model.md. · value: US$108.5M · internal derivation, arithmetic in the value model · copy/value-model.md · accessed 2026-08-22
    17. Outside proprietary markets, Amex reaches APAC through Global Network Services licensed bank issuers and acquirers chosen market by market. · value: market-by-market licensed partners · https://www.sec.gov/Archives/edgar/data/4962/000119312517047588/d321397d10k.htm · accessed 2026-08-22
    18. The Amex network partnership with Maybank in Malaysia is live. · value: Maybank (Malaysia) · https://www.americanexpress.com/en-my/network/credit-cards/maybank/ · accessed 2026-08-22
    19. Amex acceptance reached about 160 million locations globally as of mid-2025, up 16% year over year and about 5x since 2017. · value: about 16% in a year, to roughly 160 million · https://finance.yahoo.com/news/american-express-accepted-160-million-123000510.html · accessed 2026-08-22
    20. Industry-average merchant attrition is 23.6% per year (The Strawhecker Group). · value: 23.6% · https://www.digitaltransactions.net/magazine_articles/a-remedy-for-attrition-headaches/ · accessed 2026-08-22
    21. Merchant attrition costs acquirers about $2 billion a year in losses plus about $1 billion a year replacing lost merchants. · value: $2 billion a year in losses plus about $1 billion replacing lost merchants · https://www.paymentsjournal.com/solving-the-merchant-attrition-problem/ · accessed 2026-08-22
    22. The 2024 Amex SG hackathon was an uplift-modeling contest scoring 12,604,600 customer x merchant pairs on incremental activation rate against a randomized holdout. · value: 12.6 million customer x merchant pairs · https://github.com/ZzzhangXiao/Amex-AI-Hackathon · accessed 2026-08-22
    23. Cardlytics measures card-linked offer incrementality via test and control comparisons, with no published methodology partners can audit. · value: test-and-control incrementality · https://www.cardlytics.com/ · accessed 2026-08-22
    24. A published CausalML case reports uplift-targeting 30% of users matching the conversion gain of blanket promotion. · value: 30% · https://github.com/uber/causalml · accessed 2026-08-22
    25. American Express FY2025 consolidated Marketing expense was $6,252 million, per the consolidated statements of income in the FY2025 Form 10-K (FY2024 $6,040 million, FY2023 $5,213 million). · value: $6.25 billion · https://www.sec.gov/Archives/edgar/data/4962/000000496226000080/axp-20251231.htm · accessed 2026-08-25
    26. GBTA forecasts APAC business-travel spend at $700.9 billion in 2026, up 10.9% year over year, the fastest of any region. · value: $700.9 billion · https://gbta.org/asia-pacific-business-travel-to-surpass-700-billion-in-2026-leading-global-growth-amid-geopolitical-uncertainty/ · accessed 2026-08-22
    27. Visa's TREASURE transaction encoder improves abnormal-behavior detection by 111% over Visa production systems and lifts recommendation models 104% as an embedding provider. · value: +111% · https://arxiv.org/abs/2511.19693 · accessed 2026-08-22
    28. Nubank's nuFormer transaction model runs in production decision engines serving over 100 million customers, with about 1.2% average AUC lift across benchmark tasks. · value: 100+ million customers · https://arxiv.org/abs/2507.23267 · accessed 2026-08-22
    29. Revolut's PRAGMA scales an encoder-only banking-event model to 1 billion parameters over 24 billion events. · value: a billion parameters · https://philippdubach.com/posts/inside-pragma-revoluts-foundation-model-for-banking/ · accessed 2026-08-22
    30. Amex's Singapore Decision Science Center of Excellence launched in December 2022 and expanded in November 2023 into marketing and servicing models plus a generative AI R&D practice. · value: Singapore Decision Science CoE · https://www.edb.gov.sg/en/about-edb/media-releases-publications/american-express-expands-singapore-decision-science-center-of-excellence.html · accessed 2026-08-22
    31. The MAS SAFR white paper (July 3, 2026) defines four runtime components that evaluate every proposed AI-agent action before execution. · value: SAFR · https://www.mas.gov.sg/publications/monographs-or-information-paper/2026/safeguards-for-agentic-finance-at-runtime · accessed 2026-08-22
    32. SAFR v1.0 page 9 states the framework's own open problem for an agent-declared governance envelope: the envelope is treated as a document to be authenticated against its origin, because its own contents cannot be taken as proof of what the original instruction produced. · value: page 9 · https://www.mas.gov.sg/-/media/mas-media-library/development/fintech/ai-safr/safr.pdf · accessed 2026-08-23
    33. In the 25-page SAFR v1.0 PDF, both open-loop networks appear as contributors and carry published four-component case-study mappings (Mastercard 5 mentions, Visa 8). American Express appears zero times, in neither the contributor list nor the case studies. Counted by grep over the PDF text by this team on 2026-08-23. · value: 0 mentions · https://www.mas.gov.sg/-/media/mas-media-library/development/fintech/ai-safr/safr.pdf · accessed 2026-08-23
    34. Enigma Merchant Transaction Signals sells transaction-derived monthly merchant attributes for 10M+ US small businesses; US-only, raw signals rather than ranked signing propensity. · value: US merchants only · https://enigma.com/resources/blog/introducing-merchant-transaction-signals · accessed 2026-08-22
    35. Mastercard Economics Institute Travel Trends is a descriptive annual corridor report, not a forward-looking per-corridor scoring product. · value: annual report · https://www.mastercard.com/news/ap/en/newsroom/press-releases/en/2025/mastercard-economics-institute-on-travel-in-2025-asia-pacific-leads-trending-summer-destinations-for-second-year-running/ · accessed 2026-08-22
    36. PRAGMA reports a 47.1% drop on AML performance because single-user sequences miss cross-user graph structure. · value: 47.1% · https://philippdubach.com/posts/inside-pragma-revoluts-foundation-model-for-banking/ · accessed 2026-08-22
    37. MAS Veritas Toolkit 2.0 (June 2023) is the open-source assessment methodology for fairness and explainability in financial-institution AI. · value: Veritas 2.0 · https://www.mas.gov.sg/news/media-releases/2023/toolkit-for-responsible-use-of-ai-in-the-financial-sector · accessed 2026-08-22
    38. The ghost-ads design (Johnson, Lewis, Nubbemeyer, JMR 2017) measures incremental lift continuously at platform scale and runs in production at DoorDash Ads. · value: ghost ads · https://journals.sagepub.com/doi/10.1509/jmr.15.0297 · accessed 2026-08-22
    39. Amex leadership states GenAI is not used for credit line or approval decisions. · value: no credit decisions by GenAI · https://www.americanbanker.com/news/american-express-unveils-its-approach-to-generative-ai · accessed 2026-08-22
    40. ConnectChain is Amex's open-source LangChain-derived enterprise orchestration framework for GenAI. · value: ConnectChain · https://americanexpress.io/generative-ai-meets-open-source-at-american-express/ · accessed 2026-08-22
    41. Amex envelopes its AI systems in an AI firewall with orchestration layers and human-in-the-loop review. · value: AI firewall · https://venturebeat.com/ai/how-amex-uses-ai-to-increase-efficiency-40-fewer-it-escalations-85-travel-assistance-boost · accessed 2026-08-22
    42. Amex's Generative AI Council triaged about 500 identified use cases down to about 70 in implementation. · value: 500 use cases triaged to about 70 · https://www.americanbanker.com/news/american-express-unveils-its-approach-to-generative-ai · accessed 2026-08-22
    43. Amex's Agentic Commerce Experiences developer kit defines five services, with two of five still under development as of April 2026. · value: two of five services under development · https://www.americanexpress.com/en-us/company/agentic-commerce/ · accessed 2026-08-22
    44. Amex's engineering blog names traceability from final transaction back to original customer intent as an unsolved problem in agentic commerce. · value: intent-to-settlement traceability · https://americanexpress.io/shaping-the-future-of-agentic-commerce/ · accessed 2026-08-22
    45. Duncan, Keller-McNulty and Stokes, Disclosure Risk vs. Data Utility: The R-U Confidentiality Map, National Institute of Statistical Sciences Technical Report 121, 2001. Sets out the standard statistical-disclosure-control way to report a protection decision: plot disclosure risk against data utility for a family of disclosure-limiting transformations and read the trade-off off the curve rather than asserting one setting. · value: NISS Technical Report 121, 2001 · https://www.niss.org/sites/default/files/technicalreports/tr121.pdf · accessed 2026-08-24
    46. Singapore's Shared Responsibility Framework, effective 16 December 2024, covers seemingly authorised transactions, where a scammer obtains credentials and transacts, and adds an FI duty of real-time fraud surveillance for accounts rapidly drained of a material sum. Self-effected transfers by the account holder sit outside that defined scope. · value: 16 December 2024 · https://www.mas.gov.sg/regulation/guidelines/guidelines-on-shared-responsibility-framework · accessed 2026-08-23
    47. Amex EVP for global fraud Tina Eide said that as transaction controls became sophisticated enough that bad actors struggled to make money there, attacks came back around to scams, social engineering and fraudulent applications (Payments Dive interview, 19 December 2023). · value: coming back around to more of the scams · https://www.paymentsdive.com/news/amex-fraud-trends-social-engineering-payments-ai-cards-tina-eide/702911/ · accessed 2026-08-23
    48. TabBERT/TabFormer (ICASSP 2021) introduced hierarchical BERT-style pretraining over tabular transaction sequences. · value: arXiv 2011.01843 · https://arxiv.org/abs/2011.01843 · accessed 2026-08-22
    49. Amex AI Research published sequential deep learning over card-transaction histories for credit risk monitoring. · value: arXiv 2012.15330 · https://arxiv.org/abs/2012.15330 · accessed 2026-08-22
    50. NVIDIA published a transaction foundation model blueprint on 16 June 2026 that pretrains a compact Llama-style decoder-only model with a next-token (causal language modelling) objective on the same public IBM TabFormer corpus this page uses, then feeds its pooled embeddings to an XGBoost fraud classifier and reports an average-precision lift over that XGBoost baseline. · value: NVIDIA transaction foundation model blueprint, June 2026 · https://developer.nvidia.com/blog/build-your-own-transaction-foundation-model-for-financial-intelligence/ · accessed 2026-08-23
    51. The NVIDIA blueprint's transaction foundation model is a decoder-only Llama architecture of about 29 million parameters, hidden size 512, 8 layers, grouped-query attention with 8 query heads and 2 key-value heads. · value: 29M · https://developer.nvidia.com/blog/build-your-own-transaction-foundation-model-for-financial-intelligence/ · accessed 2026-08-23
    52. The NVIDIA transaction foundation model blueprint ships as an Apache-2.0 repository of notebooks plus a pretrained checkpoint, so the prior art is public and reproducible. · value: Apache-2.0 · https://github.com/NVIDIA-AI-Blueprints/transaction-foundation-model · accessed 2026-08-23
    53. The Amex Default Prediction Kaggle competition (2022, about 4,875 teams) was won by GBDT models built on per-customer temporal aggregates, ensembled with transformers. · value: per-customer temporal aggregates plus GBDT · https://www.kaggle.com/competitions/amex-default-prediction · accessed 2026-08-22
    54. Pendlebury, Pierazzi, Jordaney, Kinder and Cavallaro, TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and Time, USENIX Security 2019. Names temporal bias, from an incorrect time split of training and test data, and spatial bias, from an unrepresentative test distribution, as the two sources that inflate reported classifier performance, and treats the experimental protocol as the thing to constrain and measure. · value: USENIX Security 2019 · https://www.usenix.org/conference/usenixsecurity19/presentation/pendlebury · accessed 2026-08-24
    55. Ji, Sun, Zhang and Li, A Critical Study on Data Leakage in Recommender System Offline Evaluation, ACM Transactions on Information Systems 41(3), Article 75, 2023. Shows that a train/test split that ignores the global timeline changes the recommendation lists a model produces and moves its reported accuracy, and that the direction of the change is not predictable, so more future data in training does not reliably mean higher accuracy. · value: ACM TOIS 41(3), Article 75, 2023 · https://dl.acm.org/doi/10.1145/3569930 · accessed 2026-08-24
    56. Pre-registration of the Complete Journey replication: hypotheses, seeds, split and decision rule, committed before the run. · value: 6c9c7a9 · CJ-REPLICATION-PREREG.md · accessed 2026-08-24
    57. Pre-run amendment to the Complete Journey pre-registration: scored unit, pinned evaluation seed and unscorable floor, committed before the run. · value: 51c69e5 · CJ-REPLICATION-PREREG-AMENDMENT-1.md · accessed 2026-08-24
    58. American Express reported the lowest US credit card fraud rate among major card networks for the 19th consecutive year (July 2026). · value: 19 straight years · https://www.americanexpress.com/en-us/newsroom/articles/products-and-services/for-the-19th-year-in-a-row--american-express-has-the-lowest-u-s-.html · accessed 2026-08-22
    59. Amex's Gen X fraud stack issues about 8 billion automated risk decisions a year over more than $1 trillion of transactions with millisecond scoring. · value: 8 billion-plus automated risk decisions on over $1 trillion of volume · https://www.nvidia.com/en-us/customer-stories/american-express-prevents-fraud-and-foils-cybercrime-with-nvidia-ai-solutions/ · accessed 2026-08-22
    60. Ascarza (JMR 2018) showed in randomized field experiments that switching from response-based to treatment-effect targeting yielded up to 6.8 percentage points additional churn reduction. · value: 6.8 percentage points · https://journals.sagepub.com/doi/abs/10.1509/jmr.16.0163 · accessed 2026-08-22
    61. The X-learner meta-learner (Kunzel et al.) is provably efficient when treatment arms are unbalanced. · value: X-learner · https://arxiv.org/abs/1706.03461 · accessed 2026-08-22
    62. Pre-registered design of the Criteo uplift exhibit, naming conversion the primary outcome and visit the robustness outcome, committed in scripts/uplift_exhibit.py before the full-data run. · value: 2fd8c23 · scripts/uplift_exhibit.py · accessed 2026-08-24
    63. Foursquare OS Places is an open dataset of 106M+ points of interest under Apache-2.0, including APAC coverage. · value: Foursquare OS Places · https://opensource.foursquare.com/os-places/ · accessed 2026-08-22
    64. Pre-registration of the whitespace forward check: cutoff rule, outcome window, arms and decision rule, committed before any outcome was computed. · value: fef8af5 · WHITESPACE-TEMPORAL-PREREG.md · accessed 2026-08-24
    65. MinT trace-minimization reconciliation (Wickramasuriya, Athanasopoulos, Hyndman, JASA 2019) makes hierarchical forecasts coherent across aggregation levels. · value: grouped reconciliation · https://robjhyndman.com/publications/mint/ · accessed 2026-08-22
    66. World Bank/KNOMAD bilateral matrices allocate flows in proportion to bilateral migrant stocks and incomes, the standard gravity approach for corridor modeling. · value: migrant stocks · https://data360files.worldbank.org/data360-data/metadata/WB_KNOMAD/WB_KNOMAD_BRE.pdf · accessed 2026-08-22
    67. Pre-registration of the corridor forecast combination: fixed equal weights and decision rule, committed before any combination number was computed. · value: 418aef3 · CORRIDOR-COMBINATION-PREREG.md · accessed 2026-08-24
    68. Lambda publishes on-demand per-GPU-hour prices for its A100 and H100 instances; these are the rates the One Loop cost model assumes for those cards. · value: public on-demand GPU rates · https://lambda.ai/pricing · accessed 2026-08-24
    69. CoreWeave publishes an on-demand instance price for HGX H200, used because Lambda lists no on-demand H200 rate; CoreWeave sits above the specialist-cloud range, so the mapping overstates cost. · value: public on-demand H200 instance rate · https://www.coreweave.com/pricing · accessed 2026-08-24
    70. AWS EC2 c7i.4xlarge on-demand Linux pricing in us-east-1, read from the Vantage mirror of AWS published prices, is the assumed CPU core-hour rate in the cost model. · value: public on-demand CPU rate · https://instances.vantage.sh/aws/ec2/c7i.4xlarge · accessed 2026-08-24
    71. Orchestra is Amex's proprietary real-time relevance-prediction engine powering consumer Amex Offers personalization. · value: Orchestra · https://www.americanbanker.com/payments/news/inside-amexs-data-driven-effort-to-personalize-customer-service · accessed 2026-08-22

    incremental effect

    Whether the offer caused the visit

    The incremental effect, also called uplift or lift, is the extra outcome an action causes: visits that happened because the offer went out, not visits that would have happened anyway.

    Why it matters here. Partners pay a share only of certified incremental billed business, so the fee base is the increment and not the response. This layer exists to prevent payment for customers who would have converted anyway.

    • Uplift ranking, top tenth59.3
    • Response ranking, top tenth48.6

    Percent of every incremental site visit the campaign produced that is captured by targeting one tenth of a randomized Criteo holdout; visit endpoint, point estimates.

    Endpoint
    • Uplift ranking, top tenth59.3
    • Response ranking, top tenth48.6
    • Uplift ranking, top tenth59.2
    • Response ranking, top tenth68.6

    In this entry. On the randomized Criteo holdout of 4,193,878 rows, ranking by predicted incremental effect concentrates 5.9 times the average visit effect into the top tenth, and adds 0.011 incremental visits per customer over response ranking, interval [0.009, 0.014], entirely above zero.

    For the technical reader. The estimand is the conditional average treatment effect, CATE: treated minus control outcome rate within a group, estimated per customer by an X-learner read off a randomized holdout. The average treatment effect over the holdout is 0.010326; the top decile by predicted CATE returns 0.061241. The ratio carries no interval. The paired difference against response ranking does, and it clears zero only at the top decile, fading at twenty and thirty percent. The result depends on the choice of endpoint: on conversion the same ranking lost at every depth.

    Read from

    randomized holdout

    The customers who received no offer, and why that is by design

    A randomized holdout is a slice of customers chosen at random to receive no offer. Their outcome is what would have happened anyway, so treated minus holdout is the increment, with nothing modelled.

    Why it matters here. Every partner campaign ships with one by design, co-signed by the partner before it opens. It is the only instrument that tells a team which targeting rule produced the increment; a response score cannot provide this information.

    • Segment the response model skips5.82
    • Segment the response model mails2.31

    Measured visit uplift in points from Hillstrom's randomized arms: the segment ranked 26 of 28 moves most; the one ranked 5 cannot be told from zero.

    In this entry. A response model ranks one segment 26 of 28 and never mails it; the randomized arms measure its visit uplift at 5.82 points, interval [3.96, 7.67], clear of zero. The segment it ranks 5 and mails measures 2.31 points, interval [-1.25, 5.87], spanning zero.

    For the technical reader. The design is the ghost-ads pattern: assignment is randomized before targeting, so the holdout is exchangeable with the treated arm and the arm contrast is an unbiased estimate of the effect among the targeted. Intervals on the segment table are the normal approximation on the stored standard error. The cost is real: the control arm forgoes the effect, and the fee is weighed against that cost. The Phase 2 gate is measured incremental billed business above zero at ninety-five percent confidence, recomputable by the partner from its own arm counts.

    Read from

    Qini coefficient

    Whether the whole ranking beats random targeting

    The Qini coefficient scores an uplift ranking as a whole: the area between the cumulative incremental gain of the ranking and the straight line random targeting would draw. Higher means the increment is found sooner.

    Why it matters here. It is the curve metric on the pre-registered primary endpoint, and the uplift ranking lost it on conversion and was not separated on visit. The win claimed is therefore a depth result at the top decile rather than a whole-curve result.

    • X-learner, uplift ranking0.000199
    • Response ranking0.000345

    Qini coefficients on the conversion endpoint of the randomized Criteo holdout; the paired difference sits fully below zero, so response ranking wins the curve.

    Endpoint
    • X-learner, uplift ranking0.000199
    • Response ranking0.000345
    • X-learner, uplift ranking0.002993
    • Response ranking0.003015

    In this entry. Qini on conversion: X-learner 0.000199 against response ranking 0.000345, paired difference interval [-0.000194, -0.000089], fully below zero. On visit, 0.002993 against 0.003015, paired difference [-0.000229, 0.000159], spanning zero: the two rankings are not separated across the curve.

    For the technical reader. Qini is computed on the holdout by sorting customers on the model score and, at each depth, taking treated conversions minus control conversions rescaled to the treated count, then integrating the excess over the random diagonal. Conversion is rare here, holdout rates 0.00304 treated and 0.001947 control, so per-customer effect estimates are noise-dominated at the resolution an X-learner needs, and the effect correlates with baseline propensity, so response propensity proxies the uplift order. Paired difference intervals are stratified bootstrap. The endpoint hierarchy was fixed before the full run, at 2fd8c23.

    Read from

    PR-AUC

    Two ranking metrics for one review queue

    ROC-AUC is the chance a random fraud scores above a random non-fraud. PR-AUC is the area under precision against recall, and it moves only with how the rare positives are ranked near the top.

    Why it matters here. Fraud positives are rare enough that an uninformed ranking scores 0.00315 PR-AUC, so ROC-AUC is carried by the negatives and PR-AUC leads. The two can disagree, and on the surprise score they do.

    • Contextual surprise, the model0.1338
    • Global rarity, counts only0.1282
    • Both, unweighted sum0.3296

    PR-AUC on the same 300,000 scored synthetic transactions carrying 945 fraud positives; the combination clears both single scores on both metrics.

    Metric
    • Contextual surprise, the model0.1338
    • Global rarity, counts only0.1282
    • Both, unweighted sum0.3296
    • Contextual surprise, the model0.9022
    • Global rarity, counts only0.9367
    • Both, unweighted sum0.9497

    In this entry. Contextual surprise minus global rarity is 0.0056 on PR-AUC, interval [-0.0391, 0.0506], spanning zero, and -0.0345 on ROC-AUC, interval [-0.0523, -0.0159], entirely below zero: a tie on one metric and a loss on the other, which is why the model alone does not carry the task.

    For the technical reader. Both are threshold-free ranking metrics. ROC-AUC integrates true positive rate against false positive rate and is insensitive to class balance; PR-AUC integrates precision against recall, so every false positive near the top costs it directly, which is why it discriminates when positives are this rare. Intervals are entity-level paired bootstrap on identical rows. The split keeps every fraud positive and thins the negatives to a row cap, so PR-AUC levels compare the two arms rather than size a live queue. Neither metric was named the lead before results existed.

    Read from

    transfer

    Whether frozen embeddings help an unseen task

    Transfer is when a model trained for one job helps another it was never trained on. Here the frozen backbone's embeddings are added to a strong per-task baseline, and the gain, the delta, is what transferred.

    Why it matters here. Transfer is the whole case for one backbone behind three heads: one embedding service instead of a feature pipeline per task. It counts only where the delta clears zero, and the nulls are printed beside the gains.

    • Next category, synthetic corpus0.02180.0103 to 0.0343
    • Household transfer, real corpus-0.0160-0.0207 to -0.0113
    • Store head, real corpus0.19120.0064 to 0.3865

    Delta with embeddings over the L3 baseline, paired entity-clustered intervals: a win on synthetic TabFormer, a loss and a win on real Complete Journey, first pretraining seed.

    Corpus
    • Next category, top one0.02180.0103 to 0.0343
    • Next category, top five0.02180.0106 to 0.0336
    • Fraud PR-AUC, full corpus0.0021-0.0041 to 0.0086
    • Fraud ROC-AUC, full corpus-0.0003-0.0012 to 0.0001
    • Household transfer, first seed-0.0160-0.0207 to -0.0113
    • Household transfer, second seed-0.0147-0.0197 to -0.0101
    • Store head, first seed0.19120.0064 to 0.3865
    • Store head, second seed0.22420.0449 to 0.4343

    In this entry. On the synthetic corpus the next-category baseline scores 0.235 top-one and 0.257 with embeddings, delta interval [0.0103, 0.0343], about two more right answers in every hundred. On the real corpus the same recipe cost accuracy: -0.0160 and -0.0147, both intervals below zero.

    For the technical reader. The baseline is LightGBM on per-entity temporal aggregates, the pattern that wins on Amex-style data in public competition; the comparison adds the frozen backbone's as-of embeddings to that same baseline. The full-corpus fraud delta is a null on a saturated baseline at 0.99389 PR-AUC; the gain appears where the baseline has headroom. On the real corpus the with-embedding arm lands under the 0.5101 majority-class floor, so protocol and corpus are both variables, and real closed-loop data at Amex scale stays the untested case the Phase 1 gate exists for.

    Read from

    leakage ladder

    The share of the lift attributable to protocol

    Leakage is test-time information reaching the model through the evaluation itself. The ladder re-scores one frozen backbone under four protocols, L0 permissive to L3 strict, turning the guards on one step at a time.

    Why it matters here. Published transfer lifts are far larger than ours, and the open question is whether the gap comes from the model or from the measurement. Applied to our own claim, the ladder attributed 63.0 percent of it to protocol. It does not audit any other published result.

    • L0 permissive, shared accounts0.0590
    • L1 entity-disjoint split-0.0240-0.0373 to -0.0112
    • L2 as-of embeddings0.01590.0077 to 0.0246
    • L3 baseline aggregates, shipped0.02180.0103 to 0.0343

    Next-category top-one delta, same frozen checkpoint, synthetic corpus. L0 scores different rows and carries no paired interval; L1 sits below zero; L3 is the shipped configuration.

    Add one guard
    • L0 permissive, shared accounts0.0590
    • L0 permissive, shared accounts0.0590
    • L1 entity-disjoint split-0.0240-0.0373 to -0.0112
    • L0 permissive, shared accounts0.0590
    • L1 entity-disjoint split-0.0240-0.0373 to -0.0112
    • L2 as-of embeddings0.01590.0077 to 0.0246
    • L0 permissive, shared accounts0.0590
    • L1 entity-disjoint split-0.0240-0.0373 to -0.0112
    • L2 as-of embeddings0.01590.0077 to 0.0246
    • L3 baseline aggregates, shipped0.02180.0103 to 0.0343

    In this entry. Next-category top-one delta: 0.0590 at L0 and then -0.0240 at L1 with interval [-0.0373, -0.0112] below zero, 0.0159 at L2 and 0.0218 at L3 which is the figure shipped. The entity-disjoint split alone costs -0.0831, enough to turn a headline gain into a loss.

    For the technical reader. Rungs: L0 is a temporal-only split with the same accounts on both sides, no per-entity aggregates in the baseline, and one pooled embedding per account that can see the scored transaction; L1 adds the entity-disjoint split; L2 makes embeddings as-of, stopping strictly before the scored row; L3 adds per-entity aggregates to the baseline and reproduces the shipped table to a largest difference of 0. L1 to L2 and L2 to L3 score identical rows, so each guard carries a paired interval; L0 to L1 moves three things and carries none.

    Read from

    entity-disjoint split

    No account on both sides of the split

    An entity-disjoint split holds out whole accounts, households, stores or merchants: every row of a held-out entity is test, none is train. A time-only split lets the same account sit on both sides.

    Why it matters here. It is the guard that moved the most. With shared accounts, a pooled embedding works as a key to that account's own category mix; with accounts the model has never seen, the headline gain reverses to a loss.

    • Same accounts on both sides0.0590
    • Whole accounts held out-0.0240

    Next-category top-one delta on the synthetic corpus, rung L0 against rung L1; the L1 interval [-0.0373, -0.0112] stays below zero.

    In this entry. The shipped split holds out 400 of 2,000 accounts. Switching to it sends the delta from 0.0590 to -0.0240. On the real corpus every split was household-disjoint or store-disjoint, and pair retention was merchant-disjoint, 167,245 training merchants against 71,370 test merchants sharing 0.

    For the technical reader. The L0 to L1 step is compound: it replaces row selection, an earlier against later slice of the post-cut pool becoming whole held-out accounts drawn across the window; it changes what the test set samples; and it restricts the rows the frequency encodings are fitted on to accounts not held out. All three hang off one flag, so no paired interval is claimed on that step; the L1 interval carries the point. Bootstrap resampling is clustered on the same entity. Shared-account leakage flatters embeddings, so merchant-only-disjoint runs read as conservative.

    Read from

    confidence interval

    Whether the gain can be distinguished from zero

    A confidence interval is the range of gains the data cannot rule out. Paired means both arms are scored on identical rows and the difference is resampled, so noise common to both arms cancels out.

    Why it matters here. Where one is claimed, the verdict is read off it and not off the point estimate: above zero is a win, below zero a loss, spanning zero a null. Twenty-seven pre-registered comparisons here carry one. A few do not, and say so instead: where the arms score different rows, and where the holdout is a single short time series, a point comparison is stated with no interval claimed.

    • Comparisons carrying an interval27
    • Cleared zero7
    • Landed on the wrong side7
    • Straddle zero13

    Pre-registered comparisons on this page, one row per metric, depth, seed and arm the pre-registrations named, all drawn against one zero line in Section 5 of the entry.

    In this entry. The next-category delta on the synthetic corpus is 0.0218 with interval [0.0103, 0.0343], a win. The store head's accuracy delta in the first seed is +0.1957 with interval [-0.0217, 0.4130] spanning zero, a null on that endpoint, while the same seed clears zero on macro F-one at +0.1912.

    For the technical reader. Intervals are ninety-five percent entity-clustered paired bootstrap: resample held-out entities, accounts, households or stores, with replacement; recompute both arms' metric; take the difference; read the two-and-a-half and ninety-seven-and-a-half percentiles. Clustering on the entity keeps every row of one account together, matching the entity-disjoint split. Where the arms score different rows, as at ladder rung L0 against the rest, no paired interval is claimed; where the holdout is one short time series, as for the corridor blend, a point comparison is stated instead. Deltas ship as obtained.

    Read from

    pre-registration

    The rule was written before the result

    Pre-registration means committing the task, split, metric, decision rule and any weights to the repository before the run, so a result cannot be chosen after it is seen. The commit hash is the proof.

    Why it matters here. It is why the losses are on the page: the corridor blend, the whitespace forward check and the real-corpus replication each shipped as obtained, under a rule fixed before any number existed.

    • Real corpus replication6c9c7a9
    • Corridor blend weights418aef3
    • Whitespace forward checkfef8af5
    • Offers endpoint hierarchy2fd8c23

    Commit hashes, each committed before its run. The offers commit is page plumbing that carried the script, fixed before the full run rather than before every result.

    In this entry. The real-corpus rule reads: "pre-registered in CJ-REPLICATION-PREREG.md: positive only if both seed intervals sit above zero; anything else ships as a null or mixed result; no third seed, no protocol changes after seeing numbers" Under it the household transfer shipped negative in both seeds, the store head positive in both, and the offers head as mixed. The corridor blend weights were fixed at equal before the run at 418aef3, and no other weight was computed.

    For the technical reader. Pre-registration fixes the estimand and the analysis path, which is what makes an interval interpretable: choosing the endpoint, depth or seed after the fact inflates the chance a spurious gain clears zero. The exceptions are labelled. The protection combination was chosen after the two single scores were compared and is a reported finding; the whitespace stratified read is post hoc and counted in no tally; the PR-AUC lead on fraud was reasoned, not pre-registered. Phase 1 hands model risk the leakage rules and the pre-registrations before the build starts.

    Read from

    macro MASE

    Forecast error, scaled to the cheap baseline

    MASE divides a forecast's average absolute error on the held-out months by a naive forecast's error inside the training window, so it carries no unit. Macro MASE averages it across corridors, each weighted equally.

    Why it matters here. A MASE below one does not by itself mean beating the seasonal naive on the holdout, the false rule our own direction check catches. The corridor lane is claimed on neither view.

    • Model alone0.623
    • Seasonal naive0.53
    • Blend, equal weights0.506

    Macro MASE over 13 held-out months across twelve corridors of Singapore inbound arrivals; lower is better, a point comparison with no interval claimed.

    View
    • Model alone0.623
    • Seasonal naive0.53
    • Blend, equal weights0.506
    • Seasonal naive113211.69
    • Blend, equal weights114180.52

    In this entry. Averaging the model with the seasonal naive at weights fixed before the run lowers macro MASE from 0.53 to 0.506. Totalled in arrivals the naive is ahead, 113211.69 against 114180.52, so the blend costs 968.83 more over the holdout and wins 6 of 12 corridors.

    For the technical reader. MASE is scale-free, so corridors of different volume can be averaged; that is also why the corridor-weighted average and the arrival-weighted total can disagree. A MASE below one does not mean beating the seasonal naive on the holdout, since the scaling denominator is in-sample naive error; that false rule in a fact bundle is what produced the 5 reversed narratives the direction check catches. Grouped reconciliation moves total-level MASE from 0.904 to 0.441, still behind the naive's 0.348 on the total.

    Read from

    seasonal naive

    Same month last year, as the forecast

    The seasonal naive forecasts each month as the same month one year earlier. It costs nothing and has no parameters, and any corridor model must outperform it before it is included.

    Why it matters here. The model alone loses to it pooled. Only the pre-registered blend of model and naive beats it, and only on the corridor-weighted average, so the naive stays in the product.

    • Model, before reconciliation0.904
    • Model, reconciled0.441
    • Seasonal naive on the total0.348

    Total-level backtest MASE: grouped reconciliation pools variance across the twelve series, and the seasonal naive computed on the total still stays ahead.

    In this entry. Against a per-corridor seasonal naive over 13 held-out months, the model's MASE of 0.623 sits against 0.53 and loses pooled. Added to the naive at equal weights it helps: the blend beats the naive in 6 of 12 corridors.

    For the technical reader. The forecast for a month is the observed value twelve months earlier. It is the conventional MASE scaling denominator for monthly data and the null control in the pre-registered combination 418aef3, where model and naive are averaged at fixed equal weights with no other weight computed. Attribution explains why it is hard to beat: 82.1 percent of the model's own attribution is recent level, and seasonality is 8.1 plus 3.1 percent, so on this public proxy the forecaster is mostly persistence with a seasonal correction.

    Read from

    surprise score

    The model's surprise at a single transaction

    The label-free surprise score asks the frozen backbone how unlikely each field of a transaction is given the account's own history, then sums the surprises. No fraud label is used anywhere in building it.

    Why it matters here. Authorized scams carry no labels: 81.8% of Singapore scam cases are self-effected transfers. A detector that needs no labels can be stood up on a population where labelled data does not exist.

    • Global rarity alone0.3228
    • Rarity plus surprise0.5989

    Share of fraud positives held by the top one percent of the ranking, 300,000 synthetic transactions; the model contributes as the second term of the score rather than as the first.

    Field set
    • Global rarity alone, top slice0.3228
    • Rarity plus surprise, top slice0.5989
    • Rarity plus surprise, PR-AUC0.3296
    • Rarity plus surprise, top slice0.6783
    • Rarity plus surprise, PR-AUC0.4133

    In this entry. Of 300,000 scored transactions the top one percent is a 3,000 row queue. With the surprise score added to the counting control, that queue holds 566 of the fraud transactions instead of 305, 261 more at identical analyst capacity, on a split whose positive rate is thinned upward.

    For the technical reader. The score is the pseudo-log-likelihood of a masked language model, following the recipe of Salazar and colleagues: one field is masked at a time, the negative log-likelihood of the observed value is taken, and the values are summed across fields. Masking every field at once is out of distribution for a model pretrained at mask probability 0.15. The label file opens only after every score exists and the score matrix hash is recorded. On its own the score is below global rarity on ROC-AUC and equal to it on PR-AUC; when standardized and summed with rarity, unweighted and unfitted, the combination exceeds both. This combination is reported but was not pre-registered.

    Read from

    kill gate

    A gate we might fail, on purpose

    A kill gate is a pre-set test that stops the project when it fails: if embeddings do not beat production per-task features on at least two of three heads offline, with paired intervals, the partner phase never opens.

    Why it matters here. Feasibility is judged on the basis of the plan rather than on an assurance of success. On the one real corpus tested the gate is close: the store head passed in both seeds, household transfer failed in both, and the offers head came back mixed.

    • Store head, sales growth delta0.19120.0064 to 0.3865
    • Household transfer, top-one delta-0.0160-0.0207 to -0.0113
    • Offers head, AUC delta0.0141-0.0094 to 0.0382

    Three heads on the real Complete Journey corpus under the shipped L3 protocol, first pretraining seed; positive only if both seed intervals sit above zero.

    Pretraining seed
    • Store head, sales growth delta0.19120.0064 to 0.3865
    • Household transfer, top-one delta-0.0160-0.0207 to -0.0113
    • Offers head, AUC delta0.0141-0.0094 to 0.0382
    • Store head, sales growth delta0.22420.0449 to 0.4343
    • Household transfer, top-one delta-0.0147-0.0197 to -0.0101
    • Offers head, AUC delta0.02690.0023 to 0.0551

    In this entry. The Phase 1 gate reads: embeddings beat production per-task features on at least two of three heads, offline, with paired CIs. The real-corpus rehearsal gave the following results: store head +0.1912 and +0.2242, both above zero; household transfer -0.0160 and -0.0147, both below; offers +0.0141 null and +0.0269 positive.

    For the technical reader. Phase 1 runs inside the SG Decision Science CoE on data it already holds, months zero to six, after the leakage rules and pre-registrations are handed to model risk; the go or no-go is taken at month six against the printed gate. A failed gate still leaves the protocol ladder behind: it needs no Amex data, and it priced our own claim from 0.0590 to 0.0218. The Phase 2 gate is measured incremental billed business above zero at ninety-five percent confidence, plus signing-list conversion above the current prioritization.

    Read from

    whitespace composite

    Whether the signing score beats plain density

    The whitespace composite is the signing head's score: four weighted channels ranking pseudonymized Singapore merchant buckets. Spearman, a rank correlation, scores how closely that order matches where venues actually appeared afterwards.

    Why it matters here. Its pre-registered forward check lost: plain pre-cutoff venue density predicted where card-accepting venues appeared better than the composite. Density is the bar the composite must clear in Phase 1 against real signings.

    • The shipped composite0.5477
    • Pre-cutoff venue density0.6869
    • Equal weights, same channels0.4498

    Spearman rank correlation with venue formation over the following 24 months, buckets formed before 2024-08-10; the composite minus density difference sits entirely below zero.

    Metric
    • The shipped composite0.5477
    • Pre-cutoff venue density0.6869
    • Equal weights, same channels0.4498
    • The shipped composite0.38
    • Pre-cutoff venue density0.66
    • Equal weights, same channels0.32

    In this entry. Spearman 0.5477 for the composite against 0.6869 for density, difference -0.139 with interval [-0.178, -0.102] entirely below zero. Post hoc, inside 10 density bands, the non-density channels order formation at 0.2383 against density's own 0.1086.

    For the technical reader. Spearman's rho is Pearson correlation computed on ranks: a value of one indicates identical order, and a value of zero indicates no rank correlation. Precision at top fifty is the share of an arm's top fifty buckets that also sit in the outcome's own top fifty. The interval is a cell-clustered bootstrap. The outcome is Foursquare record creation, a proxy for opening, on venues still open at the snapshot, not signings. The backbone column is a per-category constant entering neither the tested composite nor the shipped order, so nothing about the backbone was tested. The stratified re-read is post hoc.

    Read from