Capability · Cascade
Their spreadsheet, in your fields.
A spreadsheet arrives with columns nobody agreed on. You get it in your fields the same morning, with a confidence score and the reason beside every one.
- Languages in the hints
- 5
- Per map, flat
- $0.001
- Rows sent, in schema-only
- 0
The job
Someone sent you a file, and it is not in your shape.
employee_id. Today somebody opens the file, reads the header row, and types the mapping out by hand. Next month the same sender does it again, slightly differently.What changes
What you get
Not a demo. The thing that runs every day.
Layer 1
Statistics
Auto-accept from what your workspace has already confirmed. When the evidence is overwhelming, the column is mapped before a single string is compared.
HowAUTO_ACCEPT_RULES in lib/matcher.ts is exactly {minN:100, minRatio:0.95} and {minN:20, minRatio:1.00}. Both are production-tested; below them the column falls through instead of guessing.
Layer 2
Heuristic
A normalized compare against every name a field answers to — its column, its label and each of its multilingual hints. Adding hints is the highest-leverage way to catch new vocabulary.
Hownormalize() strips accents, punctuation and whitespace, then the header is compared against column, label and every hints entry (DE / FR / IT / EN / ES). An exact hit after normalization scores 1.00, a containment hit 0.85 — the accept floor.
Layer 3
Fuzzy
Typos, abbreviations and token-order drift. “DOB”, “Date of Birth” and “birth_date” all land on the same field, with no network call and no key.
HowfuzzyScore() in lib/fuzzy.ts — token-set ratio plus Levenshtein over the normalized strings, auto-accepting at FUZZY_AUTO_ACCEPT = 0.80. Pure compute, still free.
Layer 4
Semantic
Embedding similarity between the header and the field’s meaning, for the columns the string layers cannot reach. Cheap, cached, and optional.
HowCosine over embeddings of the header versus label + hints, accepting at SEMANTIC_THRESHOLD = 0.78. It is config-gated on MAPR_EMBEDDING_API_KEY — absent, the layer is skipped and the cascade runs on without it.
Layer 5
LLM, batched once
Everything still unresolved is settled in ONE collision-aware call, constrained to your allowed column set and told which fields are already claimed. The only metered step in the cascade.
HowllmMatchHeadersBatch() sends every leftover header in a single request rather than one call per column, and it is skipped entirely for any column an earlier layer resolved — that short-circuit is the biggest cost lever in the system. Gated on MAPR_LLM_API_KEY.
Guarantee
Contests are settled by quality, never by column order
A target field, once assigned, cannot be reused by another header in the same file. Two source columns can never map to one target — and reversing your columns cannot change the answer.
Howmatch() proposes across layers 1–3, arbitrates once on (layer rank, then confidence, then column index), then runs layer 4 on what is left. Claim-on-accept used to give email to a 0.85 hit on e_mail_alt and leave the exact 1.00 match unmapped; the cascade is now order-independent end to end.
How it works
The six layers, and what each one costs
Six layers run in strict cost order and short-circuit: the moment a layer resolves a column with enough confidence the target field is claimed, and no later — more expensive — layer sees it. Layers 1–3 are free, deterministic and run in-process. Layer 3½ is a cross-encoder we fine-tuned and host ourselves, priced per column it resolves. Layer 4 is cheap and cached. Layer 5, the third-party LLM, only ever sees what none of them could place.
- 01
Parse, then clamp
The file becomes headers plus rows inside the Worker. In schema-only that grid is cut to three rows of ≤80 characters by
clampForSchemaOnly()before any layer sees it. - 02
Propose, do not claim
Layers 1–3 each score every header in pure compute. A hit is a bid, not an assignment — nothing is claimed while the walk is still running.
- 03
Arbitrate once
The bids are sorted by layer rank, then confidence, then column index, and fields are handed out top-down. A claimed field leaves the pool for everyone else.
- 04
Fall through
Only what is still unplaced reaches the cached embedding layer, and only what survives that reaches the single batched AI call. Every earlier resolution is a call not made.
- 05
Confirm, and it gets cheaper
A commit records the per-pair statistics and the whole confirmed layout, so the same file next week lands on layer 1 — or on a fingerprint lookup with no cascade at all.
| Layer | What it compares | Auto-accepts at | Cost |
|---|---|---|---|
| 1 · statistics | Confirmed header → field counts for your workspace | ≥100 @ 95% · ≥20 @ 100% | Free · deterministic |
| 2 · heuristic | Normalized header vs column, label and every hint (DE/FR/IT/EN/ES) | ≥ 0.85 | Free · deterministic |
| 3 · fuzzy | Token-set ratio + Levenshtein over normalized strings | ≥ 0.80 | Free · pure compute |
| 3½ · ranker | A fine-tuned cross-encoder scores every unresolved column against every field + “none”, then assigns one-to-one | p ≥ τ fitted at 99.5% precision | Per column resolved · our own GPU, in Switzerland |
| 4 · semantic | Embedding cosine, header vs label + hints | ≥ 0.78 | Cheap, cached · off without an embedding key |
| 5 · ai | One batched call over everything still unresolved | Model pick, constrained to the unclaimed column set | Metered · the only paid layer |
Try it
Four German headers, no AI, no key
Send headers plus up to three sample rows. The response names the layer that resolved each column, so you can see for yourself how little of the work needed a model.
curl https://api.adaptivmapr.com/v1/match \
-H "Content-Type: application/json" \
-d '{
"template_id": "patient_demographics_v1",
"headers": ["Nachname", "Vorname", "Geburtsdatum", "E-Mail"],
"sample_rows": [["Lovelace", "Ada", "10.12.1815", "ada@example.org"]]
}'{
"template_id": "patient_demographics_v1",
"matches": [
{ "source_col": "Nachname", "target_field": "last_name", "source": "heuristic", "confidence": 1 },
{ "source_col": "Vorname", "target_field": "first_name", "source": "heuristic", "confidence": 1 },
{ "source_col": "Geburtsdatum", "target_field": "date_of_birth", "source": "heuristic", "confidence": 1 },
{ "source_col": "E-Mail", "target_field": "email", "source": "heuristic", "confidence": 1 }
],
"auto_accept_threshold": [{ "minN": 100, "minRatio": 0.95 }, { "minN": 20, "minRatio": 1 }],
"cascade_layers": ["statistics", "heuristic", "fuzzy", "ranker", "semantic", "ai"],
"unmapped": []
}POST /v1/matchis public and needs no key — 100 requests an hour per IP, withsample_rowshard-clamped to 3 at the HTTP edge. It is the same endpoint theadaptivmapr.match_headersMCP tool calls, which sendsADAPTIVMAPR_API_KEYwhenever one is configured.- The metered layers are unreachable without a key: an anonymous request runs the deterministic layers 1–3 (plus layer 4 where embeddings are configured) and stops, and costs nothing. Present a Bearer key and the mapping ranker and AI fallback switch on, and the map is billed to that key’s workspace wallet — so a scraped endpoint can never run up an LLM bill.
- The response returns
cascade_layersandauto_accept_thresholdverbatim, so an importer can render the cascade-efficiency bar without hardcoding a single threshold.
The surface
Every route this page actually has.
The cascade is reachable two ways: one stateless call for a header row you already have, or the stored-upload pipeline when a human reviews the mapping before it commits.
- POST
/v1/matchHeaders + ≤3 sample rows in, per-column matches with the resolving layer out. 100 requests an hour per IP, 200 headers per call.no key - POST
/v1/uploadsMultipart file up to 10 MiB, parsed and held in Cloudflare KV under the workspace retention window (24h default, up to 30d); the response carries expires_at.transform scope - POST
/v1/uploads/{id}/matchRun the cascade over a stored upload and get the per-column proposal back.read scope - POST
/v1/uploads/{id}/match/streamThe same run as Server-Sent Events, one frame per layer — what the dashboard renders live.read scope - PATCH
/v1/uploads/{id}/mappingsConfirm or override the mapping. Returns requires_hitl on a medium- or high-risk template.commit scope - POST
/v1/uploads/{id}/commitValidate and emit the confirmed rows — and record the layout so the next identical file is a lookup.commit scope - GET
/v1/templatesThe catalogue the engine serves — 33 templates across 7 packs, with every field, hint and validator.no key
Limits & failure modes
What it refuses to do — and the code it says it with.
A gate that fails open is worse than one that is missing. Every refusal below is a specific status and code you can branch on, not a generic 400.
400 too_many_headersMore than 200 headers in one call.
A hard cap on both /v1/match and the layout routes. Split the sheet, or map it through an upload.
404 template_unknownThe template_id is not one the engine serves.
List them with GET /v1/templates, or define your own with POST /v1/schemas and pass that id instead.
429 rate_limited100 requests an hour per IP on the public route.
The shared Postgres counter is authoritative; the per-isolate counter only ever denies earlier. Authenticate for the workspace-scoped limits.
Layers 4 & 5 offNo embedding key or no LLM key is configured on the deployment.
Not an error. The cascade runs on the deterministic layers alone and returns fewer matches — quietly weaker, never quietly broken.
No AI without a keyAn anonymous call to the public /v1/match.
Layers 1–4 run and the call stops. The metered layer is unreachable without a Bearer key, so a scraped endpoint cannot run up an AI bill.
402 phi_gateway_requiredFull-data mode asked for without the PHI entitlement.
The request is refused rather than silently downgraded to schema-only, because a downgrade would misreport what left your system.
What it costs
One prepaid wallet, drawn down per call.
No free tier. Top up from $10 — the balance is shared across the phi-cloud suite — and every operation draws it down at the rate below. An optional $6/mo plan grants a credit that resets each billing cycle instead; overage falls back to the wallet.
Every map
$0.001
The flat per-map fee, drawn from the prepaid wallet — including a map that finished entirely on the free deterministic layers.
Layers 1–4
No token cost
Statistics, heuristics and fuzzy are pure compute. Embeddings are cheap and cached, and are skipped entirely when the key is unset.
Layer 5, when reached
cost × 2
The phi-cloud tokens actually consumed, at provider cost × 2 — or × 0.5 with your own LLM key. One batched call, not one per column.
A PHI/enterprise-routed run multiplies the whole charge by 1.2 (+20%), flat fee included — and only when the run genuinely got that routing. PHI stays locked until the workspace accepts the BAA in Settings → Security & Data. Full pricing
Questions
The ones asked before signing.
What it costs, what leaves your network, and the claims we will not make.
Does my data get sent to an LLM?
What makes the statistics layer safe to auto-accept?
Can two of my columns map to the same field?
What file formats can I map?
How do I teach it our in-house vocabulary?
Keep reading
The rest of the same engine.
One key, from a messy file to your schema.
Top up a $10 prepaid wallet and start mapping. In schema-only mode only your headers and three clamped rows ever leave you.