AdaptivMapr

Capability · Cascade

Their spreadsheet, in your fields.

A spreadsheet arrives with columns nobody agreed on. You get it in your fields the same morning, with a confidence score and the reason beside every one.

Languages in the hints
5
Per map, flat
$0.001
Rows sent, in schema-only
0
The six cascade layers, cheapest first — most columns resolve before any model is consulted1Statistics2Heuristic3Fuzzy3½Ranker4Semantic5LLM

The job

Someone sent you a file, and it is not in your shape.

Personalnummer. N° de dossier. E-Mail Adresse. The same six fields arrive spelled nine ways, in five languages, abbreviated by whoever built the export — and none of them is called employee_id. Today somebody opens the file, reads the header row, and types the mapping out by hand. Next month the same sender does it again, slightly differently.

What changes

Send the file and get your fields back, each one scored, each one saying which layer decided it. You correct what is wrong; the correction is remembered, so the next file from that sender is already mapped.

What you get

Not a demo. The thing that runs every day.

Layer 1

Statistics

Auto-accept from what your workspace has already confirmed. When the evidence is overwhelming, the column is mapped before a single string is compared.

HowAUTO_ACCEPT_RULES in lib/matcher.ts is exactly {minN:100, minRatio:0.95} and {minN:20, minRatio:1.00}. Both are production-tested; below them the column falls through instead of guessing.

Layer 2

Heuristic

A normalized compare against every name a field answers to — its column, its label and each of its multilingual hints. Adding hints is the highest-leverage way to catch new vocabulary.

Hownormalize() strips accents, punctuation and whitespace, then the header is compared against column, label and every hints entry (DE / FR / IT / EN / ES). An exact hit after normalization scores 1.00, a containment hit 0.85 — the accept floor.

Layer 3

Fuzzy

Typos, abbreviations and token-order drift. “DOB”, “Date of Birth” and “birth_date” all land on the same field, with no network call and no key.

HowfuzzyScore() in lib/fuzzy.ts — token-set ratio plus Levenshtein over the normalized strings, auto-accepting at FUZZY_AUTO_ACCEPT = 0.80. Pure compute, still free.

Layer 4

Semantic

Embedding similarity between the header and the field’s meaning, for the columns the string layers cannot reach. Cheap, cached, and optional.

HowCosine over embeddings of the header versus label + hints, accepting at SEMANTIC_THRESHOLD = 0.78. It is config-gated on MAPR_EMBEDDING_API_KEY — absent, the layer is skipped and the cascade runs on without it.

Layer 5

LLM, batched once

Everything still unresolved is settled in ONE collision-aware call, constrained to your allowed column set and told which fields are already claimed. The only metered step in the cascade.

HowllmMatchHeadersBatch() sends every leftover header in a single request rather than one call per column, and it is skipped entirely for any column an earlier layer resolved — that short-circuit is the biggest cost lever in the system. Gated on MAPR_LLM_API_KEY.

Guarantee

Contests are settled by quality, never by column order

A target field, once assigned, cannot be reused by another header in the same file. Two source columns can never map to one target — and reversing your columns cannot change the answer.

Howmatch() proposes across layers 1–3, arbitrates once on (layer rank, then confidence, then column index), then runs layer 4 on what is left. Claim-on-accept used to give email to a 0.85 hit on e_mail_alt and leave the exact 1.00 match unmapped; the cascade is now order-independent end to end.

How it works

The six layers, and what each one costs

Six layers run in strict cost order and short-circuit: the moment a layer resolves a column with enough confidence the target field is claimed, and no later — more expensive — layer sees it. Layers 1–3 are free, deterministic and run in-process. Layer 3½ is a cross-encoder we fine-tuned and host ourselves, priced per column it resolves. Layer 4 is cheap and cached. Layer 5, the third-party LLM, only ever sees what none of them could place.

  1. 01

    Parse, then clamp

    The file becomes headers plus rows inside the Worker. In schema-only that grid is cut to three rows of ≤80 characters by clampForSchemaOnly() before any layer sees it.

  2. 02

    Propose, do not claim

    Layers 1–3 each score every header in pure compute. A hit is a bid, not an assignment — nothing is claimed while the walk is still running.

  3. 03

    Arbitrate once

    The bids are sorted by layer rank, then confidence, then column index, and fields are handed out top-down. A claimed field leaves the pool for everyone else.

  4. 04

    Fall through

    Only what is still unplaced reaches the cached embedding layer, and only what survives that reaches the single batched AI call. Every earlier resolution is a call not made.

  5. 05

    Confirm, and it gets cheaper

    A commit records the per-pair statistics and the whole confirmed layout, so the same file next week lands on layer 1 — or on a fingerprint lookup with no cascade at all.

LayerWhat it comparesAuto-accepts atCost
1 · statisticsConfirmed header → field counts for your workspace≥100 @ 95% · ≥20 @ 100%Free · deterministic
2 · heuristicNormalized header vs column, label and every hint (DE/FR/IT/EN/ES)≥ 0.85Free · deterministic
3 · fuzzyToken-set ratio + Levenshtein over normalized strings≥ 0.80Free · pure compute
3½ · rankerA fine-tuned cross-encoder scores every unresolved column against every field + “none”, then assigns one-to-onep ≥ τ fitted at 99.5% precisionPer column resolved · our own GPU, in Switzerland
4 · semanticEmbedding cosine, header vs label + hints≥ 0.78Cheap, cached · off without an embedding key
5 · aiOne batched call over everything still unresolvedModel pick, constrained to the unclaimed column setMetered · the only paid layer
Layers run in order and short-circuit. Layers 4 and 5 are config-gated and fail soft to OFF, so an unconfigured deployment still maps on the deterministic layers rather than erroring.

Try it

Four German headers, no AI, no key

Send headers plus up to three sample rows. The response names the layer that resolved each column, so you can see for yourself how little of the work needed a model.

curl
curl https://api.adaptivmapr.com/v1/match \
  -H "Content-Type: application/json" \
  -d '{
    "template_id": "patient_demographics_v1",
    "headers": ["Nachname", "Vorname", "Geburtsdatum", "E-Mail"],
    "sample_rows": [["Lovelace", "Ada", "10.12.1815", "ada@example.org"]]
  }'
response
{
  "template_id": "patient_demographics_v1",
  "matches": [
    { "source_col": "Nachname",     "target_field": "last_name",     "source": "heuristic", "confidence": 1 },
    { "source_col": "Vorname",      "target_field": "first_name",    "source": "heuristic", "confidence": 1 },
    { "source_col": "Geburtsdatum", "target_field": "date_of_birth", "source": "heuristic", "confidence": 1 },
    { "source_col": "E-Mail",       "target_field": "email",         "source": "heuristic", "confidence": 1 }
  ],
  "auto_accept_threshold": [{ "minN": 100, "minRatio": 0.95 }, { "minN": 20, "minRatio": 1 }],
  "cascade_layers": ["statistics", "heuristic", "fuzzy", "ranker", "semantic", "ai"],
  "unmapped": []
}
→ 4 of 4 resolved on layer 2 · 0 headers reached the LLM · no API key sent
  • POST /v1/match is public and needs no key — 100 requests an hour per IP, with sample_rows hard-clamped to 3 at the HTTP edge. It is the same endpoint the adaptivmapr.match_headers MCP tool calls, which sends ADAPTIVMAPR_API_KEY whenever one is configured.
  • The metered layers are unreachable without a key: an anonymous request runs the deterministic layers 1–3 (plus layer 4 where embeddings are configured) and stops, and costs nothing. Present a Bearer key and the mapping ranker and AI fallback switch on, and the map is billed to that key’s workspace wallet — so a scraped endpoint can never run up an LLM bill.
  • The response returns cascade_layers and auto_accept_threshold verbatim, so an importer can render the cascade-efficiency bar without hardcoding a single threshold.

The surface

Every route this page actually has.

The cascade is reachable two ways: one stateless call for a header row you already have, or the stored-upload pipeline when a human reviews the mapping before it commits.

  • POST/v1/matchHeaders + ≤3 sample rows in, per-column matches with the resolving layer out. 100 requests an hour per IP, 200 headers per call.no key
  • POST/v1/uploadsMultipart file up to 10 MiB, parsed and held in Cloudflare KV under the workspace retention window (24h default, up to 30d); the response carries expires_at.transform scope
  • POST/v1/uploads/{id}/matchRun the cascade over a stored upload and get the per-column proposal back.read scope
  • POST/v1/uploads/{id}/match/streamThe same run as Server-Sent Events, one frame per layer — what the dashboard renders live.read scope
  • PATCH/v1/uploads/{id}/mappingsConfirm or override the mapping. Returns requires_hitl on a medium- or high-risk template.commit scope
  • POST/v1/uploads/{id}/commitValidate and emit the confirmed rows — and record the layout so the next identical file is a lookup.commit scope
  • GET/v1/templatesThe catalogue the engine serves — 33 templates across 7 packs, with every field, hint and validator.no key

Limits & failure modes

What it refuses to do — and the code it says it with.

A gate that fails open is worse than one that is missing. Every refusal below is a specific status and code you can branch on, not a generic 400.

400 too_many_headers

More than 200 headers in one call.

A hard cap on both /v1/match and the layout routes. Split the sheet, or map it through an upload.

404 template_unknown

The template_id is not one the engine serves.

List them with GET /v1/templates, or define your own with POST /v1/schemas and pass that id instead.

429 rate_limited

100 requests an hour per IP on the public route.

The shared Postgres counter is authoritative; the per-isolate counter only ever denies earlier. Authenticate for the workspace-scoped limits.

Layers 4 & 5 off

No embedding key or no LLM key is configured on the deployment.

Not an error. The cascade runs on the deterministic layers alone and returns fewer matches — quietly weaker, never quietly broken.

No AI without a key

An anonymous call to the public /v1/match.

Layers 1–4 run and the call stops. The metered layer is unreachable without a Bearer key, so a scraped endpoint cannot run up an AI bill.

402 phi_gateway_required

Full-data mode asked for without the PHI entitlement.

The request is refused rather than silently downgraded to schema-only, because a downgrade would misreport what left your system.

What it costs

One prepaid wallet, drawn down per call.

No free tier. Top up from $10 — the balance is shared across the phi-cloud suite — and every operation draws it down at the rate below. An optional $6/mo plan grants a credit that resets each billing cycle instead; overage falls back to the wallet.

Every map

$0.001

The flat per-map fee, drawn from the prepaid wallet — including a map that finished entirely on the free deterministic layers.

Layers 1–4

No token cost

Statistics, heuristics and fuzzy are pure compute. Embeddings are cheap and cached, and are skipped entirely when the key is unset.

Layer 5, when reached

cost × 2

The phi-cloud tokens actually consumed, at provider cost × 2 — or × 0.5 with your own LLM key. One batched call, not one per column.

A PHI/enterprise-routed run multiplies the whole charge by 1.2 (+20%), flat fee included — and only when the run genuinely got that routing. PHI stays locked until the workspace accepts the BAA in Settings → Security & Data. Full pricing

Questions

The ones asked before signing.

What it costs, what leaves your network, and the claims we will not make.

Does my data get sent to an LLM?
In schema-only mode, never the rows — only your headers and up to three sample rows, each clamped to 80 characters. Layers 1–3 resolve most columns with no AI at all, and the batched layer-5 call only fires for what nothing else could place. Row-level AI is a separate, opt-in full-data mode.
What makes the statistics layer safe to auto-accept?
It accepts only on overwhelming evidence: at least 100 prior confirmations agreeing 95% of the time, or at least 20 agreeing unanimously. Those two rules are production-tested and deliberately not loosened; below them the column falls through to the next layer rather than guessing.
Can two of my columns map to the same field?
No — it is forbidden by construction, because you cannot emit two source columns into one target. Once a field is claimed it is skipped for every other header in the same call, and a contested field goes to the strongest bid by layer rank, then confidence. Reordering your columns cannot change the result.
What file formats can I map?
CSV, TSV, Excel (multi-sheet), JSON, XML, Parquet, Word/PowerPoint tables and SQL dumps (CREATE TABLE + INSERT). PDFs, images, audio and web pages come in through the convert route. The cascade works on headers and sample values, so once a source is parsed its original format is irrelevant.
How do I teach it our in-house vocabulary?
Add hints to the field — the heuristic layer compares the header against every hint, in five languages, and a hint costs nothing at runtime. Confirmed mappings also feed the statistics layer, so the more you commit, the more resolves on layer 1 next time, with no AI.

One key, from a messy file to your schema.

Top up a $10 prepaid wallet and start mapping. In schema-only mode only your headers and three clamped rows ever leave you.

Schema mapping cascade — AdaptivMapr — AdaptivMapr