AdaptivMapr

HR & People pack

Payroll CSV import API

Import payroll runs from any CSV — gross salary, pay period, payout IBAN. PII-aware: raw salary rows never leave, only clamped samples.

30-second curl
curl -X POST https://api.adaptivmapr.com/v1/uploads \
  -H "Authorization: Bearer $ADAPTIVMAPR_API_KEY" \
  -F "template=payroll_v1" \
  -F "file=@your_data.csv"
→ 6 canonical fields · 1 validated · high risk

Canonical columns

The whole schema, printed as it ships.

Every canonical column, the type each row carries, whether it is required, the field-level validators that fire on commit, and the multilingual header hints the cascade resolves against. This is the shipped definition, not a summary of it.

payroll_v1
fields
6
required
3
validated
1
hints
28
Canonical columnTypeRequiredValidatorsHeader hints the cascade matches
employee_idstringyes—personalnummermatriculematricolaemployee idnúmero de empleado
gross_salarynumberyes—bruttolohnsalaire brutstipendio lordogross salarysalario bruto
currencystringyes—währungdevisevalutamoneda
pay_perioddate——lohnperiodepériode de paieperiodo di pagapay periodperíodo de pago
ibanstring—ibanibancomptekontobank account
tax_codestring——steuercodecode fiscalcodice fiscaletax codecódigo fiscal

Read the same definition as JSON at GET /v1/templates/payroll_v1. A hint match resolves on layer 2 — no LLM call, no token spend, just the flat per-map fee. Hover a validator id to see what it checks.

  • 6 canonical fields
  • 3 required
  • 1 validated
  • 28 header hints, 5 languages

Why it exists

Written for the file you actually receive.

The Payroll template is the canonical schema for a pay run — the file a payroll-provider export, a salary-register dump, or a finance reconciliation extract reduces to. Each row carries an employee_id (required), a gross_salary (required), a currency (required), a pay_period parsed across locales, an optional payout IBAN, and an optional tax_code. Payroll and finance teams reach for it when migrating between payroll providers, when backfilling a warehouse with historical pay runs for cost analysis, and when reconciling a payroll register against the general ledger. It is `high` risk because salary plus bank detail is among the most sensitive PII an organisation holds. Schema-only mode is therefore the default ingress: raw payroll rows never leave the customer; only headers and a handful of clamped sample cells are processed to decide the mapping.

employee_id, gross_salary, and currency are all required — a pay row without an amount or a currency is meaningless. gross_salary is number-typed and locale-tolerant (US "5,000.00", EU "5.000,00", Swiss "5'000.00"); pay_period auto-detects ISO, US, and EU date formats. IBAN runs mod-97 with a country restriction (CH, LI, DE, FR, IT, ES) so a transposed digit blocks a mispayment before it happens. tax_code is a free string because withholding-code formats vary per jurisdiction. Hints cover DE / FR / IT / ES / EN so a multilingual payroll export does not escalate to the LLM.

Migration scenarios & the foreign headers they ship

Migration scenarios for the Payroll template: porting pay-run history between providers (ADP → Deel, Swissdec-based systems → a new bureau) so the new platform has a complete register, backfilling a finance warehouse with historical runs for labour-cost analysis, reconciling the payroll register against the general ledger each cycle, and consolidating multi-entity payroll after a merger. Foreign headers we routinely see: "Personalnummer / Matricule / Matricola / Número de empleado / Bruttolohn / Salaire brut / Stipendio lordo / Salario bruto / Währung / Devise / Valuta / Moneda / Lohnperiode / Période de paie / Periodo di paga / Período de pago / IBAN / Compte / Konto / Steuercode / Code fiscal / Codice fiscale / Tax code / Código fiscal". The cascade resolves every one through the registered hints — no LLM call, no salary figures in any prompt.

The cascade

Six layers, and the cheapest one wins.

Layers run in order and stop the moment a column resolves. That is the single biggest cost lever in the system: a column caught on layer 2 never reaches the metered layer 5.

  1. L1Statisticsno LLM

    Auto-accepts a header that past confirmations already resolved the same way, at {minN:100, minRatio:0.95} or {minN:20, minRatio:1.00}.

  2. L2Heuristicno LLM

    Normalises accents, punctuation and whitespace, then compares against the column name, the label, and every registered hint (DE / FR / IT / EN / ES).

  3. L3Fuzzyno LLM

    Token-set ratio plus Levenshtein over the normalised strings. Auto-accepts at 0.80 — it absorbs typos and reordered words.

  4. L4Semanticcheap, cached

    Embedding cosine between the header and the field’s label + hints. Catches the long tail of paraphrases.

  5. L5LLMmetered

    Everything still unresolved goes up in ONE batched, collision-aware call, constrained to this template’s column set so it cannot invent a field.

Try it

One template id, two ways in.

REST for your import pipeline, MCP for your editor. Both run the same cascade and both honour the same schema-only clamp.

REST · POST /v1/uploads

Name the template; the cascade picks up the rest. The canonical definition is read-only at GET /v1/templates/payroll_v1.

bash
curl -X POST https://api.adaptivmapr.com/v1/uploads \
  -H "Authorization: Bearer $ADAPTIVMAPR_API_KEY" \
  -F "template=payroll_v1" \
  -F "file=@your_data.csv"
→ upload created · mappings ready · confirm before commit

MCP · Cursor / Claude Desktop

Drop AdaptivMapr into your editor and call the same cascade as a tool. Schema-only calls leave only column names and up to three clamped sample rows.

mcp
// In Cursor or Claude Desktop with the AdaptivMapr MCP server installed:
adaptivmapr.match_headers({
  template_id: "payroll_v1",
  headers: ["employee_id", "gross_salary", "currency", "pay_period"]
})
schema-only · headers and ≤3 rows, 80 chars each
MCP install instructions
high-risk template

The mappings response comes back flagged. PATCH /uploads/:id/mappings returns requires_hitl: true and hitl_status: "pending_review" so you can hold the commit in your own workflow — the flag is a signal, not a queue we run. Schema-only mode (headers plus at most three sample rows, each clamped to 80 characters) is a data-minimization mode enforced at the HTTP edge. Full-data mode routes the metered layer-5 call to phi-cloud in-region under a BAA, costs 20% more on the whole map charge, and stays locked until the workspace accepts the BAA/NDA in Settings → Security & Data.

How the wallet is charged

Questions

Payroll CSV import — FAQ

Does salary data leave our environment during mapping?
No. Schema-only mode is the default: only headers and three clamped sample rows (≤80 chars each) are processed to decide the mapping. Full payroll rows never touch our infrastructure unless you explicitly opt into full-data mode under an active subscription.
Is gross_salary the only amount field?
Yes — the canonical row is deliberately minimal (gross_salary plus pay_period). Net pay, deductions, and employer contributions vary too much per jurisdiction to enum-bound; add them via a workspace fork if your downstream needs the breakdown.
How does the IBAN country restriction work?
The default allow-list is CH / LI / DE / FR / IT / ES — the cluster most European payroll runs pay into. Fork the template and broaden the `countries` array for cross-border payroll; the validator otherwise rejects out-of-list IBANs at import.
Why does pay_period use a date rather than a string like "2026-06"?
The auto date parser accepts month-level and full-date inputs and normalises them, so "Juni 2026", "06/2026", and "2026-06-30" all resolve. Storing a real date lets downstream cost analysis group and sort correctly.

Map payroll in production — without shipping raw records.

Schema-only mode leaves only headers and a handful of clamped samples. Add full-data when you need row-level AI, routed in-region under a BAA.