AT0M by Pienomial

// a decision model you own, not an API you rent · v0 to evaluate, v1 to deploy

AT0M: a private, System One Decision AI that runs locally.

Give it a state and a set of questions — choice, score, or yes-no. It returns one probability per option, in a single forward pass.

No tokens and no free text, so it cannot answer outside your candidate set. It runs as a single local executable, which keeps inference inside your perimeter.

Explore the evidence
Accuracy against cost per workflow: AT0M at 78.9%, alone at the low-cost edge of the chart; every other system sits further right and lower.
Accuracy against what a decision costs. A workflow is five typed decisions about one record. Competitor accuracies and prices are as published; ours is measured over all 2,000 test decisions and priced from a laptop’s measured throughput plus its amortised system cost. The full basis is in the report.
78.9%
typed-decisions, 2,000 decisions
16 ms
response on a T4
5.6%
calibration error, nothing fitted

// how you deploy it

The executable is the product.

AT0M will ship as a single self-contained Rust binary — no Python at inference, no service in the path, nothing to rent. That is the deployment, and it will be the only thing that ends up in your perimeter. The hosted endpoint exists for one reason: so you can check every number on this page before you download anything.

Evaluate first v0 · hosted

One POST carries a record and as many questions as you like, answered together in the same pass. No SDK, no client library. Currently an evaluation build; the release will improve on it.

POST /decide/v0 — evaluation endpoint$ curl -s -X POST \
  https://at0m.pienomial.com/decide/v0 \
  -H 'Content-Type: application/json' \
  -d '{"state": "Our production API has
       been down since 6 AM.",
       "questions": {
         "queue": {"type":"choice",
                   "criteria": {...}},
         "urgent": {"type":"noul",
                    "instructions":"Needs action
                     today."}}}'
→ queue: infrastructure (p=0.999)
→ urgent: true (p=0.892)

To be released v1

Run a file of decisions, or serve the same API on your own machine. Self-hosted end to end — this is what a production deployment looks like.

./at0m — your machine, your perimeter# score a file of decisions, no server
$ ./at0m decide --in decisions.json

# or serve the same API yourself
$ ./at0m serve --port 8080 --device metal
listening on :8080 · metal · 54 dec/s
$ curl -s localhost:8080/decide \
    -d @decisions.json

→ no network, no provider, no per-decision fee

v0 is the evaluation endpoint; v1 is the executable release — the hosted service is a way to try the model, not a way to run it. Every figure on this page was measured against v0, the evaluation build, so they are the floor rather than the ceiling: v1 is expected to be more accurate, and we will republish every number on this page when it ships.

// how it works

What the model actually does.

  • Non-generative.It scores your supplied options, it doesn't write text. The output is a probability per option — so a malformed or out-of-list answer is structurally impossible, not just discouraged.
  • Calibrated by construction.Probabilities mean what they say without post-hoc temperature fitting: 5.6% ECE on typed-decisions, 3.7% on its choices, 3.7% on Enron spam. Set an escalation threshold, auto-act above it, route the rest to a human.
  • Local inference.Runs as a standalone Rust binary on CPU, Apple Metal or CUDA. Inference stays in your perimeter — no required third-party API in the path, so no per-decision network round-trip.
  • Trainable on your own decisions.Ten thousand of your own records, five questions each, is about two hours on a laptop. No datacentre GPU, no cluster time.

// the numbers, with their measurement conditions

Benchmark results.

One configuration across every public test set — no per-dataset tuning or prompt tweaks. 38,862 decisions, scored in one sweep.

Two things we'd flag before you read the numbers. The public typed-decisions benchmark's labels are noisy — they self-agree ~65% of the time against a ~73% teacher-model ceiling — so we don't read small margins near the top as meaningful, ours included. And every AT0M figure here came from a single sweep through one harness, while the competitor figures come from their own write-ups under their own conditions — so read the table as "each number with its source," not a single scoreboard.

BenchmarkAT0MLayaJevCondition
typed-decisions (2,000)0.7890.7660.740ahead of Laya by 2.3pt, marginally outside the ±1.8pt sampling error
calibration (ECE)0.0560.081*0.121–0.161*Laya after fitting a temperature, and from its specialist checkpoint; Jev is the spread across four independent studies — TypeSafe publishes none. Jev and AT0M are general models, one configuration for every row; AT0M fits no temperature
DAIR Emotion0.9210.5950.480bare labels
Banking77 (77 intents)0.8530.4250.870Jev on 72 labels — not the same label space. Ours moves between 85% and 88% with how the intents are worded, so read the margin here as a band, not a point
AG News0.8960.9500.910the one clean multi-way we trail
CLINC150 + out-of-scope0.795——151 candidates, scored in full
Enron spam / held-out phishing0.955 / 0.805——phishing never trained on — the honest transfer test
toxicity guardrail (F1)0.626——7.1% positive — answering "no" to everything scores 92.9% accuracy and catches nothing

Laya figures, and Jev's on the classification sets, are as reported in the Laya write-up. Jev's typed-decisions figure is DecisionEval's independent rerun. On calibration, TypeSafe publishes no figure at all, so Jev's range is assembled from four independent studies, none temperature-fitted.

// what we'd actually defend

The parts that don't depend on the benchmark.

The accuracy number is near the top but sits inside admitted label noise, so we don't hang the pitch on it. The claims we'd stand behind under scrutiny are architectural:

  • Bounded output — can't answer off your candidate set. (A generative model can; that's the failure mode this design removes.)
  • Calibration without temperature fitting — thresholds you can trust straight out of the model.
  • Local inference — the model is an artifact you run in your perimeter, not an API you rent, and there's no per-decision round-trip.
  • Cost structure, not a discount — nothing is generated, so there are no output tokens to pay for. The marginal cost of the next decision is electricity.
  • Deterministic — send the same request twice and you get the same probabilities, to the last digit. Two full re-runs of every set on this page, 38,862 decisions each, came back identical on every metric: accuracy, calibration error, mean confidence, F1. The headline survives load as well — typed-decisions scores 0.7885 whether you send it one request at a time or five hundred at once. The one thing that moves is the split underneath that total: batched beside different neighbours, an item sitting on a decision boundary can land the other way, so per-primitive figures wander by a decision or two in 2,000 while the total holds. Re-run it and you get our numbers.
  • Reproducibility — the endpoint that produced these numbers is the same one you can point your own eval at.

And the parts we wouldn't: it compares a value to a threshold well, but not one field of a record against another, so compute those predicates in code. Calibration decays off-distribution — 0.037 on spam it has seen, 0.171 on phishing it hasn't. Both are measured below.

// from idea to deployment

How can I build my own?

An AT0M app is a record, a set of typed questions, and the code that acts on the probabilities. Here's how teams get from one to the other.

You don't start from a blank page.Every use case in the showcase is a working set of questions — across banking, insurance, pharma, government and everyday tasks. Pick the one closest to your problem, then adapt it.

Browse use cases
  1. Find the atomic decision

    Look for a judgment your team makes hundreds of times a day, sitting between an incoming record and an action: route this request, flag this report, check this file. If you can list every acceptable answer in advance, AT0M can make the call. If the answer has to be written, it can't.

  2. Define the state

    The state is the text AT0M reads: a message, a form, a call note, a summary of sensor readings. Comparisons between fields — a lab value against baseline, an amount against a limit — are computed in your code and passed in as facts. The evaluation endpoint accepts requests up to 4,096 tokens.

  3. Write the questions

    Give each question an id your code will use, and pick one of three types. Ask as many as you need; they're answered together in one pass.

    choicePick one option. Describe each one, and include an option for "none of these".
    scoreRate on a scale, such as urgency or severity from 1 to 5.
    noulA yes-or-no question, answered with a probability.
  4. Try it on the hosted endpoint

    Write your questions and run them against the evaluation endpoint. One POST, no SDK, no client library — the API call updates as you edit, ready to copy. Every use case in the showcase has a “Use as template” button that loads it here.

    This is the same endpoint that produced every number on this page. Point your own eval at it, or re-run the public test splits, and you'll get our figures.

  5. Test it on your own records

    Label a few hundred real records and run them through. Check accuracy and calibration for each question, then refine question wording and option descriptions until the results hold up.

  6. Set thresholds and actions

    Because the probabilities are calibrated, you can act automatically above a threshold, send a middle band to a person, and log every probability for audit. Decisions about a person's rights, money or job stay with people; AT0M prepares, routes and prioritizes the work.

  7. Train on your own decisions, if you need to

    When your decisions differ from what the general model knows, train on them. About 10,000 records with five questions each takes roughly two hours on a laptop — no datacentre GPU, no cluster time.

  8. Deploy inside your perimeter

    When v1 ships, the same requests run against a single binary on your own CPU, Apple Metal or CUDA hardware, with no network dependency and no per-decision fee. The model is a fixed, versioned file you validate once. See how deployment works, or join the waitlist to hear when v1 is out.

Evaluating for your organization?

We'll help you scope a local-training or private-deployment evaluation against your own decisions — and share the full technical report.

// for engineers & procurement

The full evidence, when you want it.

A comparative evidence inventory across AT0M, Jev 1.13.0 and Laya, carried from the technical paper. It is not a harmonized three-model benchmark: figures were obtained under different conditions, and each is tagged with its source and how it was measured.

Full technical comparison: AT0M · Jev · Laya

[S1] Pienomial test report — AT0M results, measured by us against the public endpoint and reproducible by anyone: the same script run against the same endpoint returns the same figures. [L1/L2] Convai Laya model cards — vendor-reported; specialist and routed results kept distinct. [I1] DecisionEval — independent Jev 1.13.0 rerun on a frozen test. [J1/J2] TypeSafe documentation — product spec, subject to change. [B1] Public typed-decisions leaderboard.

A1 · Architecture & interface

ParameterAT0MLayaJev 1.13.0
DeveloperPienomialConvai InnovationsTypeSafe AI
Delivery formStandalone Rust binary; HTTP / CLI / JSON-fileOpen weights; Python SDK / Jev-compatible HTTP serverManaged API
BackboneNot disclosed in reportModernBERT-large 421M; mmBERT-base 322MNot disclosed
Primitiveschoice · score · noulchoice · score · noulchoice · score · noul
InferenceNon-autoregressive; no free textNon-autoregressive; no free textNon-autoregressive; no free text
Multi-question callYesYesYes
Max context32k request — the verification server caps requests at 4,096 tokens512 EN / 1,024 TD, multi 1,02432k state + longest Q / 64k request
Model selectionOne configuration across all testsRouter: EN / multilingual / specialistOne managed version

A2 · Accuracy & benchmark quality

MeasureAT0M [S1]Laya [L2]Jev [I1]
TD overall0.7890.766 specialist0.740 (0.727 older [B1])
TD choice0.7550.7330.738
TD score0.7630.7230.705
TD noul0.8570.8570.787
Invoice / service / obs / security.846 / .786 / .764 / .758.804 / .764 / .730 / .766.786 / .788 / .640 / .744
AG News0.896 described0.950 routed0.910 (secondary)
DAIR Emotion0.921 bare0.595 routed0.480 (secondary)
Banking770.853 / 77 labels0.425 / 77 labels0.870 / 72 labels ≠
Held-out phishing0.802 (excluded from training)——
CLINC150 + out-of-scope0.794 / 151 choices——
Base vs specialistOne config throughoutGeneral EN checkpoint 0.362 on TDZero-shot general weights

A3 · Calibration & reliability

MeasureAT0MLayaJev
TD ECE (10-bin)0.057 no temp fitted [S1]0.213 raw specialist [L2]0.121–0.161 across four independent studies
Per-type ECEchoice .052 · score .086 · noul .0570.081 after temperature [L1]TypeSafe publishes none
BrierNot provided0.062 specialist0.148 independent
Confidence gatingFull coverage curve not publishedNo matched curve≥0.7 → 57.6% kept @ 86.4%; ≥0.9 → 24.5% @ 93.7%
Option-order98% stableNot establishedNot established
Counterfactual13/15 directional; identifier-repeat pair fails——
Empty recordConfidence falls 0.723→0.616; still no abstentionNot measuredNot measured

A4 · Inference & throughput

MetricAT0MLayaJev
5-question, T482 ms server time84.5 ms EN / 40.1 ms multi687 ms p50, network incl. [I1]
1-question, T416 ms T4 server; 22.1 ms M1 Max39.5 ms EN / 32.8 ms multiolder 236–276 ms hosted
ThroughputM1 Max ~41→54 dps (1→32 clients)103–332 q/s batched, T4Provider-managed
77 / 151 candidates92.7 ms / 165.5 ms, M1 MaxBudget competition on optionsNot matched
HardwareCPU · Apple Metal · CUDACPU / GPU self-hostProvider-managed
NetworkServer figure excludes round-tripInference; network not included687 ms explicitly end-to-end

A5 · Enterprise ownership

ParameterAT0MLayaJev
WeightsTraining emits own weights; distribution terms per engagementApache-2.0 downloadableProprietary; no download
Self-hostingStandalone executablePython / HTTP serverNot offered publicly
Data transmissionLocal inference stays in perimeterLocal inference stays in perimeterSent to provider unless contracted otherwise
Version managementExport validation; customer handles rolloutPin model hashes / runtimeVersioned IDs / alias pinning
Audit trailRaw probabilities; app owns the logOpt-in prediction hooksVersion returned; app logs decisions

A6 · Economics

ItemAT0MLayaJev
Published priceCommercials will be released with v1. What do you think they should be? Please let us know at info@pienomial.comApache-2.0; hosting not free$0.042 / 1M input tokens; output free
Recurring costLocal hardware; no per-decision token feeSelf-host compute + opsPer-token input charges
Correct unitNormalized cost per verified useful decision = (inference HW + service ops + adaptation amortization + review + provider fees) ÷ verified useful decisions. No party publishes a hardware-normalized total — do not infer a universal lowest cost.

Sources: [S1] Pienomial AT0M benchmark report v1.0, 27 Sep 2026 · [L1/L2] Convai Laya model cards · [I1] DecisionEval Jev 1.13.0 independent rerun, 20 Sep 2026 · [J1/J2] TypeSafe documentation · [B1] public typed-decisions leaderboard. Latencies are not hardware- or network-normalized; pricing and availability may change after the report date. "Not disclosed" means the cited sources don't establish the item — not that it's unavailable.

AT0M→M0LECULE→P0LYMER

Everything complex is made of something that isn't.

An atom is one pass and no bonds — the smallest decision that can be made. M0LECULE bonds them into a sequence. P0LYMER runs the long chains. We are building up from the irreducible unit rather than carving down from a larger one.

Join the waitlist

V1 is the single binary you run on your own hardware. Leave an address and we'll tell you when it ships — nothing else, and no one else gets it.

One email when V1 ships. No newsletter.