Skip to evidence
MemCat
Evidence Logged experiments6 October 2026

Every decision,
accounted for.

A real local encoder. 3,080 labelled banking messages. The accepted decisions, the ones sent for review, and the mistakes—all from the same frozen test run.

Classify before Qwen. Measure both routes.

SDK 0.3.1 · local reference

70.5% fewer LLM calls. Accuracy changed from 69.5% to 83.4%, below our preset 90% floor. This configuration did not pass.

The frozen sample contains 770 public BANKING77 messages, ten per category. Direct Qwen classifies every message; the hybrid runs MemCat first and makes a fresh Qwen request on review or failure. Both return an exact label or review. These are classification calls; no answers are generated for a customer.

Measured outcomeQwen directlyMemCat + Qwen
Correct / 770535 / 770 (69.48%)642 / 770 (83.38%)
Wrong / review / failed218 / 17 / 0123 / 5 / 0
Actual LLM calls770227
Reported prompt + output tokens2,103,931620,541
Sum of measured query times299.67 s87.72 s
Per-message p50 / p95353.57 ms / 794.05 ms8.41 ms / 436.41 ms

MemCat accepted 543 messages locally: 527 were correct and 16 were wrong (97.05% accepted precision). On the remaining 227, the actual Qwen fallback returned 115 correct labels (50.66%), 107 wrong labels and 5 reviews. Most final errors came from the fallback route.

The paired accuracy difference was 13.90 percentage points; the one-sided 95% lower bound was 11.95 points. The interval uses 20,000 predeclared stratified paired bootstrap samples. Review and failure count as incorrect on this in-scope sample. This is an internal tolerance, not customer approval.

Target frozen before inferenceResult
All 770 cases recorded on both routesMet
At least 90% correct across all hybrid inputsNot met
Paired 95% lower bound no worse than −2 percentage pointsMet
At most 1% hybrid operational failuresMet
At least 50% fewer actual LLM requestsMet
At least 50% fewer returned prompt + output tokensMet
At least 30% less measured query timeMet

Qwen2.5 7B Q4_K_M ran locally through Ollama 0.35.1 on an Apple M4 Max with an 8,192-token context. It received the 77 labels, descriptions and one training example per category. MemCat used 32 examples plus a description per category. This comparison does not establish the best possible Qwen configuration.

Category descriptions were literal names. The upstream label get_physical_card is attached to PIN-related examples, which can mislead a prompted model. An independent source check verified every imported label against the original CSV files. The frozen prompt and this failed result are preserved; future prompt changes require separate evaluation.

Qwen ran first, then the hybrid, with the target model unloaded and the same three excluded warmups repeated between routes. Prefix caching remained active. Query time includes local classification and fallback when used; setup, warmups and receipt writing are excluded and recorded separately. Fixed route order may affect timing. The sample was previously evaluated by MemCat; it is not a fresh held-out test or customer traffic. Known mixed-intent failures below remain unresolved.

Comparison report · SHA-256582a276a3995f5c10be8553b632fe60a9b82462a108fdaeb0bde82f5c5dbd858

Public BANKING77 data: Casanueva et al. / PolyAI, CC BY 4.0. Dollar savings, total compute cost and customer adoption remain unmeasured. Sources and licenses.

What happened to the messages?

All 3,080 official BANKING77 test rows, across 77 categories. Thresholds and prototypes were frozen before this test. No test rows were removed.

English · banking support
Correct among automatic decisions97.48%

2,092 correct / 2,146 accepted

Automatically classified69.68%

2,146 accepted / 3,080 total

Incorrect automatic routes54

1.75% of all 3,080 inputs

2,092 correct54 incorrect934 review0 failed or skipped

95% Wilson interval for correctness among accepted decisions: 96.73%–98.07%. Review means the classifier declined to choose; no fallback model ran, and those messages are not counted as correct answers.

Precision and coverage belong together.

Without abstention, its top-choice accuracy is 88.25% and macro-F1 is 0.8821. The higher accepted precision comes from sending harder messages to review.

Compared with simpler options.

The same test inputs and labels. Each baseline used only training data; rejection thresholds were selected on calibration data. The local text baselines require no neural encoder.

MethodTop-choice accuracyAccepted precisionCoverageWrong routes
MemCat + MiniLMExample similarity + abstention88.25%97.48%69.68%54
Exact training lookup0.13%—0.00%0
Category word overlap36.66%—0.00%0
TF-IDF nearest centroid81.17%96.93%46.46%44
Always review——0.00%0

“—” means no automatic decisions, so precision is undefined. Exact lookup and category overlap did not meet the calibration target and accepted no test messages under their selected rejection policy. Always choosing the most common training label achieved 1.30% accuracy.

Measured speed. Defined boundaries.

Individual, sequential queries on Apple M4 Max, Node 22.22.3. These measurements include token validation, tokenization, q8 encoding, category matching, and the review decision.

Warm latency · median
1.78 msPer message; 3 warmup calls, no result cache
Warm latency · p95 / p99
2.70 ms / 3.72 ms
Model load from local disk
371.73 msExcludes first network download; not a browser cold-start figure
Sequential throughput
524 messages / secondMeasured across the full test callback loop
Example index
10.07 MB JSON / 3.67 MB gzip2,537 prototype vectors; encoder runtime and weights are additional
Decision policy
Similarity ≥ 0.70 · margin ≥ 0.08Cosine scores are not probabilities of correctness

Runtime: ONNX CPU; individual sequential queries. Browser WASM uses a different execution environment; the live demo reports its own loading and inference times. These figures do not measure an LLM workflow’s end-to-end latency or cost savings.

Run the same policy in your own project.

The Node starter includes the encoder adapter, pinned dependencies, and a replay command. We verified a fresh installation outside this repository using the published SDK: it downloaded the assets, ran local inference, and wrote inspectable results without an API key or source edits.

This is an internal integration check on Node 22 and Apple silicon. Its four authored examples are separate from the accuracy benchmark above. Setup time, first inference, asset requests, and failures are logged separately; no LLM cost savings were measured.

Where it still struggles.

Passing an in-scope benchmark does not make every decision safe. We also ran a separate unknown-topic diagnostic and an authored stress suite.

Narrow diagnostic

Distant, non-banking requests

0 / 240

Accepted as a banking category. All official CLINC test rows from eight predeclared tasks, including weather, recipes, and music. The 95% Wilson upper bound is 1.58%.

This does not establish rejection of unfamiliar requests close to banking topics.

Stress gate failed

Mixed intent, negation, and spelling

3 / 14

Authored cases missed their expected handling. One mixed request was routed automatically; two other cases were deferred instead of receiving the expected category.

These are regression cases, not a representative accuracy sample.

  1. Wrong automatic decision
    “Why was I charged twice, and how do I close my account?”
    Expected
    Review
    Returned
    transaction_charged_twice
    Case
    mixed-02
  2. Deferred · expected a category
    “I do not want to close my account. Where is my card?”
    Expected
    card_arrival
    Returned
    Review
    Case
    negation-02
  3. Deferred · expected a category
    “my caard got stoln”
    Expected
    lost_or_stolen_card
    Returned
    Review
    Case
    typo-01

The incorrect automatic decisions.

The first 10 of 54 errors in source order. Related banking categories are sometimes confused. The full report contains every result, expected label, similarity decision, and timing.

  1. “Is there a way to know when my card will arrive?”
    Expected
    card_arrival
    Returned
    card_delivery_estimate
    Source ID
    banking77-test-00004
  2. “How long does a card delivery take?”
    Expected
    card_arrival
    Returned
    card_delivery_estimate
    Source ID
    banking77-test-00012
  3. “How do I change currencies to euros?”
    Expected
    fiat_currency_support
    Returned
    exchange_via_app
    Source ID
    banking77-test-00256
  4. “Can I get my card expedited?”
    Expected
    card_delivery_estimate
    Returned
    card_arrival
    Source ID
    banking77-test-00309
  5. “How do I unblock my card using the app?”
    Expected
    card_not_working
    Returned
    pin_blocked
    Source ID
    banking77-test-00369
  6. “Can I change from one currency to another?”
    Expected
    exchange_via_app
    Returned
    fiat_currency_support
    Source ID
    banking77-test-00403
  7. “What all currencies can be exchanged?”
    Expected
    exchange_via_app
    Returned
    fiat_currency_support
    Source ID
    banking77-test-00411
  8. “Is it possible to change to another currency?”
    Expected
    exchange_via_app
    Returned
    fiat_currency_support
    Source ID
    banking77-test-00415
  9. “Can I change to another currency?”
    Expected
    exchange_via_app
    Returned
    fiat_currency_support
    Source ID
    banking77-test-00419
  10. “Change currency”
    Expected
    exchange_via_app
    Returned
    fiat_currency_support
    Source ID
    banking77-test-00426

What this result does—and doesn’t—prove.

One declared workflow

English, single-intent banking messages with 77 published labels. Public benchmark performance does not establish accuracy on your support queue, language, or category definitions.

Existing source overlap, disclosed

7 official test messages match earlier source rows after normalization. All test rows are retained. Excluding those overlaps gives 97.48% accepted precision at 69.61% coverage across 3,073 messages. Public data may also have appeared in encoder pretraining.

An existing encoder, with controls

MemCat uses MiniLM embeddings and nearest-example cosine similarity. It adds explicit thresholds, serializable indexes, evaluation, tracing, and shadow mode. This is not a newly trained foundation model.

Economics still need your workload

This local run made 0 hosted classification calls. Local compute and review work still have costs. The separate local Qwen comparison above measures real classification requests and latency. Customer savings, paid-provider bills and independent adoption remain unmeasured.

Keep the receipts.

Training, development, and calibration contained 7,963, 965, and 1,071 rows respectively. The selected model and thresholds were frozen before testing. The report records dataset hashes, source-code hashes, outcomes, baselines, and limitations.

Full report · SHA-2568d29ebcbfa71e7000b095c983fe74b5cc4e133cd4c8aaacf19844ead5e115c41

BANKING77: Casanueva et al. / PolyAI, CC BY 4.0. CLINC150: Larson et al. / Clinc, CC BY 3.0. MiniLM encoder: Apache 2.0. Full attribution and immutable source revisions are in the linked notice.