Classification at scale

Sort ten thousand items without reading ten thousand items.

Compact embedding-based classifiers label your documents, emails and tickets in bulk. Fast, cheap per item and measured against a labelled sample before anything goes live.

Accuracy measured on your dataLocal or APILow cost per item

[ 01 — WHAT YOU GET ]

A classifier you can test.

Large language models can classify, but at volume they are slow and costly. Compact models do the bulk, and a large model or a person handles the uncertain rest.

  • [ 01 ]

    Label scheme

    A clear set of categories with definitions and edge cases, agreed with the people who use the result.
  • [ 02 ]

    Labelled baseline

    A sample labelled by your experts. It is the yardstick for every version of the classifier.
  • [ 03 ]

    Compact classifier

    An embedding-based model, such as a JEPA-style model, open-source and run locally, or via a provider's API, whichever fits your data rules.
  • [ 04 ]

    Confidence and routing

    Each item gets a score. Low-confidence items go to a person or a larger model instead of being guessed.
  • [ 05 ]

    Measured accuracy

    Precision and recall per category on held-out data, and a check on fresh samples every month.

[ 02 — PROCESS ]

Label, train, measure, run.

The labelled sample decides what is good enough.

  1. 01

    Define categories

    We work with your team on a label scheme and resolve the ambiguous cases on paper.1 week
  2. 02

    Build the baseline

    Your experts label a sample. We support with tooling and check agreement between labellers.1 to 2 weeks
  3. 03

    Train and evaluate

    We compare candidates on held-out data and report accuracy per category, including where it fails.2 weeks
  4. 04

    Run and monitor

    Batch or live classification with confidence thresholds, a review queue and monthly re-checks.Ongoing

[ 03 — USE CASES ]

Volumes that fit.

Typical scenarios. None of these describe a specific client.

Shared inbox sorting

Today
Staff read every mail to decide which team handles it.
With the workflow
The classifier assigns a team and a topic. Low-confidence mails land in a review queue.

Document archive tagging

Today
Years of scans and PDFs have no consistent type or metadata.
With the workflow
A batch run tags document types once, and new documents are tagged as they arrive.

Ticket categorisation

Today
Categories are picked by hand and drift over time.
With the workflow
The classifier suggests a category with a score. Agents confirm or correct, and corrections feed the next version.
EXAMPLE — ILLUSTRATIVE DATA, NOT A CLIENT RESULT

[ 04 — SCOPE & PRICE ]

A proof of concept with a measured result.

Classification proof of concept · 4 to 6 weeks

from CHF 9,500fixed scope · excl. VAT

Running costs depend on volume and hosting and are stated before rollout. Extensions: CHF 2,200 per day.

  • Label scheme and labelled baseline
  • Candidate classifiers compared on held-out data
  • Accuracy per category, including failures
  • Confidence thresholds and review queue design

[ 05 — QUESTIONS ]

What people ask about classification.

How accurate will it be?

We do not promise a number up front. We measure it on your labelled sample and only go live if it meets the threshold you set.

How much labelled data do we need?

Often a few hundred examples per category are enough to start, fewer for simple schemes. We tell you after a first test.

Why not just use a large language model?

You can, for small volumes. At tens of thousands of items, compact classifiers are faster and cheaper, and we can still use a large model for the uncertain cases.

Can it run without sending documents to an external provider?

Yes. Open-source embedding models run on your own servers, so the content stays with you.

What happens when categories change?

We update the scheme, add labelled examples and re-evaluate. Version history and results are kept so you can compare.

Bring a sample of what needs sorting.

We test a small set and show you the accuracy before you decide.

Book an intro call