Classification at scale
Sort ten thousand items without reading ten thousand items.
Compact embedding-based classifiers label your documents, emails and tickets in bulk. Fast, cheap per item and measured against a labelled sample before anything goes live.
[ 01 — WHAT YOU GET ]
A classifier you can test.
Large language models can classify, but at volume they are slow and costly. Compact models do the bulk, and a large model or a person handles the uncertain rest.
- [ 01 ]A clear set of categories with definitions and edge cases, agreed with the people who use the result.
- [ 02 ]
Labelled baseline
A sample labelled by your experts. It is the yardstick for every version of the classifier. - [ 03 ]
Compact classifier
An embedding-based model, such as a JEPA-style model, open-source and run locally, or via a provider's API, whichever fits your data rules. - [ 04 ]
Confidence and routing
Each item gets a score. Low-confidence items go to a person or a larger model instead of being guessed. - [ 05 ]
Measured accuracy
Precision and recall per category on held-out data, and a check on fresh samples every month.
[ 02 — PROCESS ]
Label, train, measure, run.
The labelled sample decides what is good enough.
- 01
Define categories
We work with your team on a label scheme and resolve the ambiguous cases on paper.1 week - 02
Build the baseline
Your experts label a sample. We support with tooling and check agreement between labellers.1 to 2 weeks - 03
Train and evaluate
We compare candidates on held-out data and report accuracy per category, including where it fails.2 weeks - 04
Run and monitor
Batch or live classification with confidence thresholds, a review queue and monthly re-checks.Ongoing
[ 03 — USE CASES ]
Volumes that fit.
Typical scenarios. None of these describe a specific client.
Shared inbox sorting
- Today
- Staff read every mail to decide which team handles it.
- With the workflow
- The classifier assigns a team and a topic. Low-confidence mails land in a review queue.
Document archive tagging
- Today
- Years of scans and PDFs have no consistent type or metadata.
- With the workflow
- A batch run tags document types once, and new documents are tagged as they arrive.
Ticket categorisation
- Today
- Categories are picked by hand and drift over time.
- With the workflow
- The classifier suggests a category with a score. Agents confirm or correct, and corrections feed the next version.
[ 04 — SCOPE & PRICE ]
A proof of concept with a measured result.
Classification proof of concept · 4 to 6 weeks
Running costs depend on volume and hosting and are stated before rollout. Extensions: CHF 2,200 per day.
- Label scheme and labelled baseline
- Candidate classifiers compared on held-out data
- Accuracy per category, including failures
- Confidence thresholds and review queue design
[ 05 — QUESTIONS ]
What people ask about classification.
How accurate will it be?
We do not promise a number up front. We measure it on your labelled sample and only go live if it meets the threshold you set.
How much labelled data do we need?
Often a few hundred examples per category are enough to start, fewer for simple schemes. We tell you after a first test.
Why not just use a large language model?
You can, for small volumes. At tens of thousands of items, compact classifiers are faster and cheaper, and we can still use a large model for the uncertain cases.
Can it run without sending documents to an external provider?
Yes. Open-source embedding models run on your own servers, so the content stays with you.
What happens when categories change?
We update the scheme, add labelled examples and re-evaluate. Version history and results are kept so you can compare.
Bring a sample of what needs sorting.
We test a small set and show you the accuracy before you decide.