Back to Blog
    Jev

    Jev for Classification: Six Places to Use It in Real Workflows

    A researched guide to TypeSafe's Jev with six classification workflows, a Python ticket-routing example, and a practical evaluation plan.

    Bhaulik Patel·Sep 30, 2026·8 min read

    Most workflow decisions have a small answer space. Which team owns this ticket? What kind of document just arrived? Which product category fits this description?

    Jev is worth investigating when the output you need is a category your software can act on. The engineering question is whether it makes that decision accurately enough, at an acceptable cost, with a useful fallback when it gets uncertain.

    Research checked September 30, 2026. The workflow examples below are illustrative designs, not measured customer deployments.

    What Jev does

    TypeSafe introduced Jev on September 15 as its first System One model. Its documentation describes an API that accepts a state and typed questions and returns structured decisions.

    The primitives are:

    • Choice: select from categories you define.
    • Score: evaluate against an ordered rubric.
    • Noul: evaluate a statement with a value between zero and one.

    Choice and Score return probability distributions and confidence. Multiple independent questions can share one state. This makes it possible to separate department, urgency, and issue type rather than forcing all three into one label.

    TypeSafe positions Jev around speed, efficiency, and calibrated decisions. Treat those as provider claims to evaluate on your own workload. A well-formed answer can still be the wrong answer.

    Six workflow examples

    The labels and expected routes here are design examples. They are not outputs from an API run.

    WorkflowIllustrative inputQuestion or categoriesWhat the application does
    Support routing“My invoice has two charges for the same order.”Billing, delivery, product help, otherPropose the billing queue; leave refunds to the authorized process.
    Document intake“Invoice 1048: payment due within 30 days.”Invoice, purchase order, contract, otherSend an invoice to extraction and validation; review ambiguous files.
    Product tagging“Insulated stainless-steel bottle with a carry handle.”Drinkware, clothing, electronics, otherSuggest drinkware; check against catalog rules before publishing.
    Feedback triage“I wish I could export the dashboard as CSV.”Feature request, bug report, praise, otherTag a feature request and group it with similar feedback.
    Search relevanceQuery: “reset API key”; candidate: “How to rotate service credentials”Irrelevant, partially relevant, directly relevantScore retrieved candidates before selecting context for a response.
    Meeting classification“Weekly sprint planning: priorities and capacity.”Planning, standup, retrospective, otherSuggest a planning template; let the organizer correct the tag.

    The value is in the next step each decision enables. A category that nobody uses adds another field to maintain.

    Example 1: route support tickets without writing a reply

    A support team may need both a department and an urgency signal. Keep them separate: an urgent billing problem belongs to billing, not to a vague “urgent” category.

    This original Python example follows the official Choice request pattern and SDK quickstart. Install typesafe-sdk and configure TYPESAFE_API_KEY on the server before running it. It is a documentation example; we have not executed a live request.

    from typesafe_sdk import Choice, TypeSafeClient
    
    ticket = "My invoice shows duplicate charges for order 812."
    
    with TypeSafeClient() as client:
        response = client.system_one(
            state=ticket,
            questions={
                "queue": Choice(
                    instructions="Which queue should review this ticket?",
                    criteria={
                        "billing": "Charges, invoices, and payment disputes",
                        "delivery": "Shipment status and missing packages",
                        "product_help": "Questions about using the product",
                        "other": "Insufficient detail or no suitable queue",
                    },
                ),
            },
        )
    
    answer = response.answers["queue"]
    allowed = {"billing", "delivery", "product_help", "other"}
    
    # Illustrative threshold: choose it using labeled validation cases.
    if (
        answer.choice not in allowed
        or answer.choice == "other"
        or answer.confidence < 0.90
    ):
        proposed_queue = "manual_review"
    else:
        proposed_queue = answer.choice
    
    # Suggest a route; do not issue refunds or send a reply here.
    print(proposed_queue)
    

    Add request timeouts, error handling, logging, and retry limits in a production integration. An API failure should fall back to review rather than discard the ticket.

    TypeSafe's confidence documentation explains that confidence summarizes how concentrated the probability distribution is. It is not a promise that a prediction is correct. The 0.90 cutoff above is a design placeholder, not an established accuracy level or a universal recommendation.

    Example 2: sort incoming documents

    An operations inbox contains invoices, purchase orders, contracts, and unrelated attachments. Convert supported documents into text, then ask a Choice question with explicit definitions for each document type.

    Keep classification separate from extraction. Choosing “invoice” does not verify its total, vendor, authenticity, or payment instructions.

    A practical workflow could be:

    1. Extract text with provenance and check that extraction succeeded.
    2. Classify the document, with an “other” option.
    3. Route a proposed invoice to field extraction and arithmetic checks.
    4. Send ambiguous, incomplete, or mismatched results to a reviewer.

    Evaluate scanned files, mixed attachments, unusual layouts, and multilingual examples separately. If the text is missing the decisive section, fix the input pipeline before changing the classifier.

    Example 3: tag a product catalog

    For a small shop, a handful of top-level categories may be enough. A larger catalog needs a hierarchy: first decide the broad family, then evaluate the relevant subcategories.

    For example, “drinkware” might lead to bottles, mugs, and tumblers. Define what separates them and include an escape category at each stage.

    The Choice documentation describes a limit of 255 options and recommends an “other” or equivalent option when the list is incomplete. It also describes hierarchical classification through successive questions.

    A staged design introduces more calls and lets early mistakes propagate. Compare it with a single request where feasible. Evaluate rare categories and misleading titles, not just your most popular items.

    Example 4: organize customer feedback

    A sentence can contain several signals: “The dashboard freezes, and I need a CSV export.”

    Use independent questions for “reports a bug” and “requests a feature” when both can be true. A single Choice that forces those signals to compete loses information.

    Then use code to create review tags. Keep sentiment separate from severity: a polite report can describe a serious failure. Sample tagged feedback regularly so a product manager can spot drift in the label definitions.

    Example 5: score search candidates

    After retrieval, a relevance score can help decide which documents deserve a place in the context window. Include the query and candidate text in the state, then define a narrow rubric.

    Parallel's September 18 evaluation offers a useful reality check. Jev matched at least one of Parallel's custom rerankers on NDCG@10, but its cost per document was higher. Parallel's specialized systems also performed better on topic classification and query freshness classification.

    Those are results from Parallel's tasks and infrastructure, not a general ranking. The lesson for an FDE is to compare Jev with the reranker or classifier you actually operate.

    Example 6: suggest meeting templates

    Meeting titles and descriptions can supply enough context to propose a planning, standup, or retrospective template. Ambiguous invitations should keep the default template or prompt the organizer.

    This is a useful place to begin because mistakes are visible and easy to correct. Record corrections as labeled examples and review whether the classifier helps more than a few title rules would.

    When a simpler approach wins

    Try deterministic rules first for known IDs, exact codes, or fixed formats. Compare a trained classifier or embedding-based classifier when you have stable labels and enough representative data.

    A generative model may be more appropriate when the task needs an explanation, an open-ended answer, or several reasoning steps. Jev may fit a bounded semantic judgment with labels defined at request time. That is an engineering hypothesis, not a blanket preference.

    A pilot that can earn its place

    Build a labeled dataset from the workflow you intend to improve. Separate the cases used to refine label descriptions from the held-out cases used to assess the final system.

    Compare rules, your existing classifier, a small structured-output LLM, and Jev where available. Record:

    • Precision and recall for each category, including rare ones.
    • False routes and the operational cost of correcting them.
    • Review coverage and error rates at different confidence thresholds.
    • End-to-end latency, failed requests, and total cost per accepted decision.
    • Behavior on missing context, overlapping labels, and out-of-scope inputs.

    Do not select the threshold on the final test set. Choose it on validation data, then check the resulting tradeoff on held-out data and a limited production rollout.

    Version the taxonomy, question wording, model returned by the API, and evaluation cases. Recheck behavior when any of them changes. A typed classifier can make a decision easier to integrate; the workflow around it determines whether that decision is useful.

    Discuss a classification workflow · Learn to build evaluations · Request AI training for your team

    JevClassificationWorkflow AutomationAI Evals
    Share
    BP

    Bhaulik Patel

    Forward deployed AI engineer and creator of Deployed Engineer.