SCHEDULING EXPERIMENT / INDEPENDENT GUIDE
Does Jev Improve an AI Scheduling Agent? A Three-Arm Test
We added Jev to a Chinese scheduling assistant to see whether a second decision layer would prevent mistakes. In this small evaluation, improving the original model’s context worked better than our Jev integration. Here is what failed, what the numbers mean, and what we would test next.
By Try Jev AI · Published October 8, 2026 · Experiments September 22, 2026
An arrival time is not an agreed meeting time
Consider this synthetic request, translated from Chinese:
“Help me note a meeting for Thursday morning. Lin says they’ll arrive at the office at 8:00; the discussion time depends on when I’m available.”
Thursday morning is already specified. But 8:00 describes Lin’s arrival, not the start of the meeting. A useful assistant would ask, “What time on Thursday morning should I schedule the meeting?” Asking whether 8:00 means morning or evening loses context; creating an 8:00 meeting invents an agreement.
This illustrates the problem; it is not a claim that every LLM makes this mistake or that Jev reliably fixes it. Our application, a WeChat scheduling assistant called Chengcheng, had an early time guard that split text at sentence boundaries and inspected local fragments before the LLM saw the full request. Adding another model downstream could not, by itself, repair that upstream loss of context.
Three workflows, 48 independent case families
We compared three arms on September 22, 2026. The base model was GPT-5.5 with low reasoning effort; the added decision model was pinned to jev-1.13.0. All examples were synthetic and anonymized. No customer conversations were used.
- A — Existing guard. Replay the current time guard. If it does not block the request, use the existing prompt and the structured semantic probe.
- B — Better context. Use the same model, prompt and request text, with centralized context constraints and without the premature local guard.
- C-ZH — B plus Jev. Give Jev B’s candidate and the original Chinese evidence, then adopt its judgments for action, time role, time of day, relation and date basis.
The fixture contained 240 variants across 48 independent families. The live model comparison used only the 48 original sentences, one per family—not 240 independent model trials. We split families into 24 calibration and 24 holdout cases. Most paths were run once per case, so this study does not measure variation across repeated runs.
This was an offline semantic probe, not a production end-to-end test. A did not replay every production component, and B was an experimental alternative, not a deployed fix. There were no real calendar writes, database changes or WeChat messages.
Better context won; our Jev layer added errors
| Workflow | Valid results | Exact match | Action match |
|---|---|---|---|
| A: existing guard | 47/48 | 42/48 | 42/48 |
| B: better context | 48/48 | 46/48 | 46/48 |
| C-ZH: B + Jev | 48/48 | 38/48 | 39/48 |
Exact match means matching the evaluation’s expected structured labels; action match checks the action label alone. A had one connection failure, retained in the attempted-case denominator. These labels represent this project’s scheduling policy, not an objective score for every possible assistant.
On the 24 held-out cases, exact matches were A: 21/24, B: 23/24 and C-ZH: 17/24. Jev corrected none of B’s errors and introduced eight mismatches on cases B had matched. That is evidence against this integration in this test—not evidence that Jev is universally unsuitable for agents.
Those eight mismatches were not eight dangerous writes. Some concerned overlapping labels such as answer versus no action, duty versus task, or turning a query into a clarification. Other cases produced cross-user modification suggestions that the minimal guard blocked. No real user state was changed.
B and C also shared a failure: both guessed that a Friday-at-eight gathering meant evening when our policy required clarification. Agreement between two models is not independent confirmation that the missing fact exists.
What we actually tested: a specific, imperfect integration
C directly adopted five Jev judgments. We logged probabilities but did not implement a calibrated abstention or fallback rule. We also retained B’s target identifiers and evidence quotations; replacing an action or relation could therefore leave the combined object internally inconsistent.
The action taxonomy had overlapping categories, and the English question rubrics needed review against their Chinese labels. C was not an independent complete scheduling system. It was a second classifier attached to B’s output, with all the coupling that implies.
The practical lesson is to repair evidence flow and define a narrow decision before adding another model. For example: “Does this message explicitly establish an agreed meeting start time?” That question needs an unknown or not-stated outcome. Even a correct answer does not establish authorization, ownership or a valid calendar transaction.
TypeSafe’s model documentation recommends decomposing broad judgments into smaller questions. Its Jev 1.13 limitations also describe date and time comparison weaknesses. Extracting a date can be a model task; interval arithmetic and date comparisons belong in deterministic code. These documentation links were checked October 8, after the experiment.
Would translating the Chinese evidence help?
We explored a fourth path, C-EN, translating evidence into English before Jev. Both language paths already used English questions. This extension began with 12 selected cases, then added the remaining 36 after inspecting those results. B’s candidates and the criteria were frozen, but this was exploratory follow-up, not a fresh blinded holdout.
C-EN produced valid results for 45/48 attempts, with 35/48 exact matches and 36/48 action matches. Three translations failed the literal preservation protocol—for example, changing a digit to its written form or a Chinese month into its English name. Those rejections were protocol failures, not necessarily semantic errors by Jev.
Across the 45 jointly valid pairs, Chinese and English evidence each achieved 35/45 exact matches: 34 were correct in both, nine wrong in both, one improved and one regressed. The net gain was zero.
Translation also changed evidence in consequential ways. A gathering acquired a dinner implication, and a modification instruction became a description of an already completed change. In the gathering case B had already guessed evening, so translation alone cannot explain the error. These examples show why translation must be evaluated as part of the pipeline, rather than assumed to be neutral preprocessing.
The extra layer added latency and cost
| Workflow | Total p50 | Total p95 | Paired added p50 / p95 vs B |
|---|---|---|---|
| A | 3.89 | 5.71 | — |
| B | 4.32 | 7.54 | Baseline |
| C-ZH | 5.20 | 8.43 | 0.86 / 1.16 |
| C-EN | 9.17 | 12.45 | 4.80 / 6.51 |
These are local semantic-probe timings, not isolated model inference or production end-to-end latency. Added time was calculated from paired runs; it is not the difference between the percentile columns. Seven A cases stopped at the local guard with little model time. A fast but unnecessary clarification is not a user-experience win.
The recorded experiment made 230 requests: 137 to GPT and 93 to Jev. Reused B results were counted once. Under the project’s September 22 pricing snapshot and cache accounting, the estimated equivalent cost was about $1.165 for GPT and $0.0081 for Jev, or $1.173 combined. This is an estimate, not a complete provider invoice: the failed GPT connection had no usage record, and connectivity smoke tests were excluded.
The Jev estimate used 193,692 input tokens at the recorded $0.042 per million input tokens, with output tokens free. Those are historical assumptions, not a promise of today’s pricing. Adding a cheap classifier does not remove the retained B call, and translation adds another step. The useful business metric is the cost of completing the task correctly across the whole workflow.
What would justify another trial?
Our decision from this experiment was to improve the original context handling first. The study did not justify new production Jev calls or changes to the scheduling assistant. Two narrower ideas remain worth testing:
- Asynchronous quality review. Flag possible mistakes for human inspection without changing user state. A proposed starting gate is at least 80% of flagged cases confirmed as issues and at least 50% of independently labeled issues found, while reducing review time. High precision on a tiny slice alone is insufficient.
- One narrow online check. First test in shadow mode. A proposed gate is at least a 20% relative reduction in consequential errors, no increase in unnecessary clarifications, and added p95 latency no greater than one second. Authorization and transaction checks remain in code.
These are proposed project criteria, not results or industry standards. A next evaluation should freeze labels, severity and thresholds in advance, start with at least 300 new independent Chinese scenarios, and report natural request mixes separately from hard cases. Three hundred cases is a starting point, not a guarantee of statistical certainty. Report counts and uncertainty; keep the original 48 as regression cases, not unseen validation data.
The main path must remain independently functional when the extra service fails. A fallback must neither silently approve a consequential action nor trap the user in repeated questions. Any serious authorization regression should stop the experiment.
Evidence and limits
This article adapts our September 25 draft and the saved September 22 experiment report and machine-readable summary. The underlying harness and full per-case records are not published here; the tables are aggregate results, not a publicly reproducible benchmark. Small synthetic samples, project-specific labels, one-run variability and the integration defects above all limit generalization.
The result is still useful: in this setup, preserving context helped more than adding our Jev decision layer. That is a reason to refine the question and evaluation—not to turn a negative experiment into a general verdict about a model.
Explore the scheduling question yourself
Open the Jev Playground and choose the scheduling example, “Is this meeting time actually confirmed?” Change the arrival and meeting wording to explore the distinction.
This preset illustrates a decision question; it does not reproduce the 48-case experiment or its assistant workflow. Check the current mode and result label. Live requests depend on API availability and may fail; Interactive Demo output is illustrative and cannot establish model accuracy. The historical results above do not indicate current service health.
For background, read our hands-on review of Choice, Noul and Score. For integration details, see the Jev API guide. Treat a successful demo as a starting point for your own evaluation, not proof that an agent is safe to act.
FROM CONCEPT TO INTERACTION
See the decision pattern for yourself.
Explore Choice, Noul and Score examples. The playground clearly labels Live Jev and Interactive Demo results.
Try Jev AIThis guide is independent of TypeSafe AI. For current model capabilities and access, see the official documentation.