Instructions to use autotrust/JEV-35B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use autotrust/JEV-35B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="autotrust/JEV-35B")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("autotrust/JEV-35B") model = AutoModelForMultimodalLM.from_pretrained("autotrust/JEV-35B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
autotrust/JEV-35B
🔴 Live: AutoTrust/GuruSearch recommendation & search demo → news.guru.so
Try AutoTrust/GuruSearch live (new, 11 October 2026). A recommendation and search demo in which every ranking is a calibrated System 1 decision, made in about half a second:
- Recommendation: the latest headlines from 6 news feeds, ranked by importance.
- Search: results from Google News, Bing News and Yahoo News re-ranked by relevance, side by side with the search engines' own order, plus a short answer with citations.
A System 1 decision model on Qwen3.5-35B-A3B (35 B parameters, about 3 B active per token), up to 256 options in one pass
Decision Index 0.3, public suite: 59.75 (our scoring with the kit), against 53.64 for autotrust/JEV-27B-VL on the board. Approximate public vision score 72.08 (our rebuild of the vision benchmarks; JEV-27B-VL 71.67 on the same rebuild, a tie within noise). Median single-request latency 241 ms on one B200 (JEV-27B-VL 271 ms).
NVFP4 version: autotrust/JEV-35B-NVFP4, the same model with its routed experts in NVFP4: 25.6 GB download, 23 GiB of GPU memory (bf16: 72 GB, 66 GiB), so it runs on one 32–48 GB GPU. Decision Index 0.3 public 59.18 (bf16 59.75), vision rebuild 71.23 (bf16 72.08), 96.6 % of the Decision Index answers identical to bf16, same latency, same
serve.shand API.
| version | download | GPU memory (weights) | Decision Index 0.3 public | vision (rebuild) |
|---|---|---|---|---|
| autotrust/JEV-35B (bf16, this repo) | 72 GB | 66.5 GiB | 59.75 | 72.08 |
| autotrust/JEV-35B-NVFP4 | 25.6 GB | 23.3 GiB | 59.18 | 71.23 |
Two models, two organisations. TypeSafe Jev 1.13 is the hosted, closed model made by TypeSafe AI. autotrust/JEV-35B is an independent open-weights model built by AutoTrust AI; it is not affiliated with, endorsed by, or a product of TypeSafe AI.
Decision Index 0.3 (public suite)
| public index | |
|---|---|
| autotrust/JEV-35B, System 1 (our run with the kit) | 59.75 |
| autotrust/JEV-27B-VL, System 1 (board, public part) | 53.64 |
| area (skill) | Knowledge & Reasoning | Language | Retrieval & Classification | Tools & Automation | Arts & Taste |
|---|---|---|---|---|---|
| JEV-35B | 0.426 | 0.632 | 0.710 | 0.765 | 0.422 |
| JEV-27B-VL | 0.414 | 0.564 | 0.537 | 0.735 | 0.416 |
JEV-35B: all 140,178 scoreable requests of the 0.3 public suite answered (0 errors), System 1 only (thinking off),
every choice read in one pass (up to 256 options), scored with the kit's score --edition 0.3. Our scoring, not a board
entry: the board's Full score also counts private tests (80 %), which only its maintainers run. JEV-27B-VL: the board's
public numbers for JEV-27B, whose text decisions JEV-27B-VL reproduces (see the JEV-27B-VL card).
JEV-35B is ahead in all five areas and on 24 of 37 benchmarks. Largest gains (skill points): PhishNChips +42.3, HoVer +28.3, Habermas Machine +28.1, VAST +22.8, WinoGrande +18.9, BANKING77 +18.3, iSarcasmEval +17.2, When2Call +14.9, GPQA +9.5. JEV-27B-VL is ahead on POP909-CL (+25.7), GSM8K (+18.3), ANLI (+11.4), BPoMP (+10.8), NLI4CT (+6.7) and BBH (+4.8).
Every benchmark
Skill rescales the benchmark's own metric so that chance is 0 (below chance counts as 0); ★ = gold benchmark (weight 1.2); the higher score is in bold.
Knowledge & Reasoning (area skill 0.426; JEV-27B-VL 0.414)
| benchmark | JEV-35B | JEV-27B-VL |
|---|---|---|
| GPQA Diamond ★ | 0.360 | 0.265 |
| GSM8K (0.3 rebuild) | 0.359 | 0.542 |
| ChessBench | 0.127 | 0.091 |
| MuSR | 0.361 | 0.402 |
| SATA-Bench | 0.284 | 0.328 |
| CRUXEval | 0.546 | 0.583 |
| CLadder | 0.440 | 0.410 |
| HLE ★ | 0.000 | 0.000 |
| MMLU-Pro ★ | 0.642 | 0.567 |
| BBH ★ | 0.647 | 0.695 |
| WinoGrande ★ | 0.841 | 0.651 |
Language Understanding (area skill 0.632; JEV-27B-VL 0.564)
| benchmark | JEV-35B | JEV-27B-VL |
|---|---|---|
| ContractNLI | 0.732 | 0.639 |
| ANLI ★ | 0.545 | 0.658 |
| HellaSwag ★ | 0.962 | 0.900 |
| ACOS | 0.280 | 0.169 |
| FinEntity | 0.856 | 0.755 |
| iSarcasmEval | 0.426 | 0.254 |
| VAST | 0.550 | 0.322 |
| NLI4CT | 0.627 | 0.694 |
| RAGTruth | 0.666 | 0.604 |
Retrieval & Classification (area skill 0.710; JEV-27B-VL 0.537)
| benchmark | JEV-35B | JEV-27B-VL |
|---|---|---|
| BANKING77 ★ | 0.922 | 0.739 |
| CLINC150+OOS ★ | 0.943 | 0.845 |
| BRIGHT ★ | 0.421 | 0.417 |
| Amazon ESCI | 0.519 | 0.430 |
| PhishNChips | 0.661 | 0.238 |
| HoVer | 0.761 | 0.478 |
Tools & Automation (area skill 0.765; JEV-27B-VL 0.735)
| benchmark | JEV-35B | JEV-27B-VL |
|---|---|---|
| BFCL ★ | 0.950 | 0.946 |
| ToolRet | 0.600 | 0.604 |
| API-Bank ★ | 0.856 | 0.825 |
| Home appliance simulator | 0.500 | 0.523 |
| When2Call | 0.863 | 0.714 |
Arts & Human Taste (area skill 0.422; JEV-27B-VL 0.416)
| benchmark | JEV-35B | JEV-27B-VL |
|---|---|---|
| BPoMP | 0.750 | 0.859 |
| Humicroedit | 0.232 | 0.223 |
| POP909-CL | 0.132 | 0.390 |
| cfcolor | 0.256 | 0.283 |
| Habermas Machine | 0.416 | 0.135 |
| New Yorker | 0.744 | 0.607 |
Images (approximate public vision score)
The vision benchmarks of the Decision Index are not published; this is our rebuild of their public datasets (CV-Bench, BLINK, RealWorldQA, CharXiv, InfographicVQA, Mind2Web, CORD + FUNSD, Hateful Memes, R-Bench-M, MMMU-Pro vision; Winoground not included), scored with the board's chance correction and weights. The same rebuild gives JEV-27B-VL 71.67 against its board score of 71.53 on the same benchmarks.
| approximate public vision score | accuracy | ECE | |
|---|---|---|---|
| JEV-35B | 72.08 | 78.6 % | 0.050 |
| JEV-27B-VL (same rebuild) | 71.67 | 78.0 % | 0.044 |
| benchmark | JEV-35B skill | JEV-27B-VL skill |
|---|---|---|
| CV-Bench | 75.7 | 76.9 |
| BLINK | 55.6 | 56.7 |
| RealWorldQA | 68.9 | 72.2 |
| CharXiv | 80.7 | 80.7 |
| InfographicVQA | 95.0 | 95.9 |
| Mind2Web | 83.6 | 83.2 |
| KIE (CORD+FUNSD) | 98.7 | 98.6 |
| Moderation (Hateful Memes) | 47.7 | 38.0 |
| R-Bench-M | 28.1 | 30.3 |
| MMMU-Pro vision | 43.5 | 39.2 |
Overall the two models are level: the difference (+0.4) is inside the noise (paired bootstrap 95% interval −0.7 to +1.6). Only two per-benchmark differences are statistically significant (paired McNemar test, p < 0.01), both in favour of JEV-35B: moderation (+9.6) and MMMU-Pro (+4.3). The small deficits on RealWorldQA, CV-Bench, BLINK, InfographicVQA and R-Bench-M (−1 to −3) are not significant. The vision tower is Qwen3.5-35B-A3B's, unchanged; image decisions are zero-shot.
Speed
One sequential client, the same 587 rows (387 text rows sampled across the Decision Index, 200 image rows), POST
/v1/decide on serve.sh (vLLM, System 1 as a LoRA), one B200:
| median (p95) | text | image | all |
|---|---|---|---|
| JEV-35B | 215 ms (415) | 378 ms (784) | 241 ms (682) |
| JEV-27B-VL (same benchmark) | 207 ms (617) | 457 ms (910) | 271 ms (746) |
Computer use, robot arm and games
Same demo code, seeds, scenes and opponents as the JEV-27B-VL card; every
step is one System 1 decision (POST /v1/decide, thinking off), one model on one B200.
Computer use: screenshot → which element to click. A real browser (headless Chromium). Every clickable element gets a numbered box; System 1 picks the next click (or "the task is complete"), the browser clicks it, and the loop repeats.
Robot arm: pick and place from a camera image (MuJoCo). At every step System 1 answers two questions from the top camera: is the target left or right of the gripper, and above or below it? The arm halves its step whenever an answer flips.
| JEV-35B | JEV-27B-VL | |
|---|---|---|
| Computer use, numbered boxes + element text (60 tasks: shop, settings, mail) | 95% | 95% |
| Computer use, numbered boxes only (60 tasks) | 38% | 10% |
| ms per click decision, median (6 browsers in parallel) | 397 | 720 |
| Robot arm, binary-decision servo (20 scenes): pick-and-place success | 75% | 75% |
| Robot arm: median distance from the cube centre when grasping | 2.5 cm | 2.7 cm |
| Robot arm: placed in the tray, once grasped | 15/15 | 15/15 |
| Robot arm: ms per decision, median | 239 | 239 |
| Robot arm, direct choice among 8 motor actions (10 scenes) | 0% | 0% |
- Computer use with element text: 95%, the same three failures as JEV-27B-VL (shop seeds 12, 15, 20: the colour swatch carries no text and is skipped).
- Numbered boxes only (every element read from pixels): 38% against 10%. JEV-35B completes 90% of the settings tasks but only 15% of mail and 10% of shop; most failures still declare the task complete too early (24 of 37).
- Robot arm: same success rate, different scenes. Each model misses 5 of 20 grasps (both miss scenes 4 and 19), every miss 3 cm or more off the cube centre; once grasped, every cube reaches the tray.
- Choosing directly among 8 motor commands fails for both models; decompose control into simple visual questions.
Games (seed 0 for both models; board as image + text):
| game | JEV-35B | JEV-27B-VL |
|---|---|---|
| 2048: score / largest tile | 336 / 32 | 2,080 / 128 |
| Connect Four against a rule-based opponent (6 games) | 0 wins, 6 losses | 0 wins, 6 losses |
| Flappy Bird: pipes passed | 0 | 3 |
| Snake: food eaten | 11 | 21 |
| Quick, Draw! top-1 among 16 (320 sketches; 30% / 60% / 100% of strokes) | 35.0 / 57.5 / 86.3% | 41.6 / 62.2 / 87.8% |
| Chess mate in one, choice among 16 moves (200 Lichess puzzles) | 47.0% | 52.0% |
For game play, use JEV-27B-VL. Over five seeds JEV-35B's largest 2048 tile averages 64 (random play: 102), it passes
no Flappy Bird pipe and eats 3–18 pieces of food in Snake (mean 10). Per-episode results: reports/demos/.
Quick start (vLLM)
hf download autotrust/JEV-35B --local-dir JEV-35B
bash JEV-35B/serve.sh # vLLM on :8000; one GPU with 80 GB or more
On a smaller GPU, use autotrust/JEV-35B-NVFP4 (one GPU with 32 GB or
more; same serve.sh and API):
hf download autotrust/JEV-35B-NVFP4 --local-dir JEV-35B-NVFP4
bash JEV-35B-NVFP4/serve.sh
serve.sh runs serve_decide.py: the standard vLLM OpenAI server (same flags as vllm serve) with a POST /v1/decide
route. Plain requests are System 2; requests for the LoRA module jev-decision (adapter_vllm/: backbone LoRA + the
decision head as an lm_head LoRA) are System 1. Tested with a vLLM development build from September 2026.
The adapter has no LoRA on the routed experts. serve.sh sets JEV_DECIDE_NO_MOE_LORA=1 and --lora-target-modules
so that the routed experts run on vLLM's normal MoE kernels (the decision head needs --max-lora-rank 320; vLLM's
MoE-LoRA kernel supports at most rank 128).
System 1: POST /v1/decide
curl localhost:8000/v1/decide -H 'Content-Type: application/json' -d '{
"kind": "choice",
"state": "Customer message: my card was charged twice for the same order.",
"question": "Which team should handle this ticket?",
"options": ["billing", "shipping", "technical support", "account security"]}'
| field | value |
|---|---|
kind |
noul: yes/no, probabilities for ["false", "true"] · score: 0–5 · choice: your options |
state |
what the decision is about: a string, a JSON object, or a list mixing text and images |
question |
one question about the state |
options |
choice only: 2–256 strings |
thinking |
"off" (default) |
The response has options, probabilities, choice, choice_index and usage. GET /v1/decide/info lists the
defaults.
System 2
import requests
requests.post("http://localhost:8000/v1/chat/completions", json={
"model": "autotrust/JEV-35B",
"messages": [{"role": "user", "content": "In one sentence, what is safety stock?"}],
"max_tokens": 200})
If you call /v1/completions for System 1 yourself, pass top_k: 0 and top_p: 1.0 so that the returned
probabilities are not truncated.
Engine and read-out
System 1 reads the hidden state at the last token of a bare-text prompt:
[kind] choice
[state] ...
[question] ...
[options]
A) ...
B) ...
[decision]:
A 264-slot linear head (fp32) gives one logit per slot; the active slots of the question's kind are soft-maxed with a
per-kind temperature (calibration.json). In vLLM the head is expressed as an lm_head LoRA (adapter_vllm/, rank
320) and the head bias is added client-side (adapter_vllm/decision_head.json).
Limitations
- System 1 only on this card: thinking (adaptive System 2) has not been evaluated for this model.
- Knowledge & Reasoning is the weak area (GPQA 0.360, MMLU-Pro 0.642, HLE below chance); JEV-27B-VL is better on GSM8K, ANLI and BBH.
- Weaker than JEV-27B-VL at sequential game play (2048, Flappy Bird, Snake).
- The vision score is our rebuild, not the board's; image decisions are zero-shot.
- English-centric; not for high-stakes decisions without confidence gating.
Files
model.safetensors-* · config.json · tokenizer* · chat_template.jinja · generation_config.json · *_config.json
Qwen/Qwen3.5-35B-A3B, unchanged (System 2)
adapter/ System 1 LoRA (peft), rank 32
head.safetensors 264-slot decision head (fp32)
judge_config.json slot layout, verbalizer ids, read-out
calibration.json per-kind temperatures
adapter_vllm/ System 1 for vLLM: backbone LoRA + the head as an lm_head LoRA (rank 320), plus decision_head.json
serve_decide.py · serve.sh
vLLM server with POST /v1/decide next to the OpenAI endpoints
videos/ computer-use and robot-arm episodes
reports/demos/ per-episode computer-use, robot-arm and game results
NVFP4 weights (routed experts in NVFP4, everything else identical): autotrust/JEV-35B-NVFP4.
License
Apache-2.0. This repository contains the weights of Qwen/Qwen3.5-35B-A3B (Apache-2.0) unchanged, plus the System 1 adapter, decision head and calibration.
- Downloads last month
- -
Model tree for autotrust/JEV-35B
Evaluation results
- Decision Index 0.3 public index on Decision Index 0.3 public suiteself-reported59.750
