Torchcast Decision 27B

Torchcast Decision 27B: a probability for every option, every question answered together

Decision Index 0.2.1: 65.00, +7.09 over Jev 1.13. One yes/no question in 33 ms on one H100.

Torchcast Decision 27B answers typed questions about a state: pick one of several options (choice), yes or no (noul), or a rating on an ordinal scale (score). Every question comes back with a probability for each option, and all questions of a request are answered together, in a single forward pass for typical requests. Requests and responses use the TypeSafe-compatible /v1/systemone format, so clients written for the choice, noul and score questions of Jev (TypeSafe AI) should work without changes. Context up to 262,144 tokens (256K); the evaluations below are text-only.

Results

Torchcast Decision 27B
Decision Index 0.2.1 65.00
Latency, one yes/no question 33.0 ms
Latency, one request with three questions 49.0 ms
Input text or JSON, up to 262,144 tokens

Decision Index: the official kit (apolinario/decision-index, release 0.2.1, commit 87d4650) on all 155,390 suite rows, coverage 1.0, scored by the kit. The model runs on one H100 80GB; rows were answered in bf16 by our batched runner (eval/di_batch.py in the GitHub repository), with the suite split between two such GPUs, each running a full copy of the model, to halve the wall-clock time. Latency is end to end over HTTP on one H100 80GB in bf16, one request at a time, mean of 100 after 20 warm-up requests (p95 34.6 ms and 52.1 ms); the three-question request is the ticket example under Usage.

Decision Index 0.2.1: Torchcast Decision 27B 65.00, StartLux-Decision-27B 63.88, Jev 1.13 57.91, with scores by area

Against Jev 1.13: ahead on 31 of 38 index benchmarks, with the largest gaps in each direction

Torchcast Decision 27B is ahead of Jev 1.13 on 31 of the 38 index benchmarks. The largest gains are on the home appliance simulator (tool use), POP909-CL, aspect-sentiment extraction (ACOS), stance (VAST), grade-school math, fact verification and contract reasoning. Jev 1.13 remains clearly ahead on hard knowledge benchmarks (GPQA Diamond, MMLU-Pro, BBH), and the Knowledge area of Torchcast Decision 27B is 5.8 points below it.

Compared with other decision models

Model Decision Index 0.2.1 Source
Torchcast Decision 27B 65.00 this card
StartLux-Decision-27B 63.88 its model card
Clef 61.21 self-reported by Cloudflare, 2026-10-01
Jev 1.13 57.91 public board
Surogate Rune 26B-A4B v3 57.44 public board
Decider chat ยท Gemma-4-31B 57.33 public board
pplx-decider-v1-27b 56.40 public board
Winnow-12B 50.02 public board

Our score is author-run with the official kit and is not on the public board; board entries were run by the board's maintainers. Public board: Decision Index 0.2.1, snapshot of 2026-09-28. Latency is not compared across models because the published figures use different hardware and paths. On the same pipeline, the base StartLux-Decision-27B scores 63.99, within 0.11 of its published 63.88.

Fast inference

The folder ships the inference package torchcast_decision/, which is the fast path:

  • requirements.txt installs the fast kernels, flash-linear-attention and causal-conv1d, and python -m torchcast_decision.check . confirms they are active. Without them transformers falls back to a much slower path, and the server refuses to start on a GPU.
  • The questions of a typical request share one forward pass; a very long state is read once in chunks first, and a choice with more than 26 options takes extra rounds. The server records CUDA graphs at start-up and replays them for short requests: on one H100 a single yes/no question takes 33.0 ms end to end and the three-question ticket request 49.0 ms.
  • For bulk work, decide_batch batches the questions of many requests together.

Serve the model with the included package as shown under Usage. torchcast_decision/ is Apache-2.0 and derived from StartLux-Decision's inference code; see its NOTICE.

Usage

Linux with an NVIDIA GPU with at least 64 GB of memory (54.7 GB of bf16 weights).

hf download torchcast-ai/torchcast-decision-27b --local-dir torchcast-decision-27b
cd torchcast-decision-27b
pip install -r requirements.txt                          # causal-conv1d may need --no-build-isolation
python -m torchcast_decision.check .                     # must print "fast kernels: active"
python -m torchcast_decision.server --model . --port 8090
curl -s localhost:8090/v1/systemone -H 'Content-Type: application/json' -d '{
  "state": {"ticket": "I was charged twice for order #4411 and the app still shows it as unpaid."},
  "questions": {
    "team":   {"type": "choice", "instructions": "Which team should handle this ticket?",
               "criteria": {"billing": "Payments, refunds and invoices",
                            "shipping": "Delivery and tracking",
                            "technical": "App, login and account problems"}},
    "urgent": {"type": "noul", "instructions": "Should this ticket be answered today?"},
    "severity": {"type": "score", "instructions": "How severe is the impact?",
                 "criteria": ["cosmetic", "annoying", "blocks the customer"]}
  }
}'

The response (abridged) carries a probability for every option:

{"answers": {
   "team":     {"type": "choice", "choice": "billing", "confidence": 0.948,
                "probabilities": {"billing": 0.965, "shipping": 0.001, "technical": 0.034}},
   "urgent":   {"type": "noul", "noul": 0.706},
   "severity": {"type": "score", "score": 1.71, "confidence": 0.565,
                "probabilities": {"0": 0.016, "1": 0.258, "2": 0.726}}},
 "usage": {"input_tokens": 290, "output_tokens": 0},
 "model": "torchcast-decision-27b"}

Or in Python, from the same folder:

from torchcast_decision import TorchcastDecision

m = TorchcastDecision(".")
answers, usage = m.decide(state, questions)
many = m.decide_batch([(state, questions), ...])

confidence follows TypeSafe's definitions (for a choice, (p_max โˆ’ 1/n) / (1 โˆ’ 1/n)); the top probability is in probabilities. Serving and evaluation code: Torchcast-AI/torchcast-decision-27b.

Training

A LoRA fine-tune of StartLux-Decision-27B, merged into the weights. Training inputs come from train and dev splits of public datasets and benchmark-format decision data built from them, with the datasets' gold labels as targets. Some sources carry non-commercial, share-alike or research-only terms; see LICENSE.

Benchmark exposure

Training inputs came only from train and dev splits of public datasets; no Decision Index test rows were used, and all training data was screened against every test row. As expected for these datasets, about 1% of training rows share a source document or question template with a test item, and one CLINC150 utterance appears in both splits of that dataset. ForecastBench training rows all resolve before the test resolution window (which starts 2026-07-01). Some are earlier instances of recurring question series that also appear in the test set; each of those resolves before the earliest forecast date of the test questions on the same event, and prediction-market questions that appear in the test set were excluded.

Intended use

Research and non-commercial use. Not intended for safety-critical, security, medical, legal, financial or employment decisions without human review; probabilities can be miscalibrated on new domains. No warranty (see LICENSE).

Limitations

  • The scores are single-run point estimates without uncertainty intervals, and Decision Index results guided checkpoint selection, so they are not an untouched holdout.
  • The Language area is slightly below the base model (73.81 against 74.52 on the same pipeline), and the Knowledge area trails Jev 1.13.
  • Evaluated on English text only; image inputs are supported by the inherited vision tower but were not evaluated here, and probabilities can be unreliable on new domains.
  • The base model's multi-token-prediction head is not included; the weights are for decisions, not speculative decoding.

License

Model weights: CC BY-NC 4.0, inherited from StartLux-Decision-27B; non-commercial use only, with attribution. Commercial use is not permitted under this license. It would require separate permission from StartLux Labs (contact@startlux.com) for StartLux-Decision-27B and from Torchcast AI for this fine-tune, and may also be restricted by the terms of the training data. StartLux-Decision-27B is a modified version of a model released by Alibaba Cloud under the Apache License 2.0; files taken unchanged from it, such as the tokenizer files, remain under that license. torchcast_decision/ is Apache-2.0 code derived from StartLux-Decision's inference code (see torchcast_decision/NOTICE). See LICENSE and NOTICE, which carry StartLux's license and notices.

Credit Torchcast AI for this checkpoint, StartLux Labs for StartLux-Decision, and the Qwen team at Alibaba Cloud for the underlying model. This project is independent of, and not endorsed by, StartLux Labs, TypeSafe AI, Alibaba Cloud, Cloudflare, the Decision Index maintainers, or any other model provider named here. Product names are trademarks of their respective owners and are used only to identify those products.

Downloads last month
858
Safetensors
Model size
27B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for torchcast-ai/torchcast-decision-27b

Quantizations
2 models

Spaces using torchcast-ai/torchcast-decision-27b 2