larkooo commited on
Commit
e099c73
·
verified ·
1 Parent(s): f092803

Add simultaneous streaming image and video comparisons

Browse files
.gitattributes CHANGED
@@ -37,3 +37,4 @@ docs/assets/demo-poster.jpg filter=lfs diff=lfs merge=lfs -text
37
  docs/assets/demo.mp4 filter=lfs diff=lfs merge=lfs -text
38
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
39
  gemma_rlcd/static/sample-street.jpg filter=lfs diff=lfs merge=lfs -text
 
 
37
  docs/assets/demo.mp4 filter=lfs diff=lfs merge=lfs -text
38
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
39
  gemma_rlcd/static/sample-street.jpg filter=lfs diff=lfs merge=lfs -text
40
+ docs/assets/live-demo.mp4 filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -26,9 +26,11 @@ Give the model one input and a set of questions. It encodes the input once, scor
26
 
27
  This download includes the complete 4-bit multimodal checkpoint and the `gemma_rlcd` runtime. The checkpoint preserves the pinned MLX quantization; the runtime implements parallel decision scoring. [Source code on GitHub](https://github.com/Larkooo/gemma-e2b-rlcd).
28
 
29
- <video controls playsinline preload="metadata" width="100%" poster="https://huggingface.co/larkooo/gemma-e2b-rlcd/resolve/main/docs/assets/demo-poster.jpg" src="https://huggingface.co/larkooo/gemma-e2b-rlcd/resolve/main/docs/assets/demo.mp4"></video>
30
 
31
- **[Watch the demo](https://huggingface.co/larkooo/gemma-e2b-rlcd/resolve/main/docs/assets/demo.mp4)** · 28 matching outputs · 4.41 s vs 15.66 s · 3.55× faster on the support-triage workload.
 
 
32
 
33
  ## Quick start
34
 
@@ -56,6 +58,8 @@ The download is approximately 3.6 GB and includes the image and audio encoders,
56
 
57
  Open **http://127.0.0.1:8787/demo** for 32, 64, or 128 checks over an image or video. Use the included street photo or upload your own media. The parallel scorer streams completed field batches; normal Gemma streams its generated JSON. Live clocks, per-check probabilities, answer differences, and downloadable events make the comparison inspectable. Both paths start together and stream side by side, sharing the resident weights with separate processors and KV caches. Timings measure concurrent completion on one GPU, including resource contention. The playground’s ordinary comparison remains sequential for isolated timings.
58
 
 
 
59
  ## How it works
60
 
61
  ```mermaid
 
26
 
27
  This download includes the complete 4-bit multimodal checkpoint and the `gemma_rlcd` runtime. The checkpoint preserves the pinned MLX quantization; the runtime implements parallel decision scoring. [Source code on GitHub](https://github.com/Larkooo/gemma-e2b-rlcd).
28
 
29
+ <video controls playsinline preload="metadata" width="100%" src="https://huggingface.co/larkooo/gemma-e2b-rlcd/resolve/main/docs/assets/live-demo.mp4"></video>
30
 
31
+ **128 visual checks, streaming live.** In this recorded simultaneous run: **11.57 s vs 54.10 s · 4.67× faster · 120/128 matching answers**. Both paths share one GPU. The full startup pause and original elapsed time are preserved.
32
+
33
+ [Download the video](docs/assets/live-demo.mp4) · [Recorded results](reports/live-visual-demo.json)
34
 
35
  ## Quick start
36
 
 
58
 
59
  Open **http://127.0.0.1:8787/demo** for 32, 64, or 128 checks over an image or video. Use the included street photo or upload your own media. The parallel scorer streams completed field batches; normal Gemma streams its generated JSON. Live clocks, per-check probabilities, answer differences, and downloadable events make the comparison inspectable. Both paths start together and stream side by side, sharing the resident weights with separate processors and KV caches. Timings measure concurrent completion on one GPU, including resource contention. The playground’s ordinary comparison remains sequential for isolated timings.
60
 
61
+ The first answer follows input preparation and a full multimodal prefill. The UI shows these stages, first-token time, and first-decision time separately. Repeated runs reuse compiled field definitions; media, input KV state, and answers are recomputed. Use **Focus view** to see all 128 outputs together.
62
+
63
  ## How it works
64
 
65
  ```mermaid
checkpoint-provenance.json CHANGED
@@ -1,7 +1,7 @@
1
  {
2
  "repository": "larkooo/gemma-e2b-rlcd",
3
  "runtime_repository": "https://github.com/Larkooo/gemma-e2b-rlcd",
4
- "runtime_commit": "e78b79605ac4e37b776727935e2de7dc8f259a81",
5
  "checkpoint_repository": "mlx-community/gemma-4-e2b-it-4bit",
6
  "checkpoint_revision": "238767527555cb75a05732a84dff5d6ba0dd6809",
7
  "checkpoint_modified": false,
 
1
  {
2
  "repository": "larkooo/gemma-e2b-rlcd",
3
  "runtime_repository": "https://github.com/Larkooo/gemma-e2b-rlcd",
4
+ "runtime_commit": "4ae799d204d8b2f4a802b756b4236541af0de468",
5
  "checkpoint_repository": "mlx-community/gemma-4-e2b-it-4bit",
6
  "checkpoint_revision": "238767527555cb75a05732a84dff5d6ba0dd6809",
7
  "checkpoint_modified": false,
docs/architecture.md CHANGED
@@ -68,3 +68,9 @@ Video sampling can miss brief events. The limits describe the current serving co
68
  | `--backend head` | `DecisionHeadBackend` | Candidate-conditioned head; requires a checkpoint |
69
 
70
  The cached path assigns tokenizer-verified codes to options and scores per-question suffixes against shared state. Historical answer-code and catalog measurements are retained for reproducing those experiments. The [training guide](training.md) describes the head separately.
 
 
 
 
 
 
 
68
  | `--backend head` | `DecisionHeadBackend` | Candidate-conditioned head; requires a checkpoint |
69
 
70
  The cached path assigns tokenizer-verified codes to options and scores per-question suffixes against shared state. Historical answer-code and catalog measurements are retained for reproducing those experiments. The [training guide](training.md) describes the head separately.
71
+
72
+ ## Streaming startup
73
+
74
+ The visual demo starts scoring and normal generation together on separate worker streams with separate processors and KV caches. Both process the complete media and question schema before producing answers; the UI reports input preparation, prefill, first token, and first completed decision separately. End-to-end concurrent timings include GPU contention.
75
+
76
+ The scorer encodes the common prompt once for field compilation and keeps up to eight compiled schemas in an LRU cache. Cache keys include the complete prompt and candidate definitions. This cache contains only tokenized field definitions; every request recomputes media features, input KV state, and answer probabilities. Boundary-merge validation and complete-candidate scoring remain unchanged.
docs/assets/live-demo.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:803559adc8667d1cbf92070943d005ce8659a2e4e2498a83f84b74ac65d80f8b
3
+ size 1843915
gemma_rlcd/comparison.py CHANGED
@@ -144,13 +144,19 @@ def prepare_generation(backend, state: State, questions: dict) -> tuple[str, dic
144
  return prompt, inputs
145
 
146
 
147
- def generate_answers(backend, state: State, questions: dict, on_token=None) -> dict:
 
 
148
  from mlx_vlm import generate, stream_generate
149
 
150
  started = time.perf_counter()
 
 
151
  prompt, inputs = prepare_generation(backend, state, questions)
152
  budget = output_budget(backend.tokenizer, questions)
153
  prepared = time.perf_counter()
 
 
154
  options = dict(
155
  **inputs,
156
  max_tokens=budget,
@@ -166,6 +172,8 @@ def generate_answers(backend, state: State, questions: dict, on_token=None) -> d
166
  parts = []
167
  generated = None
168
  for chunk in stream_generate(backend.model, backend.processor, prompt, **options):
 
 
169
  generated = chunk
170
  parts.append(chunk.text)
171
  on_token(chunk.text, chunk.generation_tokens)
@@ -266,10 +274,21 @@ def compare(
266
  }
267
  )
268
 
 
 
 
 
 
 
 
 
 
 
 
269
  if method == "parallel":
270
  engine = DecisionEngine(worker_backend)
271
  output = (
272
- engine.system_one(state, questions, on_answer=on_answer)
273
  if emit
274
  else engine.system_one(state, questions)
275
  )
@@ -282,7 +301,9 @@ def compare(
282
  )
283
  else:
284
  output = (
285
- generate_answers(worker_backend, state, questions, on_token=on_token)
 
 
286
  if emit
287
  else generate_answers(worker_backend, state, questions)
288
  )
 
144
  return prompt, inputs
145
 
146
 
147
+ def generate_answers(
148
+ backend, state: State, questions: dict, on_token=None, on_progress=None
149
+ ) -> dict:
150
  from mlx_vlm import generate, stream_generate
151
 
152
  started = time.perf_counter()
153
+ if on_progress:
154
+ on_progress("preparing", None)
155
  prompt, inputs = prepare_generation(backend, state, questions)
156
  budget = output_budget(backend.tokenizer, questions)
157
  prepared = time.perf_counter()
158
+ if on_progress:
159
+ on_progress("prefill", int(inputs["input_ids"].shape[-1]))
160
  options = dict(
161
  **inputs,
162
  max_tokens=budget,
 
172
  parts = []
173
  generated = None
174
  for chunk in stream_generate(backend.model, backend.processor, prompt, **options):
175
+ if generated is None and on_progress:
176
+ on_progress("generating", chunk.prompt_tokens)
177
  generated = chunk
178
  parts.append(chunk.text)
179
  on_token(chunk.text, chunk.generation_tokens)
 
274
  }
275
  )
276
 
277
+ def on_progress(stage, input_tokens):
278
+ emit(
279
+ {
280
+ "type": "progress",
281
+ "method": method,
282
+ "stage": stage,
283
+ "input_tokens": input_tokens,
284
+ "seconds": media_seconds + time.perf_counter() - started,
285
+ }
286
+ )
287
+
288
  if method == "parallel":
289
  engine = DecisionEngine(worker_backend)
290
  output = (
291
+ engine.system_one(state, questions, on_answer=on_answer, on_progress=on_progress)
292
  if emit
293
  else engine.system_one(state, questions)
294
  )
 
301
  )
302
  else:
303
  output = (
304
+ generate_answers(
305
+ worker_backend, state, questions, on_token=on_token, on_progress=on_progress
306
+ )
307
  if emit
308
  else generate_answers(worker_backend, state, questions)
309
  )
gemma_rlcd/core.py CHANGED
@@ -216,7 +216,9 @@ class DecisionEngine:
216
  def decide(self, state: State, question: Question) -> dict:
217
  return self.system_one(state, {"answer": question})["answers"]["answer"]
218
 
219
- def system_one(self, state: State, questions: Mapping[str, Question], on_answer=None) -> dict:
 
 
220
  if not questions:
221
  raise ValueError("At least one question is required")
222
  jobs = []
@@ -268,11 +270,12 @@ class DecisionEngine:
268
  )
269
 
270
  if question_score is not None:
271
- scores = (
272
- question_score(state, questions, on_scores=completed_scores)
273
- if on_answer
274
- else question_score(state, questions)
275
- )
 
276
  elif batch_score is not None:
277
  scores = batch_score(state, requests)
278
  else:
 
216
  def decide(self, state: State, question: Question) -> dict:
217
  return self.system_one(state, {"answer": question})["answers"]["answer"]
218
 
219
+ def system_one(
220
+ self, state: State, questions: Mapping[str, Question], on_answer=None, on_progress=None
221
+ ) -> dict:
222
  if not questions:
223
  raise ValueError("At least one question is required")
224
  jobs = []
 
270
  )
271
 
272
  if question_score is not None:
273
+ callbacks = {}
274
+ if on_answer:
275
+ callbacks["on_scores"] = completed_scores
276
+ if on_progress:
277
+ callbacks["on_progress"] = on_progress
278
+ scores = question_score(state, questions, **callbacks)
279
  elif batch_score is not None:
280
  scores = batch_score(state, requests)
281
  else:
gemma_rlcd/json_backend.py CHANGED
@@ -2,6 +2,7 @@
2
 
3
  import math
4
  import time
 
5
 
6
  from .cached_backend import CachedMLXBackend, PreparedState
7
  from .comparison import prepare_generation
@@ -67,12 +68,28 @@ class JSONMLXBackend(CachedMLXBackend):
67
  on_scores(completed)
68
  return output, batches
69
 
70
- def score_questions(self, state, questions, on_scores=None):
71
  started = time.perf_counter()
 
 
72
  prompt, inputs = prepare_generation(self, state, questions)
73
- fields = [
74
- compile_field(self.tokenizer, field, prompt) for field in candidate_fields(questions)
75
- ]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
76
  prefix_tokens = int(inputs["input_ids"].shape[1])
77
  if any(
78
  prefix_tokens + len(field.prefix) + max(map(len, field.candidates))
@@ -84,8 +101,12 @@ class JSONMLXBackend(CachedMLXBackend):
84
  )
85
  prepared = PreparedState(inputs, [], prefix_tokens)
86
  processed = time.perf_counter()
 
 
87
  prefix_cache = self.prefill(prepared)
88
  prefilled = time.perf_counter()
 
 
89
  self.last_stats = {}
90
  scores = [None] * len(fields)
91
  single = [
@@ -136,6 +157,7 @@ class JSONMLXBackend(CachedMLXBackend):
136
  "execution": "shared_json_prefix_gpu_batched_fields",
137
  "prefix_prefills": 1,
138
  "prefix_tokens": prefix_tokens,
 
139
  "primitive_fields": len(fields),
140
  "question_suffix_tokens": [len(field.prefix) for field in fields],
141
  "candidate_token_lengths": [
 
2
 
3
  import math
4
  import time
5
+ from collections import OrderedDict
6
 
7
  from .cached_backend import CachedMLXBackend, PreparedState
8
  from .comparison import prepare_generation
 
68
  on_scores(completed)
69
  return output, batches
70
 
71
+ def score_questions(self, state, questions, on_scores=None, on_progress=None):
72
  started = time.perf_counter()
73
+ if on_progress:
74
+ on_progress("preparing", None)
75
  prompt, inputs = prepare_generation(self, state, questions)
76
+ candidates = tuple(candidate_fields(questions))
77
+ key = (prompt, candidates)
78
+ if not hasattr(self, "_field_cache"):
79
+ self._field_cache = OrderedDict()
80
+ fields = self._field_cache.get(key)
81
+ schema_cache_hit = fields is not None
82
+ if fields is None:
83
+ base_ids = self.tokenizer.encode(prompt, add_special_tokens=False)
84
+ fields = [
85
+ compile_field(self.tokenizer, field, prompt, base_ids=base_ids)
86
+ for field in candidates
87
+ ]
88
+ self._field_cache[key] = fields
89
+ if len(self._field_cache) > 8:
90
+ self._field_cache.popitem(last=False)
91
+ else:
92
+ self._field_cache.move_to_end(key)
93
  prefix_tokens = int(inputs["input_ids"].shape[1])
94
  if any(
95
  prefix_tokens + len(field.prefix) + max(map(len, field.candidates))
 
101
  )
102
  prepared = PreparedState(inputs, [], prefix_tokens)
103
  processed = time.perf_counter()
104
+ if on_progress:
105
+ on_progress("prefill", prefix_tokens)
106
  prefix_cache = self.prefill(prepared)
107
  prefilled = time.perf_counter()
108
+ if on_progress:
109
+ on_progress("scoring", prefix_tokens)
110
  self.last_stats = {}
111
  scores = [None] * len(fields)
112
  single = [
 
157
  "execution": "shared_json_prefix_gpu_batched_fields",
158
  "prefix_prefills": 1,
159
  "prefix_tokens": prefix_tokens,
160
+ "schema_cache_hit": schema_cache_hit,
161
  "primitive_fields": len(fields),
162
  "question_suffix_tokens": [len(field.prefix) for field in fields],
163
  "candidate_token_lengths": [
gemma_rlcd/json_scoring.py CHANGED
@@ -36,11 +36,12 @@ def candidate_fields(questions: dict) -> list[JSONField]:
36
  return fields
37
 
38
 
39
- def compile_field(tokenizer, field: JSONField, prompt: str) -> FieldTokens:
40
  # Whitespace is part of the model's answer context. In particular, scoring
41
  # digits immediately after ':' instead of ': ' can score a whitespace slot.
42
  opening = "".join("{" + json.dumps(key, ensure_ascii=False) + ": " for key in field.path)
43
- base_ids = tokenizer.encode(prompt, add_special_tokens=False)
 
44
  sequences = []
45
  for value in field.values:
46
  suffix = opening + json.dumps(value, ensure_ascii=False)
 
36
  return fields
37
 
38
 
39
+ def compile_field(tokenizer, field: JSONField, prompt: str, *, base_ids=None) -> FieldTokens:
40
  # Whitespace is part of the model's answer context. In particular, scoring
41
  # digits immediately after ':' instead of ': ' can score a whitespace slot.
42
  opening = "".join("{" + json.dumps(key, ensure_ascii=False) + ": " for key in field.path)
43
+ if base_ids is None:
44
+ base_ids = tokenizer.encode(prompt, add_special_tokens=False)
45
  sequences = []
46
  for value in field.values:
47
  suffix = opening + json.dumps(value, ensure_ascii=False)
gemma_rlcd/static/demo.css CHANGED
@@ -1 +1,3 @@
1
  :root{font-family:-apple-system,BlinkMacSystemFont,"Segoe UI",sans-serif;color:#1d2924;background:#f6f7f4;--green:#23674e;--muted:#69766d;--line:#dfe5dc}*{box-sizing:border-box}body{margin:0;font-size:14px;line-height:1.5}header{display:flex;align-items:center;justify-content:space-between;gap:20px;padding:20px 36px;border-bottom:1px solid var(--line);background:#fcfdfb}.brand{font-weight:650;font-size:17px;color:inherit;text-decoration:none}nav{display:flex;gap:24px;align-items:center;font-size:12px}a{color:var(--green)}#model-status{color:var(--muted)}main{max-width:1560px;margin:auto;padding:35px 36px 70px}h1,h2,h3,p{margin:0}h1{font-size:36px;letter-spacing:-1.4px;font-weight:580;line-height:1.25;margin:8px 0}h1>span{color:var(--green)}h2{font-size:15px;font-weight:650}h3{font-size:14px;font-weight:620}.eyebrow{font-size:10px;font-weight:650;letter-spacing:1.7px;color:var(--muted)}.heading{display:flex;align-items:center;justify-content:space-between;gap:20px;margin-bottom:28px}.heading p{color:var(--muted);font-size:13px}.workspace{display:grid;grid-template-columns:minmax(280px,.75fr) minmax(0,1.6fr);gap:24px;align-items:start}.input-panel,.live-panel,.wall{border:1px solid var(--line);background:white;border-radius:10px;padding:22px}.section-title{display:flex;align-items:center;justify-content:space-between;gap:12px;margin-bottom:16px}.small{font-size:11px;color:var(--muted);line-height:1.65}.preview{height:260px;background:#eff2ed;border-radius:6px;overflow:hidden;display:grid;place-items:center;cursor:pointer}.preview img,.preview video{width:100%;height:100%;max-height:260px;object-fit:contain}.preview.dragging{outline:3px solid #7aab8f}.media-controls{display:flex;gap:12px;align-items:center;margin-top:12px}.file-button{font-size:12px;font-weight:550;color:var(--green);cursor:pointer;position:relative;flex-shrink:0}.file-button input{position:absolute;inset:0;opacity:0;width:100%;cursor:pointer}#file-name{font-size:10px;color:var(--muted);overflow:hidden;text-overflow:ellipsis;white-space:nowrap}#credit{margin-top:5px;min-height:18px}label:not(.file-button){display:block;font-size:11px;font-weight:550;margin:15px 0 6px}button,input,textarea,select{font:inherit;color:inherit}button,select{cursor:pointer}button{border:1px solid transparent;border-radius:6px;padding:10px 14px;font-size:12px;font-weight:550}button:disabled{opacity:.5;cursor:default}.primary{background:var(--green);color:white}.primary:hover:not(:disabled){background:#18513b}.secondary{background:white;border-color:var(--line)}.text-button{padding:2px;background:none;color:var(--green);font-size:11px}.input-panel select,.input-panel textarea{width:100%;border:1px solid var(--line);border-radius:6px;padding:10px;background:#fcfdfb;font-size:12px}.input-panel textarea{resize:vertical;line-height:1.5;margin-bottom:9px}.actions{display:flex;gap:10px;margin:18px 0 9px}.actions .primary{flex:1}.engines{display:grid;grid-template-columns:1fr 1fr;gap:24px}.engine{min-width:0}.engine-title{display:flex;align-items:center;justify-content:space-between;gap:8px}.engine-title>span{font-size:10px;color:var(--muted)}.parallel h3,.parallel .clock{color:var(--green)}.clock{display:block;font-size:50px;font-weight:450;letter-spacing:-2px;font-variant-numeric:tabular-nums;margin-top:8px}.clock>span{font-size:15px;letter-spacing:0;color:var(--muted);margin-left:7px}.progress{height:4px;background:#edf0e9;border-radius:3px;overflow:hidden;margin-top:12px}.progress i{display:block;height:100%;width:0;background:var(--green)}.normal .progress i{background:#4b5850}.engine-stats{display:flex;justify-content:space-between;gap:5px;color:var(--muted);font-size:10px;margin-top:7px}.engine>.small{margin-top:10px}.verdict{border-top:1px solid var(--line);border-bottom:1px solid var(--line);margin:22px 0;padding:15px 0;min-height:54px;font-size:13px;color:var(--green)}.verdict strong{font-size:22px;font-weight:600;margin-right:8px}.streams{display:grid;grid-template-columns:1fr 1fr;gap:20px;min-width:0}.streams>div{min-width:0}.stream-label{display:flex;justify-content:space-between;gap:10px;font-size:11px;color:var(--muted)}.live-dot{font-size:8px;letter-spacing:1px;color:var(--green)}pre{height:235px;margin:10px 0 0;background:#f6f8f3;border:1px solid #e9eee4;border-radius:6px;padding:12px;white-space:pre-wrap;overflow-wrap:anywhere;overflow:auto;font:11px/1.7 ui-monospace,SFMono-Regular,Menlo,monospace;color:#3e6550}.normal-stream{color:#46504a}.method-note{font-size:10px;color:var(--muted);line-height:1.8;margin-top:17px}.wall{margin-top:24px}.wall-heading{display:flex;justify-content:space-between;align-items:center;gap:20px;margin-bottom:20px}.wall-heading p{margin-top:5px}.filters{display:flex;align-items:center;gap:12px;flex-shrink:0}.filters select{font-size:11px;padding:8px;border:1px solid var(--line);border-radius:6px;background:white}.group{margin-top:19px}.group:first-child{margin-top:0}.group-title{font-size:11px;color:var(--muted);font-weight:550;margin-bottom:9px}.check-grid{display:grid;grid-template-columns:repeat(8,minmax(0,1fr));gap:7px}.check{border:1px solid #e3e8de;border-radius:5px;padding:8px 9px;min-width:0;min-height:59px;transition:background .2s,border-color .2s}.check-name{display:block;font-size:10px;white-space:nowrap;overflow:hidden;text-overflow:ellipsis;color:#53624f}.values{display:flex;gap:12px;justify-content:space-between;margin-top:6px;font:10px ui-monospace,monospace;color:#8a9585}.values b{font-weight:500}.check.detected{background:#f0f7ec;border-color:#c7dcc1}.check.different{background:#fff6e9;border-color:#e9c998}.values .yes{color:#23674e}.values .no{color:#727d70}.check.flash{animation:arrive .55s ease-out}.details{font-size:11px;color:var(--muted);margin-top:20px;max-width:1000px}.details p{margin-top:10px}.details summary{cursor:pointer}#error{margin-top:12px;padding:11px;border:1px solid #eccfc4;background:#fff6f0;color:#963e2e;font-size:12px;border-radius:6px}.sr-only{position:absolute;width:1px;height:1px;overflow:hidden;clip:rect(0,0,0,0)}[hidden]{display:none!important}:focus-visible{outline:3px solid #75a58c;outline-offset:3px}@keyframes arrive{0%{background:#dceccd}100%{}}@media(prefers-reduced-motion:reduce){*{animation:none!important;transition:none!important}}@media(max-width:1200px){.check-grid{grid-template-columns:repeat(6,minmax(0,1fr))}.workspace{grid-template-columns:minmax(275px,.8fr) minmax(0,1.4fr)}.clock{font-size:43px}.engine-title{align-items:flex-start;flex-direction:column;gap:2px}}@media(max-width:850px){header{padding:17px 20px}main{padding:25px 20px 50px}.workspace{grid-template-columns:1fr}.preview{height:280px}.preview img,.preview video{max-height:280px}.input-panel{display:grid;grid-template-columns:1fr 1fr;column-gap:20px}.input-panel>*{grid-column:1/-1}.check-grid{grid-template-columns:repeat(4,minmax(0,1fr))}.engine-title{flex-direction:row}.wall-heading{align-items:flex-start;flex-direction:column}h1{font-size:32px}}@media(max-width:520px){header{align-items:flex-start}.brand{font-size:14px}nav{flex-direction:column;gap:3px;align-items:flex-end;font-size:10px}main{padding:22px 13px 40px}.heading{align-items:flex-start}.heading>button{padding:8px;font-size:10px;white-space:nowrap}h1{font-size:27px}.heading p{font-size:12px}.input-panel,.live-panel,.wall{padding:16px}.engines{gap:17px}.engine-title{align-items:flex-start;flex-direction:column}.clock{font-size:40px}.streams{gap:12px}pre{height:215px;font-size:10px;padding:9px}.check-grid{grid-template-columns:repeat(3,minmax(0,1fr))}.check{padding:7px}.values{gap:6px}.preview{height:230px}.preview img,.preview video{max-height:230px}.filters{flex-wrap:wrap}.wall-heading p{font-size:10px}}
 
 
 
1
  :root{font-family:-apple-system,BlinkMacSystemFont,"Segoe UI",sans-serif;color:#1d2924;background:#f6f7f4;--green:#23674e;--muted:#69766d;--line:#dfe5dc}*{box-sizing:border-box}body{margin:0;font-size:14px;line-height:1.5}header{display:flex;align-items:center;justify-content:space-between;gap:20px;padding:20px 36px;border-bottom:1px solid var(--line);background:#fcfdfb}.brand{font-weight:650;font-size:17px;color:inherit;text-decoration:none}nav{display:flex;gap:24px;align-items:center;font-size:12px}a{color:var(--green)}#model-status{color:var(--muted)}main{max-width:1560px;margin:auto;padding:35px 36px 70px}h1,h2,h3,p{margin:0}h1{font-size:36px;letter-spacing:-1.4px;font-weight:580;line-height:1.25;margin:8px 0}h1>span{color:var(--green)}h2{font-size:15px;font-weight:650}h3{font-size:14px;font-weight:620}.eyebrow{font-size:10px;font-weight:650;letter-spacing:1.7px;color:var(--muted)}.heading{display:flex;align-items:center;justify-content:space-between;gap:20px;margin-bottom:28px}.heading p{color:var(--muted);font-size:13px}.workspace{display:grid;grid-template-columns:minmax(280px,.75fr) minmax(0,1.6fr);gap:24px;align-items:start}.input-panel,.live-panel,.wall{border:1px solid var(--line);background:white;border-radius:10px;padding:22px}.section-title{display:flex;align-items:center;justify-content:space-between;gap:12px;margin-bottom:16px}.small{font-size:11px;color:var(--muted);line-height:1.65}.preview{height:260px;background:#eff2ed;border-radius:6px;overflow:hidden;display:grid;place-items:center;cursor:pointer}.preview img,.preview video{width:100%;height:100%;max-height:260px;object-fit:contain}.preview.dragging{outline:3px solid #7aab8f}.media-controls{display:flex;gap:12px;align-items:center;margin-top:12px}.file-button{font-size:12px;font-weight:550;color:var(--green);cursor:pointer;position:relative;flex-shrink:0}.file-button input{position:absolute;inset:0;opacity:0;width:100%;cursor:pointer}#file-name{font-size:10px;color:var(--muted);overflow:hidden;text-overflow:ellipsis;white-space:nowrap}#credit{margin-top:5px;min-height:18px}label:not(.file-button){display:block;font-size:11px;font-weight:550;margin:15px 0 6px}button,input,textarea,select{font:inherit;color:inherit}button,select{cursor:pointer}button{border:1px solid transparent;border-radius:6px;padding:10px 14px;font-size:12px;font-weight:550}button:disabled{opacity:.5;cursor:default}.primary{background:var(--green);color:white}.primary:hover:not(:disabled){background:#18513b}.secondary{background:white;border-color:var(--line)}.text-button{padding:2px;background:none;color:var(--green);font-size:11px}.input-panel select,.input-panel textarea{width:100%;border:1px solid var(--line);border-radius:6px;padding:10px;background:#fcfdfb;font-size:12px}.input-panel textarea{resize:vertical;line-height:1.5;margin-bottom:9px}.actions{display:flex;gap:10px;margin:18px 0 9px}.actions .primary{flex:1}.engines{display:grid;grid-template-columns:1fr 1fr;gap:24px}.engine{min-width:0}.engine-title{display:flex;align-items:center;justify-content:space-between;gap:8px}.engine-title>span{font-size:10px;color:var(--muted)}.parallel h3,.parallel .clock{color:var(--green)}.clock{display:block;font-size:50px;font-weight:450;letter-spacing:-2px;font-variant-numeric:tabular-nums;margin-top:8px}.clock>span{font-size:15px;letter-spacing:0;color:var(--muted);margin-left:7px}.progress{height:4px;background:#edf0e9;border-radius:3px;overflow:hidden;margin-top:12px}.progress i{display:block;height:100%;width:0;background:var(--green)}.normal .progress i{background:#4b5850}.engine-stats{display:flex;justify-content:space-between;gap:5px;color:var(--muted);font-size:10px;margin-top:7px}.engine>.small{margin-top:10px}.verdict{border-top:1px solid var(--line);border-bottom:1px solid var(--line);margin:22px 0;padding:15px 0;min-height:54px;font-size:13px;color:var(--green)}.verdict strong{font-size:22px;font-weight:600;margin-right:8px}.streams{display:grid;grid-template-columns:1fr 1fr;gap:20px;min-width:0}.streams>div{min-width:0}.stream-label{display:flex;justify-content:space-between;gap:10px;font-size:11px;color:var(--muted)}.live-dot{font-size:8px;letter-spacing:1px;color:var(--green)}pre{height:235px;margin:10px 0 0;background:#f6f8f3;border:1px solid #e9eee4;border-radius:6px;padding:12px;white-space:pre-wrap;overflow-wrap:anywhere;overflow:auto;font:11px/1.7 ui-monospace,SFMono-Regular,Menlo,monospace;color:#3e6550}.normal-stream{color:#46504a}.method-note{font-size:10px;color:var(--muted);line-height:1.8;margin-top:17px}.wall{margin-top:24px}.wall-heading{display:flex;justify-content:space-between;align-items:center;gap:20px;margin-bottom:20px}.wall-heading p{margin-top:5px}.filters{display:flex;align-items:center;gap:12px;flex-shrink:0}.filters select{font-size:11px;padding:8px;border:1px solid var(--line);border-radius:6px;background:white}.group{margin-top:19px}.group:first-child{margin-top:0}.group-title{font-size:11px;color:var(--muted);font-weight:550;margin-bottom:9px}.check-grid{display:grid;grid-template-columns:repeat(8,minmax(0,1fr));gap:7px}.check{border:1px solid #e3e8de;border-radius:5px;padding:8px 9px;min-width:0;min-height:59px;transition:background .2s,border-color .2s}.check-name{display:block;font-size:10px;white-space:nowrap;overflow:hidden;text-overflow:ellipsis;color:#53624f}.values{display:flex;gap:12px;justify-content:space-between;margin-top:6px;font:10px ui-monospace,monospace;color:#8a9585}.values b{font-weight:500}.check.detected{background:#f0f7ec;border-color:#c7dcc1}.check.different{background:#fff6e9;border-color:#e9c998}.values .yes{color:#23674e}.values .no{color:#727d70}.check.flash{animation:arrive .55s ease-out}.details{font-size:11px;color:var(--muted);margin-top:20px;max-width:1000px}.details p{margin-top:10px}.details summary{cursor:pointer}#error{margin-top:12px;padding:11px;border:1px solid #eccfc4;background:#fff6f0;color:#963e2e;font-size:12px;border-radius:6px}.sr-only{position:absolute;width:1px;height:1px;overflow:hidden;clip:rect(0,0,0,0)}[hidden]{display:none!important}:focus-visible{outline:3px solid #75a58c;outline-offset:3px}@keyframes arrive{0%{background:#dceccd}100%{}}@media(prefers-reduced-motion:reduce){*{animation:none!important;transition:none!important}}@media(max-width:1200px){.check-grid{grid-template-columns:repeat(6,minmax(0,1fr))}.workspace{grid-template-columns:minmax(275px,.8fr) minmax(0,1.4fr)}.clock{font-size:43px}.engine-title{align-items:flex-start;flex-direction:column;gap:2px}}@media(max-width:850px){header{padding:17px 20px}main{padding:25px 20px 50px}.workspace{grid-template-columns:1fr}.preview{height:280px}.preview img,.preview video{max-height:280px}.input-panel{display:grid;grid-template-columns:1fr 1fr;column-gap:20px}.input-panel>*{grid-column:1/-1}.check-grid{grid-template-columns:repeat(4,minmax(0,1fr))}.engine-title{flex-direction:row}.wall-heading{align-items:flex-start;flex-direction:column}h1{font-size:32px}}@media(max-width:520px){header{align-items:flex-start}.brand{font-size:14px}nav{flex-direction:column;gap:3px;align-items:flex-end;font-size:10px}main{padding:22px 13px 40px}.heading{align-items:flex-start}.heading>button{padding:8px;font-size:10px;white-space:nowrap}h1{font-size:27px}.heading p{font-size:12px}.input-panel,.live-panel,.wall{padding:16px}.engines{gap:17px}.engine-title{align-items:flex-start;flex-direction:column}.clock{font-size:40px}.streams{gap:12px}pre{height:215px;font-size:10px;padding:9px}.check-grid{grid-template-columns:repeat(3,minmax(0,1fr))}.check{padding:7px}.values{gap:6px}.preview{height:230px}.preview img,.preview video{max-height:230px}.filters{flex-wrap:wrap}.wall-heading p{font-size:10px}}
2
+
3
+ .view-actions{display:flex;gap:8px;flex-shrink:0}.engine>#normal-first{margin-top:2px}.focused header{padding:12px 26px}.focused main{max-width:1920px;padding:16px 26px}.focused .heading{margin-bottom:16px}.focused h1{font-size:28px}.focused .heading .eyebrow,.focused .heading p,.focused .input-panel textarea,.focused .input-panel label[for="instructions"],.focused .input-panel>p.small:not(#credit):not(#run-note),.focused .details{display:none}.focused .input-panel,.focused .live-panel,.focused .wall{padding:16px}.focused .workspace{grid-template-columns:320px minmax(0,1fr);gap:18px}.focused .preview,.focused .preview img,.focused .preview video{height:185px;max-height:185px}.focused .clock{font-size:42px;margin-top:0}.focused .verdict{margin:13px 0;padding:10px 0;min-height:48px}.focused pre{height:155px}.focused .wall{margin-top:16px}.focused .wall-heading{margin-bottom:10px}.focused #output-grid{display:grid;grid-template-columns:repeat(4,minmax(0,1fr));gap:20px}.focused .group{margin:0}.focused .check-grid{grid-template-columns:repeat(4,minmax(0,1fr));gap:5px}.focused .check{min-height:43px;padding:5px 7px}.focused .values{margin-top:4px}.focused .actions{margin-top:12px}.focused .method-note{margin-top:10px}@media(max-width:1200px){.focused #output-grid{grid-template-columns:1fr 1fr}}@media(max-width:850px){.focused .workspace{grid-template-columns:1fr}.focused #output-grid{grid-template-columns:1fr}.view-actions{flex-direction:column}.focused main{padding:16px}}
gemma_rlcd/static/demo.html CHANGED
@@ -8,7 +8,7 @@
8
  <body>
9
  <header><a class="brand" href="/demo">Gemma E2B RLCD</a><nav><span id="model-status" role="status">Connecting…</span><a href="/">Open playground ↗</a></nav></header>
10
  <main>
11
- <div class="heading"><div><span class="eyebrow">LIVE MULTIMODAL COMPARISON</span><h1>One scene. <span id="headline-count">128</span> decisions.</h1><p>Both start together. Watch batched decisions race against normal Gemma’s streamed JSON.</p></div><button id="export" class="secondary" disabled>Export run ↓</button></div>
12
  <div class="workspace">
13
  <section class="input-panel" aria-labelledby="input-title">
14
  <div class="section-title"><h2 id="input-title">The evidence</h2><button id="sample" class="text-button">Use sample photo</button></div>
@@ -25,8 +25,8 @@
25
  <section class="live-panel" aria-labelledby="live-title">
26
  <div class="section-title"><h2 id="live-title">Live output</h2><span id="run-phase" class="small">Ready</span></div>
27
  <div class="engines">
28
- <article class="engine parallel"><div class="engine-title"><h3>Parallel scorer</h3><span id="parallel-state">Waiting</span></div><strong class="clock" id="parallel-clock">0.00<span>s</span></strong><div class="progress"><i id="parallel-progress"></i></div><div class="engine-stats"><span id="parallel-count">0 / 128 decisions</span><span>GPU batches</span></div><p id="parallel-first" class="small">First answer —</p></article>
29
- <article class="engine normal"><div class="engine-title"><h3>Normal Gemma</h3><span id="normal-state">Waiting</span></div><strong class="clock" id="normal-clock">0.00<span>s</span></strong><div class="progress"><i id="normal-progress"></i></div><div class="engine-stats"><span id="normal-count">0 / 128 decisions</span><span id="token-count">0 tokens</span></div><p id="normal-first" class="small">First answer —</p></article>
30
  </div>
31
  <div id="verdict" class="verdict" aria-live="polite">Upload a scene or try the sample. Results are measured live.</div>
32
  <div class="streams">
@@ -37,7 +37,7 @@
37
  </section>
38
  </div>
39
  <section class="wall" aria-labelledby="wall-title"><div class="wall-heading"><div><h2 id="wall-title">Every decision, as it lands</h2><p class="small">A check is “yes” when visually established. P = parallel probability of yes · G = Gemma’s boolean. Streaming JSON is provisional until validated.</p></div><div class="filters"><label for="filter" class="sr-only">Filter decisions</label><select id="filter"><option value="all">All checks</option><option value="yes">Detected by either</option><option value="different">Different answers</option></select><span id="agreement" class="small"></span></div></div><div id="output-grid"></div><p id="no-matches" class="small" hidden>No completed checks match this filter.</p></section>
40
- <details class="details"><summary>What is being measured?</summary><p>Both paths process the same complete image or sampled video, instructions, and label descriptions. The parallel scorer returns probabilities as field batches finish. Normal Gemma is asked to generate one compact JSON object with boolean decisions, without explanations or probability prose.</p><p>The timers include input preparation and inference, with shared upload decoding added equally. They exclude model loading, upload transfer, and allocator reset. This is one simultaneous run per path; GPU contention and first-use effects can affect timing. The result measures completion time while both are running, not isolated throughput. Matching answers measures agreement, not correctness. Invalid normal output is shown without a speedup claim.</p><p>Video processing targets one frame per second with a 32-frame cap. Long media plus 128 questions can exceed the 8,192-token input limit; reduce the number of checks or use a shorter clip. Media is never silently trimmed.</p></details>
41
  </main>
42
  </body>
43
  </html>
 
8
  <body>
9
  <header><a class="brand" href="/demo">Gemma E2B RLCD</a><nav><span id="model-status" role="status">Connecting…</span><a href="/">Open playground ↗</a></nav></header>
10
  <main>
11
+ <div class="heading"><div><span class="eyebrow">LIVE MULTIMODAL COMPARISON</span><h1>One scene. <span id="headline-count">128</span> decisions.</h1><p>Both start together. Watch batched decisions race against normal Gemma’s streamed JSON.</p></div><div class="view-actions"><button id="focus" class="secondary" aria-pressed="false">Focus view</button><button id="export" class="secondary" disabled>Export run ↓</button></div></div>
12
  <div class="workspace">
13
  <section class="input-panel" aria-labelledby="input-title">
14
  <div class="section-title"><h2 id="input-title">The evidence</h2><button id="sample" class="text-button">Use sample photo</button></div>
 
25
  <section class="live-panel" aria-labelledby="live-title">
26
  <div class="section-title"><h2 id="live-title">Live output</h2><span id="run-phase" class="small">Ready</span></div>
27
  <div class="engines">
28
+ <article class="engine parallel"><div class="engine-title"><h3>Parallel scorer</h3><span id="parallel-state">Waiting</span></div><strong class="clock" id="parallel-clock">0.00<span>s</span></strong><div class="progress"><i id="parallel-progress"></i></div><div class="engine-stats"><span id="parallel-count">0 / 128 decisions</span><span>GPU batches</span></div><p id="parallel-first" class="small">First decision —</p></article>
29
+ <article class="engine normal"><div class="engine-title"><h3>Normal Gemma</h3><span id="normal-state">Waiting</span></div><strong class="clock" id="normal-clock">0.00<span>s</span></strong><div class="progress"><i id="normal-progress"></i></div><div class="engine-stats"><span id="normal-count">0 / 128 decisions</span><span id="token-count">0 tokens</span></div><p id="normal-token-first" class="small">First token —</p><p id="normal-first" class="small">First decision —</p></article>
30
  </div>
31
  <div id="verdict" class="verdict" aria-live="polite">Upload a scene or try the sample. Results are measured live.</div>
32
  <div class="streams">
 
37
  </section>
38
  </div>
39
  <section class="wall" aria-labelledby="wall-title"><div class="wall-heading"><div><h2 id="wall-title">Every decision, as it lands</h2><p class="small">A check is “yes” when visually established. P = parallel probability of yes · G = Gemma’s boolean. Streaming JSON is provisional until validated.</p></div><div class="filters"><label for="filter" class="sr-only">Filter decisions</label><select id="filter"><option value="all">All checks</option><option value="yes">Detected by either</option><option value="different">Different answers</option></select><span id="agreement" class="small"></span></div></div><div id="output-grid"></div><p id="no-matches" class="small" hidden>No completed checks match this filter.</p></section>
40
+ <details class="details"><summary>What is being measured?</summary><p>Both paths must read the complete image or sampled video and all question descriptions before producing an answer. The reading stage is shown live; no answers are buffered for playback. Only reusable tokenized field definitions are cached, never media features, input KV state, or answers. The parallel scorer returns probabilities as field batches finish. Normal Gemma is asked to generate one compact JSON object with boolean decisions, without explanations or probability prose.</p><p>The timers include input preparation and inference, with shared upload decoding added equally. They exclude model loading, upload transfer, and allocator reset. This is one simultaneous run per path; GPU contention and first-use effects can affect timing. The result measures completion time while both are running, not isolated throughput. Matching answers measures agreement, not correctness. Invalid normal output is shown without a speedup claim.</p><p>Video processing targets one frame per second with a 32-frame cap. Long media plus 128 questions can exceed the 8,192-token input limit; reduce the number of checks or use a shorter clip. Media is never silently trimmed.</p></details>
41
  </main>
42
  </body>
43
  </html>
gemma_rlcd/static/demo.js CHANGED
@@ -3,6 +3,7 @@ const $ = (selector) => document.querySelector(selector);
3
  const escapeHTML = (value) => String(value).replace(/[&<>"']/g, (char) => ({"&":"&amp;","<":"&lt;",">":"&gt;",'"':"&quot;","'":"&#39;"}[char]));
4
  let config, media, mediaURL, active = false, ready = false, controller, finalResult = null;
5
  let checks = [], rows = new Map(), completed = {parallel: new Map(), normal: new Map()};
 
6
  let clocks = {}, rawText = "", parallelLines = [], events = [], firstAnswer = {}, frame = null, sampleVersion = 0;
7
 
8
  function error(message) { $("#error").textContent = message; $("#error").hidden = !message; }
@@ -39,20 +40,21 @@ async function sample() {
39
  finally { $("#sample").disabled = active; }
40
  }
41
  function reset() {
42
- finalResult = null; rawText = ""; parallelLines = []; events = []; firstAnswer = {}; clocks = {};
43
  completed = {parallel: new Map(), normal: new Map()};
44
  $("#headline-count").textContent = count();
45
  $("#parallel-stream").textContent = "Answers appear as soon as each batch completes.";
46
  $("#normal-stream").textContent = "Real token output will stream here.";
47
  $("#token-count").textContent = "0 tokens";
48
  $("#run-phase").textContent = "Ready";
 
49
  $("#agreement").textContent = "";
50
  $("#verdict").textContent = "Upload a scene or try the sample. Results are measured live.";
51
  $("#export").disabled = true;
52
  for (const method of ["parallel","normal"]) {
53
  $(`#${method}-clock`).innerHTML = seconds(0);
54
  $(`#${method}-state`).textContent = "Waiting";
55
- $(`#${method}-first`).textContent = "First answer —";
56
  $(`#${method}-progress`).style.width = "0%";
57
  $(`#${method}-count`).textContent = `0 / ${count()} decisions`;
58
  }
@@ -83,7 +85,7 @@ function updateDecision(method, path, value, probability, elapsed) {
83
  row.classList.toggle("different", Boolean(parallel && normal && parallel.value !== normal.value));
84
  if (firstAnswer[method] === undefined) {
85
  firstAnswer[method] = elapsed;
86
- $(`#${method}-first`).textContent = `First answer ${elapsed.toFixed(2)} s`;
87
  }
88
  $(`#${method}-count`).textContent = `${completed[method].size} / ${count()} decisions`;
89
  $(`#${method}-progress`).style.width = `${100 * completed[method].size / count()}%`;
@@ -122,6 +124,9 @@ function handle(event) {
122
  } else if (event.type === "phase_start") {
123
  if (!clocks[method]) clocks[method] = {running:true, start:performance.now() - event.media_seconds * 1000};
124
  $(`#${method}-state`).textContent = "Running";
 
 
 
125
  } else if (event.type === "answer") {
126
  const path = event.path.join(".");
127
  const probability = event.answer.probabilities.yes;
@@ -130,6 +135,10 @@ function handle(event) {
130
  $("#parallel-stream").textContent = parallelLines.join("\n");
131
  $("#parallel-stream").scrollTop = $("#parallel-stream").scrollHeight;
132
  } else if (event.type === "token") {
 
 
 
 
133
  rawText += event.text;
134
  $("#normal-stream").textContent = rawText;
135
  $("#normal-stream").scrollTop = $("#normal-stream").scrollHeight;
@@ -193,7 +202,7 @@ async function run() {
193
  }
194
  if(pending.trim()) handle(JSON.parse(pending));
195
  if(!finalResult) throw new Error("The stream ended before the comparison completed.");
196
- finalResult = {request:spec,response:finalResult,stream_events:events,first_answer_seconds:firstAnswer};
197
  $("#run-note").textContent = "Complete. Export preserves the answers, events, and measured timings.";
198
  } catch(failure) {
199
  controller.abort();
@@ -216,6 +225,11 @@ async function status() {
216
  } catch { ready=false; $("#model-status").textContent="Server disconnected"; }
217
  syncButtons();
218
  }
 
 
 
 
 
219
  $("#run").addEventListener("click",run);
220
  $("#stop").addEventListener("click",()=>controller?.abort());
221
  $("#sample").addEventListener("click",sample);
 
3
  const escapeHTML = (value) => String(value).replace(/[&<>"']/g, (char) => ({"&":"&amp;","<":"&lt;",">":"&gt;",'"':"&quot;","'":"&#39;"}[char]));
4
  let config, media, mediaURL, active = false, ready = false, controller, finalResult = null;
5
  let checks = [], rows = new Map(), completed = {parallel: new Map(), normal: new Map()};
6
+ let firstToken = null;
7
  let clocks = {}, rawText = "", parallelLines = [], events = [], firstAnswer = {}, frame = null, sampleVersion = 0;
8
 
9
  function error(message) { $("#error").textContent = message; $("#error").hidden = !message; }
 
40
  finally { $("#sample").disabled = active; }
41
  }
42
  function reset() {
43
+ finalResult = null; firstToken = null; $("#normal-token-first").textContent = "First token —"; rawText = ""; parallelLines = []; events = []; firstAnswer = {}; clocks = {};
44
  completed = {parallel: new Map(), normal: new Map()};
45
  $("#headline-count").textContent = count();
46
  $("#parallel-stream").textContent = "Answers appear as soon as each batch completes.";
47
  $("#normal-stream").textContent = "Real token output will stream here.";
48
  $("#token-count").textContent = "0 tokens";
49
  $("#run-phase").textContent = "Ready";
50
+ $("#run-note").textContent = "The model loads once. Every comparison uses fresh input state.";
51
  $("#agreement").textContent = "";
52
  $("#verdict").textContent = "Upload a scene or try the sample. Results are measured live.";
53
  $("#export").disabled = true;
54
  for (const method of ["parallel","normal"]) {
55
  $(`#${method}-clock`).innerHTML = seconds(0);
56
  $(`#${method}-state`).textContent = "Waiting";
57
+ $(`#${method}-first`).textContent = "First decision —";
58
  $(`#${method}-progress`).style.width = "0%";
59
  $(`#${method}-count`).textContent = `0 / ${count()} decisions`;
60
  }
 
85
  row.classList.toggle("different", Boolean(parallel && normal && parallel.value !== normal.value));
86
  if (firstAnswer[method] === undefined) {
87
  firstAnswer[method] = elapsed;
88
+ $(`#${method}-first`).textContent = `First decision ${elapsed.toFixed(2)} s`;
89
  }
90
  $(`#${method}-count`).textContent = `${completed[method].size} / ${count()} decisions`;
91
  $(`#${method}-progress`).style.width = `${100 * completed[method].size / count()}%`;
 
124
  } else if (event.type === "phase_start") {
125
  if (!clocks[method]) clocks[method] = {running:true, start:performance.now() - event.media_seconds * 1000};
126
  $(`#${method}-state`).textContent = "Running";
127
+ } else if (event.type === "progress") {
128
+ const labels = {preparing:"Preparing input", prefill:`Reading ${(event.input_tokens || 0).toLocaleString()} tokens`, scoring:"Scoring fields", generating:"Generating"};
129
+ $(`#${method}-state`).textContent = labels[event.stage] || event.stage;
130
  } else if (event.type === "answer") {
131
  const path = event.path.join(".");
132
  const probability = event.answer.probabilities.yes;
 
135
  $("#parallel-stream").textContent = parallelLines.join("\n");
136
  $("#parallel-stream").scrollTop = $("#parallel-stream").scrollHeight;
137
  } else if (event.type === "token") {
138
+ if (firstToken === null && event.text) {
139
+ firstToken = event.seconds;
140
+ $("#normal-token-first").textContent = `First token ${firstToken.toFixed(2)} s`;
141
+ }
142
  rawText += event.text;
143
  $("#normal-stream").textContent = rawText;
144
  $("#normal-stream").scrollTop = $("#normal-stream").scrollHeight;
 
202
  }
203
  if(pending.trim()) handle(JSON.parse(pending));
204
  if(!finalResult) throw new Error("The stream ended before the comparison completed.");
205
+ finalResult = {request:spec,response:finalResult,stream_events:events,first_answer_seconds:firstAnswer,first_token_seconds:firstToken};
206
  $("#run-note").textContent = "Complete. Export preserves the answers, events, and measured timings.";
207
  } catch(failure) {
208
  controller.abort();
 
225
  } catch { ready=false; $("#model-status").textContent="Server disconnected"; }
226
  syncButtons();
227
  }
228
+ $("#focus").addEventListener("click",()=>{
229
+ const focused = document.body.classList.toggle("focused");
230
+ $("#focus").setAttribute("aria-pressed",String(focused));
231
+ $("#focus").textContent = focused ? "Standard view" : "Focus view";
232
+ });
233
  $("#run").addEventListener("click",run);
234
  $("#stop").addEventListener("click",()=>controller?.abort());
235
  $("#sample").addEventListener("click",sample);
reports/live-visual-demo.json ADDED
@@ -0,0 +1,23 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "source": "Live browser capture of http://127.0.0.1:8787/demo",
3
+ "presentation": "Actual browser captures at original elapsed time; 4 fps capture; both inference paths share the GPU; model already loaded; only schema tokenization may be warm.",
4
+ "capture_duration_seconds": 58.189,
5
+ "metrics_displayed": {
6
+ "error": "",
7
+ "normal-clock": "54.10s",
8
+ "normal-count": "128 / 128 decisions",
9
+ "normal-first": "First decision 7.93 s",
10
+ "normal-token-first": "First token 7.45 s",
11
+ "parallel-clock": "11.57s",
12
+ "parallel-count": "128 / 128 decisions",
13
+ "parallel-first": "First decision 7.11 s",
14
+ "run-phase": "Comparison complete",
15
+ "verdict": "4.67\u00d7 faster this run \u00b7 120 / 128 matching answers"
16
+ },
17
+ "answers_displayed": {
18
+ "normal": "```json\n{\"people\": {\"person\": true, \"crowd\": true, \"walking\": false, \"sitting\": false, \"standing\": true, \"running\": false, \"cycling\": false, \"carrying_bag\": false, \"backpack\": false, \"handbag\": false, \"hat\": false, \"sunglasses\": false, \"umbrella\": false, \"phone_in_hand\": false, \"using_camera\": false, \"raised_hand\": false, \"waving\": false, \"pointing\": false, \"eating\": false, \"drinking\": false, \"stroller\": false, \"walking_dog\": false, \"uniform\": false, \"helmet\": false, \"high_vis_vest\": false, \"face_mask\": false, \"red_clothing\": false, \"blue_clothing\": false, \"white_clothing\": false, \"striped_clothing\": false, \"group_interaction\": false, \"person_lying_down\": false}, \"objects\": {\"car\": false, \"bus\": false, \"bicycle\": false, \"motorcycle\": false, \"truck\": false, \"taxi\": false, \"boat\": false, \"train\": false, \"traffic_light\": false, \"street_sign\": true, \"billboard\": true, \"bench\": false, \"chair\": false, \"table\": false, \"trash_bin\": false, \"traffic_cone\": false, \"bollard\": false, \"fence\": false, \"streetlamp\": false, \"shop_window\": false, \"door\": false, \"stairs\": false, \"ramp\": false, \"clock\": false, \"flag\": false, \"food_stall\": false, \"bottle\": false, \"cup\": false, \"suitcase\": false, \"screen\": true, \"fire_hydrant\": false, \"scaffolding\": false}, \"nature\": {\"dog\": false, \"cat\": false, \"bird\": false, \"horse\": false, \"cow\": false, \"sheep\": false, \"goat\": false, \"skunk\": false, \"rabbit\": false, \"squirrel\": false, \"duck\": false, \"fish\": false, \"tree\": false, \"grass\": false, \"flowers\": false, \"potted_plant\": false, \"bush\": false, \"mountain\": false, \"water\": false, \"beach\": false, \"snow\": false, \"rain\": false, \"clouds\": false, \"blue_sky\": false, \"sun\": false, \"moon\": false, \"smoke\": false, \"fire\": false, \"rocks\": false, \"fallen_leaves\": false, \"puddle\": false, \"animal_group\": false}, \"scene\": {\"outdoors\": false, \"indoors\": true, \"street\": true, \"sidewalk\": true, \"crosswalk\": false, \"buildings\": true, \"high_rise\": true, \"storefront\": true, \"park\": false, \"kitchen\": false, \"office\": false, \"living_room\": false, \"daylight\": false, \"night\": true, \"artificial_lighting\": true, \"shadows\": true, \"reflections\": true, \"wet_ground\": false, \"visible_text\": true, \"advertising\": true, \"road_markings\": false, \"red_dominant_area\": true, \"blue_dominant_area\": true, \"green_dominant_area\": false, \"yellow_dominant_area\": false, \"closeup\": false, \"wide_view\": false, \"blur\": false, \"occlusion\": true, \"dense_scene\": true, \"clear_foreground\": true, \"distant_background\": true}}\n```",
19
+ "parallel": "people.person: true (99.8% yes)\npeople.crowd: false (49.3% yes)\npeople.walking: false (42.6% yes)\npeople.sitting: false (0.1% yes)\npeople.standing: true (99.9% yes)\npeople.running: false (0.0% yes)\npeople.cycling: false (0.0% yes)\npeople.carrying_bag: false (6.9% yes)\npeople.backpack: false (0.0% yes)\npeople.handbag: false (0.0% yes)\npeople.hat: false (0.0% yes)\npeople.sunglasses: false (0.0% yes)\npeople.umbrella: false (0.0% yes)\npeople.phone_in_hand: false (0.0% yes)\npeople.using_camera: false (0.0% yes)\npeople.raised_hand: false (0.0% yes)\npeople.waving: false (0.2% yes)\npeople.pointing: false (0.1% yes)\npeople.eating: false (0.0% yes)\npeople.drinking: false (0.9% yes)\npeople.stroller: false (0.0% yes)\npeople.walking_dog: false (0.0% yes)\npeople.uniform: false (0.0% yes)\npeople.helmet: false (0.0% yes)\npeople.high_vis_vest: false (0.0% yes)\npeople.face_mask: false (0.0% yes)\npeople.red_clothing: false (42.0% yes)\npeople.blue_clothing: true (99.8% yes)\npeople.white_clothing: false (11.4% yes)\npeople.striped_clothing: false (0.0% yes)\npeople.group_interaction: false (0.3% yes)\npeople.person_lying_down: false (0.0% yes)\nobjects.car: false (0.0% yes)\nobjects.bus: false (0.0% yes)\nobjects.bicycle: false (0.0% yes)\nobjects.motorcycle: false (0.0% yes)\nobjects.truck: false (0.0% yes)\nobjects.taxi: false (0.0% yes)\nobjects.boat: false (0.0% yes)\nobjects.train: false (0.0% yes)\nobjects.traffic_light: false (0.0% yes)\nobjects.street_sign: true (99.8% yes)\nobjects.billboard: true (99.9% yes)\nobjects.bench: false (0.0% yes)\nobjects.chair: false (0.0% yes)\nobjects.table: false (0.7% yes)\nobjects.trash_bin: false (0.0% yes)\nobjects.traffic_cone: false (0.0% yes)\nobjects.bollard: false (0.0% yes)\nobjects.fence: false (0.0% yes)\nobjects.streetlamp: false (3.1% yes)\nobjects.shop_window: true (96.8% yes)\nobjects.door: false (0.5% yes)\nobjects.stairs: false (0.1% yes)\nobjects.ramp: false (1.1% yes)\nobjects.clock: false (0.0% yes)\nobjects.flag: false (0.0% yes)\nobjects.food_stall: false (0.0% yes)\nobjects.bottle: false (0.1% yes)\nobjects.cup: true (89.0% yes)\nobjects.suitcase: false (0.0% yes)\nobjects.screen: true (100.0% yes)\nobjects.fire_hydrant: false (0.0% yes)\nobjects.scaffolding: false (2.7% yes)\nnature.dog: false (0.0% yes)\nnature.cat: false (0.0% yes)\nnature.bird: false (0.0% yes)\nnature.horse: false (0.0% yes)\nnature.cow: false (0.0% yes)\nnature.sheep: false (0.0% yes)\nnature.goat: false (0.0% yes)\nnature.skunk: false (0.0% yes)\nnature.rabbit: false (0.0% yes)\nnature.squirrel: false (34.9% yes)\nnature.duck: false (0.0% yes)\nnature.fish: false (0.0% yes)\nnature.tree: false (10.2% yes)\nnature.grass: false (0.0% yes)\nnature.flowers: false (0.0% yes)\nnature.potted_plant: false (15.1% yes)\nnature.bush: false (0.5% yes)\nnature.mountain: false (0.0% yes)\nnature.water: false (0.0% yes)\nnature.beach: false (0.0% yes)\nnature.snow: false (0.0% yes)\nnature.rain: false (0.0% yes)\nnature.clouds: false (0.0% yes)\nnature.blue_sky: false (0.0% yes)\nnature.sun: false (0.5% yes)\nnature.moon: false (0.0% yes)\nnature.smoke: false (0.0% yes)\nnature.fire: false (0.0% yes)\nnature.rocks: true (91.7% yes)\nnature.fallen_leaves: false (0.0% yes)\nnature.puddle: false (0.8% yes)\nnature.animal_group: false (0.0% yes)\nscene.outdoors: false (0.0% yes)\nscene.indoors: true (100.0% yes)\nscene.street: true (100.0% yes)\nscene.sidewalk: true (92.8% yes)\nscene.crosswalk: false (0.0% yes)\nscene.buildings: true (100.0% yes)\nscene.high_rise: true (100.0% yes)\nscene.storefront: true (99.8% yes)\nscene.park: false (0.0% yes)\nscene.kitchen: false (0.0% yes)\nscene.office: true (99.9% yes)\nscene.living_room: false (0.0% yes)\nscene.daylight: false (0.1% yes)\nscene.night: true (99.9% yes)\nscene.artificial_lighting: true (100.0% yes)\nscene.shadows: true (99.6% yes)\nscene.reflections: true (99.8% yes)\nscene.wet_ground: false (0.3% yes)\nscene.visible_text: true (100.0% yes)\nscene.advertising: true (100.0% yes)\nscene.road_markings: false (0.0% yes)\nscene.red_dominant_area: false (11.2% yes)\nscene.blue_dominant_area: true (98.6% yes)\nscene.green_dominant_area: false (0.1% yes)\nscene.yellow_dominant_area: false (1.5% yes)\nscene.closeup: false (7.6% yes)\nscene.wide_view: true (85.1% yes)\nscene.blur: false (1.8% yes)\nscene.occlusion: true (69.9% yes)\nscene.dense_scene: true (100.0% yes)\nscene.clear_foreground: true (78.3% yes)\nscene.distant_background: true (100.0% yes)"
20
+ },
21
+ "accuracy_note": "Answer agreement is not accuracy; no ground truth was supplied for this scene.",
22
+ "github_video_asset": "https://github.com/user-attachments/assets/2dcac6b3-4dbe-4b1a-94fe-6c015f7defb0"
23
+ }
tests/test_comparison.py CHANGED
@@ -117,7 +117,7 @@ def test_concurrent_comparison_streams_both_paths_before_either_finishes(monkeyp
117
  def symbols(self, count):
118
  return ("A", "B")
119
 
120
- def score_questions(self, state, questions, on_scores):
121
  assert generated.wait(2), "Generation never started alongside scoring"
122
  result = TokenScores((5, 0), 1, 20)
123
  on_scores([(0, result)])
@@ -126,7 +126,7 @@ def test_concurrent_comparison_streams_both_paths_before_either_finishes(monkeyp
126
 
127
  backend = Backend()
128
 
129
- def generate(worker, state, questions, on_token):
130
  assert worker.processor is not backend.processor
131
  assert worker.processor.tokenizer is not backend.processor.tokenizer
132
  on_token('{"visible":', 1)
 
117
  def symbols(self, count):
118
  return ("A", "B")
119
 
120
+ def score_questions(self, state, questions, on_scores, on_progress=None):
121
  assert generated.wait(2), "Generation never started alongside scoring"
122
  result = TokenScores((5, 0), 1, 20)
123
  on_scores([(0, result)])
 
126
 
127
  backend = Backend()
128
 
129
+ def generate(worker, state, questions, on_token, on_progress=None):
130
  assert worker.processor is not backend.processor
131
  assert worker.processor.tokenizer is not backend.processor.tokenizer
132
  on_token('{"visible":', 1)
tests/test_json_backend.py CHANGED
@@ -35,3 +35,43 @@ def test_complete_candidate_likelihood_uses_suffix_and_ignores_padding(batch_siz
35
  assert sum(batches) == 3
36
  assert max(batches) <= batch_size
37
  assert streamed == [[(0, results[0])]]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
35
  assert sum(batches) == 3
36
  assert max(batches) <= batch_size
37
  assert streamed == [[(0, results[0])]]
38
+
39
+
40
+ def test_repeated_schema_reuses_only_tokenization_and_still_prefills_each_image(
41
+ monkeypatch, tmp_path
42
+ ):
43
+ from gemma_rlcd import json_backend
44
+ from gemma_rlcd.core import Noul, State, TokenScores
45
+
46
+ encoded, prefills = [], []
47
+ backend = JSONMLXBackend.__new__(JSONMLXBackend)
48
+ backend.tokenizer = SimpleNamespace(
49
+ encode=lambda text, **kwargs: encoded.append(text) or list(map(ord, text))
50
+ )
51
+ backend.max_input_tokens = 8192
52
+ backend.last_stats = {}
53
+ monkeypatch.setattr(
54
+ json_backend,
55
+ "prepare_generation",
56
+ lambda backend, state, questions: (
57
+ "prompt",
58
+ {"input_ids": SimpleNamespace(shape=(1, 6)), "image": state.images},
59
+ ),
60
+ )
61
+ backend.prefill = lambda prepared: prefills.append(prepared.inputs["image"]) or []
62
+ backend._sequence_scores = lambda cache, tokens, fields, on_scores: (
63
+ {index: TokenScores((float(len(prefills)), 0), 1, 10) for index, _ in fields},
64
+ [2],
65
+ )
66
+ questions = {"visible": Noul("Visible?")}
67
+ paths = [tmp_path / name for name in ("one.jpg", "two.jpg")]
68
+ for path in paths:
69
+ path.touch()
70
+ first = backend.score_questions(State(images=(str(paths[0]),)), questions)
71
+ assert not backend.last_stats["schema_cache_hit"]
72
+ count = len(encoded)
73
+ second = backend.score_questions(State(images=(str(paths[1]),)), questions)
74
+ assert backend.last_stats["schema_cache_hit"]
75
+ assert len(encoded) == count
76
+ assert prefills == [(str(paths[0]),), (str(paths[1]),)]
77
+ assert first[0].logits != second[0].logits
tests/test_web.py CHANGED
@@ -474,7 +474,7 @@ def test_streaming_comparison_orders_events_and_cleans_uploads(playground, monke
474
 
475
  client, _, directory = playground
476
 
477
- def generate(backend, state, questions, on_token=None):
478
  on_token('{"animal":', 1)
479
  on_token('"cat"}', 2)
480
  return {
 
474
 
475
  client, _, directory = playground
476
 
477
+ def generate(backend, state, questions, on_token=None, on_progress=None):
478
  on_token('{"animal":', 1)
479
  on_token('"cat"}', 2)
480
  return {