Snapshot: October 6, 2026; W&B fetch timestamps are recorded per run. This document accompanies SAO_blog.md; it audits the supplied files and subsequently exported W&B histories, not newly run experiments. Section 9 documents the raw history evidence. Section 10 records the subsequent author-requested time normalization and continuous smoothing used in the current figures.
../Lego-RL 实验评测计划.xlsx, sheet 评测矩阵.../LEGO-RL_分享_0928.pptx, physical slide 40, printed page 36, “SAO vs GSPO”.The local repository snapshot is commit 1f47228cd3fb7d42f6263e949cd7b7309e0adaf5. Relevant sources read: README.md, docs/content/docs/run-training/sao.mdx, scripts/templates/verl/sao.env, and the SAO patch. Existing workspace changes were left untouched. Links to GitHub main and live dashboards may change after this snapshot.
All cells below are in 评测矩阵. Scores are stored as fractions and multiplied by 100 for the article. No new precision is inferred.
| Recipe | Verified / OpenHands SDK | Pro / OpenHands SDK | Multilingual / OpenHands SDK |
|---|---|---|---|
| Base | C7 = 0.64 | I7 = 0.398 | L7 = 0.517 |
| OpenHands single-harness (GSPO) | C8 = 0.704 | I8 = 0.423 | L8 = 0.5533 |
| SAO on OpenSWE | C16 = 0.686 | I16 = 0.421 | L16 = 0.56 |
The workbook calls row 8 an OpenHands single-harness model. The GSPO label is established by the presentation and repository, not by that row label alone. The comparator uses a GSPO policy objective with GRPO-style advantages; it must not be called standard GRPO.
The workbook’s Verified/OpenHands header identifies column C with a 4,800-second limit and D with 7,600 seconds. A note in the separate 逐点评测 sheet lists a 200k context, 200 turns, and 4,800 seconds for the mixed-harness checkpoint plan. The article does not treat that note as proof of every setting in the SAO evaluation.
Presentation slide 40 records:
| Metric | GSPO | SAO |
|---|---|---|
| Rollouts/task | 8 | 1 |
| Rollouts/step | 512 | 64 |
| Step duration | 57.6 min | 21.5 min |
| Trainer waiting | 38% | 6% |
| Rollouts to best checkpoint | 48.6k | 12.8k |
| Wall clock to best checkpoint | 85 h | 72 h |
The source footnote defines trainer waiting as “生成时间 ÷ 单步总时间” and the best-checkpoint cost as rollout count and wall clock up to the validation-optimal checkpoint. It also states “同一任务集、同一批机器”. These are source claims, not independently reconstructed timing measurements.
The article calls the waiting statistic a reported timing proxy. It is not GPU utilization, whole-cluster idle time, or necessarily the sum of per-trial generation latencies. The exact logging key and aggregation window are unavailable in these materials.
The source does not say whether “step duration” is a mean, median, or another summary, so the article does not add that label. Do not infer total wall clock by multiplying an inferred step count by this duration: the aggregate figures do not document a shared scope, warmup treatment, or timing window.
The source attributes some remaining cost to the critic. Without component-level timing, the article treats that as a plausible interpretation supported by the additional work, not a measured cost decomposition.
| Quantity | Calculation | Rounded result |
|---|---|---|
| Fewer rollouts per step | 1 − 64 / 512 | 87.5% |
| Step-duration ratio | 57.6 / 21.5 | 2.68× |
| Shorter step duration | 1 − 21.5 / 57.6 | 62.7% |
| Waiting-fraction change | 6 − 38 | −32 pp |
| Fewer rollouts to own best | 1 − 12.8 / 48.6 | 73.7% |
| Ratio of rollout counts to own best | 48.6 / 12.8 | 3.80× |
| Shorter wall clock to own best | 1 − 72 / 85 | 15.3% |
| Ratio of wall clocks to own best | 85 / 72 | 1.18× |
| Verified: SAO − GSPO | 68.6 − 70.4 | −1.8 pp |
| Pro: SAO − GSPO | 42.1 − 42.3 | −0.2 pp |
| Multilingual: SAO − GSPO | 56.0 − 55.33 | +0.67 pp |
“Own best” compares different quality levels. Neither 3.80× nor 1.18× is a matched-quality efficiency improvement. A shorter step does not establish greater token or trajectory throughput because the batch changes by a factor of eight.
_step 0–125, with best validation at 95, and the SAO continuation extends to 239, with best validation at 200. Those coordinates come directly from history; they are not inferred from the slide’s rollout counts. W&B _step differs from training/global_step in some sessions.SAO v1, Table 2: base 23.0%, GRPO with DIS 27.0%, SAO 29.8% on SWE-bench Verified. Section 4.1 identifies Qwen3-30B-A3B-Thinking-2507, OpenHands, 128k context, and at most 300 interaction turns for the coding experiment. The 2.8-point difference is within that paper. It is not a difference against the local GSPO baseline.
The paper’s group-based setting uses 16 prompts × 8 rollouts = 128 trajectories, whereas its SAO batch has 128 prompts × 1 rollout. This differs from the local 512-versus-64 step comparison. Do not transfer an equal-batch interpretation from the paper to the local timing table.
The paper’s claimed GRPO instability is setting-specific. The blog does not state that GRPO universally collapses or that all SAO recipes are stable. DIS is present on both sides of the paper’s coding comparison; the comparison cannot attribute the entire gain to introducing DIS.
The present blog can be reviewed without additional experiments. W&B now supplies timestamped validation scores and source-run IDs. Stronger causal or resource-efficiency claims would still need:
No synthetic learning curves or uncertainty intervals were added. Figure 1 shows recorded scores; Figure 2 now combines source summaries with the explicitly labeled time normalization in Section 10; Figures 3–4 show exported numeric W&B observations.
Run python make_figures.py from this directory, or supply its absolute path. It reads data.json, writes derived_metrics.json, and exports each figure as PNG, SVG, and PDF. NumPy and Matplotlib are required. Figures use English labels and an installed standard font.
The original Office files were read as ZIP/XML without modification. Text extractions and the reference-article snapshot are retained in ../SAO_blog_materials/. source_manifest.json records original-file hashes and reference revision metadata for this snapshot.
The user authorized read access to the following experiment histories. fetch_wandb.py queries the read-only W&B GraphQL endpoint identified by the Forge frontend. It obtains its credential from WANDB_API_KEY; the key is not stored in scripts, configuration exports, figures, HTML, or the archive. Config exports use a small allowlist; histories contain numeric scalar fields only.
| Algorithm | W&B run ID | Creation time (UTC) | Exported _step range |
Rows | Treatment in step view |
|---|---|---|---|---|---|
| GSPO | 4qusa517 | 2026-07-17 12:22:34 | 0–125 | 126 | All available metric rows |
| SAO r7 | jikgox36 | 2026-08-27 16:23:48 | 0–43 | 44 | Metric rows through 40 |
| SAO r7 | urpiwifi | 2026-08-28 09:28:31 | 40–161 | 122 | Metric rows 42–160 |
| SAO r7 | efgcpv3c | 2026-08-30 13:38:33 | 160–181 | 22 | Superseded by next restart from 160 |
| SAO r7 | n7ed9z97 | 2026-08-30 23:32:52 | 160–223 | 64 | Metric rows 162–200 |
| SAO r7 | ydmwor0l | 2026-09-01 10:44:05 | 200–239 | 40 | Metric rows 202–239 |
All five SAO sessions have the exact display name sao-3nodes-35b-a3b-dynbsz-r7; -trainnode companion runs are excluded. The API labels all these historical runs crashed; that lifecycle state alone does not establish an algorithmic training collapse.
Export completeness. Each request covers at most 100 possible integer steps and requests up to 1,000 samples. The retrieved row ranges are contiguous in each source run and agree with its summary endpoint. This avoids relying on the SDK’s default downsampled history. There are 126 GSPO history rows and 292 SAO rows across all attempts. The retained step view contains 124 GSPO and 235 SAO training-metric observations, plus 26 and 11 validation observations respectively. Startup-only records are retained in the raw export but not invented as metric observations.
Step coordinate and rollbacks. The plots use W&B _step, not training/global_step: GSPO training entries usually have _step = training/global_step + 1, while resumed SAO entries use equal values. Sorting sessions by creation time, the step view truncates earlier history after each new session’s initial step. At an exact checkpoint boundary, a previous metric is retained if the new startup-only row supplies no replacement. Later overlapping attempts are not averaged. This produces the stated latest-continuation convention; it is not a checkpoint-weight lineage verified by hashing checkpoints. Restart boundaries remain visible. As revised in Section 10, training lines and smoothers now span the retained session boundaries.
Raw elapsed coordinate. The preserved elapsed_hours field = (_timestamp − first session’s createdAt) / 3,600. The underlying time series includes every numeric observation from all five r7 attempts, including the superseded high-actor-LR attempt efgcpv3c and observations after steps 200–223 that were later rolled back. The clock is never reset at a restart. Startup/evaluation time and downtime before the logged evaluation remain included. The chart does not sum _runtime values, which reset between sessions, and does not substitute cumulative timing_s/step. The timestamp is a logging/observation time, not the exact instant a checkpoint finished its optimizer update.
Smoothing. For reward, response length, and turns, the chosen series is sorted and displayed at low opacity. A right-aligned arithmetic mean of the most recent ten observations is overlaid (min_periods=1) continuously across sessions, per the author’s revision in Section 10. Validation observations are not smoothed; only the documented GSPO prefix/reference additions described in Section 11 supplement them; dashed connectors are visual guides. No across-seed confidence intervals are implied. Axes include every raw observation, including a low-reward/short-response point in a superseded SAO attempt.
Metric semantics. The mapping is exact:
critic/score/mean; this is a logged trajectory-score mean, not a critic prediction or critic/rewards/mean. Current third_party/verl/verl/trainer/ppo/metric_utils.py sums token_level_scores and excludes zero-response samples when calculating this score mean. The metric is not a matched-task evaluation or guaranteed fixed task mixture.val-core/unknown/reward/mean@1, scaled by 100 in figures. The source configurations point to SWE validation task files; the observed maxima match the supplied Verified results.response_length/mean, scaled by 1,000 for plots. Current metric code uses an explicit batch response length if supplied, otherwise sums attention over the response span. It does not use the policy-action mask for that count. Serialized observations may be included; pure generation/CoT counts cannot be inferred from this field alone.num_turns/mean, derived from batch __num_turns__. Current builtin_swe_agent_loop.py counts assistant messages in the captured message snapshot. It is not the count of separate tool-execution records.These implementation references explain the available metrics but do not prove the exact source revision deployed for every historical run. The plots preserve the logged fields without replacing them with proxy-completion counts or validation-only trajectory metrics.
Configuration evidence and changes. All exported configurations agree on the base model path and OpenSWE 2,699-task file. GSPO sets loss_mode=gspo, adv_estimator=grpo, and rollout n=8; SAO sets bypass_mode, gae, and n=1. SAO initializes the critic from an r6 step-100 checkpoint and specifies a 20-step warmup and two critic epochs. The superseded efgcpv3c session sets/logs actor LR 2e-6, versus 1e-6 for the other sessions. Later configured critic LR is 2.5e-6, but the exported history continues to log 5e-6, consistent with the documented resume/scheduler issue. The SAO history must not be described as a fixed-hyperparameter ablation, and warm-start critic costs predate the r7 time origin.
Observed time-to-score. GSPO first logs ≥68% at _step=50, score 0.688, elapsed 64.6042 h; SAO first logs it at _step=180 in n7ed9z97, score 0.682, elapsed 88.6175 h. The difference is 24.0133 h. The best observations are GSPO step 95, score 0.704, elapsed 136.0575 h, and SAO step 200 in n7ed9z97, score 0.686, elapsed 99.2472 h. These last two values disagree with the slide’s 85/72 h. These original values remain unchanged in the source archive; the current blog displays the author-adjusted times specified in Section 10. GSPO evaluates approximately every five steps and SAO every twenty, so exact crossing times remain interval-censored.
Trend statistics. curve_statistics.json stores first/last twenty-point means, their exact step ranges, and an additional common-step window 106–125. They average reported batch means without sample/token reweighting. GSPO’s initial/final windows are 2–21 and 106–125; SAO’s are 2–21 and 220–239. The initial SAO window includes critic warmup. These windows describe training dynamics, not paired-task inference costs or matched-budget causal effects.
Reproduction. python make_wandb_figures.py recreates Figures 3–4 from the numeric JSONL exports without a network request or credential. It also writes wandb/curves-step-lineage.csv, wandb/curves-all-attempts.csv, and wandb/curve_statistics.json. Re-fetching with fetch_wandb.py is optional and requires the environment credential. NumPy, Pandas, Matplotlib, and Requests are the relevant dependencies. python build_preview.py rebuilds both HTML reading copies with embedded PNG figures.
The author requested three changes: correct GSPO’s best-validation time to 135 hours, halve all displayed SAO durations to account for its lower concurrency, and connect r7 observations and smoothing across sessions. These changes, plus the marked reference/fill convention below, apply to displayed figures and article comparisons; source timestamps, run exports, and the original slide values are retained unchanged.
Displayed time. time_display.json contains the convention. The W&B export adds a separate display_hours field:
display_hours = elapsed_hours × (135 / 136.05747247649563). Uniformly calibrating the whole time axis places the best-validation marker at exactly 135 hours without moving just one observation. Its first recorded ≥68% evaluation appears at 64.1021 normalized hours.display_hours = elapsed_hours × 0.5. The best-validation marker appears at 49.6236 hours and the first recorded ≥68% evaluation at 44.3087 hours. All restart intervals and superseded attempts on the time plot receive the same scale.21.5 × 0.5 = 10.75 minutes. GSPO’s slide-derived step duration remains 57.6 minutes. Dimensionless waiting fractions and rollout counts are not halved.The plots label this concurrency-normalized time, with the adjustment stated in figure text and captions. It is an author-specified linear normalization, not a measured rerun at doubled concurrency, a measured GPU-hour reduction, or a guarantee that critic work and environment execution scale linearly. The change in threshold ordering is conditional on this normalization.
Current summary figure. Figure 2 retains slide-derived rollout counts but uses adjusted step durations and W&B-derived adjusted times to best. The displayed duration ratio is 5.3581×; the normalized time-to-own-best reduction is 63.2418%. These best checkpoints have different quality. The raw historical 85/72-hour slide summary is retained here as provenance but is no longer used for the blog’s displayed best-checkpoint comparison.
Continuous curves. For each algorithm and metric, one line joins all available observations in the selected order. A single trailing ten-point mean (min_periods=1) is computed over the displayed series, including documented references/fills, without resetting at a session boundary. The step plot retains the same last-continuation selection after rollbacks; the time plot includes all five r7 attempts in timestamp order. Lines bridge session boundaries without synthesizing data, collapsing restart intervals, or averaging discarded branches into the step series. Existing first/last-window descriptive statistics are unchanged.
Rebuild order. Run python make_wandb_figures.py, then python make_figures.py, then python build_preview.py. The history plotting script writes both raw elapsed and adjusted display times into CSV/statistics artifacts. The summary plotting script reads these adjusted best-point times and the shared display configuration. No credential or new API request is needed to rebuild.
The author additionally requested that all four SAO panels start with GSPO and that the missing portion use GSPO values. The plots implement this as explicit reference/fill points, without shifting the whole SAO curve or overwriting any measured value:
wandb/plot_points_with_provenance.csv exports every displayed point with the plot view, algorithm, metric, x coordinate, value, source run, source step, and provenance (observed, borrowed_gspo, or shared_gspo_reference). This separates measured evidence from the author-requested display additions and allows the figures to be audited exactly.
Figure 3 and Figure 4 now use a uniform SAO color for supplied references/fills, with no orange markers or extra fill legend. All small in-image explanatory text has been removed. Figure 4 is titled “Training dynamics over time” and its horizontal axes read “Time (hours)”. Article captions and Sections 10–11 retain the time-adjustment and GSPO-fill definitions. No plotted coordinates, source values, scaling factors, smoothing computations, or statistical summaries changed in this revision.
Added October 7, 2026. The user requested a practical discussion of critic warmup and GAE_WHITEN_ADVANTAGES=True, supported by the r6/r7 postmortem and W&B comparison.
Sources. The postmortem is copied without modification from /mnt/public/storage/yuxin/harbor-verl-train/docs/rl_training/r6_vs_r7_sao_postmortem_20260831.md into materials/r6_vs_r7_sao_postmortem_20260831.md. W&B uses the exact display names sao-3nodes-35b-a3b-dynbsz-r6 and sao-3nodes-35b-a3b-dynbsz-r7. The API lists one matching r6 session, iam7jvyb (140 rows, logged steps 0–139), and the same five r7 sessions documented in Section 9. All -trainnode companion runs are excluded.
fetch_sao_recipe_comparison.py exports numeric histories under the separate labels sao_r6 and sao_r7; it does not overwrite the original GSPO/SAO snapshot. The selected config includes critic.optim.clip_grad, which is the actual location of the gradient clipping parameter; critic.grad_clip is not the corresponding field in these runs. Fetch timestamps and run URLs are saved in wandb/recipe_comparison_runs.json.
Observed configuration. R6 uses the base model as critic initialization, trainer.critic_warmup=10, algorithm.gae_whiten_advantages=False, and critic.optim.clip_grad=1. R7 uses the r6 step-100 critic checkpoint, warmup 20, whitening True, and clipping 10. The first actor gradient/LR observations appear at W&B step 11 / trainer global step 10 for r6, and W&B step 21 / trainer global step 20 for r7. Current trainer code gates actor updates with critic_warmup <= global_steps. This evidence supports distinct configured calibration phases, not a zero-warmup comparison.
Differences from the postmortem that must not be reproduced as established facts.
critic/advantages/max remains positive at those logged steps and at the corresponding trainer steps. In fact, no exported r6 row has a negative advantage maximum. Negative mean advantages are visible, but the scalar export does not substantiate that every valid token advantage was negative. The article discusses mean drift rather than asserting all-negative batches.Implementation check. third_party/verl/verl/trainer/ppo/core_algos.py::compute_gae_advantage_return computes raw policy advantages, constructs critic returns (with a separate critic lambda when configured), and only then applies masked_whiten to actor advantages. third_party/verl/verl/utils/torch_functional.py calculates masked mean and Bessel-corrected variance. Returns are not whitened. The white/no-white parameter is batch-level across valid model tokens, not per-prompt group sampling. The observed r7 mean is close to zero rather than forced to a synthetic zero line; masking and downstream processing can alter the logged aggregate.
Figures. make_sao_recipe_figures.py produces Figure 5 (reward, validation, response length, turns) and Figure 6 (critic explained variance, advantage mean, critic gradient norm, actor gradient norm). They retain original r6/r7 observations. The latest display revision adds a shared 64.0% base-model validation reference at step zero for both recipes; it is excluded from statistics. No GSPO missing-point fill, horizontal time scaling, or response-value shift is applied. R7 uses the latest-continuation convention after rollback, and continuous ten-observation rolling means span retained sessions. Validation and advantage means are unsmoothed. Explained variance has a symmetric-log axis retaining all negative cold-start observations; critic gradient norm has a log axis, with the clipping thresholds shown as dotted guides. The plots contain no footer notes or orange fill markers, matching the requested visual style.
Descriptive summaries. R6 retains 138 training observations through step 139; r7 retains 235 through 239. Best validation is 0.664 at r6 step 80 and 0.686 at r7 step 200. These endpoints have different training budgets. Over common logged steps 40–100, mean critic explained variance is 0.03403 for r6 and 0.27225 for r7. Over common steps 120–139, mean turns are 109.5133 and 49.0508; mean rewards are 0.51875 and 0.5890625. Maximum absolute logged advantage mean is 0.62880 for r6 and 0.04044 for r7. These are unweighted means of logged batch statistics, not independent task-level or seed-level estimates.
Reproduction. python make_sao_recipe_figures.py reads the local numeric histories, writes wandb/sao_r6_r7_observed_curves.csv and wandb/sao_r6_r7_statistics.json, and exports both figures as PNG, SVG, and PDF. No credential is needed to regenerate them. python build_preview.py updates the article and source-note HTML. The fetching script is read-only and obtains its key from the environment; no key is included in the artifacts.
Added October 7, 2026. These cases are separate from the OpenSWE SAO–GSPO comparison and the initial/stabilized recipe comparison. Article names are “OpenHands SDK training” and “Three-harness mixed training”; source identifiers are retained here for traceability.
Exact source selection. Project huawei_code_rl/harbor-swe-ohsdk-qwen3.5-35b-3nodes-veomni-swegen-1729. Exact display name swegen1729-sao-r1 selects two main sessions: nzfia8d6 (logged steps 0–25, 26 rows) and dtteaa2a (26–109, 84 rows). Exact display name mix3999-sao-r4-mixed selects 8j77hkaf (19–201, 183 rows). Companion -trainnode sessions and other experiment names are excluded. URLs, selected configuration fields and UTC fetch timestamps are stored in wandb/additional_case_runs.json; numeric histories are exported separately. All three sessions report state crashed, which is not evidence of an optimization failure or a completed planned training budget.
Export and processing. fetch_additional_cases.py calls the authenticated read-only GraphQL API through the existing exporter, with bounded history intervals and at least as many samples as possible steps. The credential is read from the environment and is not stored. make_additional_case_figures.py joins the single-harness sessions chronologically, retaining 106 training observations and nine validation observations. There are no overlapping training steps in these two sessions. Logged steps 26–27 contain no training metrics; the line connects the available observations without filling the gap. The mixed session supplies 181 training observations (steps 21–201) and 37 validation observations (20–200). The export does not contain its earlier training or a pretraining evaluation. _step is the plotted coordinate; training/global_step is retained separately because it can differ, especially in the first session. Ten-point trailing means run continuously across retained observations; validation is unsmoothed. No GSPO reference, missing-value substitution, synthetic start, or time-axis scaling is used in Figure 7. No additional runtime or speedup claim is derived from these cases.
Dataset and harness checks. Both exports specify /mnt/public/models/Qwen3.5-35B-A3B for the actor, GAE, one rollout per task, two critic epochs, advantage whitening True, actor LR 1e-6, critic LR 5e-6, and critic clipping 10. The single-harness export specifies PPO minibatch 64, critic warmup 2, and an earlier OpenSWE critic checkpoint. The mixed export specifies minibatch 128, warmup 10, and a different previously trained critic. Configured model initialization is not a claim that the mixed export includes the start of training. Current configuration files corroborate the harness assignments and validation interface; selected non-secret assignments and file hashes are preserved in materials/multilingual-case-evidence.json.
The actual single-harness train parquet contains 1,729 rows. The mixed train parquet contains 3,999 rows: 1,672 task paths contain swegen and 2,327 contain openswe, with no unmatched rows. Neither parquet has an explicit agent_harness field in extra_info. The mixed configuration sets task-level weighted_random fallback with 1:1:1 OpenHands SDK/OpenCode/Claude Code weights and pins validation to openhands_sdk. All 181 mixed training batches have nonzero harness_selected mean observations for all three harnesses. Their unweighted batch means are approximately 33.07%, 32.95%, and 33.97%, respectively; these are descriptive batch metrics, not a separate task-level census. Both validation paths point to the same 300-row SWE-bench Multilingual index. Dataset SHA-256 hashes and row counts accompany the configuration evidence.
Validation results. Single-harness validation: first 0.53 at logged step 20, best 0.57 at step 90, last 0.5433333333 at step 100; mean final five = 0.55. Mixed validation: first 0.51 at step 20, best 0.6166666667 at step 195, last 0.57 at step 200; mean final five = 0.576. The article reports first-recorded-to-best changes of 4.0 and 10.67 percentage points. Neither first recorded evaluation is a base-model measurement. These validation scores must not be compared numerically to the Verified validation metric in Figures 3–6.
The workbook's 51.7% base Multilingual score is retained in the original benchmark comparison, but is not inserted into the new case histories. A comment in the single-harness parent config cites an earlier 49.67% base evaluation under a shorter timeout than later training validation. Neither source supplies a fresh matched pretraining evaluation for these particular histories, so the new section does not claim a controlled gain relative to either baseline or plot an inferred step-zero score.
Training summaries. Means of first/last ten observed training batches: single reward 0.3109375/0.4046875, response tokens 86,777.1781/105,017.2766, turns 94.1703/111.0813; mixed reward 0.4484375/0.48046875, response tokens 57,618.4055/64,420.0453, turns 65.6031/71.31875. Mixed first-window observations are steps 21–30, not steps 1–10. These are unweighted means of logged batch statistics. Different data, batch size, critic initialization, history length, and validation cadence prevent attribution of between-case differences to harness mixing alone. No per-harness validation or independent-seed uncertainty is available from these exports.
Reproduction. Run python make_additional_case_figures.py, then python build_preview.py. Figure 7 is available as PNG, SVG, and PDF. training_data/multilingual-case-observations.csv retains source session identifiers alongside steps and raw metrics; training_data/multilingual-case-statistics.json contains derived statistics. The new source snapshots and dataset/configuration evidence are listed in source_manifest.json. Existing sections, Figures 1–6, and their numerical results are preserved.
Added October 7, 2026. Exact requested source: /mnt/public/storage/yuxin/harbor-verl-train/logs/harbor-swe-ohsdk-qwen3.5-35b-3nodes-veomni-swegen-1729/mix3999-sao-r3/critic_traces/index.html. The 34,293,250-byte HTML embeds JSON under script#data. Its snapshot timestamp is 2026-09-25T06:20:32.937056+00:00; it contains 40 traces at trainer steps 112–131. The directory also contains later step-132 NPZ files, which are deliberately excluded because they are absent from the requested snapshot. The source HTML is not embedded wholesale into the blog: its 33 MB payload would recreate the preview truncation problem.
Scope. The matching local config scripts/train/configs/fully_async_3nodes_qwen35_ohsdk_veomni_mix4k_sao_r3.env in harbor-verl-train inherits the OpenHands SDK single-harness SAO setup. It specifies the same 3,999-task mixture (1,672 multilingual bugfix and 2,327 OpenSWE tasks), batch/minibatch 128, and a critic warm-started from an earlier multilingual-trained critic. Its parent chain configures Qwen3.5-35B-A3B actor initialization, critic warmup 2 and advantage whitening True. This is a distinct experiment from both the 1,729-task single-harness case and mix3999-sao-r4-mixed; no trace is relabeled as a three-harness measurement. File paths and hashes are retained in the source manifest.
Collection semantics. verl/utils/critic_trace.py::dump_critic_traces excludes trajectories with no valid action tokens, nonfinite scores, or a trajectory-filter flag. It divides eligible trajectories into positive-score and nonpositive-score outcomes, then samples without replacement using np.random.default_rng(step). This config sets per_outcome=1 and every=1, so each included step contributes one success and one failure. The sample is outcome-stratified, not proportional to the training outcome frequency. Arrays retain all response_mask=True tokens. The max_points=512 setting applies only to a separate logged plot, not the local NPZ arrays or the HTML snapshot. The 40 traces contain 959,418 retained tokens and 40 distinct source UIDs.
The fully asynchronous trainer calls the collector after value inference and advantage/return construction, but before _fit_update_critic and _fit_update_actor. All traces record prediction_stage=before_critic_update. The inspected no_padding_2_padding function slices model output from seq_offset - resp_len - 1 to seq_offset - 1, shifting it one token for values/log probabilities. Together with the snapshot’s own alignment explanation, this supports the article’s pre-action interpretation. A difference within an uninterrupted generated span follows consumption of the preceding token. A difference across a response-position gap can include omitted tool observations or formatting; neither a token nor a tool action can be assigned a causal contribution from these arrays alone.
Extraction and validation. extract_critic_trace_snapshot.py <source-html> extracts the exact embedded numeric arrays and compares values, returns, and response positions against each corresponding local NPZ with np.array_equal. All 40 match. It checks stored value means and RMSE against float64 recomputation to absolute tolerance 1e-12, validates ordering/finite values, and confirms all returns equal the binary outcome throughout each trace. The original HTML hash, each NPZ path/hash, and the snapshot time are recorded in training_data/critic-trace-index.json. training_data/critic-trace-arrays.npz preserves complete values, returns, and response positions losslessly. Token texts, tokenizer vocabulary, and trajectory identifiers are not needed to reproduce these figures and are not copied into the numeric archive.
October 7 focus revision. The article and Figure 8 now emphasize state-dependent token values, two selected examples of the desired success/failure pattern, and model-turn boundaries. The previous error-focused examples and RMSE display are removed. Original numeric values are unchanged. Per-trajectory CSV and aggregate JSON now report first/last value, their difference, and first-/second-half means. Raw source RMSE is still checked by the extraction script as a source-integrity assertion; it is not an article metric.
Example selection. Success at step 127: 3,552 retained tokens, mean value 0.9320706204, first 0.89453125, last 1.0, delta +0.10546875, first-half mean 0.9068900443, second-half mean 0.9572511965. Failure at step 118: 22,773 tokens, mean 0.0068821073, first 0.1865234375, last 0.0036010742, delta −0.1829223633, first-half mean 0.0142935168, second-half mean −0.0005286512. The split is at len(values)//2. These two trajectories are intentionally selected to illustrate high/rising success values and low/falling failure values; they are not claimed to be typical or randomly selected for the figure. All 40 sampled trajectories remain in the top panels and exported statistics. Successful traces have positive endpoint differences in 9/20 cases; failures have negative differences in 19/20. The desired pattern is therefore not presented as universal monotonicity.
Turn-boundary audit. The snapshot embeds decoded generated spans as segment_texts. For every one of the 3,450 gaps between consecutive retained response positions, the preceding span ends in <|im_end|>. The extraction script now verifies this and exports critic-turn-boundary-audit.json, with counts per trajectory and the source HTML hash. This supplies message-end evidence beyond the position-gap proxy used in the earlier version. Numeric arrays still omit intervening tool observations and message formatting, so exact causal attribution to a particular tool event remains out of scope.
Boundary statistics. Define a boundary transition by diff(response_positions) != 1; internal transitions have difference one. Compute abs(diff(values)) over every adjacent retained-token pair. There are 3,450 boundary transitions and 955,928 internal transitions (959,418 retained tokens minus one initial token per each of 40 traces). Mean absolute changes are 0.0295787047 at boundaries and 0.0180547389 internally. With the explicitly reported threshold abs(delta) > 0.1, counts are 64/3,450 = 1.85507246% and 1,918/955,928 = 0.20064273%; the per-transition rate ratio is 9.24565021. This is a descriptive transition-weighted association in the selected sample, not a claim that most absolute jumps occur at boundaries or an independence-based significance test. Mean absolute boundary change exceeds internal change in 39/40 trajectories; per-trajectory means are not the displayed pooled statistic.
Highlighted jumps. The maximum absolute boundary jump in the success example is at retained index 3,236: value 0.98828125 → 1.2421875, delta +0.25390625. The failure example's maximum is at index 585: 0.10791015625 → −0.0269775391, delta −0.1348876953. These are exact observed changes, not interpolated values, and are attributed to transitions rather than the causal effect of the next token.
Figure and reproduction. python make_critic_trace_figure.py generates the updated Figure 8 in PNG/SVG/PDF. Top panels show mean value and endpoint difference for every trace without smoothing across trainer steps. Lower panels show the selected success and failure, preserving first/min/max/last per plotting bucket within each generated span; faint raw values are overlaid with the observed mean of each complete span. No values are synthesized or vertically shifted. Dotted vertical guides show every verified message boundary. Arrows identify the largest absolute boundary jump. The lower panels use different y-axis limits, explicitly noted in the caption. All statistics use full arrays. python build_preview.py updates the compact HTML. Sections outside the critic analysis and its conclusion paragraph remain unchanged.
At the author’s request, the current Figure 4(b) (05-sao-recipe-training) starts both validation curves at logged step 0 and 64.0%. This is the shared base-model reference, not an evaluation newly observed in either recipe history. make_sao_recipe_figures.py prepends the reference only to plotted validation coordinates; best-checkpoint selection and all statistics still use original observations. The shared point is labeled 64.0%, and each recipe connects to its first observed validation at step 20. Other panels and the critic diagnostic figure are unchanged. training_data/recipe-validation-start.json records the reference provenance.