CODEX v2 PUBLIC · TRIAL 1c42f196d immutable · Chrome 150 · 1440×900
COMPARABLE TELEMETRY v2 · PINNED AUDIT · 2026-07-17
CODEX vs FABLE 5 v2 実装・デバッグ実測比較
v2では「取得不能」を放置せず、初回の比較Pages公開検証後・finalized manifest公開前のgoal runtime checkpoint区間を保存した。 同時にFableの公開repositoryをcommit固定で監査し、正確な値、近似の手集計、 再計算できないheadlineを分ける。既存の約2,969万checksと同一Chrome 6,000 framesも消さずに残す。
既存v1記事比較の表示範囲。新しいpinned repository監査とv2 telemetryはこの切替の影響を受けない。
CODEX v2 PUBLIC · TRIAL 1c42f196d immutable · Chrome 150 · 1440×900
FABLE PUBLIC · TRIAL 1Chrome 150 · 1440×900
データ量の前に、単位と出所を揃える
Codexはgoal全体の正確なcheckpoint差分、Fableはcommit
03356fd9…
のstructured worklog監査。母集団が違う値は同じ行に置いても、倍率にはしない。
公開JSONを検証中…
400,444 → 2,548,223 / 7,655 s。final manifest公開直前の固定境界。
約148 task agent executions。normalized event logではない。
Codex atomic checks / tests / browser trialsと同じ単位ではない。
| 観測軸 | CODEX v2 | FABLE pinned main | 比較ルール |
|---|---|---|---|
| tokens | 2,147,779publication-cutoff goal delta正確・集約 | 約18.14Msubagent cumulative近似・手集計 | run scopeも集計方法も異なるため、token効率の倍率は出さない。 |
| agent grain | 14 task-path observationsroot + 13 bounded · peak 4個別token N/A | 約148 executionsidentity countではない近似・手集計 | Codexのrole / identity観測とFable task executionsを同じ「agent数」として足し引きしない。 |
| checks | 29,687,939v1 deep atomic / 0 fail再実行 | 3,012explicit numeric checksrepo監査 | 定義が一致しない。Codex v2 deepは下のfinalized JSON hookへ追加する。 |
| browser raw | 6,000 framesv1 same-Chrome direct A/B再実行 | 同じ6,000 framescurrent public artifact再実行 | Fable worklogのpass / screenshot headlineはraw trial ledgerから再計算できないため、別欄へ混ぜない。 |
二重計上しない計算式
total = input + output
cached inputはinputの部分集合、reasoning outputはoutputの部分集合なので再加算しない。取得不能はnullであり0ではない。checkpointはagent別recordへ按分しない。runのwall-clock開始境界は未取得だが、goal counter差分は初回比較Pages公開検証後・finalized manifest公開前の境界で固定した。
exec / app-server usageをexact recordへ
利用可能な実行ではturn・thread境界のusageをそのまま取り込み、input / cached input / output / reasoning outputをidentity単位でreconcileする。
時間・test・browser・deployを同じrun IDへ
OpenTelemetry spanまたはcollector eventへstart/end、command、exit、browser trial、artifact hashを記録。近似値から正確値を逆算しない。
core JSON確認中
procedural core / streaming / growth / navigation / quality。
deep core JSONpublic run確認中
Codex immutable Pages deploymentとFable public deploymentのstationary raw trialsを検証する。
raw 6,000-frame JSONdebug contract repeats確認中
fresh CDP target / debug mutation / fixed-stepのfunctional contractを確定JSONから描画する。
25-target raw ledger5-run分布を確認中
成功runだけを抜き出さず、5観測の分布と失敗gateを同じ分母で表示する。
5-run performance ledger58 tests / 13,102 expects PASS
0 fail。deep suiteやbrowser assertionへ加算しない独立した通常回帰。
engineering worklogPUBLIC Codex v2 / PUBLIC Fable stationary baseline
Codex immutable c42f196dとFable publicを、同じChrome・1440×900・cache disabled・各5試行・counterbalancedで測定。ゲーム内容量・描画内容・配信originは同一ではなく、game全体の優劣には外挿しない。
| 測定値 | CODEX v2 PUBLIC | FABLE PUBLIC |
|---|---|---|
| trials | — | — |
| measured frames | — | — |
| frame p50 | — | — |
| frame p95 | — | — |
| frame p99 | — | — |
| >50ms frames | — | — |
| draw calls / frame | — | — |
| mean heap after | — | — |
| mean resources | — | — |
| browser errors | — | — |
同じgoal、違うスタート地点
機能の有無はperformance負荷の同等性を意味しない。定数が確認できた側だけ具体値を出す。
| 機能 | SOL ROLLER v2 | FABLE |
|---|---|---|
| 舞台 | 東京 実装契約 | 東京 source確認 |
| 屋内スタート | 民家の中 実装契約 | 戦国電気・electronics store source確認 |
| 最終goal | 東京スカイツリー 634m 実装契約 | 東京スカイツリー 634m source確認 |
| BGM / SE | procedural WebAudio 実装契約 | procedural WebAudio source確認 |
| dash | 0.8s / cooldown 4s / 2.2× 定数 | Space dash gauge / Shift boost source確認 |
| 競技system | timer・rare・score・rank 実装契約 | timer・rare collection・score・rank source確認 |
| タイムアタック時計 | active foreground wall time 実装・browser検証 | fixed simulation time source確認 |
| X share | result intent URL 実装契約 | result intent URL source確認 |
時計の比較境界: rank定数に共通点があっても、Codexはforeground中の実経過時間、Fableはfixed simulation時間を採用しており、elapsed timeやrankを同一計測として直接比較しない。
約793万ケースを、約2,969万回チェック
同じframeを別シナリオとして水増しせず、入力ケースとatomic assertionを別々に保存した決定論キャンペーン。
各ケース内の境界、同一性、保存則、誤差、budgetを個別に検査。
seed・tier・chunk・trajectory・方位・品質traceの固有入力。
修正後の再実行。evidence digest c58f3b92 が一致。
Bun 1.3.11 / fixed seeds / CPU-side validation。
14,636,922 atomic
1,161,978 cases · 64 seeds × 6 tiers × 121 chunks × 24 objects
- 1,115,136 descriptors
- byte determinism
- ID collision 0
8,490,528 atomic
4,286,304 cases · 9,504 streamed worlds
- 600 hard cap
- 4,147,200 identity reuse checks
- active ID collision 0
2,709,785 atomic
369,971 cases · 4,096 complete trajectories
- 319,003 accepted events
- 30,865 rejected events
- 16,007 boundary cases
1,882,624 atomic
626,176 cases · bearing / distance / floating origin
- 622,080 vector cases
- 最大距離誤差 3.5e−10
- 4,096 landmark cases
1,968,080 atomic
1,492,520 cases · 2,048 adaptive traces
- 491,520 samples
- density / DPR budgets
- 4,095 transitions
Main + 3 specialist tracks
Main + 3 specialist tracksというv1のrole記録。unique identity数を推測しない。v2のgoal aggregate tokenも各trackへ按分していない。
設計統合、実装、修正、最終判断
約793万cases、境界bug 2件
同一Chrome 6,000 frames
35 attempts、Fable provenance監査
PUBLIC対PUBLIC、同じChrome、同じフレーム数
Codexのimmutable Pages deploymentとFable publicを同一CDP endpointで起動し、各5試行×600 framesを交互順序で測定した。両側に同じ方法の実測値がある比較軸。
| 測定値 | CODEX PUBLIC | FABLE PUBLIC | 読み方 |
|---|---|---|---|
| successful trials | 5 / 5 | 5 / 5 | 10 trialすべて起動・採取・PNG保存まで成功 |
| measured frames | 3,000 | 3,000 | 全体6,000 raw rAF intervals |
| frame p50 | 27.1 ms | 28.7 ms | 同じPUBLIC/PUBLIC stationary run内の中央値 |
| frame p95 | 40.4 ms | 44.2 ms | 同一Chrome・同一採取方法のtail指標 |
| frame p99 | 48.5 ms | 52.8 ms | 機能量非同一のためゲーム全体へ外挿しない |
| maximum frame | 92.3 ms | 259.6 ms | 単一最大値は外れ値の影響が大きい |
| frames > 50 ms | 170.57% | 421.40% | 各3,000 frames、long-frame thresholdは50ms |
| draw calls / frame | 14p50 / p95固定 | 44p50 / p95固定 | pre-navigation WebGL draw* wrapperで計測 |
| mean JS heap after | 4.70 MB | 8.73 MB | CDP Runtime.getHeapUsage、各trial終了時平均 |
| mean navigation transfer | 3.77 KB | 12.24 KB | CDP encodedDataLength |
| mean resource transfer | 168,237 B | 738,388 B | document以外。cold-cache平均 |
| browser errors | 0console / runtime / log / network | 0console / runtime / log / network | 両側ともerrorなし |
Codex / Fable。API call数でありGPU時間ではない。
ゲーム内容量・world content非同一の参考比。
今回のPUBLIC/PUBLIC cold-cache stationary run。
stationary sample内だけの比。実走性能ではない。
PUBLIC/PUBLIC固定条件: Codex https://c42f196d.sol-katamari.pages.dev/対Fable public。5 trials × 600 frames / site、合計6,000 frames。ゲーム内容量が非同一なのでstationary baseline以外の優劣は主張しない。
Fableの数字は「外部自己申告累計」として読む
現行公開物から再計測できる値と、記事で報告された開発累計を分離した。25 passes等をv1実測として扱わない。
| 指標 | CODEX | FABLE v1 | 根拠・注意 |
|---|---|---|---|
| recorded tracks (v1) | 4role記録 | 48外部 | CodexはMain + 3 specialist tracksで、unique identity数ではない。agent粒度は非同一。 |
| total tokens | 2,147,779v2 publication-cutoff goal delta正確・集約 | 約377万外部 | Codexはcheckpoint区間、Fableはversion報告。scopeが違うため効率比は出さない。 |
| elapsed | 455 sgoal runtime delta正確・集約 | 約134分外部 | 5版累計scopeでは約19時間。wall-clockとgoal runtimeを同じ時間にはしない。 |
| assertions | 29,687,939deep atomic checks再実行 | v1単体値なし外部 | pinned repo監査では5版に3,012 explicit checks。定義が違うため倍率比較禁止。 |
| same-Chrome trials | 5direct A/B再実行 | 5direct A/B再実行 | このページが同条件で採取したtrial。外部報告passとは別。 |
| same-Chrome PNG | 51 per trial再実行 | 51 per trial再実行 | 両側合計10枚。代表2枚をheroへ掲載。 |
| reported browser passes | 対象外— | raw再計算不可N/A | 記事headlineは存在するが、pinned structured rawから累計を再計算できない。 |
| reported screenshots | 対象外— | raw再計算不可N/A | immutable hash付きの完全manifestがないため、headlineを検証済み実測にしない。 |
SOL ROLLER
民家から始まるprocedural Tokyo、6 scale tiers、固定instance予算、東京スカイツリーへのnavigation。
- 世界
- home + Tokyo terrain / scenery
- scale
- HOME ROOM → SKYTREE SKYLINE
- instance
- 600 steady / 1,200 transition
- goal
- 東京スカイツリー 634m
- game systems
- dash / timer / rare / score / rank / X
FABLE KATAMARI
OSM東京をstreamingし、2cmのネジから634mのスカイツリーまで成長する情報量重視の実装。
- buildings
- 14,035 current records
- roads
- 12,929 current records
- reported buildings
- 14,563(公開資料)
- scale
- 2cm → 634m
- game systems
- timer / combo / 実況 / collection
検証量を増やした結果、2件の境界バグを発見
単にchecksを増やしたのではなく、既存suiteを通過していた失敗を再現し、修正後の回帰testまで残した。
ゴール境界のcube-root丸め
必要半径ぴったりの体積を設定しても、cube-rootの丸めで僅かに閾値を下回り、吸収可能にならない入力をdeep growth matrixが検出。
4秒 / 12秒quality閾値の分割誤差
sample時間の足し上げが閾値直前に留まり、十分な低速・高速traceでもadaptive qualityが切り替わらないケースを検出。
失敗入力をmatrixから回帰suiteへ降ろす
bun tests/deep-validation.ts
bun test
bun run test:cross
- 固定seedとinput matrix
- semantic / atomicを別集計
- raw JSONとdigestを保存
- 2件を通常suiteへ固定
多版・多agentで探索範囲を広げる
- terminal speed到達不能
- landmarkのtier gap
- 人間確認でopening occlusion
- OSM / Overpass外部検算
上記は記事報告。今回のrepositoryからFableの全suiteを再実行した結果ではない。
参照ページの「30指摘・21確定・9棄却」から計算すると棄却率は30%。36%表記には別分母が必要なため、このページでは30%とする。
2条件・45 attempts — 37 pass / 8 failure
主campaignとfresh-profile concurrent-load stressを混ぜずに併記し、合計でも失敗runを分母から除外しない。
結果ファイルを確認中
未確定値はheadline・比較表・結論へ加算しません。
失敗を除外しない
いずれもtimeoutではなく、公開版を同じChromeで反復した際に観測された不安定性またはperformance gate超過。
入力を送った観測窓で描画frameが進まず、before / afterの移動量が0。25 smoke中の1 failure。
同trialで4,065.7msのp99 outlierを観測。performance runを成功扱いにしない。
chunk-crossing soakのp95 budgetを超過。10 performance runs中の2件目のfailure。
集計契約: 35 attemptsを分母に維持し、3 failuresを削除しない。compact summaryを表示元とし、sanitized full campaign JSONも公開対象に残す。
追加10 attempts — smoke 5 / 5、performance 0 / 5
全件がframe-time gate超過。一方で、instance pool・draw-call上限・geometry/texture plateau・heap構造budgetは維持された。負荷下の時間感度と構造上限を分けて読む。
v2で量は増えた。比較単位はあえて混ぜない
deep 2 suite、25回browser campaign、更新6,000-frame A/B、goal token checkpointを追加した。一方、重複可能性・機能非同等・Fable手集計のprovenanceは消えない。
Codexで確証できたこと
- v2 depth: feature matrix 5,834,730 cases / 42,969,756 checks、core 7,917,462 / 31,769,395を別表示、双方0 fail
- v2 unit: 58 tests / 13,102 expects、0 fail
- v2 browser: 25 / 25 fresh-CDP contract repeats、1,550 / 1,550 exact assertions、62 unique contracts
- v2 performance repeat: 4 / 5 PASS(80%)。
streaming.heap-plateau1件を分母に残す - v2 direct A/B: 6,000 frames、10 / 10 trials、error 0 / 0
- usage: publication cutoff 2,147,779 tokens / 7,655 secondsをexact aggregate deltaとして保存
- orchestration: root + 13 bounded task paths、peak concurrency 4。個別tokenは取得不能
- v1 continuity: 29,687,939 checks、45-attempt campaign、境界bug 2件の証拠も保持
Fableで確認・報告されたこと
- real data: 現行14,035 buildings / 12,929 roads
- breadth: pinned repo報告 約148 task executions / 約18.14M subagent tokens
- explicit QA: 3,012 checks。browser / screenshot headlineはraw再計算不可
- game richness: timer・combo・実況・collection・OSM東京
このデータから言えないこと
Codexがゲーム全体でFableより何倍優秀、token効率が何倍、同機能で何倍高速、とは言えない。Codex tokenはgoal区間のaggregate、Fable tokenは累計の近似手集計でscopeが違う。直接性能表は静止baseline、deep matrixとFable explicit checksも別単位だからだ。
今回は「Codex側のデータが少ない比較」から、2系統のdeep matrix、25回のfresh-tab browser reliability、5回のperformance分布、更新same-Chrome A/B、exact goal usage checkpointを持つ比較へ更新した。合算できない値を合算せず、失敗runも分母に残すことが検証結果の一部である。
raw evidenceまで辿れる
v1実測生成日は2026-07-13、v2 telemetry / pinned repository監査は2026-07-17。近似の手集計は正確値から分ける。
- 01Codex deep validation — full JSON
- 02PUBLIC/PUBLIC Same-Chrome — raw 6,000 frames
- 03Legacy v1 Same-Chrome capture — retained history
- 04Repeated E2E campaign — final compact summary
- 05Repeated E2E campaign — sanitized full report
- 06Fresh-profile concurrent-load stress — summary
- 07Fresh-profile concurrent-load stress — sanitized full report
- V2Repeated Chrome performance — five-run ledger
- 08Codex engineering worklog
- 09Codex performance contract
- 10Fable 5 implementation report
- 11Fable current Tokyo manifest
- 12Reference: Sonnet 5 vs Fable 5
再実行実測repository script / current public artifact / raw JSON
repository記録test output・Worklog・architecture
外部自己申告Fable記事・参照比較ページの累計
比較不能母集団またはtelemetryが揃わない値