# FIB Preliminary Scientific Laboratory — 2026-08-24

## Status

**Purpose:** pre-implementation scientific validation.
**Production changes:** NONE.
**Commit to FIB repository:** NONE (the local Windows repository is not mounted in this chat).
**Decision:** do not integrate TMO/RCB/X into production yet.

## Source data

| Game | SHA-256 |
|---|---|
| Melate | `6cb1dc5b5d6fa630dcb1fb273e92f11e7e0efa202fcaaffdce44f1e18e886775` |
| Revancha | `0971fd724f0344492ac978f1fa1669a73ab859d59d62e285c6a7ad2f1bc25346` |
| Revanchita | `db3b2234603e8d360a6941aa548fc9f10bf584c94d717b500597bc6cb8fb7157` |

All target predictions were calculated using only draws strictly prior to each target draw.

## Experiment A — three-game blind laboratory

Current 56-number regime. Evaluation targets: contests **2691..4256** (1566 targets per game).

Temporal split:
- Discovery: 2691..3190 (500/game)
- Confirmation: 3191..3690 (500/game)
- Final Blind: 3691..4256 (566/game)

Candidate search:
- 3,000 TMO formulas
- 3,000 RCB formulas
- 2,000 X formulas

This produced approximately **12,095,094 model-target forecast evaluations** before auxiliary controls, corresponding to about **677 million number-score evaluations**.

### Final Blind absolute performance

Uniform expectations:
- Top-6 expected hits: 0.642857
- Top-28 expected hits: 3.000000

| Candidate | Top-6 mean | Lift vs Uniform | p one-sided | Top-28 mean | Lift vs Uniform | p one-sided |
|---|---:|---:|---:|---:|---:|---:|
| TMO | 0.641343 | -0.24% | 0.534 | 3.016490 | +0.55% | 0.280 |
| RCB | 0.633687 | -1.43% | 0.700 | 3.010012 | +0.33% | 0.362 |
| X | 0.655477 | +1.96% | 0.236 | 2.987044 | -0.43% | 0.676 |

**Conclusion:** none demonstrates absolute predictive advantage over Uniform in Final Blind.

### Time-shuffle temporal control

When the same final predictions were paired with the same outcomes after destroying their temporal ordering:

- TMO Top-28: empirical p ≈ 0.003
- RCB Top-28: empirical p ≈ 0.0017
- X Top-28: empirical p ≈ 0.441

Interpretation: TMO/RCB show evidence of some temporal alignment relative to their own shuffled baselines, but their absolute performance remains approximately Uniform. This is diagnostic information, not predictive authority.

## Experiment B — direct FrEWBayesian protocol reproduction

Melate only, exact frozen splits:
- train: 2500..3300
- tuning: 3301..3600
- validation: 3601..4251

FrEWBayesian v0.6-A was independently reimplemented from its documented equations:
- EW single and pair counts
- residual pair field
- double residualization
- sequential sampling without replacement
- exact 720-order unordered-set probability via 64-subset DP
- JLS against `1/C(56,6)`

### Reproduction

| Metric | Reproduced here | Frozen project reference |
|---|---:|---:|
| Train JLS | +0.000165245 | +0.000165 |
| Tuning JLS | +0.001462702 | +0.001463 |
| Validation JLS | +0.001178895 | +0.001179 |

This reproduces FrEWBayesian v0.6-A to the expected precision.

## Experiment C — standalone challengers under FrEW protocol

Each candidate was selected only from train/tuning and evaluated on untouched validation.
A Plackett-Luce sequential joint distribution was used so each challenger could be scored with the same JLS concept as FrEW.

| Candidate | Selected temperature | Validation JLS | Validation blocks positive | Verdict |
|---|---:|---:|---:|---|
| TMO | 0.025 | +0.000886 | 4/7 | INCONCLUSIVE |
| RCB | 0 | 0.000000 | 0/7 | REJECT as standalone current formulation |
| X | 0.025 | -0.000887 | 3/7 | REJECT/INCONCLUSIVE |

TMO bootstrap 20-draw 95% CI for validation mean: approximately **[-0.00252, +0.00347]**. It includes zero.

## Experiment D — FrEW + challenger hybrid

The frozen FrEW v0.6-A pair interaction remains intact:

`score(i | A,t) = 0.05 * sum_j phiPerp[i,j]`

A challenger can contribute only an additive marginal evidence term:

`score_hybrid(i | A,t) = 0.05 * sum_j phiPerp[i,j] + eta * challenger_score[i]`

`eta` was chosen using train/tuning only by maximizing the minimum train/tuning JLS.

### Results

| Hybrid | eta | Train JLS | Tuning JLS | Validation JLS | Delta vs FrEW |
|---|---:|---:|---:|---:|---:|
| FrEW reference | 0 | +0.000165 | +0.001463 | +0.001179 | 0 |
| FrEW + TMO | +0.025 | +0.001770 | +0.001549 | **+0.002061** | **+0.000883** |
| FrEW + RCB | 0 | +0.000165 | +0.001463 | +0.001179 | 0 |
| FrEW + X | +0.025 | +0.001215 | +0.001864 | +0.000290 | -0.000889 |

Geometric probability-ratio interpretation:
- FrEW validation JLS +0.001179 → about **+0.118%** geometric likelihood vs Uniform.
- FrEW+TMO +0.002061 → about **+0.206%** vs Uniform.
- Incremental FrEW+TMO delta +0.000883 → about **+0.088%** geometric likelihood over FrEW.

However, FrEW+TMO is **not confirmed**:
- only 4/7 validation blocks positive;
- HAC-style paired delta test remains non-significant (p ≈ 0.55);
- moving-block bootstrap 95% CI for incremental delta ≈ **[-0.00246, +0.00353]**.

## Experiment E — external transport 2200..2499

The fixed FrEW+TMO configuration was transported to an older block not used in its selection.

- FrEW mean JLS: -0.000962
- FrEW+TMO mean JLS: +0.001470
- Incremental delta: +0.002432
- Delta bootstrap 95% CI ≈ **[-0.00112, +0.00630]**

Direction is encouraging but uncertainty still includes zero.

## Experiment F — post-protocol targets 4252..4256

No parameter repair was permitted.

| Contest | FrEW JLS | FrEW+TMO JLS | Delta |
|---|---:|---:|---:|
| 4252 | -0.035002 | -0.058740 | -0.023738 |
| 4253 | -0.063433 | -0.114107 | -0.050674 |
| 4254 | +0.033674 | +0.050091 | +0.016416 |
| 4255 | -0.036327 | -0.010416 | +0.025911 |
| 4256 | -0.024865 | -0.022729 | +0.002136 |

Mean:
- FrEW: **-0.025191**
- FrEW+TMO: **-0.031180**
- incremental: **-0.005990**

Geometric interpretation of the five-target mean:
- FrEW: about **-2.49%** vs Uniform
- FrEW+TMO: about **-3.07%** vs Uniform
- TMO incremental effect: about **-0.60%** relative to FrEW

Therefore the apparently stronger validation result of FrEW+TMO **did not survive the first five post-protocol targets as an average improvement**.

## Current scientific verdict

### FIB-Metric
**PASS as a research/evaluation framework.**

It correctly distinguishes:
- selection quality;
- concentration/assembly;
- absolute skill vs Uniform;
- temporal-shuffle behavior;
- model selection vs final-blind generalization;
- JLS-compatible comparison.

It is not a predictor and should not receive production authority.

### TMO — Temporal Multiscale Oscillation
**RESEARCH_ONLY / INCONCLUSIVE.**

Evidence:
- no standalone absolute advantage in three-game Final Blind;
- some temporal alignment under shuffle control;
- FrEW+TMO improved historical validation JLS;
- incremental improvement not significant or stable by blocks;
- five post-protocol targets worsened on average.

Do not integrate into FrEW production. Retain as a shadow hypothesis only.

### RCB — Recency-Cycle Balance
**CURRENT FORMULATION REJECTED as predictive addition.**

The robust hybrid selection chose `eta = 0`, meaning the evidence preferred FrEW without RCB.

RCB may remain a conceptual **concentration/composer research role**, but the present formula has no empirical right to influence FrEW.

### X — cross-game context
**CURRENT FORMULATION NOT SUPPORTED.**

It did not beat Uniform robustly and reduced FrEW validation JLS in the hybrid test. It should remain out of production.

### FrEWBayesian
Still the strongest reproducible scientific baseline among the tested structured models, but its frozen project verdict remains **MODEL_ADVANTAGE = NOT_PROVEN**.

## Formula weights used in the direct FrEW-comparable experiment

### TMO
[
  [
    "freq10",
    -0.10034502297639847
  ],
  [
    "freq40",
    -0.38830873370170593
  ],
  [
    "trend40",
    -0.1894286870956421
  ],
  [
    "freq320",
    -0.03609587252140045
  ],
  [
    "trend320",
    0.28582167625427246
  ]
]

### RCB
[
  [
    "trend40",
    -0.1466040164232254
  ],
  [
    "overdue",
    -0.16480468213558197
  ],
  [
    "recency20",
    0.08583652228116989
  ],
  [
    "recency40",
    0.34993723034858704
  ],
  [
    "global",
    -0.0015016569523140788
  ],
  [
    "cycle_ratio",
    0.2513158917427063
  ]
]

### X
[
  [
    "spread_freq10",
    -0.08292818069458008
  ],
  [
    "other_freq40",
    0.34480756521224976
  ],
  [
    "other_trend40",
    -0.20204290747642517
  ],
  [
    "other_exp80",
    0.3702213168144226
  ]
]

## Integration recommendation

Do **not** modify FIB production or FrEWBayesian formulas from these results.

The safe next implementation should add only a **read-only / shadow scientific laboratory seam** that can:
1. ingest frozen historical snapshots;
2. run registered candidate formulas;
3. run FIB-Metric;
4. run a Skeptic/adversarial review;
5. emit human and machine-readable research reports;
6. propose promotion without applying it.

Any website/UI echo should consume research status through a read-only contract. No UI component should alter model authority, parameters, champion state, prediction store, or official evidence.

## Correction to earlier exploration

An earlier exploratory discussion referred to a TMO candidate with a stronger 4256 Revanchita coverage result. Under the stricter reproducible candidate-selection protocol used in this report, that exploratory result is **not carried forward as evidence**. Only the results in this document should be used for the Codex implementation decision.
