The strongest result is a targeted improvement in afternoon demand forecasting and uncertainty estimation. The project has not established profitable trading. Every historical comparison below uses revised EIA data, with assumed publication clocks; F002 and F003 reuse an already inspected development year.
1. Much of the first improvement is simple bias correction
F001 compared three fixed methods on 357 common days in 2025.
| Forecast | MAE, MW | Signed forecast bias, MW | MAE improvement versus operator |
|---|---|---|---|
| Published operator peak | 1,626.71 | +905.13 | — |
| Mean earlier-error correction | 1,485.61 | +317.03 | 8.67% |
| Fixed ridge correction | 1,425.53 | +148.98 | 12.37% |
The simple correction accounts for 70.14% of ridge's total MAE reduction. Ridge improves a further 4.04% over that stronger simple comparator. This is a useful constraint on the story: systematic operator errors mattered, and a large share of the gain did not require a complex model.
Paired ridge-minus-operator absolute error averaged −201.18 MW, with adjusted descriptive interval [−311.34, −101.91] MW. Eight evaluation dates remain unavailable: four fail raw-data/time checks and four later dates lack the exact seven-day lag. They are retained in the calendar, not replaced or concealed. F001 report
2. Afternoon information helps; blindly carrying the full error does not
F002 compared four distributions at all three clocks, on the same 361 scoreable days per clock.
| Chicago clock | Calibrated operator MAE | Full update MAE | Half update MAE | Ridge MAE |
|---|---|---|---|---|
| Noon | 1,192.29 | 1,428.91 | 1,203.06 | 1,265.70 |
| 14:00 | 1,184.06 | 1,267.07 | 1,125.06 | 1,006.11 |
| 15:00 | 1,165.96 | 1,096.41 | 1,031.09 | 830.25 |
At 15:00, ridge improves MAE by 19.48%, CRPS by 16.78%, and Brier by 15.00% versus half update. At noon, ridge loses to calibrated operator on all three metrics. Full update also performs poorly early in the day: the latest observed forecast error need not persist unchanged through every remaining hour.
The 19.48% comparison is descriptive. Ridge-versus-half was not one of F002's registered interval contrasts, so no new confidence interval was chosen for it after inspection. Registered ridge-versus-full CRPS contrasts favor ridge at 14:00 and 15:00; the noon interval includes zero.
The calibrated operator is not the raw operator point forecast: it already uses a separate residual library and the revised observed maximum. The unadjusted operator peak has MAE 1,615.77 MW on this panel. Comparing against the stronger distribution baseline avoids assigning its calibration benefit to ridge. F002 report
3. Better average accuracy does not guarantee reliable tails
F002's 15:00 ridge covered only 86.43% of outcomes with its nominal 90% interval. F003 tested one specific response: scale complete residual vectors using a separate tuning window.
These are matched F003 controls, not a comparison between F002 and F003:
| Clock | Ridge 90% interval score, unchanged → scaled | Improvement | Coverage, unchanged → scaled | Width, unchanged → scaled, MW |
|---|---|---|---|---|
| Noon | 7,524.10 → 7,269.77 | 3.38% | 82.27% → 90.03% | 4,100.49 → 5,162.36 |
| 14:00 | 6,910.77 → 5,871.16 | 15.04% | 81.72% → 90.30% | 3,555.81 → 4,376.23 |
| 15:00 | 6,195.67 → 5,524.88 | 10.83% | 81.16% → 88.37% | 2,959.05 → 3,761.27 |
At 14:00, the paired interval-score difference is −1,039.61 MW, with adjusted descriptive interval [−2,107.47, −204.62] MW. The score charges for width and missed outcomes, so the improvement is not a reward for unlimited widening. The noon gain falls below the predeclared 5% practical threshold; noon and 15:00 difference intervals include zero.
Matched CRPS improvements are smaller: 2.42% at 14:00 and 1.25% at 15:00. Both adjusted ridge-versus-ridge CRPS intervals include zero. F003 does not establish that dispersion improves every probability metric.
Three further details matter:
- At 15:00, median MAE worsens 1.18% despite the better interval score. Scaling preserves hourly residual means; it need not preserve the median of the daily maximum.
- At 14:00, scaled ridge's nominal 80% interval covers 78.39%, and its nominal 95% interval covers 94.74%. At 15:00, nominal 95% coverage is only 91.97%. Good performance at one interval level does not certify the full distribution.
- The tuning procedure selected no widening in six of twelve 14:00 ridge months and seven of twelve 15:00 ridge months. The result is not “always use a larger error bar.”
All 24 registered contrasts, five variants and all three clocks remain in the F003 report. F003 reserved an extra 60-day tuning window, making its fit/library data older than F002's; matched controls share that cost. Cross-experiment before/after figures would confound it with dispersion.
4. Repeated outputs are not independent outcomes
F002 generated 4,344 forecasts and 4,332 qualified scores. F003 generated 5,430 forecasts and 5,415 scores. Both represent 361 distinct scoreable operating days per clock, not thousands of independent trials. Methods and clocks share the same daily maximum; thresholds are nested claims on that same outcome.
Every 365-day calendar and 1,095 clock decisions remain visible. December 3 received predictions, but its later quality-qualified label was ineligible. December 4 lacked the observed prefix; December 5–6 lacked forecast paths. A retrospective label exclusion is not a simulated real-time abstention.
Monthly chronological fits prevent direct training on the evaluated month's labels. They do not make the research sequence an untouched holdout: F001 exposed 2025 before F002/F003. Shared seven-calendar-day bootstrap blocks and multiple-comparison adjustment address specific statistical problems; they do not erase development selection or supply original data vintages. Methods
5. A narrow negative trading result is still useful
E-H06 screened adjacent threshold packages in recorded books. For , one YES() plus one NO() pays at least , with an extra dollar when the peak lies in the middle band. Its unconditional gross surplus is
There were zero strictly positive values among 127 eligible pair observations, with a maximum of $0.0000 before fees. Another 1,853 pair observations lacked a required purchasable side. The observations cover one operating date and repeat three pair identities.
That rejects these particular static packages as guaranteed-profit candidates under nonnegative costs. It does not reject a probabilistic band bet, missing observations, other portfolios or future opportunities. No fee/fill simulation was needed to reject a package already at or above its minimum payout, and no “zero-profit trade” was booked. E-H06 report
6. Source freshness is measurable, but not yet guaranteed
An audit of 30 retained EIA responses found full required observation prefixes at three examined evening clocks: 17/17, 18/18 and 19/19 hours. Those are overlapping prefixes of one operating day, not the pilot's noon/14:00/15:00 clocks.
Two post-start observations first appeared 46.20 and 41.20 minutes after their hour endings. The newest usable demand cell ranged from 41.20 to 101.18 minutes old across responses: a new HTTP response is not necessarily new demand information. Initial history already present at startup was separated from newly detected data.
These are detection delays under a five-minute recorder. Neither original publication time nor exact post-write persistence was captured. The sample cannot promise future receipt timing or prove equality with the settlement source. Availability evidence
7. What has not been demonstrated
As of the retained September 7, 2026 09:35 UTC paper checkpoint, there were zero completed decisions, zero orders and zero settled events. All 18 alternative accounts remained at $200. They are different hypothetical policies, not an invested $3,600 portfolio. The initial experiment contains one daily event and cannot identify a profitable winner.
There is no measured energy trading ROI, win rate, Sharpe ratio, annual bankroll curve or funded execution. P002 market capture stopped earlier that morning under a separate diagnosis; continuous coverage must not be inferred from model or listener status. The latest saved paper audit at the time was the September 6 startup audit, not a fresh audit of hypothetical future fills.
The next decisive evidence is whether a frozen forecast contains information beyond contemporaneous prices, survives real receipt and entry delays, and produces favorable conditional paper cash flows after the declared costs and missingness rules. Actual orders would require separate execution evidence. The forecast results do not answer those questions.
What the engineering evidence supports
Independent arithmetic checks reconstructed 1,071 F001 predictions, all 4,344 F002 outputs and 4,332 scores, and all 5,430 F003 predictions, 5,415 scores, 17,280 tuning forecasts and 24 bootstrap contrasts. F003 checked 342 artifacts and 10,000 draws through 1,420,372 numerical comparisons. F001/F002 bootstrap artifacts were checked by hash, not independently recalculated.
These audits support confidence that the reported calculations match retained artifacts. They do not verify original EIA editions, establish an exact ERCOT target, or prove future profitability. The project combines established statistical methods with explicit controls on causality, revisions, missing data, costs and failed attempts; no claim of a newly invented forecasting algorithm is needed.
This is a curated energy-only history of completed source checks, forecast experiments and the initial paper setup. It records negative findings and unresolved gates alongside improvements. The publication checkpoint is September 7, 2026, before the first permitted paper decision.
1. Correct the contract and operating-day definition
The initial research request named ERCOT NP6-345-CD. Inspection of actual contract rules corrected the target to NP6-346-CD, Actual System Load by Forecast Zone, and specifically the maximum unrounded TOTAL in the first complete daily CSV. Later revisions are ignored for that contract. The date-like event ticker also pointed to the day after the ruled operating date; dates therefore come from rules and metadata, not ticker parsing.
Official ERCOT product/schema documentation was available, but direct report acquisition encountered HTTP 403 responses and a certificate-authentication redirect. The restriction was not bypassed. Exact historical first-edition forecast and settlement data remained unavailable. ERCOT product definition, official API specifications
Decision: separate a public-data forecasting diagnostic from any claim of exact settlement or historical trading replication.
2. Establish a usable, explicitly revised EIA proxy
The official EIA grid dashboard supplied the full ERCO workbook. A separately bounded acquisition retained 73,883,200 bytes, received September 6 at 21:57:33 UTC. Earlier timeouts and a size-cap refusal were preserved; the larger download was a declared new attempt rather than a silent retry.
Schema inspection found raw hourly demand and forecasts, separate imputed/adjusted fields, explicit hour-ending clocks and a quality sheet. Normalization preserved Chicago 23/24/25-hour days and exact UTC identity. The 2023–2025 hourly panel retained all 1,096 calendar days, including missing and flagged hours. Experiments use raw fields; substituted imputed demand could mechanically agree with a forecast used to create it.
The workbook is revised history received in 2026. It does not demonstrate that any historical value was available at a simulated decision time. EIA dashboard, data definitions, workbook endpoint
3. F001: test fixed peak-residual models
One registered comparison used operator peak, mean earlier residual and fixed ridge correction, with monthly chronological fits and complete common evaluation support. Ridge reduced MAE by 12.37% versus operator on 357 days; most of the reduction was already obtained by mean error correction.
All omissions and both registered comparisons were retained. An independent saved-coefficient check reconstructed 1,071 predictions without refitting. This supported a conditional forecasting lead, not a historical trading strategy. F001 report
Next question: does the day's arriving load error help predict its remaining path, rather than only a corrected daily peak?
4. F002: test noon and afternoon residual updates
Four fixed methods—operator, full update, half update and ridge—were evaluated at noon, 14:00 and 15:00 Chicago. Separate historical donor days supplied whole residual paths, preserving hourly dependence before computing the daily maximum.
The afternoon ridge improved average point and distribution scores; at noon it lost to calibrated operator. At 15:00, its MAE improvement over half update was 19.48%, while its nominal 90% interval covered only 86.43% of outcomes. That combination identified a specific weakness: sharper forecasts could remain overconfident.
Every calendar decision remained in the report. The independent arithmetic audit reconstructed all 4,344 forecast outputs and 4,332 qualified scores; it did not independently rerun bootstrap arithmetic. Since F001 had already exposed 2025, this and later comparisons are reused development evidence. F002 report
5. Check the proxy against finalized contract values
A separate fixed nine-day source comparison found zero exact scalar matches between EIA daily peaks and Kalshi final values, despite agreement on all 113 listed binary threshold outcomes. Absolute differences ranged from 1.70 to 18.53 MW in that small sample.
All 113 outcomes share only nine daily peaks. Agreement at those strikes does not make EIA the exact source, nor establish a bound on future error near a strike. Independently retrieving the first complete ERCOT CSV remained unresolved. Source-header and narrow rule-text parsing failures were retained and addressed through declared amendments rather than changes to prior evidence.
Decision: use EIA as a clearly labeled signal proxy; use validated exchange-finalized payouts for prospective paper cash, without claiming independent ERCOT reconciliation.
6. Prepare a frozen prospective paper family
F004 prepared operator, half-update and ridge models for each of three future decision clocks using history already received by preparation time. The paper family froze source selection, forecast semantics, integer sizing, fees, delayed execution, unresolved-state handling and payout validation before future predictions.
T002's single operating day is September 7, 2026, with decisions at 17:00, 19:00 and 20:00 UTC—noon, 14:00 and 15:00 Chicago. Three models, two risk fractions and three depth/slippage scenarios produce 18 alternative $200 accounts. They share one underlying daily peak and do not constitute 18 independent trials.
The separate operational auditor checks receipt selection, cash and fill arithmetic, first-attempt failures, inconsistent settlements and retained failures. It cannot turn displayed quotes into actual fills. F003 was deliberately not inserted into this already frozen family. T002 protocol, independent audit protocol
7. F003: isolate the dispersion problem
One separate registration tested four fixed dispersion scales for half update and ridge, with an older training window, a disjoint residual library and another disjoint tuning window. Every unchanged control used the same windows, making the cost of reserving tuning data visible.
At 14:00, scaled ridge's 90% interval score improved 15.04% and coverage rose from 81.72% to 90.30%. The adjusted descriptive interval favored it. Matched CRPS gains were smaller and their intervals included zero; noon did not meet the predeclared 5% practical score improvement. Every clock, scale-selection result and all 24 comparisons were retained.
The independent auditor reconstructed all predictions, tuning choices, scores and bootstrap contrasts using retained artifacts, without refitting. This was a targeted development result, not a change to the prospective trading policy. F003 report, frozen protocol
8. E-H06: reject the observed static packages early
A registered screen checked a necessary condition for adjacent-threshold packages: could buying YES at a lower threshold and NO at the next higher threshold cost strictly less than their $1 minimum payout?
It found zero strictly positive gross-surplus packages among 127 eligible observations; 1,853 observations lacked a required purchasable side. One initial snapshot lacked eligible metadata. The 127 observations repeat three pair identities on one operating day, during a late-afternoon/evening window rather than the planned forecast clocks.
Because no eligible package passed even before fees, a favorable fee/fill assumption could not rescue that unconditional bound under nonnegative costs. No trade or zero-profit account entry was manufactured. This rejects only those observed static packages, not all directional or probability-based strategies. E-H06 report, screen protocol
9. Measure actual EIA receipt behavior
A finite audit of 30 recorded responses checked the complete observation prefixes at three evening clocks. Required coverage was 17/17, 18/18 and 19/19 hours. Two newly appearing hours were detected 46.20 and 41.20 minutes after hour ending; 47 hours already present at startup were not falsely classified as fresh arrivals.
The newest usable cell aged from 41.20 to 101.18 minutes across successive responses. Polling success is therefore not a substitute for information freshness. The archive's declared availability field was assigned from receipt time, and the ingestion timestamp preceded append; exact post-write persistence was not recorded. The audit reports those limits rather than converting them into a publication-latency claim.
The three evening prefixes overlap and do not prove noon/14:00/15:00 availability on the next day. No model or market price was scored in this audit. EIA availability evidence
10. Gas storage: access works, target vintages still matter
A bounded official EIA probe acquired revision/history workbooks and selected dated gas updates. It made 15 requests including ordinary redirects, retained 1,719,225 bytes, and fitted no model or return strategy.
The original_data sheet contains 565 weekly stock observations through April 10, 2026; latest revised history contains 870 weeks through August 28. Sixty-nine overlapping stock levels differ. An archived December 12, 2025 storage report states a 167 Bcf withdrawal, while the latest history reports 166 Bcf. A first-announced net change therefore cannot be replaced by the latest revised change.
Nor does differencing adjacent original stock levels necessarily recover the announced change: the previous week's comparison stock can be revised in the current release. All 869 comparable latest national changes equal adjacent latest stock differences in the acquired workbook, but this is consistency of that edition, not a universal rounding rule or first-release proof. Retain the explicitly published change, its contemporaneous current/comparison stocks, revisions and reclassifications separately. Net stock change can also differ from implied physical flow when gas is reclassified between working and base gas. Revision policy, revision workbook, revised history
Release dates and expectation availability introduced further traps. A January 10 commentary page discusses a release scheduled for January 8 at noon Eastern. Two sampled updates reproduce survey medians after publishing the outcome; they do not establish pre-release receipt. A JSON date at midnight is not necessarily the actual release time. Official exception schedule, January update, April update
Disposition: partial source feasibility, no complete first-release surprise backtest, no validated executable instrument-price panel, and no gas profit claim. Work stopped at the data gate instead of fitting many models to a misidentified target.
Publication checkpoint and remaining work
At September 7 09:35 UTC, retained paper progress showed zero completed decisions, zero orders and zero settled events. All accounts remained at their initial $200. P002 market capture had stopped at 05:17 UTC under a separate diagnosis, so uninterrupted recording is not claimed. A future decision still requires its own complete, timely evidence; missing inputs cannot be reconstructed into a favorable historical intent.
The next economic test is the frozen source → forecast → intent → first delayed book → conditional fill → finalized payout chain. Longer confirmation would need a new predeclared horizon based on a useful effect size and day-level variability, not a winning week or a selected account. Forecast improvement, successful arithmetic replay, and proven trading value remain separate milestones.
The five files under research/energy/ are exact copies of historical protocol/proposal bytes. Their original status language and source-workspace references are preserved for hash identity. Current interpretation is provided here and in Methods, Findings and Reproducibility; old dates are not authorization to resume or extend an experiment.