What is the probability actually about?
The homepage always reserves a place for the probability percentage. It estimates the chance of another automatic Codex allowance reset in the next 24, 48 or 72 hours. These are cumulative windows from the latest data refresh. A banked reset credit is a separate event, and your own weekly timer is separate too.
Our historical labels are public announcement times, used as a proxy for reset timing. We cannot measure when a reset reaches an individual account. An announced reset stays “awaiting rollout” until a newer completed entry appears in the public feed; a deadline passing is never treated as completion.
A percentage and a Pacific countdown
NR-1.2 keeps the percentage visible alongside an announced delivery window. The main estimate covers the next 24 hours. The countdown shows time remaining to the end of the interpreted window, using Pacific Time (America/Los_Angeles, including daylight saving). It is not an exact rollout appointment: a reset may arrive sooner.
For an explicit, source-linked promise whose deadline is still ahead and falls inside the forecast horizon, we assign a 95% announcement prior. We combine it with the history baseline as P = 0.95 + 0.05 × P(history), capped at 98%. For example, a 19% historical baseline gives about 96%. This is a declared modeling assumption, not a measured 95% fulfillment rate or a calibrated prediction. The calculation assumes the historical baseline applies if the promised delivery does not happen.
The prior does not increase just because the countdown is shorter. If the deadline falls outside a horizon, is unspecified, or has passed without confirmation, that horizon uses history alone; we do not invent an early or late delivery distribution. An overdue window stays “awaiting confirmation.” Fresh history is still required for a numerical estimate; unavailable data leaves the percentage slot marked “—%.”
The history model and weaker signals
- History sets the starting point. NR-1 uses 45 automatic or combined announcements (44 completed intervals) from our selected archive. Banked-only entries are excluded. The ongoing wait counts as time without an observed event, not as a completed interval.
- Recent history matters more. Past intervals receive an exponential weight with a 90-day half-life. We estimate a daily event rate in waiting-time bands of 0–2, 2–5, 5–10, 10–20 and 20+ days. Each band is smoothed toward the weighted overall rate with 14 prior exposure-days. This avoids extreme predictions from one or two observations.
- Public evidence adjusts the odds. Vague hints and explicit negative statements adjust the odds when no direct commitment is active. A timed commitment adds the announcement prior for horizons containing its deadline, and a countdown. A recent Codex or ChatGPT Work service recovery contributes a small, fading adjustment. The same recovery is not counted again when the reset promise already discusses fixes.
rate = (weighted events + 14 × overall rate)
/ (weighted exposure days + 14)
P(reset within h) = 1 − exp(−integrated rate)Integrate the rate from the current waiting age to that age plus the selected horizon.
SIGNAL FUSIONadjusted odds = history odds × Tibo weight × context weight
probability = adjusted odds / (1 + adjusted odds)“pp” on the homepage means percentage points. Each contribution is the change after that adjustment, in the displayed order.
Signal weights you can inspect
| Signal | Odds multiplier | Rule |
|---|---|---|
| Explicit timed reset promise | 95% prior | P = 95% + 5% × history baseline, only when the future deadline falls inside the horizon. Otherwise history alone. |
| Vague hint | Up to 2× | Must mention a reset. Unrelated product posts do not qualify. |
| Explicit denial / cancellation | Down to 0.4× | Uses the newest qualifying signal, replacing an older promise. |
| Recent relevant recovery | Up to 1.15× | One resolved incident in the last 24 hours; fades to neutral. Not stacked with an explicit reset commitment. |
| Banked credit / unrelated post | 1× | No effect on the automatic reset forecast. |
These signal weights are editorial assumptions, not learned coefficients. The signal-adjusted probability has not been calibrated against a historical tweet dataset. Weak signals expire after at most 36 hours and fade during their last 12 hours. A direct commitment switches the display mode and stops statistical probability output. A newer recorded reset supersedes the older announcement. Estimates are capped at 98%; this cap is not evidence of accuracy.
What the historical backtest says
The history calculation is unchanged in NR-1.2. These results do not validate the 95% announcement prior.
We replayed 162 daily forecast origins, 2026-03-28 through 2026-09-05, fitting only announcements dated before each origin. Each forecast has a full 72-hour outcome window. At least 12 completed intervals are required before evaluation starts.
| Window | NR-1 Brier score | Constant-rate baseline | Recorded positives |
|---|---|---|---|
| 24 hours | 0.1622 | 0.1616 | 31 / 162 |
| 48 hours | 0.2517 | 0.2532 | 58 / 162 |
| 72 hours | 0.2820 | 0.2899 | 79 / 162 |
Brier score measures squared probability error; lower is better, zero is perfect. The baseline uses the same recency-weighted history at a single constant rate. The 24-hour model does not beat that baseline in this replay. The longer windows show small improvements, without evidence of a reliable predictive edge. Historical predictions also tend to understate the observed event frequency.
This is a retrospective, incomplete catalogue. We do not have historical discovery timestamps, so the replay cannot account for reporting delays or entries that were missed. Overlapping 48/72-hour outcomes are correlated. This is a history-only backtest; it does not validate today’s tweet-adjusted probability.
Download every backtest prediction ↗
Where the live evidence comes from
The forecast service checks Codex Resets’ public API for newer reset entries, scheduled announcements and active watches, retaining Tibo’s original X links. We use the source text, not that site’s probability. This is a curated reset-signal feed, not complete access to Tibo’s X timeline. A missing signal does not prove he has posted nothing.
Service context comes from the official OpenAI status page. General API or ChatGPT incidents without Codex / ChatGPT Work scope do not change the estimate. Community likes, repost totals and anonymous rumors do not receive numerical weight.
Requests are cached for up to five minutes. An open, visible homepage refreshes every five minutes; there is no background account monitoring. Every response includes source checks and an expiry time. If the history feed fails, the live percentage is unavailable. If another source fails, the estimate uses only available evidence and says so. Expired snapshots keep the percentage slot but replace its value with “—%.” If an announced delivery window passes without confirmation, we show “awaiting confirmation” instead of inventing a reset.
We checked Tibo’s original announcement: “a reset is also landing by midnight today.” The source does not specify a timezone; the public feed interprets the deadline as September 12, 00:00 PDT (Pacific Time). That interpretation is not an official exact delivery time.
What improves this model next?
A dated, prospective log of predictions, public signals and eventual outcomes is needed before fitting tweet weights or claiming calibration. NR-1 is a first transparent model, not a proven advantage. Model versions and the current retrospective backtest are published here so the claims can be checked.