← Back to blog

My pipeline said OK. The data had been wrong since January.

This is the message Telegram sent me every weeknight when the STAIR pipeline finished:

✅ STAIR pipeline OK — 2026-08-14. Precio: 776.340027, VIX: ok, News: ok, Macro: pendiente (FRED).

Green. Exit code 0. The external monitor got its ping. All good.

Meanwhile, one of the features my models train on had 147 wrong values in the database, from 2 January to 7 August 2026. No test caught it, no pipeline validation, no monitor. I found it by accident, while building a tool for something else.

If you came from the video, this is the long version: what exactly went wrong, why nothing fired, and what would have had to exist to see it sooner.

One clarification first, because it matters. This is more than seven months of corrupted data, not seven months of a bug running. STAIR hasn't been in production that long. The pipeline ingests history, so a configuration error can write months of bad rows without having been alive for those months.

The feature

The feature is yield_spread_10y3m: the difference between the yield on the 10-year US Treasury note and the 3-month bill. It's one of the classic macro inputs in any market model, because it compresses into one number what the bond market expects from the economy. In my system it's a core feature: if it's missing from a row, that row doesn't make it into training.

Both legs come from FRED, the St. Louis Fed's database. The 10-year is the DGS10 series, daily. For the 3-month bill, FRED offers two series that look very much alike:

  • DGS3MO: 3-month constant maturity yield, daily.
  • TB3MS: 3-month bill rate on the secondary market, monthly.

Same tenor, nearly the same name, different frequency. My pipeline asked for the second one.

The bug: one line of config

The list of series the pipeline pulls from FRED lives in settings.yaml. It said TB3MS. According to git log -S, that line came in with the commit that assembled the full pipeline, on 22 April 2026.

The irony is that the code did know the right series. In daily_pipeline.py there was a block with DGS3MO, commented "daily 3M CMT (matches historical base)". But it was the default value of a .get() that only applies if the YAML doesn't define the series. And the YAML did define them. The correct version was written, documented and unreachable: two sources of truth, and the wrong one won.

A monthly series inside a daily table raises no error. It gives you a value that repeats day after day until the next month. Here are the first two weeks of January in my database versus FRED (in percentage points):

Date In my database FRED Difference
2026-01-02 0.64 0.54 +0.10
2026-01-05 0.64 0.53 +0.11
2026-01-06 0.64 0.55 +0.09
2026-01-07 0.64 0.53 +0.11
2026-01-08 0.64 0.57 +0.07
2026-01-09 0.64 0.56 +0.08
2026-01-12 0.64 0.52 +0.12
2026-01-13 0.64 0.51 +0.13
2026-01-14 0.64 0.48 +0.16
2026-01-15 0.64 0.49 +0.15

The real value moves between 0.48 and 0.57. Mine is 0.64 every single day. None of those numbers is absurd: a spread of 0.64 is perfectly plausible. That's the problem.

Why nothing fired

I went through every layer that should have caught it, and each one had a good reason not to.

The tests. The macro feature tests run on synthetic series: they check that the columns exist, that the values make sense, that the ranges are reasonable. None of them depends on which series ID is requested from FRED. A test that makes up its own data can't find out that the real source is a different one.

Validation. Any check that looks at a value in isolation — does it exist, is it the right type, is it in a reasonable range — accepts 0.64. Because it is fine, as a number. What's wrong isn't the value, it's that it doesn't belong to that day.

The monitor. This is the part that stings most. The Telegram message did say something: "Macro: pendiente (FRED)", pending. But the script treated it as information rather than a warning, because FRED publishes with a one-to-two-day lag and I had decided that was normal. The warning showed up every night. And a warning that shows up every night stops being read.

The common pattern is that every check compared the system with itself. Did the flow finish? Is today's row there? Does the value look right? None of them asked the only question that mattered: does what I've stored match what the source says?

How it surfaced

I didn't find it by looking for it. In August the same bug started showing up more visibly: from 10 August the spread was being written as NULL. The pipeline log said it plainly:

FRED [TB3MS] → 'rate_3m': 0 non-null obs, 100.0% NaN

The daily cron only asks FRED for the last few days, and in a window that short a monthly series has no observations at all.

A NULL in a core feature has a quiet side effect: the data loader drops any incomplete row, so the training window was getting shorter with nothing to say so. Chasing those gaps led me to settings.yaml, and I switched the series to DGS3MO. To fill the NULLs I wrote a tool that compares market_features against FRED row by row, with a tolerance of 0.005.

On 24 August I ran it over the whole history, from 2000 onwards. Besides the gaps, it returned this:

VALORES NO-NULL QUE NO COINCIDEN CON FRED (solo informe — NO se tocan)
  [yield_spread_10y3m] 147 filas | 2026-01-02 -> 2026-08-07

(Non-null values that don't match FRED, report only: 147 rows.)

The NULLs were just the visible tip. Before it started leaving gaps, the config had spent months writing plausible, wrong values. The same tool over 2024 and 2025 found not a single discrepancy: everything before 2026 matches FRED within tolerance.

I haven't reconstructed with certainty why the boundary falls exactly at the turn of the year, and I'd rather say so than make up an explanation.

The fix, and the fix I didn't apply

Fixing new values meant changing one line. Fixing the 147 historical ones had a catch.

I already had a repair script that recomputes macro history from FRED. Its dry run showed it didn't just recompute the spread: it cascaded through twelve columns, including VIX-derived features that had nothing to do with this bug, and left thirty new gaps in another core feature that was complete at the time. Fixing 147 cells at the cost of breaking others was a bad trade, so I didn't run it.

Instead I added an option to the backfill script that applies exactly what the reconciliation had only been reporting, restricted to that one column. Dry run first: 147 corrections planned, no other column touched. Then for real. On 29 August the audit over all of 2026 came back "Sin discrepancias", no discrepancies, and January was moving day to day again: 0.54, 0.53, 0.55…

What I didn't measure

I didn't measure how much the models changed because of this bug. The data is fixed and any later retrain starts from the corrected base, but I didn't run a before-and-after comparison, so I'm not going to imply an impact I don't have.

It's also worth saying what this bug isn't. It isn't the story of a bug that wrecked some brilliant results: STAIR's central finding is already that no model robustly beats the baseline. It's the story of a bug nobody would ever have seen, because nothing in the system was designed to see it.

What I'm taking away

Validate against the source, not against yourself. Tests, validation and monitoring all checked internal consistency. Only a comparison with the original source could spot a value that was plausible and false. That comparison now exists as a tool in the repository.

A permanent warning isn't a warning. If an alert appears every night and you've decided it's normal, you've built a filter, not a monitor. "Pending (FRED)" should have had an expiry: pending for a day, fine; pending for several sessions in a row, alarm.

Zero rows isn't success. A download that returned zero observations finished exactly like one that returned a hundred. An empty response from a source that should have data is a failure, even if nothing throws.

One source of truth for configuration. The right value was in the code, in a branch that never ran. A default that contradicts the real configuration isn't a safety net: it's false documentation.

A plausible wrong value is worse than a gap. The NULL in August broke something, and that's why I saw it. The 147 values before it broke nothing. Bugs that make noise get fixed; quiet ones stay.


STAIR is a research platform built as a final degree project at the Universidade da Coruña. It is not financial advice. The figures in this article come from the pipeline's real output and from the reconciliation tool run against FRED.