
First published 2026-07-30 · updated 2026-07-30 · translated from the Finnish original
Can AI Make Things Cheaper—and Governments Riskier?
How Sami Miettinen used Gemini 3.6 Thinking, Codex GPT-5.6 SOL, Python and public FRED data to turn an intuition about AI, wages, tax bases and sovereign risk into a falsifiable working paper—and why the first test did not confirm the thesis.
How one economic intuition became a tested, criticized and deliberately inconclusive working paper with Gemini, Codex, Python, FRED and the Tekoälytyö bots.
I was one of the quantitative high performers at Helsinki School of Economics. A professor once suggested that I continue into academia. I chose investment banking and entrepreneurship instead.
This week I briefly returned to the path not taken. I started with a hypothesis that sounded plausible but was too loose to be useful:
If AI substitutes for human labour quickly enough, it may reduce wages and the prices of labour-intensive services. If governments still rely heavily on labour taxation, their tax base may weaken at the same time as transfers rise. Inflation expectations could then fall while fiscal risk rises.
The important words are may and if. AI can also raise productivity, profits, consumption and new forms of employment. Those effects can expand other tax bases and dominate the negative channel. The hypothesis is not “AI automatically causes deflation and default.” It is a conditional chain whose links should be observable—and breakable.
The result is the working paper “AI Deflation, Sovereign Risk, and the TIPS Signal.” You can download the full 15-page working paper as a PDF.
The first empirical verdict is simple: the public-data exercise did not confirm the AI-risk thesis. It did something more useful than manufacturing a significant result. It showed where the measurement and identification chain fails.
From a narrative to code
Gemini 3.6 Thinking deserves real credit for the early work. It helped turn the verbal idea into an initial research architecture: a dated configuration, synthetic macro data, event-study features, a Newey–West interrupted-time-series estimator and a corporate-bond panel difference-in-differences skeleton. That was a strong first translation from story to executable model.
I then moved the project into OpenAI Codex using GPT-5.6 SOL in Extra High reasoning mode. Codex converted the prototypes into an ai_tips_analysis Python package, added data validation and a complete pytest suite, repaired the event-study normalization and fixed-effect specification, ran public FRED diagnostics, scanned alternative break dates and rebuilt the paper around what the data could actually support.
The software checks matter because synthetic data can create false confidence. The tests therefore verify mechanics rather than the economic thesis:
- the omitted event month is correctly fixed at
event_-1; - every event-by-exposure term equals its event indicator times labour intensity;
- the macro models return HAC/Newey–West covariance estimates;
- the corporate panel absorbs entity and month fixed effects without rank failure;
- calibrated synthetic data recover the deliberately imposed coefficient directions;
- missing observations fail cleanly instead of silently changing the sample.
The final suite contains 14 tests, all passing. That proves the implementation behaves as specified. It does not prove the hypothesis. Synthetic data recover the signs because the signs were placed in the data-generating process.
What the FRED exercise tested
The public diagnostic used three monthly series:
- BAA10Y, the Moody’s Baa corporate bond yield spread over the ten-year Treasury;
- T5YIE, the five-year TIPS breakeven inflation rate;
- STLFSI4, the St. Louis Fed Financial Stress Index.
The macro model was an interrupted time series estimated with HAC/Newey–West standard errors. January 2024 was the focal break. Because choosing one convenient date can manufacture a story, Codex also fitted every plausible monthly break over the preceding four years.
The scan produced 44 candidate breaks; 38 had at least 12 post-break months. The relative slope changes had the narrative’s signs—credit up relative to its old trend, breakeven down relative to its old trend—on 35 of those 38 dates. Both relative changes survived false-discovery-rate adjustment on only three dates, clustered in September–November 2023. That sounds less impressive once the stronger test is applied.
The strong prediction requires the actual post-break credit slope to be positive while the actual post-break breakeven slope is negative. That occurred on only one of 38 adequately powered dates. It occurred on zero of 38 dates with both legs statistically robust after false-discovery-rate adjustment.
At the focal January 2024 break, the slope-change coefficients pointed in the hypothesized relative directions, but both fitted total post-break slopes were positive. The stronger joint prediction therefore failed even in the focal specification.
The joint best-fitting break by BIC was September 2022, before ChatGPT’s public launch. That is much more consistent with a pre-existing inflation, monetary-policy or credit regime than with a frontier-model event.
So the honest summary is:
The relative trend changes are easy to find. The hypothesized joint post-break state is not. The timing points away from an AI interpretation.
This is why a four-year break scan is an audit, not an AGI detector. Model-release dates are annotations. AI adoption diffuses across firms, occupations and countries; it is not a single clean macro shock.
The TIPS lesson
The original intuition gave TIPS too much work to do.
A breakeven rate is approximately the nominal Treasury yield minus the real TIPS yield. It is not a pure inflation forecast. It also reflects inflation-risk compensation, relative liquidity, the TIPS deflation floor and other market conventions. More importantly, a credit component common to nominal Treasuries and TIPS should largely cancel in the difference. A breakeven can reveal differential treatment between the two instruments, but it is not a direct meter of U.S. sovereign default probability.
BAA10Y has the opposite problem: it is a corporate credit spread, not a sovereign one. The two public proxies are useful for testing code and inspecting macro co-movement, but they do not measure the core causal variables in the hypothesis.
That distinction became sharper after discussions with my former classmate Tuomo Vuolteenaho, who earned his PhD in finance at the University of Chicago and later served as an economics professor at Harvard. His research and career have focused on how information about cash flows and risk is incorporated into market prices.
Tuomo was my guest in Neuvottelija episode 280, “Velkakriisiin keulaportti auki”. We discussed U.S. debt, real rates, TIPS, inflation and the limits of financial repression. His criticism of this project was direct: U.S. nominal Treasuries and TIPS do not offer a meaningful default-risk wedge for this purpose; normal-times TIPS liquidity is unlikely to explain an economically large effect; and the component that can actually be measured is the inflation-risk premium. It is also quite possible that my thesis is simply wrong.
That is exactly the kind of criticism a working paper needs. Tuomo gave his response in five minutes and will study the paper in greater depth when he has more time.
The Tekoälytyö review room
I also sent the paper to the bots in our Tekoälytyö WhatsApp group. The review request named Samantha; Amon Ra, Lauri’s bot; Donna, Ville’s bot; Emelia, Eemeli’s bot; and Jessie, Jussi’s bot. Jessie delegated the deeper search review to Hakutaku and explicitly gave it much of the credit.
The comments converged on a better research design.
Samantha argued that the central estimand should be a triple interaction: pre-measured AI exposure × realised adoption × dependence on labour taxation. She also insisted on measuring the transmission chain in order—wage bill, net tax revenue, primary balance, then direct credit risk.
Donna’s compressed verdict was excellent: the proxy failed; the theory has not yet. A corporate spread and a raw breakeven cannot settle a sovereign-risk hypothesis.
Emelia proposed a gated empirical programme. First establish an effect on wages, hours or employment. Then look for the tax-and-transfer effect. Only after those links appear should the analysis move to CDS or cash–OIS spreads. She also warned against placing post-treatment mediators into one final regression and pretending the resulting coefficient is causal.
Amon Ra identified the deepest problem as the AI shock itself. Pre-AI occupational exposure can correlate with education, services intensity, digitalisation and institutions. Realised adoption is endogenous too: weak firms may automate faster. A credible study needs a pre-specified causal graph, pre-trends, placebo exposures, sector outcomes and a defensible source of adoption variation.
Jessie and Hakutaku added the most important competing explanation. Falling inflation expectations and widening credit spreads are also the textbook pattern of an ordinary demand shock or recession fear. A distinguishing prediction is therefore essential. Supply-side AI deflation should combine lower unit labour costs with stable or rising output and margins in exposed sectors. Recessionary deflation should not.
What should be tested next
The next study should be pre-specified as a country or euro-area panel, not another search for the most attractive U.S. break date.
The first gate is the labour market: do highly exposed occupations experience weaker real wage bills, hours, employment, vacancies or worker flows after measured adoption? The second gate is fiscal transmission: do labour-tax receipts fall, transfers rise and primary balances weaken relative to credible controls? The third gate is heterogeneity: are the effects larger where labour taxation and debt-service burdens are high and the ability to tax capital income is weak?
Only then should the project test direct sovereign-risk measures such as liquid sovereign CDS, cash–OIS or asset-swap spreads, with currency regime, redenomination risk, liquidity and monetary capacity handled explicitly.
Inflation should become a separate branch of the design. It should combine swaps, surveys and the distribution implied by caps and floors rather than asking a raw breakeven to identify everything. Tuomo’s question—how large is the inflation-risk premium?—belongs in the measurement table, with estimates shown across models rather than assumed away.
The euro area may offer the closest thing to a laboratory: shared monetary policy and partly shared inflation conditions, but country-specific fiscal structures and sovereign spreads. At the corporate level, the paper’s near-term test is simpler: compare bond-spread changes across firms with pre-measured AI exposure and realised adoption.
The rejection rule must be equally clear. If adoption does not reduce the wage bill of exposed groups—or if higher productivity, profits, consumption, new work and tax adaptation compensate for it—the proposed sovereign-risk channel should be rejected or narrowed. AI may still be transformative or deflationary. This particular fiscal mechanism would have lost.
What this experiment actually demonstrated
AI did not replace scientific judgement here. It made a rigorous loop faster:
intuition → formal hypothesis → code → tests → public data → robustness scan → hostile review → narrower claim.
Gemini 3.6 Thinking did valuable early modelling work. Codex GPT-5.6 SOL supplied the engineering discipline and statistical audit. The Tekoälytyö bots generated unusually substantive peer-style criticism. Tuomo supplied the expert challenge that prevents an elegant TIPS narrative from surviving on intuition alone.
The paper does not prove that AI-driven wage deflation will raise sovereign risk. It does not yet disprove the broader conditional mechanism either. It turns the idea into something more valuable: a claim that can be measured in stages and genuinely rejected.
That is progress.
Paper, episode and data
- Download “AI Deflation, Sovereign Risk, and the TIPS Signal” (PDF)
- Watch Neuvottelija episode 280 with Tuomo Vuolteenaho
- FRED: BAA10Y corporate credit spread
- FRED: T5YIE five-year breakeven inflation
- FRED: STLFSI4 financial stress index
Disclosure: this is a non-peer-reviewed working paper and a methods experiment, not investment advice. The FRED exercise is descriptive and model-dependent; it does not identify a causal AI effect.
Download the working paper (PDF)
Markdown mirror: index.md