<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Carlos Chavez]]></title><description><![CDATA[Peruvian Economist.]]></description><link>https://carloschavezp29.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!zh0V!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffff9b4a5-53ee-4548-85da-563173f13249_2666x2666.jpeg</url><title>Carlos Chavez</title><link>https://carloschavezp29.substack.com</link></image><generator>Substack</generator><lastBuildDate>Thu, 03 Sep 2026 14:35:37 GMT</lastBuildDate><atom:link href="/__u/carloschavezp29.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Carlos Chavez]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[carloschavezp29@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[carloschavezp29@substack.com]]></itunes:email><itunes:name><![CDATA[Carlos Chavez]]></itunes:name></itunes:owner><itunes:author><![CDATA[Carlos Chavez]]></itunes:author><googleplay:owner><![CDATA[carloschavezp29@substack.com]]></googleplay:owner><googleplay:email><![CDATA[carloschavezp29@substack.com]]></googleplay:email><googleplay:author><![CDATA[Carlos Chavez]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Generalized Method of Moments is more influential than I thought.]]></title><description><![CDATA[What modern econometrics inherited from GMM, and what it dropped.]]></description><link>https://carloschavezp29.substack.com/p/the-generalized-method-of-moments</link><guid isPermaLink="false">https://carloschavezp29.substack.com/p/the-generalized-method-of-moments</guid><dc:creator><![CDATA[Carlos Chavez]]></dc:creator><pubDate>Mon, 31 Aug 2026 08:46:30 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!jjmZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a37299-7ca0-43eb-8f1c-cf448cf94db4_672x314.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>First of all, apologies for not publishing as often as before. I have been in the middle of some deadlines on the projects I have been working on and dealing with some personal issues, so I had to decide to stop the Substack essays for a while.</em></p><p>This essay came out of a conversation I had the other day with a friend of mine about GMM (the Generalized Method of Moments), and whether this econometric tool is still relevant nowadays. I have used it in some published papers, especially Arellano-Bond methods, but my prior was that it was not, because credibility revolution techniques are now used in so many places.</p><p>In fact, I have already finished <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6788719">a paper about the returns to the credibility revolution</a>, and how these techniques can increase the likelihood of publishing and affect where you publish, but this is another discussion. The point is that after the talk with my friend, I started to do a literature review about the history and current state of GMM, and what I found was interesting, and the opposite of what I expected.</p><p>The first thing that made me reconsider my position was <a href="https://arxiv.org/abs/2505.09942">a recent paper on triple differences</a> that I had read a few weeks earlier for a completely different reason, and that came back to me while we were talking. There is a section in that a paper where the authors justify their doubly robust estimator, and the argument looked familiar because they take the recentered influence functions of the component estimators, use the fact that influence functions have mean zero, and stack them into a vector of moment conditions whose common solution is the parameter of interest, after which they choose an appropriate weighting matrix. In other words, the estimator can be written as an optimal GMM estimator, but the authors do not make a big deal about it and simply mention it in a remark before moving on.</p><p>Then I started to find the same thing in other papers, and at that point I thought that maybe I had been wrong about GMM.</p><p>For example, <a href="https://arxiv.org/abs/2102.09948">Egami and Yamauchi</a> build their double-DID estimator using a GMM criterion, and they show that standard, extended, and sequential DiD estimators can all be obtained by changing the weighting matrix. <a href="https://jrgcmu.github.io/2sdd_gtty.pdf">Gardner, Thakral, T&#244;, and Yap</a> use the same type of machinery to derive the variance of two-stage difference-in-differences, while double machine learning is based on Neyman-orthogonal moment conditions combined with cross-fitting. Automatic debiased machine learning goes even further because it constructs the moment function itself by estimating a Riesz representer instead of deriving the correction by hand.</p><p>Of course, we would not normally call any of these &#8220;GMM papers,&#8221; and the authors are not trying to hide anything either, because GMM is simply part of the language they are using. This made me curious about how a method that I had associated with older econometrics had ended up being used inside some of the most modern parts of the literature, so I decided to go further back and look at where the method came from. The sequence turned out to be short enough to summarize before going into it, and what struck me while putting it together is that the method was criticized and set aside once, rebuilt decades later for a different problem, criticized again, and then absorbed into a literature that stopped using its name.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!VxRq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47f74f16-9426-4b3f-9d00-4d777ae22aed_690x274.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!VxRq!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47f74f16-9426-4b3f-9d00-4d777ae22aed_690x274.png 424w, /__u/substackcdn.com/image/fetch/$s_!VxRq!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47f74f16-9426-4b3f-9d00-4d777ae22aed_690x274.png 848w, /__u/substackcdn.com/image/fetch/$s_!VxRq!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47f74f16-9426-4b3f-9d00-4d777ae22aed_690x274.png 1272w, /__u/substackcdn.com/image/fetch/$s_!VxRq!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47f74f16-9426-4b3f-9d00-4d777ae22aed_690x274.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!VxRq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47f74f16-9426-4b3f-9d00-4d777ae22aed_690x274.png" width="690" height="274" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/47f74f16-9426-4b3f-9d00-4d777ae22aed_690x274.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:274,&quot;width&quot;:690,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:38865,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/213513034?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47f74f16-9426-4b3f-9d00-4d777ae22aed_690x274.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!VxRq!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47f74f16-9426-4b3f-9d00-4d777ae22aed_690x274.png 424w, /__u/substackcdn.com/image/fetch/$s_!VxRq!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47f74f16-9426-4b3f-9d00-4d777ae22aed_690x274.png 848w, /__u/substackcdn.com/image/fetch/$s_!VxRq!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47f74f16-9426-4b3f-9d00-4d777ae22aed_690x274.png 1272w, /__u/substackcdn.com/image/fetch/$s_!VxRq!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47f74f16-9426-4b3f-9d00-4d777ae22aed_690x274.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>What Fisher Did to the Method of Moments</h2><p>I knew that Karl Pearson had introduced the method of moments and that Ronald Fisher had criticized it, but I did not know how strong the disagreement had been. <a href="https://doi.org/10.1098/rsta.1894.0003">Pearson&#8217;s 1894 paper</a> used the method of moments to estimate a mixture of two normal distributions, and the basic idea was to set sample moments equal to their population counterparts and solve the resulting equations. At the time this made a lot of sense because more complicated estimation methods were difficult to implement, but Fisher did not think that computational convenience was enough to justify the method.</p><p>He started criticizing it in his first paper in 1912, when he was still an undergraduate, and <a href="https://www.stats.org.uk/statistical-inference/Fisher1922.pdf">by 1922</a> he had developed a much more general argument in the same paper where he developed the modern statistical ideas of consistency, efficiency, sufficiency, and likelihood. Fisher&#8217;s main concern was efficiency, and his criticism was much stronger than simply saying that moment estimators could sometimes perform a little worse, because he produced examples within Pearson&#8217;s own family of curves in which the lower bound on the efficiency of moment-based estimates was zero.</p><p>The disagreement continued for years, which is also one of the things I found entertaining while reading about it. There were disputes over what Fisher could publish in <em>Biometrika</em>, Fisher eventually stopped submitting his work there, and in 1937 he published an article called &#8220;<a href="https://onlinelibrary.wiley.com/doi/10.1111/j.1469-1809.1937.tb02149.x">Professor Karl Pearson and the Method of Moments</a>,&#8221; which gives you a good idea of how friendly the disagreement had become.</p><p>However, the important part for the history of GMM is that Fisher&#8217;s efficiency criticism was never really overturned, because what changed later was the estimation problem itself. Hansen himself explains that the original appeal of the method of moments was partly computational, because it provided an easy way to estimate parameters when likelihood-based methods were difficult to use, but as numerical methods improved, maximum likelihood and Bayesian methods became easier to implement and that particular advantage became less important.</p><p>This is where I had misunderstood the history, because I had implicitly thought of Hansen&#8217;s GMM as a rehabilitation of the old method of moments, while the real story is a little different. GMM did not come back because somebody proved that Fisher had been wrong, but because economists started dealing with models where the full likelihood was not available in the first place. Once the question changes from &#8220;Why should I use moments if I know the full probability model?&#8221; to &#8220;What can I do if the theory only gives me a few restrictions on the data?&#8221;, Fisher&#8217;s criticism is no longer the end of the story. This distinction seems obvious now, but I had never really thought about it before doing this literature review.</p><h2>Estimating a Model You Cannot Write Down</h2><p>The change came largely from macroeconomics and, more specifically, from Euler equations, which was another connection that I had somehow never made. I use Euler equations in some of my own work, but I had always thought about them mainly as theoretical conditions rather than as the starting point of an estimation problem. An Euler equation says that some function of observables and structural parameters has conditional expectation zero given an agent&#8217;s information set, so the economic model gives you a set of restrictions on moments of the data.</p><blockquote><p>1 = E&#8348;[ &#946; (C&#8348;&#8330;&#8321; / C&#8348;)&#8315;&#7518; R&#8348;&#8330;&#8321; ]</p></blockquote><p>which is the same as</p><blockquote><p>E&#8348;[ &#946; (C&#8348;&#8330;&#8321; / C&#8348;)&#8315;&#7518; R&#8348;&#8330;&#8321; &#8722; 1 ] = 0</p></blockquote><p>In many models, that is basically the empirical content of the theory, because the theory does not tell you the complete distribution of every variable that appears in the model. If you want to write a likelihood, however, you need to specify all of those missing pieces, which means introducing distributions and parameters that may not come from the economic theory itself. After <a href="https://doi.org/10.1016/S0167-2231%2876%2980003-6">Lucas (1976)</a>, simply filling in those parts for convenience was not an especially attractive way to proceed.</p><p>Hansen describes this problem very clearly when he explains that the parameters economists care about are often not sufficient to write down a likelihood function, because the economic model is only partially specified. <a href="https://www.jstor.org/stable/1912775">Hansen (1982)</a> gives a way to work in exactly this environment, because instead of specifying the parts of the model that we do not know, we can estimate the parameters using the orthogonality conditions that the theory actually gives us.</p><blockquote><p>E[ g(Z&#7522;, &#952;&#8320;) ] = 0</p></blockquote><p>The estimator is then the value of &#952; that makes the sample version of those conditions as close to zero as possible, in the metric given by a weighting matrix W&#8345;.</p><blockquote><p>&#7713;&#8345;(&#952;) = (1/n) &#931;&#7522; g(Z&#7522;, &#952;) and Q&#8345;(&#952;) = &#7713;&#8345;(&#952;)&#8242; W&#8345; &#7713;&#8345;(&#952;)</p></blockquote><p>We can then look at the family of estimators generated by different combinations of those moments and choose the weighting that gives the best asymptotic performance within that family. From this perspective, GMM does not really contradict Fisher, because if we actually knew the correct likelihood, using only a few moments would often throw away information. The interesting case is when the moments are all we are willing to assume, and then the question becomes how much we can learn from them without pretending that we know the rest of the model.</p><p>There were also a couple of things about Hansen&#8217;s paper that I did not know before reading more about it. One is that Hansen is very explicit about Sargan&#8217;s earlier contribution, because <a href="https://doi.org/10.2307/1907619">Sargan</a> had already developed the theory of overidentified instrumental variables and specification testing in a series of papers in the late 1950s, and some of the notation and representation that Hansen later generalized came from that work. I have occasionally seen the story told as if Hansen had somehow taken the idea from a forgotten LSE econometrician, but that is not the story Hansen himself tells, and it is hard to square with the references in the original paper.</p><p>The other thing is that <em>Econometrica</em> did not publish many of the formal proofs in the 1982 article, and Hansen published them thirty years later in the <em>Journal of Econometrics</em>, including a uniform law of large numbers for stationary ergodic processes. I found this detail funny because many of the formal proofs omitted from one of Hansen&#8217;s most influential papers did not appear in print until 2012.</p><p>In his <a href="https://www.nobelprize.org/uploads/2018/06/hansen-lecture.pdf">Nobel lecture</a>, Hansen describes the tradition he was working in as &#8220;doing something without having to do everything,&#8221; and I think this captures the motivation for GMM much better than the usual textbook presentation. We normally learn GMM by starting with a vector of moment conditions, introducing a weighting matrix, and minimizing an objective function, which makes the method look like another estimation formula that has to be memorized. The economic idea behind it is simpler and, in my opinion, much more interesting, because sometimes we know something important about the data-generating process without knowing everything, and we should be able to use what we know without having to invent the rest.</p><h2>The First Application Rejected the Model</h2><p>Another thing that changed after going back to the original papers was my understanding of the first empirical applications. The version I had in my head was that <a href="https://larspeterhansen.org/wp-content/uploads/2016/11/Generalized-Instrumnetal-Variables-Estimation-of-Nonlinear-Rational-.pdf">Hansen and Singleton</a> used the J test in 1982 to reject the consumption CAPM, and that the equity premium puzzle followed naturally from that result, but the actual history is not that clean. Hansen treats the 1982 and <a href="https://doi.org/10.1086/261141">1983</a> papers together and describes the evidence as showing empirical shortcomings of macroeconomic models with power utility preferences, rather than presenting one single test as a decisive rejection of the model.</p><p>He is also careful about separating the work he did with Singleton from the &#8220;equity premium&#8221; label that became popular later. What he emphasizes is that their contribution provided a statistically rigorous way to characterize the anomaly, and he also argues that the problem was broader than simply explaining the difference between the expected return on stocks and bonds. Later research complicated the interpretation even more because economists found that Euler-equation models can suffer from weak identification, which means that both rejection and non-rejection can be difficult to interpret when the moments do not strongly identify the parameters. This is important because the attraction of a specification test is that it seems to give us a clear answer, but that answer is only as useful as the identification behind it.</p><p>The limitation that I found most interesting, however, is one that Hansen himself points out. Suppose we specify a parametric family of stochastic discount factors and find that no value of the parameters satisfies all of the pricing restrictions. The overidentification test gives us a formal way to say that the model does not fit the data, which is already useful, but it does not tell us what the model is being rejected against or what kind of alternative model we should consider. In practice, the conclusion is something like &#8220;this family is wrong,&#8221; and then we are left to decide what comes next.</p><p>This is where <a href="https://larspeterhansen.org/lph_research/implications-of-security-market-data-for-models-of-dynamic-economies/">Hansen and Jagannathan (1991)</a> becomes particularly interesting, because instead of choosing another parametric family immediately, they characterize the larger set of stochastic discount factors that could satisfy the pricing equations and ask what properties those discount factors must have. The result is a bound rather than a point estimate, and when I read the paper from this perspective, it looked surprisingly close to what we would now call a partial-identification argument.</p><h2>Where the Reputation Came From</h2><p>At this point I should also say that my original impression of GMM was not completely unreasonable, because there are good reasons why the method developed a bad reputation in parts of applied economics. A lot of this happened during the 1990s, when GMM became part of the standard econometric infrastructure. <a href="https://academic.oup.com/restud/article-abstract/58/2/277/1563354">Arellano and Bond (1991)</a> turned dynamic-panel estimation into a system of stacked moment conditions that use lagged levels as instruments, and the method became extremely popular in growth regressions, corporate finance, and firm-level empirical work.</p><p><a href="https://www.its.caltech.edu/~mshum/gradio/papers/BerryLevinsohnPakes1995.pdf">Berry, Levinsohn, and Pakes (1995)</a> also made GMM part of the standard approach to estimating demand systems, so by the middle of the decade the method was being used in many areas where economists had relatively complicated models and many moment conditions.</p><p>The problem was that asymptotic efficiency did not necessarily translate into good performance in the samples economists actually had, and the July 1996 issue of the <em>Journal of Business and Economic Statistics</em> is very interesting in this respect. <a href="https://larspeterhansen.org/lph_research/finite-sample-properties-of-some-alternative-gmm-estimators/">Hansen, Heaton, and Yaron</a> studied the finite-sample behavior of alternative GMM estimators for asset-pricing models and introduced the continuous-updating estimator, while <a href="https://doi.org/10.1080/07350015.1996.10524661">Altonji and Segal</a> showed that the asymptotically optimal weighting matrix could produce substantial bias when estimating covariance structures in realistic samples.</p><p><a href="https://www.tandfonline.com/doi/abs/10.1080/07350015.1996.10524659">Christiano and Den Haan</a> reached a similar conclusion for postwar quarterly US data, because the available asymptotic approximations were not always good guides for samples of the size that macroeconomists actually use. What makes these papers particularly important is that the criticism was coming from inside the GMM literature itself, and Hansen was one of the people documenting the problem.</p><p>At around the same time, the identification literature started showing that there were other problems beyond finite samples. <a href="https://stock.scholars.harvard.edu/publications/instrumental-variables-regression-weak-instruments">Staiger and Stock (1997)</a> formalized the weak-instrument problem, while <a href="https://doi.org/10.1111/1468-0262.00151">Stock and Wright (2000)</a> brought weak identification directly into GMM and developed inference that does not depend on the usual local identification condition.</p><p>Dynamic-panel applications later developed another practical problem because researchers could generate very large numbers of instruments, and <a href="https://doi.org/10.1111/j.1468-0084.2008.00542.x">Roodman (2009)</a> showed that instrument proliferation can seriously weaken the overidentification test and generate implausibly high p-values. When the instrument count becomes too large, the Hansen test can lose much of its ability to detect misspecification, which is a strange outcome for a test whose purpose is supposed to be detecting problems with the model. Recent work continues to show similar tradeoffs, and there are still cases in which the estimator that looks more sophisticated asymptotically can behave worse than a simpler alternative.</p><p>While all of this was happening, applied economics was also changing in a very different direction because the credibility revolution was rewarding research designs that made identification more transparent. Difference-in-differences, regression discontinuity, randomized experiments, and instruments based on institutional variation became associated with credible empirical work, while GMM became increasingly associated with structural models, dynamic panels, many instruments, and difficult-to-interpret specification tests.</p><p>This is basically where my own prior came from, and I still think there was some logic behind it. If somebody gives you the choice between a transparent research design and an Arellano-Bond regression with a huge instrument matrix and a Hansen test that never rejects, most applied economists are going to prefer the first one. Where I was wrong was in taking this particular criticism of applied GMM and concluding that the method itself had become irrelevant.</p><h2>What the Literature Kept, and What It Dropped</h2><p>What I had missed is that GMM did not really disappear, because many of its ideas were absorbed into modern econometrics and the acronym became less important. <a href="https://academic.oup.com/ectj/article/21/1/C1/5056401">Chernozhukov and coauthors (2018)</a>, for example, organize double machine learning around moment functions that are orthogonal to first-order errors in nuisance estimates, and cross-fitting makes it possible to estimate those nuisance functions flexibly without affecting the asymptotic behavior of the parameter of interest. <a href="https://doi.org/10.3982/ECTA18515">Chernozhukov, Newey, and Singh (2022)</a> go even further because they automate the construction of the debiased moment by estimating a Riesz representer.</p><p>The modern difference-in-differences literature also uses moment representations to derive estimators and variances, so a student today can learn a large amount of GMM algebra in courses on causal inference, semiparametric estimation, or machine learning without necessarily hearing the acronym very often.</p><p>There is, however, one part of the older GMM framework that seems to have survived much less clearly, and this is the part I kept thinking about after finishing the literature review. In Hansen&#8217;s original framework, overidentifying restrictions give the model a way to be wrong that the data can detect. If the model gives us more moment restrictions than there are parameters to estimate, the system is overidentified, and the extra restrictions give us a way to ask whether the model can satisfy all of the moment conditions simultaneously. This is where the J test enters, because after estimating the parameters there are still q &#8722; p overidentifying restrictions that the model has to explain.</p><blockquote><p>q moment restrictions &gt; p parameters &#10233; q &#8722; p restrictions left to test</p></blockquote><p>This distinction between fitting a model and confronting it with additional evidence also appears in <a href="https://www.aeaweb.org/articles?id=10.1257/jep.10.1.87">Hansen and Heckman&#8217;s (1996)</a> discussion of calibration, where they emphasize the difference between matching selected features of the data and establishing the empirical credibility of a model.</p><p>Many of the modern estimators I was reading, particularly canonical double machine-learning estimators, are designed for a different purpose, because the score is typically constructed around the target parameter rather than around a large collection of overidentifying restrictions, with orthogonality chosen to make that parameter robust to errors in the estimation of nuisance functions. This is extremely useful if the objective is estimation, but it also means that there may be no overidentifying restrictions left for a Hansen-J-type specification test. I do not think this is a problem with double machine learning or with any particular modern method, because the moment conditions are simply being asked to do a different job. The modern literature often designs moments to protect the estimator from nuisance estimation, while part of the older GMM literature also cared about building restrictions that allowed the economic model itself to fail. In that sense, modern econometrics kept a lot of the algebra but seems to have moved away from some of the specification-testing philosophy. The contrast is easier to see side by side, and what matters is not that the vocabulary changed, because most of these rows are continuity under a new name. The last three are where the two frameworks actually come apart.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!jjmZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a37299-7ca0-43eb-8f1c-cf448cf94db4_672x314.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!jjmZ!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a37299-7ca0-43eb-8f1c-cf448cf94db4_672x314.png 424w, /__u/substackcdn.com/image/fetch/$s_!jjmZ!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a37299-7ca0-43eb-8f1c-cf448cf94db4_672x314.png 848w, /__u/substackcdn.com/image/fetch/$s_!jjmZ!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a37299-7ca0-43eb-8f1c-cf448cf94db4_672x314.png 1272w, /__u/substackcdn.com/image/fetch/$s_!jjmZ!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a37299-7ca0-43eb-8f1c-cf448cf94db4_672x314.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!jjmZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a37299-7ca0-43eb-8f1c-cf448cf94db4_672x314.png" width="672" height="314" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/54a37299-7ca0-43eb-8f1c-cf448cf94db4_672x314.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:314,&quot;width&quot;:672,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:58811,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/213513034?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a37299-7ca0-43eb-8f1c-cf448cf94db4_672x314.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!jjmZ!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a37299-7ca0-43eb-8f1c-cf448cf94db4_672x314.png 424w, /__u/substackcdn.com/image/fetch/$s_!jjmZ!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a37299-7ca0-43eb-8f1c-cf448cf94db4_672x314.png 848w, /__u/substackcdn.com/image/fetch/$s_!jjmZ!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a37299-7ca0-43eb-8f1c-cf448cf94db4_672x314.png 1272w, /__u/substackcdn.com/image/fetch/$s_!jjmZ!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54a37299-7ca0-43eb-8f1c-cf448cf94db4_672x314.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>I also do not want to push this argument too far because Hansen himself has warned against saying that &#8220;everything is GMM,&#8221; and I think he is right. In his Nobel lecture, he argues that treating GMM as simply one special case of a much larger family of extremum estimators misses some important features of the original framework, especially the connection between statistical efficiency and restrictions that actually come from an economic model.</p><p>After all, almost anything can be written as a moment condition if we are willing to manipulate it enough, so the interesting point is not that modern estimators can be expressed using moments. What matters is that many economic models naturally give us a relatively small number of moment restrictions while saying very little about the rest of the data-generating process, and these restrictions can still be enough to estimate something useful.</p><p>So I finished this literature review almost in the opposite place from where I started. I thought GMM was an older econometric method that had survived mainly in structural estimation and dynamic panels, while the credibility revolution had largely replaced it in modern applied work. Instead, I found that a substantial part of modern econometrics still works through moment conditions, even when nobody calls the resulting estimator &#8220;GMM.&#8221; The terminology has changed, the estimators are more sophisticated, and the applications look different, but the reason for working with moments is surprisingly close to the one Hansen gave in 1982, because most economic models are only partially specified and we often know some useful things about the data-generating process without knowing enough to write down the entire likelihood.</p><p>The history therefore ended up being much more circular than I expected when I started reading about it. Pearson introduced the method of moments more than 130 years ago, Fisher showed why relying only on moments could be extremely inefficient when a correctly specified likelihood was available, and Hansen brought the idea back for a very different situation in which economists did not want to pretend that they knew the full likelihood in the first place.</p><p>Today the same logic appears inside causal inference, semiparametric econometrics, and machine learning, although the papers often do not call themselves GMM papers and sometimes mention the connection only in passing. The strange part is that the piece of the old framework that seems to have faded is not moment-based estimation itself, but the part that was explicitly designed to tell us when the economic model was wrong.</p><h2>Essential Reading</h2><p>Fisher, R. A. (1922), &#8220;<a href="https://www.stats.org.uk/statistical-inference/Fisher1922.pdf">On the Mathematical Foundations of Theoretical Statistics</a>,&#8221; <em>Philosophical Transactions of the Royal Society A</em> 222, 309-368. Introduces consistency, efficiency, sufficiency, and likelihood, and includes the efficiency argument against moment estimators.</p><p>Hansen, L. P. (1982), &#8220;<a href="https://www.jstor.org/stable/1912775">Large Sample Properties of Generalized Method of Moments Estimators</a>,&#8221; <em>Econometrica</em> 50(4), 1029-1054. The foundational GMM paper and still one of the clearest places to understand why partial specification is central to the method.</p><p>Hansen, L. P. and K. J. Singleton (1982), &#8220;<a href="https://larspeterhansen.org/wp-content/uploads/2016/11/Generalized-Instrumnetal-Variables-Estimation-of-Nonlinear-Rational-.pdf">Generalized Instrumental Variables Estimation of Nonlinear Rational Expectations Models</a>,&#8221; <em>Econometrica</em> 50(5), 1269-1286. One of the first major empirical applications of the framework, together with the 1984 erratum in <em>Econometrica</em> 52(1), 267-268.</p><p>Hansen, L. P. and R. Jagannathan (1991), &#8220;<a href="https://larspeterhansen.org/lph_research/implications-of-security-market-data-for-models-of-dynamic-economies/">Implications of Security Market Data for Models of Dynamic Economies</a>,&#8221; <em>Journal of Political Economy</em> 99(2), 225-262. Particularly interesting when read as a response to model rejection, because the paper asks what can still be learned from the pricing restrictions without committing to one parametric family.</p><p>Hansen, L. P. (2014), &#8220;<a href="https://www.nobelprize.org/uploads/2018/06/hansen-lecture.pdf">Nobel Lecture: Uncertainty Outside and Inside Economic Models</a>,&#8221; <em>Journal of Political Economy</em> 122(5), 945-987. The best starting point of everything on this list, because it is Hansen&#8217;s own account of where GMM came from, what problem it was designed to solve, and why he is skeptical of interpretations that make the framework too broad.</p><p>Chernozhukov, V., D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W. Newey, and J. Robins (2018), &#8220;<a href="https://academic.oup.com/ectj/article/21/1/C1/5056401">Double/Debiased Machine Learning for Treatment and Structural Parameters</a>,&#8221; <em>Econometrics Journal</em> 21(1), C1-C68. Develops double machine learning using orthogonal moment conditions and cross-fitting.</p><p>Ortiz-Villavicencio, M. and P. H. C. Sant&#8217;Anna (2025), &#8220;<a href="https://arxiv.org/abs/2505.09942">Better Understanding Triple Differences Estimators</a>,&#8221; working paper. The paper that started this essay, and the remark noting that the doubly robust estimator can be read as optimal GMM on recentered influence functions.</p><h2>Further Reading</h2><p>Pearson, K. (1894), &#8220;<a href="https://doi.org/10.1098/rsta.1894.0003">Contributions to the Mathematical Theory of Evolution</a>,&#8221; <em>Philosophical Transactions of the Royal Society A</em> 185, 71-110. The origin of the method of moments, applied to separating a mixture of two normal distributions.</p><p>Fisher, R. A. (1937), &#8220;<a href="https://onlinelibrary.wiley.com/doi/10.1111/j.1469-1809.1937.tb02149.x">Professor Karl Pearson and the Method of Moments</a>,&#8221; <em>Annals of Eugenics</em> 7(4), 303-318. Worth reading for the history of the disagreement between Fisher and Pearson, which had lasted for more than two decades by this point.</p><p>Sargan, J. D. (1958), &#8220;<a href="https://doi.org/10.2307/1907619">The Estimation of Economic Relationships Using Instrumental Variables</a>,&#8221; <em>Econometrica</em> 26(3), 393-415. Develops overidentified instrumental variables and the specification test that later became an important part of the GMM framework.</p><p>Sargan, J. D. (1959), &#8220;<a href="https://doi.org/10.1111/j.2517-6161.1959.tb00317.x">The Estimation of Relationships with Autocorrelated Residuals by the Use of Instrumental Variables</a>,&#8221; <em>Journal of the Royal Statistical Society Series B</em> 21(1), 91-105. The second of the two Sargan papers Hansen names as the basis for his treatment of the family of estimators.</p><p>Lucas, R. E. (1976), &#8220;<a href="https://doi.org/10.1016/S0167-2231%2876%2980003-6">Econometric Policy Evaluation: A Critique</a>,&#8221; <em>Carnegie-Rochester Conference Series on Public Policy</em> 1, 19-46. The argument that made economists reluctant to fill in parts of a model that the theory does not deliver.</p><p>Hansen, L. P. and K. J. Singleton (1983), &#8220;<a href="https://doi.org/10.1086/261141">Stochastic Consumption, Risk Aversion, and the Temporal Behavior of Asset Returns</a>,&#8221; <em>Journal of Political Economy</em> 91(2), 249-265. The companion application, which Hansen later treats together with the 1982 paper as a single body of evidence.</p><p>Chamberlain, G. (1987), &#8220;<a href="https://doi.org/10.1016/0304-4076%2887%2990015-7">Asymptotic Efficiency in Estimation with Conditional Moment Restrictions</a>,&#8221; <em>Journal of Econometrics</em> 34(3), 305-334. Derives efficiency bounds for models defined by conditional moment restrictions and characterizes how those bounds depend on the conditional moments.</p><p>Arellano, M. and S. Bond (1991), &#8220;<a href="https://academic.oup.com/restud/article-abstract/58/2/277/1563354">Some Tests of Specification for Panel Data: Monte Carlo Evidence and an Application to Employment Equations</a>,&#8221; <em>Review of Economic Studies</em> 58(2), 277-297. The dynamic-panel estimator that helped make GMM a standard tool in applied empirical work.</p><p>Berry, S., J. Levinsohn, and A. Pakes (1995), &#8220;<a href="https://www.its.caltech.edu/~mshum/gradio/papers/BerryLevinsohnPakes1995.pdf">Automobile Prices in Market Equilibrium</a>,&#8221; <em>Econometrica</em> 63(4), 841-890. The paper that made moment-based estimation standard in empirical industrial organization.</p><p>Hansen, L. P., J. Heaton, and A. Yaron (1996), &#8220;<a href="https://larspeterhansen.org/lph_research/finite-sample-properties-of-some-alternative-gmm-estimators/">Finite-Sample Properties of Some Alternative GMM Estimators</a>,&#8221; <em>Journal of Business and Economic Statistics</em> 14(3), 262-280. Studies finite-sample problems in GMM and introduces the continuous-updating estimator.</p><p>Altonji, J. G. and L. M. Segal (1996), &#8220;<a href="https://doi.org/10.1080/07350015.1996.10524661">Small-Sample Bias in GMM Estimation of Covariance Structures</a>,&#8221; <em>Journal of Business and Economic Statistics</em> 14(3), 353-366. Shows how the asymptotically efficient weighting matrix can create important finite-sample bias.</p><p>Christiano, L. J. and W. J. Den Haan (1996), &#8220;<a href="https://www.tandfonline.com/doi/abs/10.1080/07350015.1996.10524659">Small-Sample Properties of GMM for Business-Cycle Analysis</a>,&#8221; <em>Journal of Business and Economic Statistics</em> 14(3), 309-327. Concludes that the available asymptotic theory is not a good guide at the sample sizes macroeconomists actually work with.</p><p>Hansen, L. P. and J. J. Heckman (1996), &#8220;<a href="https://www.aeaweb.org/articles?id=10.1257/jep.10.1.87">The Empirical Foundations of Calibration</a>,&#8221; <em>Journal of Economic Perspectives</em> 10(1), 87-104. Discusses what matching moments can establish and why verification requires evidence beyond the moments already used to construct a model.</p><p>Staiger, D. and J. H. Stock (1997), &#8220;<a href="https://stock.scholars.harvard.edu/publications/instrumental-variables-regression-weak-instruments">Instrumental Variables Regression with Weak Instruments</a>,&#8221; <em>Econometrica</em> 65(3), 557-586. The formalization of the weak-instrument problem.</p><p>Stock, J. H. and J. H. Wright (2000), &#8220;<a href="https://doi.org/10.1111/1468-0262.00151">GMM with Weak Identification</a>,&#8221; <em>Econometrica</em> 68(5), 1055-1096. Develops inference for GMM when the usual local identification conditions are weak or fail.</p><p>Newey, W. K. and R. J. Smith (2004), &#8220;<a href="https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1468-0262.2004.00482.x">Higher Order Properties of GMM and Generalized Empirical Likelihood Estimators</a>,&#8221; <em>Econometrica</em> 72(1), 219-255. Compares GMM with generalized empirical likelihood and shows advantages of the latter in certain higher-order properties, especially bias.</p><p>Roodman, D. (2009), &#8220;<a href="https://doi.org/10.1111/j.1468-0084.2008.00542.x">A Note on the Theme of Too Many Instruments</a>,&#8221; <em>Oxford Bulletin of Economics and Statistics</em> 71(1), 135-158. The reference for understanding instrument proliferation in dynamic-panel GMM and what it does to specification tests.</p><p>Hansen, L. P. (2012), &#8220;<a href="https://www.sciencedirect.com/science/article/abs/pii/S0304407612001200">Proofs for Large Sample Properties of Generalized Method of Moments Estimators</a>,&#8221; <em>Journal of Econometrics</em> 170(2), 325-330. Contains formal proofs that were not included in the original 1982 <em>Econometrica</em> paper and were published thirty years later.</p><p>Chernozhukov, V., W. K. Newey, and R. Singh (2022), &#8220;<a href="https://doi.org/10.3982/ECTA18515">Automatic Debiased Machine Learning of Causal and Structural Effects</a>,&#8221; <em>Econometrica</em> 90(3), 967-1027. Automates the construction of debiased moment functions through estimation of the Riesz representer.</p><p>Egami, N. and S. Yamauchi (2023), &#8220;<a href="https://arxiv.org/abs/2102.09948">Using Multiple Pretreatment Periods to Improve Difference-in-Differences and Staggered Adoption Designs</a>,&#8221; <em>Political Analysis</em> 31(2), 195-212. Combines several DiD estimators through a GMM criterion, so that the familiar estimators appear as particular choices of the weighting matrix.</p><p>Gardner, J., N. Thakral, L. T. T&#244;, and L. Yap (2026), &#8220;<a href="https://jrgcmu.github.io/2sdd_gtty.pdf">Two-Stage Differences in Differences</a>,&#8221; working paper. Conducts inference for two-stage difference-in-differences within a conventional GMM asymptotic framework.</p>]]></content:encoded></item><item><title><![CDATA[Sufficient for What?]]></title><description><![CDATA[A Silent Revolution in Macroeconomics]]></description><link>https://carloschavezp29.substack.com/p/sufficient-for-what</link><guid isPermaLink="false">https://carloschavezp29.substack.com/p/sufficient-for-what</guid><dc:creator><![CDATA[Carlos Chavez]]></dc:creator><pubDate>Mon, 06 Jul 2026 21:47:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!9XhY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60854446-cf86-41f0-b1b9-65b96a31de30_643x363.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Suppose you want to know how much a productivity shock to one particular industry, like semiconductors, electricity, or trucking, moves the GDP of the country. A priori, it seems like you cannot answer this question without knowing, for instance, who buys inputs from whom, how easily firms substitute when a supplier gets expensive, which sectors are bottlenecks for the economy, and how the shock travels through a network of thousands of buyers and sellers.</p><p>Charles Hulten, in 1978, proposed that, to a first approximation, you would need to know only the value of the industry&#8217;s sales as a share of GDP. A shock to a giant retailer and a shock to the electrical grid matter, to first order, in proportion to their sales shares, and nothing else. What about the rest of the network, such as who supplies whom and how easily inputs substitute? None of it is necessary once you know those shares.</p><p>In simple terms, this is what a <strong>sufficient statistic</strong> promises to deliver. You keep a few numbers instead of proposing a full model. The results are often good ones. But the concern about sufficient statistics is usually in the fine print. The whole argument of this essay is that a sufficient statistic is never sufficient for the world. It is sufficient only for one counterfactual, embedded in one class of models, for one set of experiments. So the question worth asking is: sufficient for what? Some researchers add that even when the number is exactly right, it can still be the wrong answer.</p><h2>Where the bargain was first struck</h2><p>The phrase comes from statistics. Ronald Fisher, in 1922, called a statistic T(X) <em>sufficient</em> for a parameter &#952; if, once you know T(X), the raw data contain no further information about &#952;. Formally, the likelihood factors as</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!uaLq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffa425a8-9463-4147-9e36-96c8580e521e_246x58.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!uaLq!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffa425a8-9463-4147-9e36-96c8580e521e_246x58.png 424w, /__u/substackcdn.com/image/fetch/$s_!uaLq!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffa425a8-9463-4147-9e36-96c8580e521e_246x58.png 848w, /__u/substackcdn.com/image/fetch/$s_!uaLq!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffa425a8-9463-4147-9e36-96c8580e521e_246x58.png 1272w, /__u/substackcdn.com/image/fetch/$s_!uaLq!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffa425a8-9463-4147-9e36-96c8580e521e_246x58.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!uaLq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffa425a8-9463-4147-9e36-96c8580e521e_246x58.png" width="246" height="58" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ffa425a8-9463-4147-9e36-96c8580e521e_246x58.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:58,&quot;width&quot;:246,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4360,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/205673266?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffa425a8-9463-4147-9e36-96c8580e521e_246x58.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!uaLq!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffa425a8-9463-4147-9e36-96c8580e521e_246x58.png 424w, /__u/substackcdn.com/image/fetch/$s_!uaLq!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffa425a8-9463-4147-9e36-96c8580e521e_246x58.png 848w, /__u/substackcdn.com/image/fetch/$s_!uaLq!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffa425a8-9463-4147-9e36-96c8580e521e_246x58.png 1272w, /__u/substackcdn.com/image/fetch/$s_!uaLq!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffa425a8-9463-4147-9e36-96c8580e521e_246x58.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p>which just says that &#952; only ever &#8220;touches&#8221; the data through T. The sample mean is sufficient for the mean of a normal distribution. Knowing every individual observation buys you nothing beyond knowing their average. Fisher&#8217;s notion is a statement about information.</p><p>Economics kept Fisher&#8217;s word and changed its object, from parameters to counterfactuals. When an economist says &#8220;sufficient statistic,&#8221; she usually means sufficient for a <em>welfare calculation</em> or a <em>policy counterfactual</em>. The modern usage was codified by Raj Chetty in a 2009 review article with a subtitle that tells you exactly what the method is for: &#8220;A Bridge Between Structural and Reduced-Form Methods.&#8221; The idea had been building in public finance for years. Martin Feldstein had shown in the 1990s that to compute the deadweight loss of the income tax you do not need to separately model how people adjust their hours, their effort, their occupation, or their tax avoidance. You need a single number, the elasticity of taxable income with respect to the net-of-tax rate. All the behavioral margins that matter for efficiency are already summarized in how taxable income responds. Earlier still, Martin Baily (1978) had shown that the optimal level of unemployment insurance could be written in terms of just the drop in consumption a worker suffers on losing a job and the elasticity of unemployment duration with respect to benefits.</p><p>The structure of all these results is the same, and it is worth writing down because every macro result later in this essay is a variation on it:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Nv6e!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56d45399-81f2-4a63-9745-96b7cd4c803d_391x96.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Nv6e!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56d45399-81f2-4a63-9745-96b7cd4c803d_391x96.png 424w, /__u/substackcdn.com/image/fetch/$s_!Nv6e!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56d45399-81f2-4a63-9745-96b7cd4c803d_391x96.png 848w, /__u/substackcdn.com/image/fetch/$s_!Nv6e!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56d45399-81f2-4a63-9745-96b7cd4c803d_391x96.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Nv6e!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56d45399-81f2-4a63-9745-96b7cd4c803d_391x96.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Nv6e!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56d45399-81f2-4a63-9745-96b7cd4c803d_391x96.png" width="391" height="96" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/56d45399-81f2-4a63-9745-96b7cd4c803d_391x96.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:96,&quot;width&quot;:391,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:8469,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/205673266?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56d45399-81f2-4a63-9745-96b7cd4c803d_391x96.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Nv6e!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56d45399-81f2-4a63-9745-96b7cd4c803d_391x96.png 424w, /__u/substackcdn.com/image/fetch/$s_!Nv6e!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56d45399-81f2-4a63-9745-96b7cd4c803d_391x96.png 848w, /__u/substackcdn.com/image/fetch/$s_!Nv6e!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56d45399-81f2-4a63-9745-96b7cd4c803d_391x96.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Nv6e!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56d45399-81f2-4a63-9745-96b7cd4c803d_391x96.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>The welfare effect of a policy change, on the left, equals some formula, on the right, involving a small number of statistics you can estimate from data: elasticities, covariances, consumption drops. The structural primitives (utility functions, production technologies, the whole apparatus) have vanished from the right-hand side. That disappearance is what the method is selling.</p><p>But notice what Chetty himself was careful to say, and what a decade of enthusiasm sometimes forgot: the approach is <em>not model-free</em>. A different sufficient-statistic formula must be derived for <em>each</em> question, and deriving it requires taking a stand on the structure of the model. The formula F is only valid inside the class of models it was derived for. The primitives disappear from the formula because you <em>assumed</em> a class of environments in which their details cancel out. They still sit inside the derivation, doing work you have agreed not to examine. That qualification is the heart of the essay.</p><h2>Three macro instances</h2><p>Much of the canonical early sufficient-statistics work was developed in relatively partial-equilibrium policy settings. The hard and interesting move of the last fifteen years has been to carry the idea into full <em>general</em> equilibrium, into a setting where everything feeds back on everything else, and where the temptation to summarize is greater and more dangerous. Let me take three of the best examples, because seeing the pattern three times is what makes it visible.</p><h3>One: aggregation, and Hulten&#8217;s revenge</h3><p>Return to Hulten&#8217;s theorem, which I opened with. In an efficient economy, the elasticity of aggregate output to a productivity shock in sector i is that sector&#8217;s <strong>Domar weight</strong>, sales over GDP: </p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!yufC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07a72ced-e064-49ce-b242-a111e279591e_200x69.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!yufC!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07a72ced-e064-49ce-b242-a111e279591e_200x69.png 424w, /__u/substackcdn.com/image/fetch/$s_!yufC!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07a72ced-e064-49ce-b242-a111e279591e_200x69.png 848w, /__u/substackcdn.com/image/fetch/$s_!yufC!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07a72ced-e064-49ce-b242-a111e279591e_200x69.png 1272w, /__u/substackcdn.com/image/fetch/$s_!yufC!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07a72ced-e064-49ce-b242-a111e279591e_200x69.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!yufC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07a72ced-e064-49ce-b242-a111e279591e_200x69.png" width="200" height="69" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/07a72ced-e064-49ce-b242-a111e279591e_200x69.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:69,&quot;width&quot;:200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4910,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/205673266?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07a72ced-e064-49ce-b242-a111e279591e_200x69.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!yufC!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07a72ced-e064-49ce-b242-a111e279591e_200x69.png 424w, /__u/substackcdn.com/image/fetch/$s_!yufC!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07a72ced-e064-49ce-b242-a111e279591e_200x69.png 848w, /__u/substackcdn.com/image/fetch/$s_!yufC!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07a72ced-e064-49ce-b242-a111e279591e_200x69.png 1272w, /__u/substackcdn.com/image/fetch/$s_!yufC!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07a72ced-e064-49ce-b242-a111e279591e_200x69.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>Notice what is <em>absent</em> from the formula. Nothing about the network, nothing about substitution or market power. One point will matter later: the Domar weight is itself an equilibrium object. It already encodes gross sales, input linkages, and intermediate flows. The network still matters in a deep sense. To first order, though, it does not enter <em>separately</em> once you know the sales shares, it is already folded into them. Those shares are a sufficient statistic for how micro shocks aggregate. Xavier Gabaix (2011) used a close cousin of this logic to argue that idiosyncratic shocks to the very largest firms, a &#8220;granular residual,&#8221; can drive a large share of aggregate volatility, because the firm-size distribution is so skewed that big firms never wash out in the average.</p><p>Now the bill. In 2019, David Baqaee and Emmanuel Farhi published a paper whose title is the warning: &#8220;The Macroeconomic Impact of Microeconomic Shocks: Beyond Hulten&#8217;s Theorem.&#8221; Their point is that Hulten&#8217;s sufficiency is a strictly <em>first-order</em> result, exact only for Cobb-Douglas economies, and only as a local approximation otherwise. Go to second order and every ingredient Hulten let you ignore returns: elasticities of substitution, the shape of the network, returns to scale, the extent to which factors can reallocate. And the second-order terms matter. They make the distribution of output asymmetric and fat-tailed <em>even when the underlying shocks are symmetric and thin-tailed</em>. Negative shocks get amplified, positive ones attenuated, so that recessions are endogenously deeper than booms are high. In their calibration, the output losses from business-cycle fluctuations come out roughly an order of magnitude larger than the near-zero welfare cost Robert Lucas made famous.</p><p>So the Domar weight is sufficient, for the linear response, in an efficient economy, to a small shock. Change any of those qualifiers and it is not. The sales share tells you the slope. It is silent about the curvature, and the curvature is where the disasters live. Figure 1 draws that gap: the tangent line is all Hulten gives you, and the true response bends away beneath it, further on the downside than the up.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!9XhY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60854446-cf86-41f0-b1b9-65b96a31de30_643x363.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!9XhY!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60854446-cf86-41f0-b1b9-65b96a31de30_643x363.png 424w, /__u/substackcdn.com/image/fetch/$s_!9XhY!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60854446-cf86-41f0-b1b9-65b96a31de30_643x363.png 848w, /__u/substackcdn.com/image/fetch/$s_!9XhY!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60854446-cf86-41f0-b1b9-65b96a31de30_643x363.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9XhY!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60854446-cf86-41f0-b1b9-65b96a31de30_643x363.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!9XhY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60854446-cf86-41f0-b1b9-65b96a31de30_643x363.png" width="643" height="363" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/60854446-cf86-41f0-b1b9-65b96a31de30_643x363.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:363,&quot;width&quot;:643,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:70544,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/205673266?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60854446-cf86-41f0-b1b9-65b96a31de30_643x363.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!9XhY!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60854446-cf86-41f0-b1b9-65b96a31de30_643x363.png 424w, /__u/substackcdn.com/image/fetch/$s_!9XhY!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60854446-cf86-41f0-b1b9-65b96a31de30_643x363.png 848w, /__u/substackcdn.com/image/fetch/$s_!9XhY!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60854446-cf86-41f0-b1b9-65b96a31de30_643x363.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9XhY!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60854446-cf86-41f0-b1b9-65b96a31de30_643x363.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Two: transmission, and the moment that matters</h3><p>Second example. When a central bank cuts rates, or a government sends out checks, how much does spending respond? The textbook representative-agent answer uses one economy-wide marginal propensity to consume. But households differ, and which feature of that distribution is sufficient depends on the question you ask.</p><p>Adrien Auclert&#8217;s 2019 work on the &#8220;redistribution channel&#8221; of monetary policy makes this concrete. The extra consumption response, beyond the representative-agent benchmark, is captured by a set of <strong>covariances</strong> between MPCs and household balance-sheet exposures. The <em>average</em> MPC alone misses it. The clearest is the interest-rate exposure channel:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!LbJ4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36ee8192-800e-4682-9b96-1ca99e902fcc_235x51.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!LbJ4!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36ee8192-800e-4682-9b96-1ca99e902fcc_235x51.png 424w, /__u/substackcdn.com/image/fetch/$s_!LbJ4!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36ee8192-800e-4682-9b96-1ca99e902fcc_235x51.png 848w, /__u/substackcdn.com/image/fetch/$s_!LbJ4!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36ee8192-800e-4682-9b96-1ca99e902fcc_235x51.png 1272w, /__u/substackcdn.com/image/fetch/$s_!LbJ4!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36ee8192-800e-4682-9b96-1ca99e902fcc_235x51.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!LbJ4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36ee8192-800e-4682-9b96-1ca99e902fcc_235x51.png" width="235" height="51" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/36ee8192-800e-4682-9b96-1ca99e902fcc_235x51.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:51,&quot;width&quot;:235,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3878,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/205673266?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36ee8192-800e-4682-9b96-1ca99e902fcc_235x51.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!LbJ4!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36ee8192-800e-4682-9b96-1ca99e902fcc_235x51.png 424w, /__u/substackcdn.com/image/fetch/$s_!LbJ4!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36ee8192-800e-4682-9b96-1ca99e902fcc_235x51.png 848w, /__u/substackcdn.com/image/fetch/$s_!LbJ4!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36ee8192-800e-4682-9b96-1ca99e902fcc_235x51.png 1272w, /__u/substackcdn.com/image/fetch/$s_!LbJ4!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36ee8192-800e-4682-9b96-1ca99e902fcc_235x51.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>where URE&#7522;, the <em>unhedged interest rate exposure</em>, is the difference between a household&#8217;s maturing assets and maturing liabilities. The logic is that a rate change redistributes between borrowers and savers, and this only moves aggregate spending if the winners and losers have <em>different</em> propensities to spend. Empirically the covariance is negative (high-MPC households tend to be borrowers), so heterogeneity <em>amplifies</em> monetary policy. Two more covariances, one for inflation exposure and one for income, complete the picture. Auclert traces the lineage of this move explicitly back to Harberger (1964) and Chetty (2009). It is the public-finance bargain, transplanted into monetary economics.</p><p>The deeper version of this idea is the &#8220;intertemporal Keynesian cross&#8221; of Auclert, Rognlie, and Straub (2024), where the sufficient statistic for the output response to fiscal policy grows into an entire matrix of <strong>intertemporal MPCs</strong>, how much you spend today out of income you expect to receive at each horizon in the future. Their empirical estimates are inconsistent with representative-agent and two-agent models, but can be matched by heterogeneous-agent ones, and they imply deficit-financed spending multipliers above one. Here the &#8220;few numbers&#8221; have become a whole object, but the philosophy is unchanged: pin down that object from micro data and you can compute the macro counterfactual without solving the full model.</p><p>The lesson repeats. The average MPC is sufficient for one question and insufficient for another. What you need is the <em>joint</em> distribution of propensities and exposures. The marginals will lie to you.</p><h3>Three: policy evaluation, and the Lucas critique on a leash</h3><p>Third example. Set welfare and aggregation aside: was a given policy decision actually <em>good</em>? The usual way to answer is to build a full structural model, which means taking a stand on almost everything. Barnichon and Mesters argued in 2023 that you need just two objects, and neither requires solving a fully specified structural model: the <strong>impulse responses</strong> of the objectives (say, inflation and unemployment) to policy shocks, and the central bank&#8217;s own <strong>conditional forecasts</strong> of those objectives.</p><p>Here is the reasoning. At an optimum, there is no predictable way to do better, so the gradient of the policymaker&#8217;s loss function with respect to a small policy change should be zero. That gradient is a weighted product of the two statistics, which makes the optimality condition a kind of <em>orthogonality</em>: your forecast of where the economy is headed should be uncorrelated with your ability to move it. When they correlate, you are leaving something on the table. Applied to US monetary policy, the optimal adjustments are usually small, around 25 basis points, with a few exceptions, the zero lower bound chief among them.</p><p>That is a lot of mileage from two estimable objects. But the catch is right there, and it is an old objection. The whole thing works only if the impulse responses you estimated under <em>past</em> policy behavior remain valid under the <em>alternative</em> policy you&#8217;re contemplating, that is, only if the private sector doesn&#8217;t re-optimize in response to the new rule. This is the Lucas critique, and the entire semi-structural literature (Barnichon and Mesters, but also McKay and Wolf (2023) on time-series counterfactuals, and Wolf&#8217;s (2023) &#8220;missing intercept&#8221;) is best understood as an effort to put the Lucas critique on a leash rather than pretend it has been solved. The move is to state the invariance assumption <em>explicitly</em> and only compute counterfactuals within its reach. That is real progress.</p><h2>The pattern, and the sharper failure underneath it</h2><p>Step back and the three examples rhyme. In each, a formidable structural object collapses to a few measurable numbers, and in each, that collapse is licensed by an assumption about the class of environments you&#8217;re willing to consider:</p><ul><li><p>Hulten&#8217;s Domar weights are sufficient <em>for the first-order response in an efficient economy</em>.</p></li><li><p>Auclert&#8217;s covariances are sufficient <em>for the first-order consumption response, within the class of models where those balance-sheet channels operate</em>.</p></li><li><p>Barnichon and Mesters&#8217; impulse responses are sufficient <em>for policy counterfactuals in which the private sector is invariant to the perturbation</em>.</p></li></ul><p>The same shape appears far outside macro proper. In trade, Arkolakis, Costinot, and Rodr&#237;guez-Clare (2012) showed that the welfare gains from trade, across a range of models with different micro foundations, depend on just two numbers: the share of spending on domestic goods and the trade elasticity. Their title asks, drily, &#8220;New Trade Models, Same Old Gains?&#8221; and the answer is yes: conditional on the trade data, the micro details cancel. This is Hulten&#8217;s irrelevance wearing a different suit, and just as conditional. The moment you allow variable markups or non-CES demand, the two statistics stop being sufficient, and a literature grows up to handle each escape.</p><p>This much is the familiar half of the story. That a sufficient statistic is model-relative is what Chetty conceded and what Henrik Kleven (2021) drove home when he revisited the program. The selling point was &#8220;no structural model required.&#8221; Kleven shows how much that oversells. You can write sufficient-statistics formulas under very general conditions, but as you relax the assumptions the estimation requirements explode, and the feasible implementations turn out to be structural approaches in all but name. The theory has not gone away. It has been moved from the estimation stage, where you would have to defend it, to the derivation stage, where it sits silently inside the formula. An honest sufficient-statistic result always ships with its model class attached, the way a drug ships with its contraindications.</p><p>That is the familiar critique, and it does not go deep enough.</p><h2>When sufficiency masks identification</h2><p>There is a failure mode underneath model-relativity, and it is the one I find most worth dwelling on, because it survives even when you have done everything right. The sufficient statistic can be right, and well estimated. The trouble is that a single statistic can collapse several structurally different worlds into one number, worlds that agree on the number and disagree on the counterfactual you actually care about. The problem is a <em>many-to-one</em> mapping from structure to summary. It has nothing to do with noise, and measuring the summary more precisely does nothing to break the tie. The statistic can be accurate and still hide the thing you need. Call that <strong>masking</strong>. Figure 2 is that collapse in one picture: several worlds funnel into one number, which fans back out to several answers.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!vmFb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ffe647b-ea7a-4b2a-ab26-683cf4c3b22b_639x337.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!vmFb!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ffe647b-ea7a-4b2a-ab26-683cf4c3b22b_639x337.png 424w, /__u/substackcdn.com/image/fetch/$s_!vmFb!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ffe647b-ea7a-4b2a-ab26-683cf4c3b22b_639x337.png 848w, /__u/substackcdn.com/image/fetch/$s_!vmFb!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ffe647b-ea7a-4b2a-ab26-683cf4c3b22b_639x337.png 1272w, /__u/substackcdn.com/image/fetch/$s_!vmFb!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ffe647b-ea7a-4b2a-ab26-683cf4c3b22b_639x337.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!vmFb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ffe647b-ea7a-4b2a-ab26-683cf4c3b22b_639x337.png" width="639" height="337" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8ffe647b-ea7a-4b2a-ab26-683cf4c3b22b_639x337.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:337,&quot;width&quot;:639,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:67642,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/205673266?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ffe647b-ea7a-4b2a-ab26-683cf4c3b22b_639x337.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!vmFb!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ffe647b-ea7a-4b2a-ab26-683cf4c3b22b_639x337.png 424w, /__u/substackcdn.com/image/fetch/$s_!vmFb!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ffe647b-ea7a-4b2a-ab26-683cf4c3b22b_639x337.png 848w, /__u/substackcdn.com/image/fetch/$s_!vmFb!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ffe647b-ea7a-4b2a-ab26-683cf4c3b22b_639x337.png 1272w, /__u/substackcdn.com/image/fetch/$s_!vmFb!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ffe647b-ea7a-4b2a-ab26-683cf4c3b22b_639x337.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Two things here are easy to conflate. The first is plain <em>insufficiency</em>: the statistic you reached for does not carry enough information, and a richer one would. Most of this essay has been about insufficiency, and insufficiency is the benign case. Baqaee and Farhi&#8217;s second-order terms are of this kind. The Domar weight is silent about curvature, so you go and measure elasticities of substitution as well. Auclert&#8217;s lesson has the same shape. The average MPC is too coarse, so you measure its covariance with exposure instead. In both cases the fix is to <em>measure more</em>, and the program absorbs the correction without complaint.</p><p>Masking is worse, because measuring more of the <em>same</em> statistic does not help. Here the statistic can be point-identified, estimated without error, and still fail to pin down the counterfactual, because the map from the structural world to the statistic is many-to-one. Two different economies produce the identical number and would respond differently to the policy. This is an <em>identification</em> failure sitting on top of a clean <em>estimation</em>. You can be perfectly right about the present and still wrong about the counterfactual, and no confidence interval will warn you, because the confidence interval is about the number, and the number is not where the ambiguity lives.</p><p>The sharpest case is hidden in one we have already met. Return to the policy-evaluation result. Barnichon and Mesters recover a set of impulse responses and treat them as sufficient for judging whether policy was optimal. Estimate those responses without error, and they are still consistent with several structural worlds that differ only in how much the private sector re-optimizes when the rule changes. Under the rule that generated the data, those worlds are observationally identical. Under the alternative rule you are weighing, they disagree. That is the Lucas critique, read as masking. The impulse response is a many-to-one shadow of the structures that could have produced it, and no confidence interval around it tells you which one you are in. Only a restriction does, the assumption that the private sector is invariant to the rule you are considering, and that restriction is exactly the structure the sufficient statistic promised to let you skip. Figure 3 shows the collision: the two responses coincide to the left of the rule change and separate to the right, one estimate, two counterfactuals.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!48QF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3bbe656-8077-427f-9186-559186ea3dba_643x349.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!48QF!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3bbe656-8077-427f-9186-559186ea3dba_643x349.png 424w, /__u/substackcdn.com/image/fetch/$s_!48QF!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3bbe656-8077-427f-9186-559186ea3dba_643x349.png 848w, /__u/substackcdn.com/image/fetch/$s_!48QF!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3bbe656-8077-427f-9186-559186ea3dba_643x349.png 1272w, /__u/substackcdn.com/image/fetch/$s_!48QF!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3bbe656-8077-427f-9186-559186ea3dba_643x349.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!48QF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3bbe656-8077-427f-9186-559186ea3dba_643x349.png" width="643" height="349" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e3bbe656-8077-427f-9186-559186ea3dba_643x349.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:349,&quot;width&quot;:643,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:59293,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/205673266?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3bbe656-8077-427f-9186-559186ea3dba_643x349.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!48QF!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3bbe656-8077-427f-9186-559186ea3dba_643x349.png 424w, /__u/substackcdn.com/image/fetch/$s_!48QF!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3bbe656-8077-427f-9186-559186ea3dba_643x349.png 848w, /__u/substackcdn.com/image/fetch/$s_!48QF!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3bbe656-8077-427f-9186-559186ea3dba_643x349.png 1272w, /__u/substackcdn.com/image/fetch/$s_!48QF!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3bbe656-8077-427f-9186-559186ea3dba_643x349.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The first two examples are milder failures: the missing object can be named, measured, and folded back into the statistic. The policy case is different. There, precision on the original number buys you nothing. Separating the structures takes new identifying restrictions, or new data that tell them apart. A tighter estimate of the same impulse response will not do it.</p><h2>What to keep</h2><p>I don&#8217;t want any of this to read as a debunking, because the sufficient-statistics program reshaped empirical macroeconomics for the better. Before it, the choice felt binary: either write down a full structural model and defend every assumption, or run a reduced-form regression that couldn&#8217;t answer the welfare or policy question you cared about. The sufficient-statistics move built a bridge, and it disciplined the conversation between the micro evidence economists can credibly estimate and the macro questions they actually want to answer. That is a considerable achievement, and the researchers I&#8217;ve cited are among the most careful in the field <em>because</em> they state their assumptions out loud.</p><p>The same warning applies wherever an exposure, an elasticity, or an impulse response is asked to do the work of a model.</p><p>A sufficient statistic is a <strong>bargain</strong>. You hand over the burden of estimating an entire structural model, and in return you get a formula in a few numbers you can measure. What you pay, always, is a set of assumptions about the class of models and the range of experiments over which that formula holds, and the price is easy to forget because it doesn&#8217;t appear on the right-hand side of the equation. Sometimes the price is that the formula holds only to first order. Sometimes it is that the same number could have come from a different world that answers your question differently. So when you meet a new sufficient statistic in the wild, the useful reflex is a question, the one this essay is named after, and the one that a perfectly estimated number can still fail to answer: <em>sufficient for what?</em></p><h2>Further reading</h2><p><strong>Chetty (2009), &#8220;Sufficient Statistics for Welfare Analysis: A Bridge Between Structural and Reduced-Form Methods.&#8221; </strong><em><strong>Annual Review of Economics</strong></em><strong> 1, 451-488.</strong> The paper that named the method and made the case for it, and the one to read first, if only because Chetty states the model-relativity caveat more honestly than much of what followed.</p><p><strong>Hulten (1978), &#8220;Growth Accounting with Intermediate Inputs.&#8221; </strong><em><strong>Review of Economic Studies</strong></em><strong> 45(3), 511-518.</strong> The original first-order irrelevance result, sales shares are all you need, and the seed of the entire aggregation literature.</p><p><strong>Gabaix (2011), &#8220;The Granular Origins of Aggregate Fluctuations.&#8221; </strong><em><strong>Econometrica</strong></em><strong> 79(3), 733-772.</strong> Why idiosyncratic shocks to a few very large firms do not wash out, and why the size distribution is itself a sufficient statistic for a chunk of aggregate volatility.</p><p><strong>Baqaee and Farhi (2019), &#8220;The Macroeconomic Impact of Microeconomic Shocks: Beyond Hulten&#8217;s Theorem.&#8221; </strong><em><strong>Econometrica</strong></em><strong> 87(4), 1155-1203.</strong> The essential correction, first-order sufficiency is exact only for Cobb-Douglas, and the second-order world is asymmetric, fat-tailed, and far more dangerous.</p><p><strong>Auclert (2019), &#8220;Monetary Policy and the Redistribution Channel.&#8221; </strong><em><strong>American Economic Review</strong></em><strong> 109(6), 2333-2367.</strong> The cleanest port of the method into monetary macro, with three covariances between MPCs and balance-sheet exposures doing all the work.</p><p><strong>Auclert, Rognlie and Straub (2024), &#8220;The Intertemporal Keynesian Cross.&#8221; </strong><em><strong>Journal of Political Economy</strong></em><strong> 132(12), 4068-4121.</strong> Where the sufficient statistic becomes a matrix of intertemporal MPCs, and where the empirical version rules out representative-agent and two-agent models outright.</p><p><strong>Barnichon and Mesters (2023), &#8220;A Sufficient Statistics Approach for Macro Policy.&#8221; </strong><em><strong>American Economic Review</strong></em><strong> 113(11), 2809-2845.</strong> Two objects, impulse responses and forecasts, are enough to judge and correct policy, provided you are willing to hold the private sector fixed.</p><p><strong>Arkolakis, Costinot and Rodr&#237;guez-Clare (2012), &#8220;New Trade Models, Same Old Gains?&#8221; </strong><em><strong>American Economic Review</strong></em><strong> 102(1), 94-130.</strong> The same pattern in trade, two numbers pin down the welfare gains across a zoo of models, with the same conditionality the moment you leave the class.</p><p><strong>Kleven (2021), &#8220;Sufficient Statistics Revisited.&#8221; </strong><em><strong>Annual Review of Economics</strong></em><strong> 13, 515-538.</strong> The reassessment to read against all of the above, arguing that once the formulas are pushed to full generality, the feasible implementations are structural approaches in disguise.</p>]]></content:encoded></item><item><title><![CDATA[On Good (Bad) Controls]]></title><description><![CDATA[Good controls can move your estimate, Bad controls can leave it unchanged]]></description><link>https://carloschavezp29.substack.com/p/on-good-bad-controls</link><guid isPermaLink="false">https://carloschavezp29.substack.com/p/on-good-bad-controls</guid><dc:creator><![CDATA[Carlos Chavez]]></dc:creator><pubDate>Thu, 04 Jun 2026 00:44:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!XKPH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f8dc040-0e4a-45e5-8a6c-38352195d747_848x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>When a coefficient barely moves after you pile in a set of controls, what has it actually shown?</p><p>Suppose you present a baseline regression in a presentation, the coefficient on X that the paper is about, and a commentator ask whether it holds up once you control for things. Covariates go in, the coefficient stays put, and the result &#8220;survives.&#8221;</p><p>The ritual, as it is usually run, shows close to nothing. Robustness checks are useful. The problem is the sentence attached to them. &#8220;The coefficient is stable when I add controls&#8221; sounds like a statement about whether an estimate is credible. Some people even <a href="/__u/carloschavezp29.substack.com/p/the-transportation-problem-in-causality">would claim causality</a>. But, actually it is a statement about something else, and the distance between the two is where applied work goes quietly wrong.</p><p>The claim, stated plainly:</p><blockquote><p>Coefficient movement is not a diagnostic of whether a control is good. It only records how the regression, and what it estimates, change once you condition.</p></blockquote><p>Three consequences follow, and they are the whole essay:</p><ol><li><p>A good control can move the coefficient a lot.</p></li><li><p>A bad control can move the coefficient very little.</p></li><li><p>Coefficient stability is not identification.</p></li></ol><p>The literature has a name for part of this, &#8220;bad controls,&#8221; though the classroom shorthand often hides more than it shows. What follows is an attempt to untangle it. I find this worthy because I have had this issue before in my own work.</p><h2>Two reasons we add controls</h2><p>Two separate questions hide inside any decision to include a covariate W. The first is structural. Does conditioning on W help recover the causal effect of X on Y? This is a question about the data-generating process, about which arrows point where, and it has the same answer in a sample of fifty or fifty million. The second is statistical. When W enters this regression, does &#946;&#8203; move, and by how much? This one is about the correlations in the sample at hand.</p><p>The structural question is whether conditioning changes the causal object being estimated. The statistical question is how much the coefficient changes in this particular sample. The first is answered by assumptions. The second is answered by algebra. The regression table can answer the second. It cannot touch the first.</p><p>Applied work runs the two together. Controls go in &#8220;to see if the result holds,&#8221; the coefficient&#8217;s movement gets read as a signal about whether those controls belonged, and a steady estimate passes for a clean bill of health. The questions are independent, though. A control can be the right thing to include and still swing the coefficient hard, which is what a confounder does once it is finally adjusted for. A control can wreck identification and leave the estimate almost untouched.</p><p>First, then, the structural picture.</p><h2>A taxonomy of controls</h2><p>Causal diagrams make the structural question concrete. That&#8217;s the reason I find very useful DAGs. The back-door criterion is the working rule: to identify the effect of X on Y, every non-causal path between them has to be blocked, without opening a new one and without blocking the causal path. Whether a given  W helps depends on where it sits.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!XKPH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f8dc040-0e4a-45e5-8a6c-38352195d747_848x528.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!XKPH!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f8dc040-0e4a-45e5-8a6c-38352195d747_848x528.png 424w, /__u/substackcdn.com/image/fetch/$s_!XKPH!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f8dc040-0e4a-45e5-8a6c-38352195d747_848x528.png 848w, /__u/substackcdn.com/image/fetch/$s_!XKPH!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f8dc040-0e4a-45e5-8a6c-38352195d747_848x528.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XKPH!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f8dc040-0e4a-45e5-8a6c-38352195d747_848x528.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!XKPH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f8dc040-0e4a-45e5-8a6c-38352195d747_848x528.png" width="848" height="528" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7f8dc040-0e4a-45e5-8a6c-38352195d747_848x528.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:528,&quot;width&quot;:848,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:123080,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/200537878?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f8dc040-0e4a-45e5-8a6c-38352195d747_848x528.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!XKPH!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f8dc040-0e4a-45e5-8a6c-38352195d747_848x528.png 424w, /__u/substackcdn.com/image/fetch/$s_!XKPH!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f8dc040-0e4a-45e5-8a6c-38352195d747_848x528.png 848w, /__u/substackcdn.com/image/fetch/$s_!XKPH!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f8dc040-0e4a-45e5-8a6c-38352195d747_848x528.png 1272w, /__u/substackcdn.com/image/fetch/$s_!XKPH!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7f8dc040-0e4a-45e5-8a6c-38352195d747_848x528.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>The six cases. The red node is the variable being conditioned on. Solid arrows are causal. The dashed red arc is the spurious association that conditioning opens up.</em></p><p>What conditioning does in each case:</p><ul><li><p><strong>Confounder.</strong> Z causes both X and Y, so it sits on the back-door path X &#8592; Z &#8594; Y. Adjusting for it is correct, the textbook good control.</p></li><li><p><strong>Mediator.</strong> M sits on the causal channel itself, X &#8594; M &#8594; Y. Conditioning on it strips out the part of the effect that runs through M. That can be appropriate if the estimand is a controlled direct effect, but it is no longer the total effect, and most papers that do it are still reporting the total.</p></li><li><p><strong>Collider.</strong> C is a common effect of X and Y. Left alone it blocks a spurious path. Conditioning on it opens a path that was closed, inventing an association between X and Y. The canonical bad control.</p></li><li><p><strong>M-bias.</strong> The case that trips people up. In the canonical M-structure, Z comes before X in time, yet it is a collider on a path running through two unobserved causes, so adjusting for it biases the estimate anyway. The rule &#8220;control for anything measured pre-treatment&#8221; does not hold.</p></li><li><p><strong>Bias amplification.</strong> Z behaves like an instrument: it moves X and reaches Y only through X. When there is unmeasured confounding U, adding Z leaves the bias in place and makes it bigger. The mistake is not that instruments are bad. It is treating an instrument-like variable as an ordinary control in a confounded OLS regression, which is a different thing from using it as an instrument in IV.</p></li><li><p><strong>Predictor of Y only.</strong> P affects Y but not X, so it cannot confound the relationship. It is not needed for identification, but it soaks up residual variance and improves efficiency. A good control, for a reason that has nothing to do with bias.</p></li></ul><p>The lesson is not that controls are good or bad by nature. The same act, conditioning, can block confounding, block the causal effect, open a collider path, amplify existing bias, or improve precision. The full catalog is in Cinelli, Forney, and Pearl (2024). The point that collider, confounding, and overcontrol bias are three distinct problems is Elwert and Winship (2014). The pre-treatment version that ruins experiments is Montgomery, Nyhan, and Torres (2018), after Rosenbaum (1984).</p><h2>The same move, three meanings</h2><p>Take three simulated worlds where the true effect of X on Y is known. In each, the regression runs with and without &#8220;the control.&#8221;</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!RKsj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90652114-302d-4fd1-b0f5-7c55b68c61bc_871x453.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!RKsj!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90652114-302d-4fd1-b0f5-7c55b68c61bc_871x453.png 424w, /__u/substackcdn.com/image/fetch/$s_!RKsj!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90652114-302d-4fd1-b0f5-7c55b68c61bc_871x453.png 848w, /__u/substackcdn.com/image/fetch/$s_!RKsj!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90652114-302d-4fd1-b0f5-7c55b68c61bc_871x453.png 1272w, /__u/substackcdn.com/image/fetch/$s_!RKsj!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90652114-302d-4fd1-b0f5-7c55b68c61bc_871x453.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!RKsj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90652114-302d-4fd1-b0f5-7c55b68c61bc_871x453.png" width="871" height="453" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/90652114-302d-4fd1-b0f5-7c55b68c61bc_871x453.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:453,&quot;width&quot;:871,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:51789,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/200537878?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90652114-302d-4fd1-b0f5-7c55b68c61bc_871x453.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!RKsj!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90652114-302d-4fd1-b0f5-7c55b68c61bc_871x453.png 424w, /__u/substackcdn.com/image/fetch/$s_!RKsj!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90652114-302d-4fd1-b0f5-7c55b68c61bc_871x453.png 848w, /__u/substackcdn.com/image/fetch/$s_!RKsj!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90652114-302d-4fd1-b0f5-7c55b68c61bc_871x453.png 1272w, /__u/substackcdn.com/image/fetch/$s_!RKsj!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90652114-302d-4fd1-b0f5-7c55b68c61bc_871x453.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><ul><li><p><strong>Confounder world.</strong> True effect 1.0. Without the control the estimate is 1.73, biased upward because the omitted Z drives both variables. Add Z and it lands at 1.00. The control corrected it.</p></li><li><p><strong>Mediator world.</strong> Total effect 1.2. Without the control the estimate is 1.20, which is right when the total effect is the target. Add the mediator and it drops to 0.49. No bias was removed. The estimand changed.</p></li><li><p><strong>Collider world.</strong> True effect 1.0. Without the control the estimate is already 1.00. Add the collider and it falls to 0.00. Conditioning broke an estimate that started out correct.</p></li></ul><p>Every time, &#946;&#770; moved when a control went in. The first movement was the control doing its job. The second was a change of estimand presented as a robustness check. The third was bias created out of thin air. The magnitude and direction say nothing about which of the three it was.</p><p>Stated plainly: the movement of &#946;&#770; does not diagnose whether W belongs in the regression. It records what happened after conditioning. If W is a confounder, movement may be the correction of bias. If W is a mediator, movement may be a change in the estimand. If W is a collider, movement may be bias the regression manufactured. And if the coefficient barely moves, none of this is settled. A bad control can be quiet. A good control can be loud.</p><p>That last point is the one the three-cases picture cannot show on its own. The full grid can. Columns are how much the coefficient moves. Rows are whether the control is good or bad. All four cells exist.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!6MKa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1163bc6-914c-45d0-b733-5417b4027733_859x517.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!6MKa!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1163bc6-914c-45d0-b733-5417b4027733_859x517.png 424w, /__u/substackcdn.com/image/fetch/$s_!6MKa!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1163bc6-914c-45d0-b733-5417b4027733_859x517.png 848w, /__u/substackcdn.com/image/fetch/$s_!6MKa!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1163bc6-914c-45d0-b733-5417b4027733_859x517.png 1272w, /__u/substackcdn.com/image/fetch/$s_!6MKa!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1163bc6-914c-45d0-b733-5417b4027733_859x517.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!6MKa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1163bc6-914c-45d0-b733-5417b4027733_859x517.png" width="859" height="517" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a1163bc6-914c-45d0-b733-5417b4027733_859x517.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:517,&quot;width&quot;:859,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:64663,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/200537878?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1163bc6-914c-45d0-b733-5417b4027733_859x517.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!6MKa!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1163bc6-914c-45d0-b733-5417b4027733_859x517.png 424w, /__u/substackcdn.com/image/fetch/$s_!6MKa!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1163bc6-914c-45d0-b733-5417b4027733_859x517.png 848w, /__u/substackcdn.com/image/fetch/$s_!6MKa!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1163bc6-914c-45d0-b733-5417b4027733_859x517.png 1272w, /__u/substackcdn.com/image/fetch/$s_!6MKa!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1163bc6-914c-45d0-b733-5417b4027733_859x517.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Top row, two good controls: the confounder swings the estimate from 1.73 to 0.99, while a pure predictor of Y leaves &#946;&#770; at 1.00 and just halves the standard error. Bottom row, two bad controls: the strong collider crashes the estimate to 0.00, while a weak collider nudges it from 1.00 to 0.99. The bottom-right cell is the dangerous one: a 1% move reads as &#8220;robust,&#8221; and the control is still illegitimate. Stability did not rescue it.</em></p><p>The collider mechanism is the hardest to believe the first time, so it earns its own picture. Consider two independent variables, talent and beauty, and a third that both of them cause, like making a living as an actor. In the whole population talent and beauty are uncorrelated. Restrict to working actors and a negative correlation appears: among the people who made it, someone short on talent has to be unusually good-looking to have gotten there, and the other way around.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!JjJg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e53ca-d418-47b8-a677-fe0d2c6c8551_866x392.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!JjJg!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e53ca-d418-47b8-a677-fe0d2c6c8551_866x392.png 424w, /__u/substackcdn.com/image/fetch/$s_!JjJg!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e53ca-d418-47b8-a677-fe0d2c6c8551_866x392.png 848w, /__u/substackcdn.com/image/fetch/$s_!JjJg!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e53ca-d418-47b8-a677-fe0d2c6c8551_866x392.png 1272w, /__u/substackcdn.com/image/fetch/$s_!JjJg!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e53ca-d418-47b8-a677-fe0d2c6c8551_866x392.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!JjJg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e53ca-d418-47b8-a677-fe0d2c6c8551_866x392.png" width="866" height="392" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c65e53ca-d418-47b8-a677-fe0d2c6c8551_866x392.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:392,&quot;width&quot;:866,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:161005,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/200537878?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e53ca-d418-47b8-a677-fe0d2c6c8551_866x392.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!JjJg!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e53ca-d418-47b8-a677-fe0d2c6c8551_866x392.png 424w, /__u/substackcdn.com/image/fetch/$s_!JjJg!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e53ca-d418-47b8-a677-fe0d2c6c8551_866x392.png 848w, /__u/substackcdn.com/image/fetch/$s_!JjJg!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e53ca-d418-47b8-a677-fe0d2c6c8551_866x392.png 1272w, /__u/substackcdn.com/image/fetch/$s_!JjJg!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e53ca-d418-47b8-a677-fe0d2c6c8551_866x392.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>+0.01 in the full population, &#8722;0.53 once the collider is conditioned on. The variables did not change. Selecting on the collider did all of it.</em></p><p>The actor story is toy language for a problem that turns up constantly: conditioning on admission, survival, employment, program take-up, presence in the administrative records, or any sample restriction defined after treatment. Each of those conditions on a collider, and the bias is baked into the sample before a model runs.</p><h2>The robustness ritual, and where Gelbach comes in</h2><p>Coefficient movement, then, does not reveal whether a control belongs. A quieter problem waits even when every control does belong, in a clean confounding world where adjusting is the right call. The usual way of reporting the movement still does not hold together.</p><p>The ritual runs like this: a regression of Y on X, then controls added one at a time, with the path of the coefficient narrated. &#8220;It drops forty percent with demographics, then settles.&#8221; The catch is that when controls are correlated with each other, the incremental contribution of each one depends on the order they enter. Whichever control goes in first gets the credit. The same control entered last, after the others have absorbed the shared variation, looks inert. Same variables, same data, a story set by typing order.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!rUp_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8cf03ee-e7fb-4287-9d5c-b1268ca2e365_882x373.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!rUp_!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8cf03ee-e7fb-4287-9d5c-b1268ca2e365_882x373.png 424w, /__u/substackcdn.com/image/fetch/$s_!rUp_!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8cf03ee-e7fb-4287-9d5c-b1268ca2e365_882x373.png 848w, /__u/substackcdn.com/image/fetch/$s_!rUp_!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8cf03ee-e7fb-4287-9d5c-b1268ca2e365_882x373.png 1272w, /__u/substackcdn.com/image/fetch/$s_!rUp_!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8cf03ee-e7fb-4287-9d5c-b1268ca2e365_882x373.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!rUp_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8cf03ee-e7fb-4287-9d5c-b1268ca2e365_882x373.png" width="882" height="373" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b8cf03ee-e7fb-4287-9d5c-b1268ca2e365_882x373.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:373,&quot;width&quot;:882,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:50961,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/200537878?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8cf03ee-e7fb-4287-9d5c-b1268ca2e365_882x373.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!rUp_!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8cf03ee-e7fb-4287-9d5c-b1268ca2e365_882x373.png 424w, /__u/substackcdn.com/image/fetch/$s_!rUp_!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8cf03ee-e7fb-4287-9d5c-b1268ca2e365_882x373.png 848w, /__u/substackcdn.com/image/fetch/$s_!rUp_!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8cf03ee-e7fb-4287-9d5c-b1268ca2e365_882x373.png 1272w, /__u/substackcdn.com/image/fetch/$s_!rUp_!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8cf03ee-e7fb-4287-9d5c-b1268ca2e365_882x373.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The left panel is that problem. The credit each control gets for moving &#946;&#770; flips when the entry order switches from W&#8321;, W&#8322;, W&#8323; to the reverse. The sequential story has no fact of the matter in it.</figcaption></figure></div><p>Jonah Gelbach settled this in &#8220;When Do Covariates Matter? And Which Ones, and How Much?&#8221; (Journal of Labor Economics, 2016). The fix falls out of the omitted-variables-bias formula. The full model is</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!HvU0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197a17b-29a8-45cc-9a32-7bacf9afc707_421x109.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!HvU0!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197a17b-29a8-45cc-9a32-7bacf9afc707_421x109.png 424w, /__u/substackcdn.com/image/fetch/$s_!HvU0!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197a17b-29a8-45cc-9a32-7bacf9afc707_421x109.png 848w, /__u/substackcdn.com/image/fetch/$s_!HvU0!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197a17b-29a8-45cc-9a32-7bacf9afc707_421x109.png 1272w, /__u/substackcdn.com/image/fetch/$s_!HvU0!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197a17b-29a8-45cc-9a32-7bacf9afc707_421x109.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!HvU0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197a17b-29a8-45cc-9a32-7bacf9afc707_421x109.png" width="421" height="109" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c197a17b-29a8-45cc-9a32-7bacf9afc707_421x109.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:109,&quot;width&quot;:421,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7304,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/200537878?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197a17b-29a8-45cc-9a32-7bacf9afc707_421x109.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!HvU0!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197a17b-29a8-45cc-9a32-7bacf9afc707_421x109.png 424w, /__u/substackcdn.com/image/fetch/$s_!HvU0!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197a17b-29a8-45cc-9a32-7bacf9afc707_421x109.png 848w, /__u/substackcdn.com/image/fetch/$s_!HvU0!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197a17b-29a8-45cc-9a32-7bacf9afc707_421x109.png 1272w, /__u/substackcdn.com/image/fetch/$s_!HvU0!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197a17b-29a8-45cc-9a32-7bacf9afc707_421x109.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>Omit the W&#8342; and the short-regression coefficient is</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!-PKG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea350129-c2f3-454f-b1c6-b8b5fe8787ec_372x137.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!-PKG!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea350129-c2f3-454f-b1c6-b8b5fe8787ec_372x137.png 424w, /__u/substackcdn.com/image/fetch/$s_!-PKG!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea350129-c2f3-454f-b1c6-b8b5fe8787ec_372x137.png 848w, /__u/substackcdn.com/image/fetch/$s_!-PKG!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea350129-c2f3-454f-b1c6-b8b5fe8787ec_372x137.png 1272w, /__u/substackcdn.com/image/fetch/$s_!-PKG!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea350129-c2f3-454f-b1c6-b8b5fe8787ec_372x137.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!-PKG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea350129-c2f3-454f-b1c6-b8b5fe8787ec_372x137.png" width="372" height="137" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ea350129-c2f3-454f-b1c6-b8b5fe8787ec_372x137.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:137,&quot;width&quot;:372,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7480,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/200537878?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea350129-c2f3-454f-b1c6-b8b5fe8787ec_372x137.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!-PKG!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea350129-c2f3-454f-b1c6-b8b5fe8787ec_372x137.png 424w, /__u/substackcdn.com/image/fetch/$s_!-PKG!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea350129-c2f3-454f-b1c6-b8b5fe8787ec_372x137.png 848w, /__u/substackcdn.com/image/fetch/$s_!-PKG!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea350129-c2f3-454f-b1c6-b8b5fe8787ec_372x137.png 1272w, /__u/substackcdn.com/image/fetch/$s_!-PKG!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea350129-c2f3-454f-b1c6-b8b5fe8787ec_372x137.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>where, in the simplest one-treatment case, &#955;&#8342; is the slope from regressing the covariate W&#8342; on X alone, and &#947;&#8342; is its coefficient in the full model. (With baseline controls, fixed effects, or weights, the same logic runs on residualized regressions.) Each product &#955;&#8342; &#947;&#8342; is control k&#8217;s contribution to the gap between the base and full estimates. Because &#955;&#8342; uses the unconditional auxiliary regression and &#947;&#8342; comes from the full model, the decomposition is unique, it ignores entry order, and the pieces sum to &#946;&#770;_base &#8722; &#946;&#770;_full.</p><p>The right panel shows it. &#946;&#770;_base = 1.395 falls to &#946;&#770;_full = 1.004, a gap of 0.391. Gelbach assigns +0.363 of it to W&#8321;, &#8722;0.184 to W&#8322;, and +0.212 to W&#8323;, and those add up to 0.391 to within floating-point error. (The decomposition also nests the Oaxaca and Blinder decomposition and extends to IV. Gelbach derives the standard errors, and the bars carry bootstrap intervals.)</p><p>That answers &#8220;which controls move the coefficient, and by how much.&#8221; It removes the order dependence and gives every covariate a fair share. It decomposes the gap between the short and long regression, nothing more.</p><p>And here is the limit. Every W&#8342; in this simulation is a real confounder. The decomposition reports how much each one moved &#946;&#770; and says nothing about whether moving &#946;&#770; was warranted.</p><blockquote><p><em>A note on that claim.</em> It is the unconditional one. Without knowing a control&#8217;s structural role, the size of the movement does not reveal whether conditioning helped or hurt. Conditional on already knowing that W is a legitimate confounder, the magnitude is informative: a large move signals severe confounding, which is exactly what Oster&#8217;s bound exploits. What carries no structural information is the movement taken on its own.</p></blockquote><p>Pointed at the collider world, the same machinery reports how much the collider &#8220;explains,&#8221; handing back a bias dressed as an explanation. Gelbach answers the statistical question with precision. The structural question, whether a variable is a confounder to adjust for or a collider that is fooling you, is not in the data and never will be. It comes from the diagram. (This is separate from the high-dimensional selection problem of Belloni, Chernozhukov, and Hansen (2014): machine learning can choose which predictors to keep, but it cannot decide whether a variable is a confounder, a mediator, or a collider.)</p><h2>The honest version of &#8220;is it stable?&#8221;</h2><p>If coefficient stability is to carry weight, say as an argument that unmeasured confounding is unlikely to flip a result, there is a disciplined way to make the case. It is Oster&#8217;s (Journal of Business &amp; Economic Statistics, 2019), building on Altonji, Elder, and Taber (2005). The idea: if selection on observables runs proportional to selection on unobservables, then coefficient movement together with R&#178; movement, as controls are added, bounds the bias from what went unmeasured.</p><p>In a simulation with one confounder left unobserved, the estimate moves from 1.89 with no controls to 1.57 with the observed ones. Oster&#8217;s logic returns a &#948; around 1.64: under proportional selection and the assumed maximum R&#178;, the unobservables would have to generate selection about 1.6 times as strong as everything that was measured before the true effect collapses to zero. That is closer to the number a referee should be asking for than &#8220;the coefficient held steady.&#8221; The companion bias-adjusted &#946;* depends heavily on the assumed maximum R&#178;. In this run it undershoots the true 1.0, which is why &#948; is the statistic to report, with &#946;* treated as a bound that carries the assumption with it.</p><p>The scope is narrow even here. Oster turns the stability heuristic into a real bound under a stated assumption, and it still takes for granted that the controls being moved toward are confounders rather than colliders. The point is not that coefficient stability is useless. It is meaningful once the control set has been justified structurally, and not before.</p><h2>Keeping the two questions apart</h2><p>So, there are two questions that stay separate. The structural question comes first, before the data. It is answered by the diagram, even a rough one: each candidate control is a confounder (adjust), a mediator (adjust only for the controlled direct effect, and label it as such), a collider or descendant of the outcome (never), or a pure predictor of Y (adjust for precision). Coefficient movement has no vote in it. For anyone trained on potential outcomes rather than graphs, Imbens (2020) is the bridge, and Cunningham&#8217;s Mixtape and Huntington-Klein&#8217;s The Effect are the easiest ways in.</p><p>The statistical question comes second. Once the control set has a structural justification, Gelbach reports how much each control moves the estimate without the answer riding on entry order, and Oster bounds the exposure to the confounders that went unmeasured.</p><p>This also suggests a better sentence than the usual one. In place of:</p><blockquote><p>&#8220;The estimate is robust to adding controls.&#8221;</p></blockquote><p>something like:</p><blockquote><p>We adjust for the variables our causal model marks as confounders, and we keep post-treatment variables out. The estimate holds within that set, and it would take unobserved selection stronger than everything we could measure to push it to zero.</p></blockquote><p>The first sentence reports a coincidence of the sample. The second states the assumptions, the magnitude, and the exposure to what could not be seen. &#8220;My result is robust to adding controls&#8221; is doing less work than it sounds like: it answers the easy question in the voice of the hard one. The coefficient moved, or it did not. That was never the test. The test is whether the causal diagram justifies the control set, and no amount of coefficient-watching can stand in for it. The test was on the whiteboard, before the regression ran.</p><h2>References</h2><ul><li><p>Altonji, J., Elder, T., &amp; Taber, C. (2005). Selection on observed and unobserved variables. Journal of Political Economy.</p></li><li><p>Belloni, A., Chernozhukov, V., &amp; Hansen, C. (2014). Inference on treatment effects after selection among high-dimensional controls. Review of Economic Studies, 81(2), 608 to 650.</p></li><li><p>Cinelli, C., Forney, A., &amp; Pearl, J. (2024). A crash course in good and bad controls. Sociological Methods &amp; Research, 53(3), 1071 to 1104.</p></li><li><p>Elwert, F., &amp; Winship, C. (2014). Endogenous selection bias: the problem of conditioning on a collider variable. Annual Review of Sociology, 40, 31 to 53.</p></li><li><p>Gelbach, J. B. (2016). When do covariates matter? And which ones, and how much? Journal of Labor Economics, 34(2), 509 to 543.</p></li><li><p>Imbens, G. (2020). Potential outcome and directed acyclic graph approaches to causality. Journal of Economic Literature.</p></li><li><p>Montgomery, J., Nyhan, B., &amp; Torres, M. (2018). How conditioning on posttreatment variables can ruin your experiment. American Journal of Political Science, 62(3), 760 to 775.</p></li><li><p>Oster, E. (2019). Unobservable selection and coefficient stability. Journal of Business &amp; Economic Statistics, 37(2), 187 to 204.</p></li><li><p>Rosenbaum, P. (1984). The consequences of adjustment for a concomitant variable that has been affected by the treatment. JRSS-A.</p></li><li><p>Angrist, J., &amp; Pischke, J.-S. (2009). Mostly Harmless Econometrics, section 3.2.3.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[The Transportation Problem In Causality]]></title><description><![CDATA[Why Causal Inference is Harder to Find Than You Can Think]]></description><link>https://carloschavezp29.substack.com/p/the-transportation-problem-in-causality</link><guid isPermaLink="false">https://carloschavezp29.substack.com/p/the-transportation-problem-in-causality</guid><dc:creator><![CDATA[Carlos Chavez]]></dc:creator><pubDate>Mon, 25 May 2026 23:58:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!BpLx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea208c98-119d-44b6-9c24-37caf64e756d_873x356.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Two variables plotted against each other (Panel A), a clean upward sweep, often a trend line, sometimes a caption that invites a conclusion. The visual is compelling enough that the conclusion feels inevitable. </p><p>The temptation gets sharper when the chart shows time (Panel B). A line moving up, an arrow marking a policy change, the line continuing to climb after the arrow. The natural reading is that the policy worked. Sometimes the policy did work. Sometimes the trend was already there and the policy added nothing.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!BpLx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea208c98-119d-44b6-9c24-37caf64e756d_873x356.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!BpLx!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea208c98-119d-44b6-9c24-37caf64e756d_873x356.png 424w, /__u/substackcdn.com/image/fetch/$s_!BpLx!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea208c98-119d-44b6-9c24-37caf64e756d_873x356.png 848w, /__u/substackcdn.com/image/fetch/$s_!BpLx!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea208c98-119d-44b6-9c24-37caf64e756d_873x356.png 1272w, /__u/substackcdn.com/image/fetch/$s_!BpLx!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea208c98-119d-44b6-9c24-37caf64e756d_873x356.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!BpLx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea208c98-119d-44b6-9c24-37caf64e756d_873x356.png" width="873" height="356" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ea208c98-119d-44b6-9c24-37caf64e756d_873x356.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:356,&quot;width&quot;:873,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:55719,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/199155042?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea208c98-119d-44b6-9c24-37caf64e756d_873x356.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!BpLx!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea208c98-119d-44b6-9c24-37caf64e756d_873x356.png 424w, /__u/substackcdn.com/image/fetch/$s_!BpLx!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea208c98-119d-44b6-9c24-37caf64e756d_873x356.png 848w, /__u/substackcdn.com/image/fetch/$s_!BpLx!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea208c98-119d-44b6-9c24-37caf64e756d_873x356.png 1272w, /__u/substackcdn.com/image/fetch/$s_!BpLx!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea208c98-119d-44b6-9c24-37caf64e756d_873x356.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Most readers will infer causality in both charts and will share them trying to prove a point that, most of the time, they already believe.</p><p>The real problem is that none of those charts can tell you which is which. In a <a href="/__u/carloschavezp29.substack.com/p/on-causality">previous post</a> I traced how economics came to think about causation, from Hume to Fisher to the credibility revolution. That essay was about the history of the question, but it also showed different forms of causality. This post is about why it is harder to establish causality between two variables.</p><p>The hardness comes from a single impossibility. To know whether the tax cuts in the second chart worked, you would need to see the same country at the same time, with the cuts and without them. To know whether immigration causes prosperity in the first chart, you would need to see each country with two different levels of immigration, everything else held equal. You never can. The world reveals one history per unit. The other history is what we wish we could see, and what no dataset, however large, will ever provide. This is pretty well known as the fundamental problem of causal inference. It means you cannot see the counterfactual. You don&#8217;t know what would have happened if a specific policy had not been applied.</p><p>However, the other history can be found, in some sense, in other unit, other place or under other circumstance. The work for economists, or in a broader sense, for social scientists, is to find this counterfactual and compare both outcomes: the outcome with the policy and the outcome without it.</p><p>So, I like to think that causal inference in social sciences is a problem about transportation. Every method we use, every assumption we make, every estimator we run, is a way of moving information from where we have it to where we need it. Some methods transport across units. Some across persons. Some across regimes. Some across contexts. The hardness of causal inference is the hardness of those moves. They are never automatic, they are never free, and they are never complete because there are always trade-offs.</p><p>This essay is a short guide to transportation problems. Holland&#8217;s fundamental problem, Heckman&#8217;s three tasks, SUTVA, the Roy model, the Lucas critique, external validity: each appears in the literature as a separate obstacle. Each, I will argue, is a special case of one underlying problem.</p><h2>What Holland actually showed</h2><p>Paul Holland&#8217;s 1986 paper can be summarized in one sentence: the causal effect for a unit is the difference between two outcomes, one of which is unobservable. That sentence is correct, and it has organized empirical methodology for forty years. But it also obscures what the paper is really arguing.</p><p>Holland&#8217;s contribution was the taxonomy of escapes. He argued that there are two escapes, each corresponding to a different way of borrowing information from elsewhere.</p><p><strong>The scientific solution.</strong> Substitute the unobserved counterfactual for unit u with an observed outcome from another unit, or the same unit at another time, that we are willing to treat as exchangeable. A physicist studying gravity treats one rock as informative about another. A chemist studying reactions treats yesterday&#8217;s experiment as informative about today&#8217;s. The information is transported across units through assumptions of homogeneity, and across time through assumptions of stability. These assumptions are what make the lab sciences work.</p><p><strong>The statistical solution.</strong> Substitute the unobserved counterfactual for unit u with the average outcome of a group of units who did not receive treatment, under the condition that those units are comparable to u in expectation. This is the route most empirical economics takes. We do not pretend to know what would have happened to Jane specifically. We claim that the average outcome among people like Jane is a defensible substitute for what would have happened to Jane on average.</p><p>Both escapes are transport. The first transports individual histories across units assumed identical. The second transports aggregated histories across populations assumed comparable. Holland did not use the word transport, but the structure is the same. The fundamental problem is solved by finding the outcome elsewhere and justifying the move, not by a further analysis of what is missing. In practice, most modern empirical papers do exactly this. They make assumptions, justify them, propose an empirical strategy, and present an identification section.</p><h2>Three tasks, one example</h2><p>The scientific and statistical solutions name what we are doing. They do not tell us how to think about what we are doing. For that, the cleanest framework comes from a series of papers by Heckman and coauthors, who have argued that empirical work involves three distinct tasks that we routinely conflate.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!rYdq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec929dff-5770-4dca-94e2-56c958bf7f41_812x368.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!rYdq!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec929dff-5770-4dca-94e2-56c958bf7f41_812x368.png 424w, /__u/substackcdn.com/image/fetch/$s_!rYdq!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec929dff-5770-4dca-94e2-56c958bf7f41_812x368.png 848w, /__u/substackcdn.com/image/fetch/$s_!rYdq!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec929dff-5770-4dca-94e2-56c958bf7f41_812x368.png 1272w, /__u/substackcdn.com/image/fetch/$s_!rYdq!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec929dff-5770-4dca-94e2-56c958bf7f41_812x368.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!rYdq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec929dff-5770-4dca-94e2-56c958bf7f41_812x368.png" width="812" height="368" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ec929dff-5770-4dca-94e2-56c958bf7f41_812x368.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:368,&quot;width&quot;:812,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:51542,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/199155042?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec929dff-5770-4dca-94e2-56c958bf7f41_812x368.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!rYdq!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec929dff-5770-4dca-94e2-56c958bf7f41_812x368.png 424w, /__u/substackcdn.com/image/fetch/$s_!rYdq!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec929dff-5770-4dca-94e2-56c958bf7f41_812x368.png 848w, /__u/substackcdn.com/image/fetch/$s_!rYdq!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec929dff-5770-4dca-94e2-56c958bf7f41_812x368.png 1272w, /__u/substackcdn.com/image/fetch/$s_!rYdq!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec929dff-5770-4dca-94e2-56c958bf7f41_812x368.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>Each task has its own logic, and each can fail in its own way. To see why the distinction matters, consider a question that seems simple but is arguably the most contentious and relevant question in labor market debates, especially in developing countries: what is the effect of raising the minimum wage?</p><p>The question hides at least four distinct counterfactuals, each defining a different parameter of interest.</p><p><strong>Static partial equilibrium.</strong> Hold firms, workers, prices, and entry decisions constant. Raise the wage floor. Ask how employment changes among affected firms. This is the counterfactual implicit in most reduced-form difference-in-differences papers, including the canonical Card-Krueger comparison of New Jersey and Pennsylvania fast-food restaurants. It is well-defined, and it answers a specific question. It abstracts from the margins that minimum-wage debates sometimes emphasize.</p><p><strong>Static general equilibrium.</strong> Allow prices, employment, and firm entry to adjust within one labor market, holding technology and capital fixed. The counterfactual now includes substitution between low-wage and higher-wage workers, exit of marginal firms, and price pass-through to consumers. The same intervention identifies a different parameter, and the policy implications shift.</p><p><strong>Long-run general equilibrium.</strong> Allow capital, occupational sorting, and technology adoption to adjust. The counterfactual now includes automation responses, training decisions, and reallocation across sectors. A minimum-wage increase that looks neutral in the short run can have substantial effects here, and vice versa. Short-run estimates speak to long-run consequences only under strong assumptions.</p><p><strong>A policy rule versus a one-time shock.</strong> A permanent change in how minimum wages are set is a different counterfactual from a one-time legislative increase. Agents respond differently to rules than to one-off events. The same observed wage change identifies different objects depending on which counterfactual the analyst has in mind.</p><p>Each of these is a legitimate object of interest. Each requires a different model. Each demands different identification assumptions. None of them is &#8220;the&#8221; effect of the minimum wage.</p><p>This is what Task 1 means in practice. It is the work of explicitly naming which of these counterfactuals the paper is about. The complaint that people like Heckman make is that modern applied work often skips this step. Modern papers estimate a treatment effect, the reader is left to figure out which counterfactual the estimate corresponds to, and the policy implication is asserted as if the counterfactual were obvious. And it can be worse than that. If the authors don&#8217;t state what their counterfactual is, readers can be wrong about the interpretation of the results, about the conclusions, and therefore about the policy recommendation. I have written about this in a <a href="/__u/carloschavezp29.substack.com/p/why-dont-we-talk-enough-about-marginal">previous piece</a>.</p><p>Task 1 is where the economics happens, and Task 1 is where empirical papers often go quiet.</p><h2>The transportation problem</h2><p>If Holland&#8217;s escapes are forms of transport, and Heckman&#8217;s tasks are stages in the transport process, the obstacles in causal inference can be reorganized around a single question. Between what and what is the information being moved?</p><p>Four answers, each corresponding to an obstacle the literature treats separately.</p><h3>Across units</h3><p>The Stable Unit Treatment Value Assumption, SUTVA, says that the treatment received by one unit does not affect the outcomes of other units. It is the assumption that makes individual-level transport possible. If your training does not change my wage, then your treated outcome and my untreated outcome can be compared without contamination.</p><p>Economics violates this assumption by construction. Labor markets clear through prices. If training increases the supply of skilled workers, wages for skilled workers adjust, and wages for everyone else adjust in response. The treatment one worker receives reaches every worker in the same market. Cr&#233;pon, Duflo, Gurgand, Rathelot, and Zamora (2013) showed this experimentally in a French job-placement program. The displacement effects were large and central to the policy interpretation. Miguel and Kremer (2004) showed it for deworming. The externality across schoolchildren was not a contamination of the treatment effect. It was a major part of the treatment effect.</p><p>In both cases, the problem is transport. The untreated unit is no longer a clean counterfactual for the treated unit, because the untreated unit has itself been partially treated through the equilibrium.</p><h3>Across persons</h3><p>The Roy model, which I have discussed at length <a href="/__u/carloschavezp29.substack.com/p/the-roy-model">elsewhere</a>, formalizes selection from comparative advantage. People sort into treatments based on expected gain. The treated and the untreated differ in whether they received treatment, and they differ in what they expected to get from it. Conventional selection bias is one symptom.</p><p>The real issue is that the marginal complier moved by an instrument is not the average citizen. Different instruments move different margins, and the local average treatment effect they identify reflects the people the instrument moved. Card&#8217;s distance-to-college instrument identifies a return to schooling for a subset of low-income, geographically constrained students. Angrist and Krueger&#8217;s quarter-of-birth instrument identifies a return to schooling for students at the compulsory-attendance margin. These are estimates of different parameters that sit on the same curve, the one formalized by the Marginal Treatment Effects approach.</p><p>Transport across persons fails because persons are not interchangeable in their response to treatment, and the instruments that identify effects do so by selecting non-representative slices of the population.</p><h3>Across regimes</h3><p>The Lucas critique is, in the framework I am proposing, the temporal analogue of SUTVA. SUTVA says estimates do not transport across units in the same period because units interact. The Lucas critique says estimates do not transport across policy regimes for the same population because behavior is forward-looking and parameters are equilibrium objects.</p><p>The textbook example is the empirical relationship between inflation and unemployment. Estimated on data through the 1960s, it was stable enough to support policy advice. Tolerate higher inflation, the models said, and unemployment will fall. The 1970s showed otherwise. Once policymakers committed to higher inflation, agents adjusted their expectations, and the trade-off vanished. The estimated parameters had reflected the equilibrium of a particular policy regime. Changing the regime changed the parameters.</p><p>The Lucas critique is not specific to macro. Every time an applied paper extrapolates from a local treatment effect to a national policy, every time we assume that what worked in a pilot will work at scale, we are asserting parameter invariance across regimes. The assertion is sometimes defensible but rarely tested.</p><h3>Across contexts</h3><p>External validity is the most familiar transportability problem and the least formally treated in mainstream economics. A deworming RCT in Western Kenya identifies a parameter. Whether that parameter applies in Madhya Pradesh, in the Peruvian Amazon, or in rural Mississippi is a separate question that the RCT itself cannot answer. The standard response, replication in new contexts, is honest but expensive. The other response, structural extrapolation under assumptions, is faster but stronger.</p><p>Deaton and Cartwright (2018) made this critique sharply. A randomized experiment with internal validity tells you the effect for the units in the experiment, in the time and place of the experiment, under the conditions of the experiment. The move to any other unit, time, place, or condition is an extrapolation, and the extrapolation is precisely what the design was not built to do.</p><h3>Four faces of one problem</h3><p>These four problems are usually treated as distinct. However, they are not. SUTVA, Roy selection, the Lucas critique, and external validity are four faces of the same underlying difficulty: the move from where we have evidence to where we want to make a claim. The first is transport across units, the second across persons, the third across regimes, the fourth across contexts. Each requires assumptions the data cannot test directly. Each fails silently when the assumptions fail.</p><p>When an applied paper concludes with a policy recommendation, it has made at least one transport claim, often several. The recommendation is only as strong as the weakest move. Most papers do not name the moves they are making, perhaps because economists are eager to apply the latest credibility-revolution techniques without first understanding the statistical foundations of those tools.</p><h2>Two traditions, two trade-offs</h2><p>If causal inference is the discipline of transport, the major traditions in empirical economics can be read as different choices about which moves to defend and which to assume away. Neither choice is obviously right. Each has costs.</p><p>The structural tradition refuses to give up Task 1. It writes down a model. The model defines counterfactuals over its primitives: preferences, technologies, constraints, equilibrium concepts. Identification is conducted within the model. The parameters recovered are interpretable as the deep objects the model names, and counterfactual simulation under alternative policies is straightforward in principle. The strength is that the policy question and the estimated parameter speak the same language. The cost is that identification depends on assumptions the data cannot adjudicate: functional forms, distributional shapes, behavioral primitives that are stipulated, not estimated. A skeptical reader can always doubt the model, and the doubt is not statistical. A bad specification of the model can really screw up the results of a paper with an actual policy-relevant question.</p><p>The design-based tradition refuses to give up Task 2. It looks for variation that is plausibly exogenous and identifies whatever parameter that variation can identify. The strength is that identification is defensible on its own terms, with minimal modeling assumptions. The cost is that the parameter identified is local to the variation: local to the instrument, local to the cutoff, local to the compliers, local to the period of the natural experiment. The original policy question and the identified parameter often do not match, and bridging them requires assumptions the design did not impose. So, policymakers reading papers based on this tradition cannot really understand what the paper is actually estimating.</p><p>Neither tradition has solved the transportation problem. Each has chosen which dimension of transport to defend. The structural approach defends transport across counterfactuals at the cost of credibility about the model. The design-based approach defends credibility within a specific variation at the cost of transport beyond it.</p><p>The credibility revolution shifted the field toward the second tradition, and the gains are real. Identification became transparent. Assumptions became visible. Debates became disciplined. We learned, collectively, what a clean identification argument looks like. That is a permanent contribution to the field, and dismissing it would be a mistake. I always say there is beauty in finding a good instrument to estimate a LATE. It is a kind of art.</p><p>But what the shift costs is less visible. Estimands became smaller. The distance between the parameter estimated and the policy question asked often grew. When the questions a society wants answered are about general equilibrium responses, long-run adjustments, or counterfactual policy regimes, a credible LATE on a specific margin is informative but not sufficient. Bridging the gap requires assumptions of the kind the structural tradition has always made. Whether to make those assumptions explicitly, inside a model that can be debated, or implicitly, in the move from estimate to recommendation, is itself a methodological choice.</p><p>Reading a structural paper, you ask whether the model is right. Reading a design-based paper, you ask whether the estimand is the one you wanted. Both questions are legitimate. Neither admits a clean answer. The trade-off at this point is clear: cleaner identification on a narrower estimand, or messier identification on a more ambitious one. There is no algorithm that resolves which side of the trade-off to take. It depends on the question, the data, and what you are willing to defend.</p><p>For questions where the variation is clean and the population of interest matches the compliers, the design-based approach is preferable. For questions where the counterfactual is policy-defined and the available variation does not match it, structural modeling is the only route, and the right response to its assumptions is to argue about them, not to dismiss them. Hopefully, as I argue in this essay, macro and micro are converging given the scope of their own type of questions. The interesting empirical work in the coming decade will, I suspect, increasingly combine the two.</p><h2>A note on frameworks</h2><p>There is a long-running argument between two formal languages for causal inference. The potential outcomes framework, developed by Donald Rubin and elaborated by Guido Imbens, and the structural causal model framework, developed by Judea Pearl and elaborated by his collaborators. Within economics, potential outcomes dominates. Within computer science and parts of epidemiology, the Pearl framework dominates. The two have argued about which is better for over thirty years.</p><p>The right reading, I think, is that each tradition saw a different piece of the transportation problem.</p><p>The potential outcomes framework saw selection. Its core question is how to identify treatment effects when units are not randomly assigned, and its tools, propensity scores, conditioning, instrumental variables, were built for that purpose. It is excellent at what it was built for, and it has shaped empirical economics accordingly.</p><p>The structural causal model framework saw transportability. Its core question is how to combine information across data sources, populations, and regimes, and its tools, do-calculus, graphical surgery, transportability theorems, were built for that purpose. The development by Bareinboim and Pearl on data fusion is the most sustained formal treatment of the move-across-contexts problem in the literature.</p><p>These are complementary specializations. Imbens (2020) concluded much the same. The frameworks are tools for different questions. The economics profession has been slower to absorb the Pearl tradition&#8217;s contributions on transport, partly out of disciplinary inertia, partly because the formal apparatus is unfamiliar.</p><h2>What this means for the rest of the series</h2><p>Every method I have discussed in several posts is a partial answer to the transportation problem.</p><p>Instruments transport to the population of compliers, under the assumption that the instrument is exogenous and affects the outcome only through treatment. Regression discontinuity transports to the cutoff, under the assumption that units on either side are comparable in the limit. Difference-in-differences transports across time, under the assumption of parallel trends. Synthetic control transports across units, under the assumption that a weighted combination of donors approximates the treated unit&#8217;s counterfactual. Structural estimation transports across counterfactuals arbitrarily, under the assumption that the model is right.</p><p>No method transports for free. Every method makes a move, the move depends on assumptions, and the assumptions are sometimes credible and sometimes not. The reader&#8217;s job is to figure out which assumptions are being made and whether they are defensible for the question at hand. The author&#8217;s job is to make those assumptions visible to the reader, especially if the paper presents causal effects on a policy-relevant question.</p><p>Causal inference is hard because every applied conclusion is a transport claim, and most transport claims are made implicitly. The discipline of making them explicit is what separates an empirical paper that informs a debate from one that decorates it.</p><p>The methods we will examine over the next essays are tools for making the moves. None of them solves the transportation problem. All of them illuminate one corner of it. That is, I think, the most honest description of where empirical economics actually stands.</p><h2>Some Worth Readings</h2><p><strong>Bareinboim, E., and Pearl, J. (2016).</strong> &#8220;Causal Inference and the Data-Fusion Problem.&#8221; <em>Proceedings of the National Academy of Sciences</em>, 113(27), 7345&#8211;7352. The most sustained formal treatment of transport across contexts.</p><p><strong>Cr&#233;pon, B., Duflo, E., Gurgand, M., Rathelot, R., and Zamora, P. (2013).</strong> &#8220;Do Labor Market Policies Have Displacement Effects? Evidence from a Clustered Randomized Experiment.&#8221; <em>Quarterly Journal of Economics</em>, 128(2), 531&#8211;580. The cleanest experimental demonstration that general equilibrium effects matter and can be measured.</p><p><strong>Deaton, A., and Cartwright, N. (2018).</strong> &#8220;Understanding and Misunderstanding Randomized Controlled Trials.&#8221; <em>Social Science &amp; Medicine</em>, 210, 2&#8211;21. The sharpest available statement of the external validity problem.</p><p><strong>Heckman, J. (2008).</strong> &#8220;Econometric Causality.&#8221; <em>International Statistical Review</em>, 76(1), 1&#8211;27. The clearest statement of the three-task framework and of the relationship between the structural and the treatment-effect traditions.</p><p><strong>Heckman, J., and Pinto, R. (2022).</strong> &#8220;Causality and Econometrics.&#8221; NBER Working Paper 29787. A more recent treatment, in dialogue with the DAG literature.</p><p><strong>Holland, P. W. (1986).</strong> &#8220;Statistics and Causal Inference.&#8221; <em>Journal of the American Statistical Association</em>, 81(396), 945&#8211;960. The founding statement of the fundamental problem and of the scientific and statistical solutions. Read it.</p><p><strong>Imbens, G. W. (2020).</strong> &#8220;Potential Outcome and Directed Acyclic Graph Approaches to Causality: Relevance for Empirical Practice in Economics.&#8221; <em>Journal of Economic Literature</em>, 58(4), 1129&#8211;1179. The state of the dialogue between economics and the Pearl tradition.</p><p><strong>Lucas, R. E. (1976).</strong> &#8220;Econometric Policy Evaluation: A Critique.&#8221; <em>Carnegie-Rochester Conference Series on Public Policy</em>, 1, 19&#8211;46. Short, deep, and still relevant.</p><p><strong>Miguel, E., and Kremer, M. (2004).</strong> &#8220;Worms: Identifying Impacts on Education and Health in the Presence of Treatment Externalities.&#8221; <em>Econometrica</em>, 72(1), 159&#8211;217. The canonical demonstration of SUTVA violations in a development context.</p>]]></content:encoded></item><item><title><![CDATA[The Cowles Commission in Chicago]]></title><description><![CDATA[The War for Econometrics]]></description><link>https://carloschavezp29.substack.com/p/the-cowles-commission-in-chicago</link><guid isPermaLink="false">https://carloschavezp29.substack.com/p/the-cowles-commission-in-chicago</guid><dc:creator><![CDATA[Carlos Chavez]]></dc:creator><pubDate>Sun, 17 May 2026 16:16:02 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!zh0V!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffff9b4a5-53ee-4548-85da-563173f13249_2666x2666.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Every empirical economist works inside a vocabulary they did not invent. Identification. Exogeneity. Structural parameters. Exclusion restrictions. Reduced form. The conceptual scaffolding behind every IV regression, every structural model, every credibility-revolution paper. Economists learn the vocabulary in the first year of graduate school and use it for the rest of their careers.</p><p>It came from sixteen years of work in one building on the south side of Chicago.</p><p>The Cowles Commission for Research in Economics was based at the University of Chicago from 1939 to 1955. In that period it built the foundations of modern econometrics, fought and won the most consequential methodological debate of the twentieth century, and was then pushed out of the department that hosted it. The intellectual victory and the institutional defeat happened in the same room, at the same time, and were not unrelated. The standard accounts treat the 1955 move to Yale as an administrative event. Better terms, family ties, Connecticut tax law. The real story is that two visions of empirical economics had been competing for the same department, and only one of them was going to survive.</p><p>This is the story of those sixteen years. It is also, indirectly, the story of why the credibility revolution thirty years later spoke in Cowles vocabulary while pursuing the opposite program.</p><h2>The arrival: Cowles before it was Cowles</h2><p>The Commission was founded in Colorado Springs in 1932 by Alfred Cowles III, an investment manager who wanted to know whether professional stock forecasters could beat random guessing. His 1933 <em>Econometrica</em> paper, &#8220;Can Stock Market Forecasters Forecast?&#8221;, answered the question with a single word: no. The Commission was the institutional home of that finding, and for its first seven years it was a small, well-funded operation devoted mostly to financial economics. Cowles himself was a Yale graduate. The detail matters later.</p><p>In 1939, looking for university affiliation and a research community, Cowles moved the Commission to the University of Chicago. Theodore Yntema served as the first research director from 1939 to 1942. Henry Schultz, Chicago&#8217;s specialist in demand estimation, was the natural patron on the Chicago side. Then Schultz died in a car accident, Yntema left for the Committee for Economic Development during the war, and the Commission was without leadership.</p><p>It got Jacob Marschak in 1943.</p><p>Marschak is the underappreciated architect of what Cowles became. A former Menshevik, refugee from Berlin via Oxford, fluent in five languages and active in five fields, he had directed the Oxford Institute of Statistics in the 1930s. He arrived at Chicago with a clear program: economic theory, probability theory, and statistical inference welded into a single methodology. Under his five-year directorship the Commission turned from a financial research institute into something more ambitious, a research program organized around a unified approach to empirical economics.</p><p>The unifying piece was Haavelmo.</p><h2>The Haavelmo turn</h2><p>Trygve Haavelmo&#8217;s 1944 monograph <em>The Probability Approach in Econometrics</em>, published as a supplement to <em>Econometrica</em> and written largely while he was visiting Cowles, did something philosophically radical that is now so deeply embedded in econometric practice that we forget how strange it once seemed.</p><p>Until 1944, most empirical work in economics treated stochastic disturbances as nuisances. Measurement errors. Small omitted influences. Statistical noise tacked onto otherwise deterministic relationships. Haavelmo argued this was the wrong starting point. Economics is stochastic from the ground up. The relationships we try to estimate are inherently probabilistic. Not deterministic laws with errors attached, but actual probability distributions. And the moment you accept this, the entire machinery of statistical inference becomes available. Estimators, standard errors, hypothesis tests, identification analysis. None of these tools make sense for deterministic relationships with noise on top. They make complete sense for probability models of behavior.</p><p>The move sounds technical. The implications were not. Once economic relationships are probability statements, you can ask formal questions about them. When does the data identify the underlying parameters? Under what assumptions are your estimators consistent? What does it mean for two structures to be observationally equivalent? These questions had not been asked rigorously before because the framework for asking them did not exist.</p><p>After Haavelmo, the framework existed. Cowles spent the next eleven years building inside it.</p><p>The team Marschak assembled was extraordinary. Between 1943 and 1955, the Commission&#8217;s research associates and visitors included Haavelmo, Tjalling Koopmans, Leonid Hurwicz, Lawrence Klein, Kenneth Arrow, Herbert Simon, Franco Modigliani, Don Patinkin, James Tobin, and G&#233;rard Debreu. Most of them would win the Nobel Prize. Marschak himself was elected president of the American Economic Association in 1977, the year he died. The concentration of talent in one building on the south side of Chicago has few parallels in the history of economics.</p><p>By 1947, Cowles had a research program. It also had an enemy.</p><h2>Measurement Without Theory</h2><p>In 1946 the National Bureau of Economic Research published <em>Measuring Business Cycles</em>, by Arthur Burns and Wesley Mitchell. The book was the culmination of nearly three decades of NBER work in the Mitchell tradition: a meticulous taxonomy of cyclical fluctuations in dozens of economic time series, with careful methodology for dating peaks and troughs, identifying reference cycles, and decomposing fluctuations into specific and reference components. It was empirical economics at its most disciplined and least theoretical. Collect everything. Classify. Describe. Let the data speak.</p><p>In August 1947, Tjalling Koopmans reviewed it in the <em>Review of Economics and Statistics</em>. The review ran twelve pages. Its title was &#8220;Measurement Without Theory.&#8221;</p><p>The opening paragraph compared Burns and Mitchell to Tycho Brahe and Johannes Kepler. The comparison was not flattering. Brahe accumulated decades of careful planetary measurements. Kepler extracted regularities from those measurements. Neither could explain <em>why</em> the planets moved as they did. That had to wait for Newton, for an underlying theory of forces and masses from which the observed regularities could be derived. Burns and Mitchell, Koopmans argued, were stuck in the Brahe-Kepler stage. They had regularities, measured in any amount of detail, with no explanatory power. They could describe business cycles forever without ever understanding them.</p><p>The substantive critique came in three layers.</p><p><strong>Variable selection without theory is arbitrary.</strong> Without an economic model specifying which variables matter and why, the NBER&#8217;s choice of which series to include and which to exclude had no principled basis. The selection was justified by intuition and accumulated practice, neither of which is a theory.</p><p><strong>Estimated parameters without structure have no economic meaning.</strong> The NBER methodology could establish that variable XX X tends to lead variable YY Y by two quarters on average. It could not say whether this reflected a causal mechanism, a common driver, or coincidence of timing. Without a structural interpretation, the estimates were descriptions rather than explanations.</p><p><strong>Policy analysis requires structure.</strong> To know what would happen if monetary policy tightened by one hundred basis points, you need a model of how economic agents respond to monetary policy. That requires theory. The NBER approach offered no path from its measurements to policy-relevant predictions.</p><p>The attack landed hard. The NBER had been founded in 1920 and had defined the empirical mainstream of American economics for a generation. Koopmans was declaring that mainstream methodologically bankrupt. The publication of &#8220;Measurement Without Theory&#8221; marked the moment when the Cowles approach went from being one school of thought to being the school of thought that intended to displace everything else.</p><p>The reply came two years later, and it came from someone almost no one reads today.</p><h2>What Vining saw</h2><p>In May 1949, Rutledge Vining responded in the same journal. His piece, &#8220;Koopmans on the Choice of Variables to Be Studied and the Methods of Measurement,&#8221; is one of the most underrated documents in the history of econometric methodology. It contains, in 1949, most of the critiques that would be aimed at structural econometrics forty years later, when the credibility revolution finally caught up.</p><p>Vining&#8217;s argument went like this. Koopmans assumes we know the structure of the economic system well enough to write down its equations, classify variables as endogenous or exogenous, specify functional forms, and impose the exclusion restrictions needed to identify parameters. For business cycles, and for most macroeconomic questions of interest, we know almost none of this. The structural approach is not &#8220;theory first.&#8221; It is <em>a particular theory first</em>, imposed by the researcher, and that theory may be wrong in ways that contaminate every estimate it produces. The NBER&#8217;s atheoretical regularities, Vining argued, are useful precisely because they do not pre-commit to a model class. They describe what is in the data. The Cowles approach describes what the data would look like if a particular theoretical structure were true.</p><p>Koopmans replied. Vining rejoined. The full exchange ran to roughly forty pages and is collected in Hendry and Morgan&#8217;s <em>Foundations of Econometric Analysis</em>.</p><p>Two things about the debate deserve more attention than they get.</p><p>Vining was largely right about something Koopmans was largely wrong about. The structural approach requires identifying assumptions that are typically harder to defend than its practitioners acknowledged. The parameters Koopmans wanted to recover from systems of simultaneous equations were only as credible as the exclusion restrictions used to identify them, and those restrictions were often imposed by statistical convenience, not derived from anything we actually know about economic behavior. When Christopher Sims wrote &#8220;Macroeconomics and Reality&#8221; in 1980, the founding document of the VAR program, his central complaint was Vining&#8217;s complaint: structural macroeconometrics is built on identifying restrictions that nobody really believes. The profession took thirty years to admit this, and when it did, it spoke in Vining&#8217;s vocabulary without crediting him.</p><p>The second thing is institutional. Beatrice Cherrier&#8217;s archival work showed that Milton Friedman helped Vining draft his reply, but asked not to be acknowledged. Friedman had returned to Chicago in 1946 and was already methodologically opposed to Cowles. He chose to fight by proxy. The Methodenstreit was not a distant fight between Cowles and the NBER. It was a fight between two factions inside the same university, with one faction ghost-writing for the side that was based elsewhere.</p><p>The seminar wars on the Midway in the late 1940s were not abstract methodological exercises. They were a slow-motion divorce.</p><h2>Monograph 10 and the birth of identification</h2><p>While the methodological battle was being fought in journals, the technical work inside Cowles was racing forward. In 1950, the Commission published Monograph 10, <em>Statistical Inference in Dynamic Economic Models</em>, edited by Koopmans. Monograph 10 is the document in which <em>identification</em> became a formal object in econometrics.</p><p>The intuitive idea had been around for a while. Working&#8217;s 1927 paper on supply and demand had shown that if you observe equilibrium prices and quantities, you cannot in general separate the supply curve from the demand curve, because both shift simultaneously. The structural parameters of one equation are confounded with the structural parameters of the other. The Cowles contribution was to state the problem generally, for any system of equations, and to derive necessary and sufficient conditions for identification.</p><p>The definition that emerged is the one that still appears in first-year graduate econometrics textbooks. A structure is identified if and only if no other structure observationally equivalent to it satisfies the same a priori restrictions. The order condition, that the number of excluded exogenous variables in an equation must be at least as large as the number of included endogenous variables minus one, was derived in Monograph 10. The rank condition, that the matrix of restrictions must have full rank, was derived in Monograph 10. The distinction between exactly identified, overidentified, and underidentified equations was made precise in Monograph 10.</p><p>Around the identification framework, the Commission built the rest of the technical apparatus. Full-information maximum likelihood for simultaneous equations was developed by Koopmans, Rubin, and Leipnik. Limited-information maximum likelihood was developed by Anderson and Rubin. The formal definition of exogeneity, the analysis of measurement error in structural equations, and the foundations of what is now called the Cowles Commission approach to macroeconometric modeling were all in place by 1953.</p><p>By 1951, if you were a serious econometrician anywhere in the world, you read Cowles monographs. By 1955, the Klein-Goldberger model, a structural macroeconometric model of the U.S. economy estimated by Cowles methods and used for policy simulation, had demonstrated the program worked at scale. The structural-econometrics paradigm Cowles built would dominate empirical macroeconomics for the next twenty-five years.</p><p>Cowles had won the argument.</p><p>It had also lost Chicago.</p><h2>The Chicago that was emerging</h2><p>While Cowles was fighting the NBER, the Chicago economics department was becoming something else. Milton Friedman returned in 1946. Theodore Schultz became department chair the same year, a position he would hold until 1961. Aaron Director arrived in 1946. Gregg Lewis was already there. Within a few years, the department had a distinctive intellectual identity that had almost nothing to do with the Cowles program.</p><p>The biographical detail about Friedman matters here. He had been a student of Arthur Burns and Wesley Mitchell at Columbia in the 1930s, and Mitchell had been his senior colleague at the National Bureau. The methodology Koopmans attacked in 1947 was the methodology Friedman had been trained in. The fight between Cowles and the NBER was, for him, also a defense of his teachers.</p><p>The emerging Chicago tradition was Marshallian rather than Walrasian. Partial-equilibrium analysis of specific markets rather than general-equilibrium modeling of whole systems. Tightly focused applied microeconomic studies rather than economy-wide structural models. The quantity theory of money rather than Keynesian aggregate demand management. Predictive adequacy as the standard of methodological success rather than structural fidelity to underlying behavior.</p><p>Friedman&#8217;s 1953 essay &#8220;The Methodology of Positive Economics&#8221; is often read as a general statement about instrumentalism in economic theory. Read in its institutional context, it is also a position paper in the war with Cowles. Friedman argued that the realism of assumptions does not matter. What matters is whether the model&#8217;s predictions are accurate. Cowles, in its commitment to identifying <em>the</em> structural parameters of the economy, the actual underlying behavioral relationships from which observed correlations emerge, was implicitly making the opposite claim. Getting the structure right is what allows you to do policy analysis. Friedman&#8217;s essay said the project was overdetermined. You did not need to know the structure. You needed to know what works.</p><p>Two visions of empirical economics were operating in the same building. The Cowles seminar and the Friedman price-theory workshop were different research programs, but they were also different epistemological cultures. One was Walrasian, formal, system-oriented, Keynesian-friendly, comfortable with applications to central planning. Linear programming, activity analysis, optimal allocation. The other was Marshallian, partial, problem-focused, monetarist, skeptical of macroeconometric modeling, and openly hostile to anything that smelled like a tool for technocratic management of the economy.</p><p>Friedman did not confine the disagreement to journals. He attended Cowles seminars and argued against the papers presented there, often in detail. In December 1950, he published a 28-page defense of his old teacher in the <em>Journal of Political Economy</em>, titled &#8220;Wesley C. Mitchell as an Economic Theorist.&#8221; Mitchell had died two years earlier. The piece was part eulogy, part counterattack on Koopmans, and part position paper on what empirical economics should look like.</p><p>The institutional friction was constant. Hiring decisions involving Cowles researchers were contentious. The Commission wanted joint appointments for its senior people. The department resisted. Schultz, as chair, presided over a department that had become methodologically uncomfortable with the visitors in the same building.</p><p>The clearest moment came in 1953.</p><h2>The Tobin episode</h2><p>Koopmans was scheduled for a sabbatical in 1954 to write <em>Three Essays on the State of Economic Science</em>. The Commission needed a new research director. Marschak, Koopmans, and Alfred Cowles approached James Tobin at Yale and invited him to move to Chicago to take the position.</p><p>The Cowles directorship came bundled with a professorship in the Chicago economics department. Tobin, before accepting, asked Schultz a direct question: would the department be interested in him independently of the Cowles affiliation? Schultz said no.</p><p>Tobin declined the offer.</p><p>The exchange has been reported by multiple historians. Christ&#8217;s 1994 <em>Journal of Economic Literature</em> article, Hildreth&#8217;s 1986 history of the Commission, and most recently Dimand&#8217;s work on the Cowles archives. The detail that matters is what Schultz was implicitly saying. The Chicago economics department did not want what Cowles wanted. It would tolerate the Commission. It was housed in the same building, it had been a courteous host since 1939. But it would not adopt Cowles as part of its own intellectual identity. The professorship attached to the Cowles directorship was a courtesy. Without the Cowles affiliation, the department would not have made the offer.</p><p>When Tobin called Koopmans to decline, Koopmans asked whether Yale might be interested in the entire Cowles operation. The answer was yes.</p><h2>The exodus</h2><p>The move was negotiated through 1954. Alfred Cowles, as a Yale alumnus with family connections to New Haven, was sympathetic to the destination. Yale required an endowment commitment in place of the annual gifts Cowles had been making, but the institutional fit was clear. In 1955, the Cowles Commission became the Cowles Foundation for Research in Economics at Yale. Tobin assumed the directorship he had previously declined when it required moving to Chicago.</p><p>The Chicago offices in the Social Sciences Research Building emptied out. The plaques came down. The monograph series continued under Yale imprint. The structural-econometrics research program that Cowles had built between 1944 and 1955 left town.</p><p>Friedman, writing in his joint memoir with Rose Friedman much later, attributed the departure primarily to financial inducements from Yale and to Alfred Cowles&#8217;s alumni loyalty. The fuller archival record, documented in Hildreth (1986), Christ (1994), Cherrier (2011), and Dimand&#8217;s more recent work, tells a different story. The financial considerations were real. The alumni ties were real. The binding constraint was that the Chicago economics department had stopped wanting the Commission as part of its intellectual identity, and the Commission&#8217;s leadership had stopped wanting to fight for the position.</p><p>Chicago kept the building. Yale got the program.</p><p>And Chicago became <em>Chicago</em>, the Chicago of Friedman, Stigler, Becker, Lucas, partly because Cowles had left. The department that took shape in the 1960s, the one we still mean when we say &#8220;Chicago School,&#8221; is the department that emerged after the Walrasian, general-equilibrium, large-scale-modeling tradition had been escorted to the train station.</p><h2>What Cowles won and what it lost</h2><p>The temptation is to read the story as a tragedy. Cowles built the foundations of modern econometrics, fought a clean methodological war, won, and was thrown out of its home at the moment of victory. There is truth in the reading. The longer arc is stranger.</p><p>For about twenty-five years after the move, the structural-econometrics paradigm was the empirical mainstream of macroeconomics. The large Keynesian macroeconometric models, Brookings, Wharton, DRI, the Federal Reserve&#8217;s MPS model, were the institutional descendants of Klein-Goldberger, which was the institutional descendant of Cowles Monograph 10. The Cowles approach defined what serious empirical macro looked like.</p><p>Then, between 1973 and 1980, the paradigm cracked.</p><p>Lucas&#8217;s 1976 critique showed that the structural parameters Cowles practitioners thought they were estimating were not invariant to policy regimes. The entire policy-evaluation use case rested on a confusion. Sims&#8217;s 1980 VAR paper attacked the identification restrictions used to pin down those parameters as &#8220;incredible.&#8221; His exact charge was that the restrictions &#8220;are mainly simplifications, chosen empirically so that they do not conflict with the data,&#8221; and therefore could not identify the underlying structure. He was careful about what this implied. Large-scale models, he conceded, &#8220;do perform useful forecasting and policy-analysis functions despite their incredible identification.&#8221; The methodology worked. The methodological claims behind it did not. And the credibility revolution in microeconometrics, beginning with Card, Krueger, Angrist, and Imbens in the late 1980s and 1990s, took the next step. Stop trying to estimate the structure of an entire economy. Focus instead on identifying one well-defined causal parameter under transparent assumptions, using research designs that approximate randomized experiments.</p><p>That program was a partial vindication of Vining&#8217;s 1949 reply. The credibility revolution accepted that we typically do not know the structure of the economy well enough to write it down. It accepted that pre-committing to a particular model class is dangerous. It shared Vining&#8217;s skepticism toward heavily imposed structural systems and echoed aspects of the design-based empirical work the NBER tradition had championed. The difference was a sharper formal apparatus for causal inference and an explicit theory of identification that Burns and Mitchell never had.</p><p>But the vocabulary in which the credibility revolution conducted its program was entirely Cowles&#8217;s gift.</p><p><em>Identification.</em> <em>Exogeneity.</em> <em>Structural parameters.</em> <em>Instrumental variables.</em> <em>Reduced form.</em> <em>Exclusion restriction.</em> <em>Endogeneity.</em> Every one of these concepts was sharpened, formalized, and given its modern meaning in the Cowles monographs of 1944&#8211;1953. Angrist and Imbens did not abandon Cowles methodology so much as inherit its conceptual core while discarding its ambition of full-system estimation. The credibility revolution speaks Cowles&#8217;s language while pursuing Vining&#8217;s program.</p><p>Both sides won different wars.</p><p>Cowles won the conceptual war. Identification, structural inference, and the probability approach are permanent features of the discipline. Anyone who runs an IV regression today is working inside a framework that was built between Marschak&#8217;s arrival in 1943 and the publication of Monograph 10 in 1950.</p><p>Vining won the methodological war, at least in part. The unifying Cowles ambition, that one large structural system could deliver policy answers for the whole economy from observational data, did not survive. Structural estimation is still alive and central in macro, IO, trade, quantitative spatial work, and search models. But it is structural estimation tailored to specific markets and specific questions, not the comprehensive economy-wide system Klein-Goldberger pointed toward. In applied micro, the large macroeconometric models of the 1960s have been replaced by smaller, design-based studies whose epistemic claims are local and modest.</p><p>The institutional victors were Friedman, Schultz, Director, and the price-theory coalition forming around them. They won a war whose consequences are still shaping the discipline. The Chicago blend of price theory, monetarism, applied microeconomics, and methodological caution defined the center of gravity of American economics for the second half of the twentieth century. Friedman was the loudest voice in that coalition and the one who corresponded with Vining behind the scenes. But the institutional outcome required Schultz&#8217;s chairmanship, Director&#8217;s intellectual brokerage, and a department that had already chosen its direction by 1953.</p><h2>What remains</h2><p>The Cowles Commission&#8217;s sixteen years on the Midway are now mostly remembered as a list of names that later won Nobel Prizes. The list is impressive. It is not the most interesting thing about what happened in that building.</p><p>What happened was the construction of the conceptual scaffolding that every empirical economist still uses, every day, whether they know it or not. The Commission lost Chicago. It did not lose the discipline.</p><p>The deeper lesson is harder to state and worth sitting with. The Cowles program won the conceptual battle and lost the methodological one. Its categories survive. Its ambition does not. The credibility revolution adopted Cowles vocabulary while rejecting the project Cowles was using that vocabulary to pursue. Vining was right about what could be defended empirically. Koopmans was right about what concepts the discipline would need. Both were partially correct, in ways that were not obvious in 1949 and that took the field three more decades to recognize.</p><p>The exodus of 1955 is sometimes treated as a footnote in the history of the Chicago department. It is closer to a load-bearing wall. The department that became <em>Chicago</em> could not have become <em>Chicago</em> with Cowles inside it. And the Cowles program, which had to leave to survive, could not have produced the work it produced without first being hosted by a department that would eventually evict it. The Midway in 1943-1955 was a room with too many ideas in it. Two of those ideas now run American economics. They both came out of that building.</p><p>The plaques are gone. The building is still there. Every regression with an instrumental variable is a Cowles monograph in different clothing.</p><h2>Some worth readings</h2><h4>The Cowles monographs themselves</h4><p>Haavelmo, Trygve. 1944. &#8220;The Probability Approach in Econometrics.&#8221; <em>Econometrica</em> 12 (Supplement): iii-vi, 1-115. The manifesto. The single most important econometric document of the twentieth century.</p><p>Koopmans, Tjalling C., ed. 1950. <em>Statistical Inference in Dynamic Economic Models</em>. Cowles Commission Monograph No. 10. New York: John Wiley &amp; Sons. Where identification becomes a formal object. Order and rank conditions, FIML, the modern framework.</p><p>Hood, William C., and Tjalling C. Koopmans, eds. 1953. <em>Studies in Econometric Method</em>. Cowles Commission Monograph No. 14. New York: John Wiley &amp; Sons. The applied companion to Monograph 10. How to use the methods in practice.</p><h4>The Methodenstreit</h4><p>Koopmans, Tjalling C. 1947. &#8220;Measurement Without Theory.&#8221; <em>The Review of Economics and Statistics</em> 29 (3): 161-172. The attack. Read it for the prose. It is one of the great polemics in economics.</p><p>Vining, Rutledge. 1949. &#8220;Koopmans on the Choice of Variables to Be Studied and the Methods of Measurement.&#8221; <em>The Review of Economics and Statistics</em> 31 (2): 77-86. The reply. Underrated. Almost no one reads it. It is forty years ahead of the credibility revolution.</p><p>Koopmans, Tjalling C. 1949. &#8220;A Reply.&#8221; <em>The Review of Economics and Statistics</em> 31 (2): 86-91. The rejoinder. Worth reading next to the Vining piece.</p><p>Hendry, David F., and Mary S. Morgan, eds. 1995. <em>The Foundations of Econometric Analysis</em>. Cambridge: Cambridge University Press. Reprints the entire Koopmans-Vining exchange in one place.</p><h4>History of the Commission</h4><p>Hildreth, Clifford. 1986. <em>The Cowles Commission in Chicago, 1939-1955</em>. Berlin: Springer-Verlag. The standard institutional history. Hildreth interviewed many of the original participants. Indispensable for the Tobin episode and the negotiations with Yale.</p><p>Christ, Carl F. 1994. &#8220;The Cowles Commission&#8217;s Contributions to Econometrics at Chicago: 1939-1955.&#8221; <em>Journal of Economic Literature</em> 32 (1): 30-59. The standard intellectual history. By an economist who was there.</p><p>Dimand, Robert W. 2019. &#8220;The Cowles Commission and Foundation for Research in Economics.&#8221; <em>Cowles Foundation Discussion Paper No. 2207</em>. Recent reassessment with newer archival material.</p><p>Cherrier, Beatrice. 2011. &#8220;The Lucky Consistency of Milton Friedman&#8217;s Science and Politics, 1933-1963.&#8221; In <em>Building Chicago Economics</em>, edited by Robert Van Horn, Philip Mirowski, and Thomas Stapleford, 335-367. Cambridge: Cambridge University Press. Documents Friedman&#8217;s ghost-writing assistance to Vining. Archival.</p><p>D&#252;ppe, Till, and E. Roy Weintraub. 2014. <em>Finding Equilibrium: Arrow, Debreu, McKenzie and the Problem of Scientific Credit</em>. Princeton: Princeton University Press. Useful context on the Cowles general-equilibrium work and the move to Yale.</p><h4>What came next</h4><p>Sims, Christopher A. 1980. &#8220;Macroeconomics and Reality.&#8221; <em>Econometrica</em> 48 (1): 1-48. Vining&#8217;s vocabulary, three decades later, with a formal apparatus.</p><p>Lucas, Robert E. 1976. &#8220;Econometric Policy Evaluation: A Critique.&#8221; <em>Carnegie-Rochester Conference Series on Public Policy</em> 1: 19-46. The critique that broke the Cowles macroeconometric program from inside the structural tradition.</p><p>Friedman, Milton. 1953. &#8220;The Methodology of Positive Economics.&#8221; In <em>Essays in Positive Economics</em>. Chicago: University of Chicago Press. The Chicago position paper, read in context as a methodological alternative to what Cowles was doing one floor away.</p>]]></content:encoded></item><item><title><![CDATA[The Roy Model: Why People Sort]]></title><description><![CDATA[Selection, Comparative Advantage, and the Foundation of Treatment Effect Heterogeneity]]></description><link>https://carloschavezp29.substack.com/p/the-roy-model-why-people-sort</link><guid isPermaLink="false">https://carloschavezp29.substack.com/p/the-roy-model-why-people-sort</guid><dc:creator><![CDATA[Carlos Chavez]]></dc:creator><pubDate>Thu, 02 Apr 2026 22:09:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!P4pi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57f76d18-18f5-49d4-a47d-e560bbe586fd_891x532.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A.D. Roy&#8217;s 1951 paper, &#8220;Some Thoughts on the Distribution of Earnings,&#8221; appeared in the <em>Oxford Economic Papers</em>. It was not about program evaluation or causal inference. It was about why the distribution of earnings looks the way it does, why some people earn a lot and others earn little, and why the overall distribution has the shape it has.</p><p>Roy posed the problem with an example. Suppose there are two occupations: hunting and fishing. Each person has a productivity in each occupation, but these productivities differ across people and are not perfectly correlated. Some people are good hunters and poor fishers. Some are the reverse. Some are good at both; some are good at neither.</p><p>If people choose freely, each person will select the occupation where their earnings are higher. A person who is a better hunter than fisher will hunt. A person who is a better fisher than hunter will fish. This sorting has consequences. The observed distribution of earnings among hunters is not the distribution of hunting productivity in the population. It is the distribution of hunting productivity among people who chose to hunt, people who are, by construction, better at hunting than at fishing. The observed distribution is selected.</p><p>Roy showed that this selection could generate a wide variety of earnings distributions depending on the underlying joint distribution of productivities and the correlation between them. If hunting and fishing skills are positively correlated (people who are good at one tend to be good at the other), the selection is less severe. If they are negatively correlated (specialists), the selection is more severe. The shape of the observed earnings distribution, its mean, variance, skewness, depends on the selection process, not just on the underlying distribution of abilities.</p><p>The paper was about income distribution, not causal inference. Roy never used the word &#8220;treatment.&#8221; But the logic he introduced, that observed outcomes reflect selected subpopulations, and that the selection depends on potential outcomes in multiple states, is exactly the logic that would later define the evaluation problem.</p><h2>The Evaluation Problem</h2><p>Reframe Roy&#8217;s setup as a treatment effect problem. Let Y&#8321; be the outcome if treated (say, earnings with a college degree) and Y&#8320; be the outcome if untreated (earnings without a college degree). Each person has both potential outcomes, but we observe only one: Y&#8321; for those who attend college, Y&#8320; for those who do not. The treatment effect for person i is Y&#8321;&#7522; &#8722; Y&#8320;&#7522;, but this is never observed directly because we never see the same person in both states.</p><p>The average treatment effect, ATE, is E[Y&#8321; &#8722; Y&#8320;], the average of the individual treatment effects across the population. The average treatment effect on the treated, ATT, is E[Y&#8321; &#8722; Y&#8320; | D = 1], the average effect among those who actually received treatment. These are different objects, and they differ precisely because of selection.</p><p>In Roy&#8217;s framework, people choose treatment when they expect it to benefit them. If expected gains from treatment are positively correlated with actual gains, then the people who select into treatment are those with above-average treatment effects. The ATT exceeds the ATE. This is selection on gains. It is not the only form of selection, people might also select based on their baseline outcomes, their costs, or their information, but it is the form that Roy&#8217;s model makes precise and that matters most for evaluation.</p><p>Consider education. The raw wage gap between college graduates and high school graduates in the United States is roughly 80 percent: workers with a bachelor&#8217;s degree earn nearly double what workers with only a high school diploma earn. But nobody believes that sending a randomly chosen high school graduate to college would double their earnings. The people who went to college chose to go, and they chose because they expected to benefit.</p><p>The numbers tell the story. Card&#8217;s (1995) OLS estimate of the return to an additional year of schooling was about 7 percent. His IV estimate, using college proximity as an instrument, was 13 percent. Angrist and Krueger&#8217;s (1991) OLS estimate was about 7 percent; their IV estimate, using quarter of birth, was also about 7 percent. Same question, same country, same decades, different answers.</p><p>The puzzle dissolves once you see it through Roy&#8217;s lens. Card&#8217;s instrument, proximity to a college, moves students who are cost-sensitive. These are students from lower-income families, students who would not have attended a distant college, students for whom the financial and logistical barriers matter. These students, it turns out, have <em>higher</em> returns to schooling than the average college attendee. They are precisely the students the market underinvests in. The IV estimate exceeds OLS because the compliers, the students moved by proximity, gain more from college than the average student who attends.</p><p>Angrist and Krueger&#8217;s instrument, quarter of birth interacting with compulsory schooling laws, moves a different population entirely. These are students at the dropout margin, students who would have left school at 16 if the law allowed but were forced to stay. They are not high-return students blocked by costs. They are students at the bottom of the schooling distribution, and their returns to an additional year of high school are roughly average.</p><p>Two instruments, two estimates, neither one wrong. They are measuring different things because they are moving different people.</p><p>This is why OLS is biased, but the direction of the bias depends on the selection. The OLS estimate of the return to schooling compares the earnings of college graduates to the earnings of non-graduates. But these groups differ in their potential outcomes even before the treatment. If graduates would have earned more even without college (selection on levels), OLS is biased upward. If graduates gain more from college than non-graduates would have (selection on gains), the bias could go either direction depending on which selection dominates. OLS conflates the treatment effect with the selection, and the conflation is specific to the population being compared.</p><h2>Heckman&#8217;s Formalization</h2><p>James Heckman recognized that Roy&#8217;s model was not just about income distribution. It was about selection, and selection was everywhere in economics: in labor supply, in migration, in program participation, in every setting where individuals choose rather than being assigned.</p><p>Heckman&#8217;s contributions came in two waves. The first, in the 1970s, developed methods to correct for selection bias when the selection process is observed only indirectly. The second, in the 1980s and beyond, developed methods to estimate the full structure of the Roy model, recovering the joint distribution of potential outcomes and the selection rule.</p><p>The selection correction, introduced in Heckman&#8217;s 1979 <em>Econometrica</em> paper, addressed a specific problem: estimating a wage equation when wages are observed only for people who work. This is sample selection bias. The people who work are not a random sample of the population; they are people who chose to work, and that choice depends on their potential wages. Heckman showed that if the selection into work follows a probit model (a latent variable crosses a threshold), the bias in the wage equation can be corrected by including a control function, the inverse Mills ratio, computed from the selection equation, as an additional regressor in the wage equation.</p><p>The selection correction was immediately useful and became one of the most cited papers in economics. But it was also limited. It corrected for selection on observables and, through the control function, for selection on a single unobservable that affected both selection and outcomes. It did not fully recover the Roy model&#8217;s structure.</p><p>The full structural approach came in Heckman&#8217;s work with collaborators in the 1980s. In a series of papers, most notably Heckman and Sedlacek (1985), Heckman estimated the complete Roy model: the joint distribution of potential outcomes in both sectors, the selection rule, and the parameters governing sorting. This required strong assumptions, typically joint normality of the unobservables, but it allowed the researcher to recover objects that the selection correction could not: the distribution of potential outcomes for the entire population, the average treatment effect, the effect on any subgroup, and counterfactual outcomes under alternative selection rules.</p><p>The structural Roy model is demanding. It requires specifying the functional forms of the outcome equations and the selection rule, the distributional assumptions on the unobservables, and the exclusion restrictions that identify the selection equation separately from the outcome equations. Critics argued that these assumptions were too strong, that the results were sensitive to specification, and that the credibility purchased by structural modeling was illusory. These critiques were not wrong, but they were not the whole story either.</p><h2>Willis and Rosen: The Schooling Application</h2><p>The most influential early application of the Roy model to treatment effects was Robert Willis and Sherwin Rosen&#8217;s 1979 paper on education. They asked the obvious question: what explains who goes to college? And their answer changed how economists think about selection forever.</p><p>Willis and Rosen modeled earnings in two sectors: the college sector and the high school sector. Each person has a potential wage in each sector, but they choose the sector that yields higher expected utility (accounting for the costs of college, including foregone earnings and tuition). The selection rule is:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!pAp2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f7f11-b3ea-403e-883f-e70fc0719ebd_258x48.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!pAp2!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f7f11-b3ea-403e-883f-e70fc0719ebd_258x48.png 424w, /__u/substackcdn.com/image/fetch/$s_!pAp2!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f7f11-b3ea-403e-883f-e70fc0719ebd_258x48.png 848w, /__u/substackcdn.com/image/fetch/$s_!pAp2!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f7f11-b3ea-403e-883f-e70fc0719ebd_258x48.png 1272w, /__u/substackcdn.com/image/fetch/$s_!pAp2!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f7f11-b3ea-403e-883f-e70fc0719ebd_258x48.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!pAp2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f7f11-b3ea-403e-883f-e70fc0719ebd_258x48.png" width="258" height="48" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/724f7f11-b3ea-403e-883f-e70fc0719ebd_258x48.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:48,&quot;width&quot;:258,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3519,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/193012648?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f7f11-b3ea-403e-883f-e70fc0719ebd_258x48.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!pAp2!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f7f11-b3ea-403e-883f-e70fc0719ebd_258x48.png 424w, /__u/substackcdn.com/image/fetch/$s_!pAp2!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f7f11-b3ea-403e-883f-e70fc0719ebd_258x48.png 848w, /__u/substackcdn.com/image/fetch/$s_!pAp2!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f7f11-b3ea-403e-883f-e70fc0719ebd_258x48.png 1272w, /__u/substackcdn.com/image/fetch/$s_!pAp2!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f7f11-b3ea-403e-883f-e70fc0719ebd_258x48.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>where W&#8321; is the log wage with college, W&#8320; is the log wage without, and C is the (monetized) cost of attending. People attend college when the wage gain exceeds the cost.</p><p>They estimated this model using data on schooling and earnings from the NBER-Thorndike sample, a longitudinal survey of World War II veterans. What they found was not just statistically significant. It was devastating for naive policy analysis.</p><p>College attendees had higher potential earnings <em>with</em> college than non-attendees would have had with college. That part was expected, the able go to college. But here is the kicker: college attendees also had <em>lower</em> potential earnings <em>without</em> college than non-attendees. The people who went to college were not simply &#8220;better&#8221; in some general sense. They were people whose skills paid off specifically in the college sector. They had comparative advantage in college-requiring jobs, not absolute advantage in everything.</p><p>This is not selection on levels. This is selection on gains. The students who attend college are precisely those for whom the <em>gap</em> between college earnings and high school earnings is largest. They are specialists, not generalists.</p><p>Now consider what this means for policy.</p><p>Every proposal to expand college access, subsidies, free tuition, outreach programs, aims to push more students into college. The students who respond to these policies are students who were on the margin. They were not attending before, either because the costs were too high or because they were not sure it was worth it. These are students for whom the gains just barely exceeded the costs, or did not quite exceed them.</p><p>In Roy&#8217;s framework, these marginal students have lower returns to college than the students who were already attending. They have to. If they had high returns, they would have found a way to attend already. The marginal student is marginal precisely because their gains are close to their costs.</p><p>Here is the punchline, and it should hit like a hammer: <strong>the students we push into college are exactly the students who will benefit least from going.</strong></p><p>This is not an argument against expanding college access. There may be good reasons to do it anyway, equity, externalities, the option value of education, the non-monetary benefits of learning. But it is an argument against using the observed return to schooling among current graduates to predict the return for new entrants. The graduates are selected. The new entrants would be selected differently. The marginal student is not the average student, and treating them as interchangeable is the mistake that the Roy model was built to expose.</p><p>Willis and Rosen&#8217;s paper framed the problem that the LATE framework would later formalize. If treatment effects are heterogeneous and selection is on gains, the average treatment effect for compliers, the people moved by a particular policy or instrument, depends on where those compliers sit in the distribution of gains. Different policies move different margins. Different instruments identify different effects. The disagreement is information about the shape of returns.</p><h2>The Roy Model and LATE</h2><p>The connection between Roy and LATE was made explicit in the 1990s, though the two literatures developed semi-independently.</p><p>Imbens and Angrist&#8217;s 1994 paper did not cite Roy. Their framework, potential outcomes, monotonicity, the LATE theorem, was developed in the language of the Rubin causal model, which emphasized randomization and design rather than structural selection models. But the two frameworks describe the same phenomenon from different angles.</p><p>The LATE framework says: with treatment effect heterogeneity, IV identifies the effect for compliers. The Roy model explains why there are compliers, and why they differ from the rest of the population. In Roy&#8217;s terms, compliers are people at the margin of selection. The instrument shifts the threshold, lowering the cost, increasing the incentive, and the compliers are those who were just below the threshold before and just above it after. They are marginal in the precise sense that their gains from treatment are close to their costs. If the gain-cost distribution has shape, the marginal people differ from the average people.</p><p>Consider college proximity as an instrument for schooling, as in Card (1995). Growing up near a college reduces the cost of attending. The compliers are students who would not have attended college if they lived far away but who attend because a college is nearby. Who are these people? Card characterizes them: they are disproportionately from lower-income families, disproportionately Black, disproportionately from families where neither parent attended college. They are students for whom the financial and logistical costs of college matter, students at the margin where a nearby college tips the decision.</p><p>In Roy&#8217;s framework, these are students for whom the wage gain from college is positive but modest, not so high that they would attend regardless of cost, not so low that even zero cost would not induce them. And yet Card&#8217;s IV estimate (13 percent per year of schooling) exceeds the OLS estimate (7 percent). The compliers have <em>higher</em> returns than the average, not lower. Why? Because these are the students the market fails: talented students blocked by costs, students whose potential is unrealized because they cannot afford to move away for college. The instrument reveals a population that policy should care about.</p><p>Different instruments move different margins. A lottery that pays tuition moves students for whom the wage gain exceeds the non-financial costs but not the financial costs. A compulsory schooling law that raises the dropout age moves students who would have dropped out at 16 but not those who would have dropped out at 14 or graduated anyway. Angrist and Krueger&#8217;s quarter-of-birth compliers are students at the very bottom of the educational distribution, students who would have left high school the moment the law allowed. Their IV estimate (about 7 percent, roughly equal to OLS) suggests these students have average returns, not the above-average returns of Card&#8217;s compliers.</p><p>Each instrument defines a different complier population, and each complier population sits at a different part of the MTE curve.</p><p>This is why different instruments give different IV estimates, even when all of them are valid. They are not estimating &#8220;different causal effects&#8221; in some vague sense. They are estimating the average effect for different subpopulations, and the subpopulations differ because selection depends on gains.</p><h2>The MTE Synthesis</h2><p>Heckman and Vytlacil&#8217;s marginal treatment effect framework, developed in papers from 1999 to 2007 and consolidated in their 2005 <em>Econometrica</em> article and 2007 <em>Handbook</em> chapter, unified the Roy model with the LATE framework.</p><p>They parameterized the Roy model in terms of a single index: the unobserved resistance to treatment, U&#7472;. This variable captures everything unobserved that affects the selection decision. People with low U&#7472; select into treatment easily; people with high U&#7472; resist. The propensity score, P(Z), gives the probability of treatment given the instrument, and the marginal person, the one just indifferent between treatment and control, is the person for whom U&#7472; = P(Z).</p><p>The marginal treatment effect, MTE(x, u), is the expected treatment effect for a person with observables x and unobserved resistance u:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Cdvt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e787547-e3a5-4aab-8a26-0ac41fbdc117_362x65.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Cdvt!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e787547-e3a5-4aab-8a26-0ac41fbdc117_362x65.png 424w, /__u/substackcdn.com/image/fetch/$s_!Cdvt!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e787547-e3a5-4aab-8a26-0ac41fbdc117_362x65.png 848w, /__u/substackcdn.com/image/fetch/$s_!Cdvt!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e787547-e3a5-4aab-8a26-0ac41fbdc117_362x65.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Cdvt!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e787547-e3a5-4aab-8a26-0ac41fbdc117_362x65.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Cdvt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e787547-e3a5-4aab-8a26-0ac41fbdc117_362x65.png" width="362" height="65" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6e787547-e3a5-4aab-8a26-0ac41fbdc117_362x65.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:65,&quot;width&quot;:362,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5610,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/193012648?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e787547-e3a5-4aab-8a26-0ac41fbdc117_362x65.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Cdvt!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e787547-e3a5-4aab-8a26-0ac41fbdc117_362x65.png 424w, /__u/substackcdn.com/image/fetch/$s_!Cdvt!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e787547-e3a5-4aab-8a26-0ac41fbdc117_362x65.png 848w, /__u/substackcdn.com/image/fetch/$s_!Cdvt!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e787547-e3a5-4aab-8a26-0ac41fbdc117_362x65.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Cdvt!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e787547-e3a5-4aab-8a26-0ac41fbdc117_362x65.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>This is the effect for the person who is exactly at the margin when the propensity score equals u. If the propensity score increases slightly, this person switches from untreated to treated. The MTE is the return to treatment for that person.</p><p>Here is what makes this useful: every treatment effect parameter, ATE, ATT, LATE, or any policy-relevant treatment effect, is a weighted average of the MTE curve</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!g4Ji!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7992581d-1b9c-41aa-90e2-99222ee503c5_654x261.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!g4Ji!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7992581d-1b9c-41aa-90e2-99222ee503c5_654x261.png 424w, /__u/substackcdn.com/image/fetch/$s_!g4Ji!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7992581d-1b9c-41aa-90e2-99222ee503c5_654x261.png 848w, /__u/substackcdn.com/image/fetch/$s_!g4Ji!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7992581d-1b9c-41aa-90e2-99222ee503c5_654x261.png 1272w, /__u/substackcdn.com/image/fetch/$s_!g4Ji!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7992581d-1b9c-41aa-90e2-99222ee503c5_654x261.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!g4Ji!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7992581d-1b9c-41aa-90e2-99222ee503c5_654x261.png" width="654" height="261" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7992581d-1b9c-41aa-90e2-99222ee503c5_654x261.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:261,&quot;width&quot;:654,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:20298,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/193012648?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7992581d-1b9c-41aa-90e2-99222ee503c5_654x261.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!g4Ji!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7992581d-1b9c-41aa-90e2-99222ee503c5_654x261.png 424w, /__u/substackcdn.com/image/fetch/$s_!g4Ji!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7992581d-1b9c-41aa-90e2-99222ee503c5_654x261.png 848w, /__u/substackcdn.com/image/fetch/$s_!g4Ji!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7992581d-1b9c-41aa-90e2-99222ee503c5_654x261.png 1272w, /__u/substackcdn.com/image/fetch/$s_!g4Ji!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7992581d-1b9c-41aa-90e2-99222ee503c5_654x261.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The weights differ. The ATE weights all margins equally. The ATT puts more weight on low-U&#7472; people (those who selected in). The LATE for a particular instrument puts weight on the part of the MTE curve that the instrument&#8217;s variation illuminates, the compliers, in Imbens-Angrist language.</p><p>In the Roy model with selection on gains, the MTE curve slopes. People who are hard to move into treatment (high U&#7472;) differ systematically from people who are easy to move (low U&#7472;). If the hard-to-move people have higher treatment effects (they resist because the costs are high relative to typical gains, but when they do switch, the gains are large), the curve slopes upward. If the easy-to-move people have higher treatment effects (they switch because they know they will benefit), the curve slopes downward. The sign and shape of the slope is an empirical question, and it determines whether different treatment effect parameters agree or disagree.</p><p>What does this look like with actual numbers? Carneiro, Heckman, and Vytlacil (2011) estimated the MTE curve for college attendance in the United States using NLSY data. They found that the return to college at the margin, the MTE, ranges from about 25 percent at the left of the curve (for people who are easy to push into college) to about 5 percent at the right (for people who resist college strongly). The curve slopes steeply downward. The people who would attend college no matter what have returns around 20 percent. The people who would never attend have returns close to zero or negative.</p><p>The policy implications are immediate. The ATE, the effect if we randomly assigned everyone to college, integrates across the whole curve and is about 14 percent. The ATT, the effect for those who actually attend, weights the left side more heavily and is about 18 percent. A policy that subsidizes tuition would move students in the middle of the curve, with returns around 12 percent. A policy that aggressively recruits students who would otherwise never attend would move students at the far right, with returns near zero. The same intervention looks like a triumph or a waste depending on where it bites.</p><p>Chart 1 illustrates this for the schooling example. The horizontal axis is unobserved resistance to college attendance. The vertical axis is the marginal return to schooling. In the canonical case, selection on gains, the curve slopes downward: people who are easy to push into college have the highest returns. Proximity-based instruments weight the left side of the curve, where returns are highest. Tuition subsidies weight a different segment. Compulsory schooling laws weight the far right, where returns may be lowest. The disagreement among IV estimates is not a sign that some instruments are invalid. It is information about the shape of the MTE curve</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!P4pi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57f76d18-18f5-49d4-a47d-e560bbe586fd_891x532.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!P4pi!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57f76d18-18f5-49d4-a47d-e560bbe586fd_891x532.png 424w, /__u/substackcdn.com/image/fetch/$s_!P4pi!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57f76d18-18f5-49d4-a47d-e560bbe586fd_891x532.png 848w, /__u/substackcdn.com/image/fetch/$s_!P4pi!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57f76d18-18f5-49d4-a47d-e560bbe586fd_891x532.png 1272w, /__u/substackcdn.com/image/fetch/$s_!P4pi!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57f76d18-18f5-49d4-a47d-e560bbe586fd_891x532.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!P4pi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57f76d18-18f5-49d4-a47d-e560bbe586fd_891x532.png" width="891" height="532" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/57f76d18-18f5-49d4-a47d-e560bbe586fd_891x532.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:532,&quot;width&quot;:891,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:76763,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/193012648?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57f76d18-18f5-49d4-a47d-e560bbe586fd_891x532.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!P4pi!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57f76d18-18f5-49d4-a47d-e560bbe586fd_891x532.png 424w, /__u/substackcdn.com/image/fetch/$s_!P4pi!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57f76d18-18f5-49d4-a47d-e560bbe586fd_891x532.png 848w, /__u/substackcdn.com/image/fetch/$s_!P4pi!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57f76d18-18f5-49d4-a47d-e560bbe586fd_891x532.png 1272w, /__u/substackcdn.com/image/fetch/$s_!P4pi!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57f76d18-18f5-49d4-a47d-e560bbe586fd_891x532.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>.</p><h2>Essential Heterogeneity</h2><p>Heckman introduced a distinction that clarifies what the Roy model contributes: the distinction between essential and accidental heterogeneity.</p><p>Accidental heterogeneity is variation in treatment effects that is unrelated to selection. Suppose people have different returns to schooling, but nobody knows their own return when they make the schooling decision. Everyone decides based on the average return plus noise. In this case, the people who go to college are not selected on gains, they are just people who faced low costs or had optimistic expectations or random shocks. With accidental heterogeneity, selection is not on gains, and the ATT may equal the ATE.</p><p>Essential heterogeneity is the opposite. People know their own returns, at least partially, and they select based on that knowledge. The returns to schooling are part of the information set when the schooling decision is made. In this case, selection is on gains, the MTE curve slopes, and different treatment effect parameters diverge.</p><p>The Roy model is a model of essential heterogeneity. It assumes that people sort based on their comparative advantage, which requires that they have information about their potential outcomes in different states. This is a strong assumption. It rules out ignorance and mistakes. But in many economic settings, people do know something about their own abilities, preferences, and prospects, and they act on that knowledge.</p><p>Whether heterogeneity is essential or accidental in any particular application is an empirical question. The distinction matters because it determines what IV identifies. With accidental heterogeneity, LATE equals ATT equals ATE (at least in expectation), and the choice of instrument is less consequential. With essential heterogeneity, they diverge, and the choice of instrument determines what you learn.</p><p>Tests for essential heterogeneity exist. Heckman, Schmierer, and Urzua (2010) developed a test based on comparing IV estimates across instruments with different strengths. If the MTE curve is flat, IV estimates should not vary with instrument strength; if it slopes, they should. The test has power against alternatives where selection on gains is present, and it can detect the direction of the slope.</p><h2>The Generalized Roy Model</h2><p>The Roy model as described so far is a two-sector model: treatment or control, college or high school, hunting or fishing. The generalized Roy model extends this to multiple treatments, continuous treatments, and dynamic selection over time.</p><p>Multiple treatments are common in practice. A worker chooses not just between working and not working but among occupations. A student chooses not just whether to attend college but which college, which major, and how long to stay. A patient chooses among treatments, including the option of no treatment. The generalized Roy model allows for arbitrary numbers of alternatives, each with its own potential outcome, and selection based on expected utility across all options.</p><p>The econometrics becomes harder. With K alternatives, there are K potential outcomes to model, K(K-1)/2 selection margins to consider, and a more complex structure of comparative advantage. But the logic is the same: selection into any particular alternative is not random, and the people observed in that alternative are selected based on their expected outcomes there relative to elsewhere.</p><p>Dynamic selection extends the Roy model over time. The decision to enter a program, persist in treatment, and eventually exit involves a sequence of choices, each made with updated information. The treatment effect depends not just on whether someone enters but on when they enter, how long they stay, and when they leave. The people who complete a program are selected differently than the people who start it, and both are selected differently than the population.</p><p>Heckman and Navarro (2007) and related work developed methods for estimating dynamic treatment effects with Roy-type selection. The hard part is that future choices are anticipated, and current choices are made in light of expected future options. The selection is forward-looking, not just contemporaneous.</p><h2>Two Traditions Reading Roy</h2><p>Both the structural and design-based traditions in econometrics accept the Roy model, but they read it differently.</p><p>The structural tradition, associated with Heckman, reads Roy as a model to be estimated. If you know the joint distribution of potential outcomes and the selection rule, you can recover any treatment effect parameter you want. The goal is to estimate the primitives, the distribution of (Y&#8321;, Y&#8320;, C), the parameters of the selection equation, the shape of the MTE curve, and then compute treatment effects as functions of those primitives. This requires strong assumptions about functional forms and distributions, but it delivers more: the ability to extrapolate to different policies, different populations, and different counterfactuals.</p><p>The design-based tradition, associated with Angrist and Imbens, reads Roy as an explanation for why LATE differs from ATE. The model clarifies the stakes: with selection on gains, IV identifies a local effect, and that effect depends on the instrument. The response is not to estimate the full Roy structure but to be honest about what you have identified. Report the LATE. Characterize the compliers. Do not claim more than the instrument supports.</p><p>These positions are not contradictory. They are different responses to the same tradeoff between credibility and ambition. The structural approach asks for more from the data (and from the researcher&#8217;s assumptions) but delivers more. The design-based approach asks for less but delivers less. The choice depends on the question and on the researcher&#8217;s tolerance for model dependence.</p><p>In practice, the two traditions have converged. Modern MTE methods, as developed by Heckman, Urzua, and Vytlacil and implemented in software like the <code>mtefe</code> package, can be applied with varying degrees of structure. With a continuous instrument that has wide support, the MTE curve can be traced nonparametrically over the range of the instrument. This is quasi-structural: it uses the Roy framework to interpret the object being estimated but does not require the full joint distribution to be specified. The LATE is simply one weighted average of the MTE curve, recoverable as a special case.</p><p>The punchline is that Roy and LATE are two languages for the same thing. Roy describes the economic behavior that generates selection. LATE describes what IV identifies given that selection. The MTE connects them: a single function that encodes the heterogeneity, from which all other parameters can be computed.</p><h2>What Roy Teaches</h2><p>The Roy model does not solve the evaluation problem. It tells you why the problem exists and why it will not go away.</p><p>Selection is not noise to be corrected. It is behavior. People sort into treatments, occupations, locations, and programs because they expect to benefit, and they are often right. The world is full of sorted populations: the employed, the educated, the migrated, the incarcerated. Every one of them is a selected sample. Every comparison of outcomes across groups is contaminated by the fact that the groups formed through choice, not assignment.</p><p>Different instruments identify different effects because different instruments move different people. Card&#8217;s proximity instrument moves cost-sensitive students with high returns. Angrist and Krueger&#8217;s compulsory schooling instrument moves reluctant students at the dropout margin. A tuition subsidy moves yet another group; a lottery-based scholarship moves another. None of these estimates is &#8220;the&#8221; return to schooling. Each is the return for a specific margin, defined by a specific policy lever, in a specific context. The disagreement among estimates is not a sign that someone made a mistake. It is the signature of selection on gains.</p><p>The marginal person is not the average person. This is the sentence that should be tattooed on the forehead of every policy analyst. The students who respond to a tuition cut are not the students who were already attending. The workers who take a job training program when offered are not the workers who would have enrolled anyway. The patients who comply with a treatment recommendation are not the patients who refuse. Every margin has its own population, and the population determines the effect.</p><p>Policy evaluation without the Roy model is astrology. You can compute a number, but the number does not mean what you think it means. You can observe that college graduates earn more than high school graduates, but you cannot conclude that sending more people to college will close the gap. You can find that job training raises earnings for participants, but you cannot conclude that expanding the program will raise earnings for new participants. The treated are selected. The marginal entrants would be selected differently. The extrapolation fails because the effect depends on who is being treated.</p><p>And yet.</p><p>The Roy model is also a source of hope. If the MTE curve can be traced, if we have instruments with enough variation, enough support, enough power to illuminate multiple margins, we can learn the shape of returns across the population. We can see where the curve is high and where it is low. We can compute the effect for any policy, any expansion, any contraction, by integrating the curve with the right weights. The Roy model does not make extrapolation impossible. It makes extrapolation honest. It tells you what you need to know and whether you know it.</p><p>Here is what it comes down to: every treatment effect estimate is a statement about who was moved, not about what the treatment does in general. The effect does not live in the treatment. It lives in the interaction between the treatment and the person. Some people gain a lot; some gain a little; some lose. The average depends on who you average over, and who you average over depends on the margin of selection. Change the margin, change the average.</p><p>This is why Roy&#8217;s 1951 paper, written about hunting and fishing in a world before randomized trials and natural experiments and LATE, still matters. He saw that people choose, that choice depends on expected outcomes, and that the distribution of observed outcomes is a shadow of the distribution of choices. Every evaluation problem is, at root, a selection problem. The Roy model is the language for saying what that means.</p><h2>Where to Start</h2><p>Roy, A.D. (1951). &#8220;Some Thoughts on the Distribution of Earnings.&#8221; <em>Oxford Economic Papers</em> 3(2): 135-146. The original paper. Short and readable. Roy uses the hunting-fishing example to show how selection on comparative advantage shapes the observed distribution of earnings.</p><p>Heckman, James J. (1979). &#8220;Sample Selection Bias as a Specification Error.&#8221; <em>Econometrica</em> 47(1): 153-161. The selection correction. Wages are observed only for workers, and workers are selected. Heckman shows how to correct for this bias using a control function derived from the selection equation. The most cited paper in economics for decades.</p><p>Willis, Robert J., and Sherwin Rosen (1979). &#8220;Education and Self-Selection.&#8221; <em>Journal of Political Economy</em> 87(5): S7-S36. The Roy model applied to schooling. Willis and Rosen estimate the joint distribution of college and high school wages and find evidence of comparative advantage: college attendees are both more productive with college and less productive without it than non-attendees.</p><p>Heckman, James J., and Bo E. Honor&#233; (1990). &#8220;The Empirical Content and Identification of the Roy Model.&#8221; <em>Econometrica</em> 58(5): 1121-1149. What can you learn from the Roy model without strong distributional assumptions? The paper establishes identification results for Roy selection under minimal parametric restrictions.</p><p>Heckman, James J., and Edward Vytlacil (2005). &#8220;Structural Equations, Treatment Effects, and Econometric Policy Evaluation.&#8221; <em>Econometrica</em> 73(3): 669-738. The MTE synthesis. Every treatment parameter is a weighted average of the marginal treatment effect curve. The paper that unified the Roy model with LATE and showed how to recover policy-relevant effects from IV variation.</p><h2>Some Additional Readings</h2><p><strong>On the Roy model and its extensions</strong></p><p>Heckman, James J., and Guilherme Sedlacek (1985). &#8220;Heterogeneity, Aggregation, and Market Wage Functions: An Empirical Model of Self-Selection in the Labor Market.&#8221; <em>Journal of Political Economy</em> 93(6): 1077-1125. The full structural Roy model estimated. Heckman and Sedlacek recover the joint distribution of potential wages in multiple sectors and show how self-selection affects the observed wage distribution.</p><p>French, Eric, and Christopher Taber (2011). &#8220;Identification of Models of the Labor Market.&#8221; <em>Handbook of Labor Economics</em> 4A: 537-617. A modern survey of identification in labor market models, including Roy and its generalizations. Technical but comprehensive.</p><p><strong>On the connection between Roy and LATE</strong></p><p>Heckman, James J. (1997). &#8220;Instrumental Variables: A Study of Implicit Behavioral Assumptions Used in Making Program Evaluations.&#8221; <em>Journal of Human Resources</em> 32(3): 441-462. Heckman&#8217;s critique of the LATE framework, arguing that the Roy model reveals the limitations of instrument-specific identification.</p><p>Heckman, James J., and Edward Vytlacil (1999). &#8220;Local Instrumental Variables and Latent Variable Models for Identifying and Bounding Treatment Effects.&#8221; <em>Proceedings of the National Academy of Sciences</em> 96(8): 4730-4734. The bridge paper. Introduces the MTE and shows how LATE relates to the structural Roy framework.</p><p>Heckman, James J., and Edward Vytlacil (2007). &#8220;Econometric Evaluation of Social Programs, Part II: Using the Marginal Treatment Effect to Organize Alternative Econometric Estimators to Evaluate Social Programs, and to Forecast Their Effects in New Environments.&#8221; <em>Handbook of Econometrics</em> 6B: 4875-5143. The definitive treatment. Long, technical, and complete.</p><p>Carneiro, Pedro, James J. Heckman, and Edward Vytlacil (2011). &#8220;Estimating Marginal Returns to Education.&#8221; <em>American Economic Review</em> 101(6): 2754-2781. The empirical application that makes the MTE curve concrete. Estimates that college returns range from 25 percent at the margin of easy selection to 5 percent at the margin of strong resistance. The paper that puts numbers on the theory.</p><p><strong>On essential heterogeneity and testing</strong></p><p>Heckman, James J., Daniel Schmierer, and Sergio Urzua (2010). &#8220;Testing the Correlated Random Coefficient Model.&#8221; <em>Journal of Econometrics</em> 158(2): 177-203. How to test whether heterogeneity is essential (correlated with selection) or accidental. The test compares IV estimates across instruments with different identifying power.</p><p><strong>On implementation</strong></p><p>Cornelissen, Thomas, Christian Dustmann, Anna Raute, and Uta Sch&#246;nberg (2016). &#8220;From LATE to MTE: Alternative Methods for the Evaluation of Policy Interventions.&#8221; <em>Labour Economics</em> 41: 47-60. A practical guide to MTE estimation with applications to parental leave policy. Accessible and applied.</p><p>Andresen, Martin Eckhoff (2018). &#8220;Exploring Marginal Treatment Effects: Flexible Estimation Using Stata.&#8221; <em>Stata Journal</em> 18(1): 118-158. The documentation for the <code>mtefe</code> command. Shows how to estimate MTE curves, compute treatment effects, and visualize results.</p><p>This is the fourth essay in a series on instrumental variables and econometric methods. The first essay covered the history of instrumental variables from Wright (1928) to the credibility revolution. The second traced the two-stage architecture from Heckman's selection correction to double machine learning. The third mapped the family of instruments and asked what each one identifies.</p>]]></content:encoded></item><item><title><![CDATA[Why Christopher Sims Was So Great?]]></title><description><![CDATA[VARs, Bayesian methods, fiscal theory, and rational inattention: four tools that reshaped a field]]></description><link>https://carloschavezp29.substack.com/p/why-christopher-sims-was-so-great</link><guid isPermaLink="false">https://carloschavezp29.substack.com/p/why-christopher-sims-was-so-great</guid><dc:creator><![CDATA[Carlos Chavez]]></dc:creator><pubDate>Mon, 23 Mar 2026 10:40:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!uuj3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17207582-95c7-41fc-a3b4-86ffc7362bac_905x705.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Christopher Sims transformed how macroeconomists think about data, uncertainty, inflation, and the limits of human attention. Across four decades, he kept asking the same question: what is this model assuming away? And then he built the tools to address it.</p><p><a href="/__u/carloschavezp29.substack.com/p/why-macro-never-had-a-credibility?r=7es0y">I have written before about Sims&#8217; critique of large-scale macroeconometric models and about why macroeconomics never had a credibility revolution</a>. Today I want to focus on what Sims built, not what he tore down. His contributions span four areas: vector autoregressions replaced incredible structural models with transparent empirical tools. Bayesian methods brought intellectual honesty to estimation under uncertainty. The fiscal theory of the price level challenged the conventional separation of monetary and fiscal policy. And rational inattention reimagined how economic agents process information. In each case, the pattern was the same: identify what was being swept under the rug, and build a framework that forced it into the open.</p><p>The figure below traces this intellectual arc across four decades</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!uuj3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17207582-95c7-41fc-a3b4-86ffc7362bac_905x705.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!uuj3!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17207582-95c7-41fc-a3b4-86ffc7362bac_905x705.png 424w, /__u/substackcdn.com/image/fetch/$s_!uuj3!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17207582-95c7-41fc-a3b4-86ffc7362bac_905x705.png 848w, /__u/substackcdn.com/image/fetch/$s_!uuj3!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17207582-95c7-41fc-a3b4-86ffc7362bac_905x705.png 1272w, /__u/substackcdn.com/image/fetch/$s_!uuj3!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17207582-95c7-41fc-a3b4-86ffc7362bac_905x705.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!uuj3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17207582-95c7-41fc-a3b4-86ffc7362bac_905x705.png" width="905" height="705" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17207582-95c7-41fc-a3b4-86ffc7362bac_905x705.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:705,&quot;width&quot;:905,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:124198,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/191848699?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17207582-95c7-41fc-a3b4-86ffc7362bac_905x705.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!uuj3!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17207582-95c7-41fc-a3b4-86ffc7362bac_905x705.png 424w, /__u/substackcdn.com/image/fetch/$s_!uuj3!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17207582-95c7-41fc-a3b4-86ffc7362bac_905x705.png 848w, /__u/substackcdn.com/image/fetch/$s_!uuj3!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17207582-95c7-41fc-a3b4-86ffc7362bac_905x705.png 1272w, /__u/substackcdn.com/image/fetch/$s_!uuj3!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17207582-95c7-41fc-a3b4-86ffc7362bac_905x705.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>I. The Language of Macroeconomic Data</h2><p>Before Sims, macroeconomists argued about models. After Sims, they argued about facts.</p><p>The dominant empirical approach of the 1970s was the large-scale structural model. These models, built at institutions like the Federal Reserve, the Brookings Institution, and the Cowles Commission, contained dozens or hundreds of equations representing consumption, investment, money demand, labor supply, price formation. The ambition was to capture the entire macroeconomic system in simultaneous equations that could be estimated, simulated, and used for policy analysis.</p><p>The problem was that each equation required assumptions about which variables appeared and which did not. A variable was declared exogenous because the model required it. A coefficient was set to zero because it simplified estimation. The models were impressive engineering, but their empirical content rested on exclusion restrictions that were difficult to defend on independent grounds.</p><p>By the mid-1970s, the limitations of this approach were widely recognized. Robert Lucas&#8217; famous critique (Lucas, 1976) showed that estimated parameters could shift when policy changed, because agents adjust their behavior to new regimes. Lucas and others proposed building models from deeper primitives, preferences, technology, rational expectations, that would remain stable across policy changes.</p><p>Sims took a different path. Rather than replacing one set of strong assumptions with another, he asked a more basic question: what can we learn from the data with minimal assumptions?</p><p>The answer was the vector autoregression (Sims, 1980). Take a set of macroeconomic variables, say, output, inflation, the interest rate, and the money supply. Instead of assuming a structural model that specifies which variable affects which, let every variable depend on its own past values and the past values of every other variable:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!O8LN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e6d7c2d-2189-4782-9de5-dd40f8e28785_384x48.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!O8LN!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e6d7c2d-2189-4782-9de5-dd40f8e28785_384x48.png 424w, /__u/substackcdn.com/image/fetch/$s_!O8LN!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e6d7c2d-2189-4782-9de5-dd40f8e28785_384x48.png 848w, /__u/substackcdn.com/image/fetch/$s_!O8LN!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e6d7c2d-2189-4782-9de5-dd40f8e28785_384x48.png 1272w, /__u/substackcdn.com/image/fetch/$s_!O8LN!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e6d7c2d-2189-4782-9de5-dd40f8e28785_384x48.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!O8LN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e6d7c2d-2189-4782-9de5-dd40f8e28785_384x48.png" width="384" height="48" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9e6d7c2d-2189-4782-9de5-dd40f8e28785_384x48.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:48,&quot;width&quot;:384,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4622,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/191848699?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e6d7c2d-2189-4782-9de5-dd40f8e28785_384x48.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!O8LN!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e6d7c2d-2189-4782-9de5-dd40f8e28785_384x48.png 424w, /__u/substackcdn.com/image/fetch/$s_!O8LN!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e6d7c2d-2189-4782-9de5-dd40f8e28785_384x48.png 848w, /__u/substackcdn.com/image/fetch/$s_!O8LN!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e6d7c2d-2189-4782-9de5-dd40f8e28785_384x48.png 1272w, /__u/substackcdn.com/image/fetch/$s_!O8LN!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e6d7c2d-2189-4782-9de5-dd40f8e28785_384x48.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>where Y_t &#8203; is a vector of macroeconomic variables and u_t&#8203; is a vector of reduced-form residuals. No economic theory is imposed. The VAR simply describes how the variables move together over time.</p><p>To give these co-movements a causal interpretation, the researcher must decompose the residuals into orthogonal structural shocks. This requires restrictions, but the restrictions are stated explicitly. The simplest approach is a recursive ordering, a Cholesky decomposition that assumes some variables respond to shocks within the period while others do not. Place the interest rate last, and you are assuming that monetary policy responds within the period to all other variables, but output and prices respond only with a lag. This is an assumption, but it is an assumption you can see, debate, and test.</p><p>Once you have identified structural shocks, you trace their effects forward through impulse response functions. What happens to output over the next twelve quarters after a monetary policy shock? The impulse response function became the signature output of the VAR literature, and it changed the terms of debate. Instead of arguing about which theoretical assumptions were correct, macroeconomists could point to empirical facts and ask: can your model replicate these?</p><p>The toolkit expanded rapidly. Blanchard and Quah (1989) introduced long-run restrictions. Uhlig (2005) proposed sign restrictions. Romer and Romer (2004) used narrative methods. Each approach made different assumptions. Each was explicit about what it required. This is the Sims legacy in empirical macro: not a specific method, but a methodological standard. State your assumptions, defend them, and show that your results survive scrutiny.</p><h2>II. Taking Uncertainty Seriously</h2><p>VARs are high-dimensional. A VAR with six variables and eight lags has hundreds of parameters. With typical macroeconomic samples of a few hundred quarterly observations, classical frequentist methods struggle. Sims&#8217; response was to turn to Bayesian methods, and in doing so he became a central figure in bringing Bayesian econometrics into mainstream macroeconomics.</p><p>The key insight was that priors are a feature, not a bug. In high-dimensional models, you need to regularize. Frequentist methods do this implicitly through model selection. Bayesian methods do it explicitly through the choice of prior.</p><p>The most influential application was the Minnesota prior (Doan, Litterman, and Sims, 1984), developed by Sims with Robert Litterman and Thomas Doan in the early 1980s. The prior embodies a simple belief: each variable in a VAR is likely to behave roughly like a random walk. It shrinks all coefficients toward zero except the first own lag of each variable, which is shrunk toward one. If you have no information about the relationship between industrial production and unemployment at various lags, start from the assumption that each variable mostly follows its own recent trend. The data are free to update this wherever the evidence is strong enough.</p><p>The Bayesian VAR became a workhorse forecasting tool. The Federal Reserve, the European Central Bank, the Bank of England, all widely use variants of BVARs for macroeconomic forecasting.</p><p>But Sims&#8217; advocacy went beyond forecasting. He argued that uncertainty about the model itself, not just uncertainty about parameters within a model, was a first-order concern for policymakers. Central banks do not know the true model of the economy. A Bayesian framework lets you assign probabilities to different models, update them as evidence accumulates, and make policy decisions that account for this uncertainty. The Smets and Wouters (2007) model, estimated with Bayesian methods and widely used at central banks, is a direct descendant of this insistence.</p><h2>III. The Fiscal Theory of the Price Level</h2><p>VARs and Bayesian methods are primarily methodological. The fiscal theory of the price level is different. It is a substantive claim about how inflation works, and one that many macroeconomists still reject.</p><p>The standard view is that inflation is fundamentally a monetary phenomenon. Sims proposed an alternative. His argument is that the price level is also determined by the government&#8217;s intertemporal budget constraint:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!tyZs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f9b89e-19d8-4e37-a5ee-5f9a1c1b53a3_191x86.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!tyZs!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f9b89e-19d8-4e37-a5ee-5f9a1c1b53a3_191x86.png 424w, /__u/substackcdn.com/image/fetch/$s_!tyZs!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f9b89e-19d8-4e37-a5ee-5f9a1c1b53a3_191x86.png 848w, /__u/substackcdn.com/image/fetch/$s_!tyZs!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f9b89e-19d8-4e37-a5ee-5f9a1c1b53a3_191x86.png 1272w, /__u/substackcdn.com/image/fetch/$s_!tyZs!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f9b89e-19d8-4e37-a5ee-5f9a1c1b53a3_191x86.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!tyZs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f9b89e-19d8-4e37-a5ee-5f9a1c1b53a3_191x86.png" width="191" height="86" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/44f9b89e-19d8-4e37-a5ee-5f9a1c1b53a3_191x86.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:86,&quot;width&quot;:191,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4666,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/191848699?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f9b89e-19d8-4e37-a5ee-5f9a1c1b53a3_191x86.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!tyZs!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f9b89e-19d8-4e37-a5ee-5f9a1c1b53a3_191x86.png 424w, /__u/substackcdn.com/image/fetch/$s_!tyZs!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f9b89e-19d8-4e37-a5ee-5f9a1c1b53a3_191x86.png 848w, /__u/substackcdn.com/image/fetch/$s_!tyZs!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f9b89e-19d8-4e37-a5ee-5f9a1c1b53a3_191x86.png 1272w, /__u/substackcdn.com/image/fetch/$s_!tyZs!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f9b89e-19d8-4e37-a5ee-5f9a1c1b53a3_191x86.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>where B_t&#8203; is nominal government debt, P_t is the price level, and &#964;_{t+s} represents future real primary surpluses. In the standard view, fiscal policy adjusts to satisfy this constraint. In the fiscal theory, when fiscal authorities show no inclination to adjust, the price level does the work instead. If the present value of future surpluses falls short of the real value of debt, inflation rises to reduce debt&#8217;s real value until the equation is satisfied.</p><p>This means a central bank cannot always independently control inflation. Its power depends on whether fiscal authorities cooperate, on whether fiscal policy is &#8220;Ricardian&#8221; (adjusting surpluses to satisfy the constraint) or &#8220;non-Ricardian&#8221; (leaving the price level to do it).</p><p>Sims applied this framework to the European Monetary Union in a prescient paper (Sims, 1999) titled &#8220;The Precarious Fiscal Foundations of EMU.&#8221; He argued that a monetary union without fiscal integration was unstable. A decade later, the Greek fiscal crisis conformed remarkably well to the scenario Sims had warned about.</p><p>The theory remains contested. Many macroeconomists, including Sims&#8217; Nobel co-laureate Thomas Sargent, have questioned whether the government budget constraint should be interpreted as an equilibrium condition that determines the price level or simply as a constraint that fiscal authorities must satisfy. But the core insight, that fiscal and monetary policy cannot be analyzed independently, has become part of how macroeconomists think about the world. The global accumulation of government debt since 2020 has only made these questions more urgent.</p><h2>IV. Rational Inattention</h2><p>The first three contributions deal with how economists analyze data and think about policy. The fourth deals with how we model people.</p><p>Standard models assume agents are fully informed. Rational expectations theory takes this further: agents use all available information optimally. Sims wanted to extend this framework by asking what happens when information processing itself has a cost. People do not track every price in the economy. Firms do not continuously update their pricing decisions. The question was how to model this limited capacity in a disciplined way.</p><p>His answer drew on Claude Shannon&#8217;s information theory. In his influential 2003 paper (Sims, 2003), Sims proposed treating economic agents as finite-capacity channels. An agent observes a multidimensional world, prices, incomes, interest rates, news, but can only process a limited amount of information per unit of time, measured in Shannon bits. The agent must choose what to pay attention to. The formal constraint is:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!qxCU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25bd8ba7-ac88-468b-8d95-3d4804b74831_140x47.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!qxCU!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25bd8ba7-ac88-468b-8d95-3d4804b74831_140x47.png 424w, /__u/substackcdn.com/image/fetch/$s_!qxCU!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25bd8ba7-ac88-468b-8d95-3d4804b74831_140x47.png 848w, /__u/substackcdn.com/image/fetch/$s_!qxCU!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25bd8ba7-ac88-468b-8d95-3d4804b74831_140x47.png 1272w, /__u/substackcdn.com/image/fetch/$s_!qxCU!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25bd8ba7-ac88-468b-8d95-3d4804b74831_140x47.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!qxCU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25bd8ba7-ac88-468b-8d95-3d4804b74831_140x47.png" width="140" height="47" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/25bd8ba7-ac88-468b-8d95-3d4804b74831_140x47.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:47,&quot;width&quot;:140,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2576,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/191848699?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25bd8ba7-ac88-468b-8d95-3d4804b74831_140x47.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!qxCU!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25bd8ba7-ac88-468b-8d95-3d4804b74831_140x47.png 424w, /__u/substackcdn.com/image/fetch/$s_!qxCU!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25bd8ba7-ac88-468b-8d95-3d4804b74831_140x47.png 848w, /__u/substackcdn.com/image/fetch/$s_!qxCU!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25bd8ba7-ac88-468b-8d95-3d4804b74831_140x47.png 1272w, /__u/substackcdn.com/image/fetch/$s_!qxCU!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25bd8ba7-ac88-468b-8d95-3d4804b74831_140x47.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>where I(C;W) is the mutual information between the agent&#8217;s actions and the true state, and &#954; is the channel&#8217;s capacity.</p><p>What does this imply? Agents&#8217; responses to external changes are sluggish and noisy, not because of mechanical frictions, but because they rationally choose not to track every fluctuation precisely. Prices update slowly not because of menu costs or Calvo contracts, but because firms have finite capacity to monitor their environment. The degree of inertia is endogenous: when inflation becomes more volatile, agents allocate more capacity to tracking it, and the economy&#8217;s response to shocks changes.</p><p>The framework also explains why individual behavior is noisy even among agents facing identical fundamentals, and why central bank communication matters. If agents have finite capacity, the way information is presented affects how it is processed.</p><p>The rational inattention literature has grown substantially. Ma&#263;kowiak and Wiederholt (2009) showed that firms pay more attention to idiosyncratic shocks than to aggregate ones, explaining why aggregate prices are stickier than individual prices. Mat&#283;jka and McKay (2015) proved that the multinomial logit model emerges as the optimal choice rule under rational inattention with Shannon entropy costs, an unexpected bridge between information theory and discrete choice.</p><h2>The Unifying Thread</h2><p>In October 2011, the Royal Swedish Academy of Sciences awarded the Nobel Memorial Prize in Economic Sciences jointly to Christopher Sims and Thomas Sargent. The citation recognized their &#8220;empirical research on cause and effect in the macroeconomy.&#8221; The pairing was apt. Sargent had pursued rational expectations from within, building structural models with forward-looking agents. Sims had pursued a parallel path from the outside, building empirical tools that could discipline theory without assuming it. They had argued for forty years, productively and with mutual respect.</p><p>The deeper unity across Sims&#8217; work is the pattern I described at the beginning. Where large-scale structural models relied on untested restrictions, Sims built VARs that made assumptions transparent. Where classical estimation treated model choice as settled, Sims introduced Bayesian methods that acknowledged uncertainty honestly. Where monetary theory treated inflation as the central bank&#8217;s domain alone, Sims showed that fiscal foundations matter. Where standard models assumed perfect information processing, Sims modeled the limits of attention.</p><p>Each contribution began with an acknowledgment that something important was being swept under the rug, and proceeded to build a framework that took it seriously. The data come first. Theory must explain the data, not the other way around. Sims was a theorist of the highest caliber, but he insisted that theory earns its place by matching empirical regularities, not by logical elegance alone.</p><p>During his freshman year at Harvard, a teaching assistant had told Sims that as a mathematician, he would never be able to change the world. He took the advice seriously. He finished the mathematics degree, then turned to economics. The teaching assistant was right about one thing: pure mathematics, in the abstract, might not change the world. But mathematics in the service of honest empirical inquiry can. Sims died on March 14, 2026. He was 83.</p><p>There is one part of the Sims story I have deliberately left out: his early work on causality. In &#8220;Money, Income, and Causality&#8221; (Sims, 1972), he applied Granger&#8217;s framework to test whether money causes income or vice versa, and his results reshaped the monetarist-Keynesian debate. That work, and the broader question of what &#8220;causality&#8221; means in time series, deserves its own treatment. I will return to it in a future essay.</p><h2>Some Worth Readings</h2><p>Sims, C.A. (1972). &#8220;Money, Income, and Causality.&#8221; <em>American Economic Review</em>, 62(4): 540&#8211;552. The paper that tested whether money causes income or the reverse, using Granger&#8217;s temporal causality framework. It sided with the monetarists, but the method mattered more than the answer.</p><p>Lucas, R.E. (1976). &#8220;Econometric Policy Evaluation: A Critique.&#8221; <em>Carnegie-Rochester Conference Series on Public Policy</em>, 1: 19&#8211;46. The paper that identified a fundamental limitation of large-scale econometric models: if agents adjust expectations when policy changes, the estimated parameters shift too. Sims built on this insight but took a different path forward.</p><p>Sims, C.A. (1980). &#8220;Macroeconomics and Reality.&#8221; <em>Econometrica</em>, 48(1): 1&#8211;48. The paper that introduced VARs into macroeconomics and called the identifying assumptions of large-scale models &#8220;incredible.&#8221; One of the most cited empirical macro papers ever written, with over 13,000 citations.</p><p>Blanchard, O.J. &amp; D. Quah (1989). &#8220;The Dynamic Effects of Aggregate Demand and Supply Disturbances.&#8221; <em>American Economic Review</em>, 79(4): 655&#8211;673. Extended the VAR framework by using long-run restrictions to separate supply and demand shocks. A permanent effect on output means it was supply; a transitory effect means it was demand.</p><p>Sims, C.A. (1999). &#8220;The Precarious Fiscal Foundations of EMU.&#8221; <em>De Economist</em>, 147(4): 415&#8211;436. Sims warned that a monetary union without fiscal integration was unstable. The eurozone crisis a decade later confirmed the warning.</p><p>Sims, C.A. (2003). &#8220;Implications of Rational Inattention.&#8221; <em>Journal of Monetary Economics</em>, 50(3): 665&#8211;690. The paper that imported Shannon&#8217;s information theory into economics. Agents have finite capacity to process information, and this alone generates the sluggish, noisy behavior we see in the data.</p><p>Romer, C.D. &amp; D.H. Romer (2004). &#8220;A New Measure of Monetary Shocks: Derivation and Implications.&#8221; <em>American Economic Review</em>, 94(4): 1055&#8211;1084. Narrative identification of monetary policy shocks. Read the Fed&#8217;s internal forecasts, strip out the predictable component, and what remains is a genuine policy surprise.</p><p>Uhlig, H. (2005). &#8220;What Are the Effects of Monetary Policy on Output? Results from an Agnostic Identification Procedure.&#8221; <em>Journal of Monetary Economics</em>, 52(2): 381&#8211;419. Sign restrictions as an alternative to recursive orderings. Instead of assuming a causal ordering, impose only that a contractionary shock does not raise output or lower rates on impact.</p><p>Smets, F. &amp; R. Wouters (2007). &#8220;Shocks and Frictions in US Business Cycles: A Bayesian DSGE Approach.&#8221; <em>American Economic Review</em>, 97(3): 586&#8211;606. The benchmark Bayesian DSGE model used at central banks worldwide. A direct descendant of Sims&#8217; insistence on taking both theory and uncertainty seriously.</p><p>Ma&#263;kowiak, B. &amp; M. Wiederholt (2009). &#8220;Optimal Sticky Prices under Rational Inattention.&#8221; <em>American Economic Review</em>, 99(3): 769&#8211;803. Showed that firms rationally pay more attention to their own costs than to aggregate conditions. This explains why individual prices move a lot but the aggregate price level is sticky.</p><p>Mat&#283;jka, F. &amp; A. McKay (2015). &#8220;Rational Inattention to Discrete Choices: A New Foundation for the Multinomial Logit Model.&#8221; <em>American Economic Review</em>, 105(1): 272&#8211;298. Proved that the multinomial logit, one of the most used models in applied economics, emerges naturally from rational inattention with Shannon entropy costs. An unexpected bridge between information theory and discrete choice.</p>]]></content:encoded></item><item><title><![CDATA[The Missing Sort Command]]></title><description><![CDATA[How a Single Line of Code Undid a Top Paper, and What AI Means for the Future of Replication]]></description><link>https://carloschavezp29.substack.com/p/the-missing-sort-command</link><guid isPermaLink="false">https://carloschavezp29.substack.com/p/the-missing-sort-command</guid><dc:creator><![CDATA[Carlos Chavez]]></dc:creator><pubDate>Fri, 20 Mar 2026 06:55:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!zh0V!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffff9b4a5-53ee-4548-85da-563173f13249_2666x2666.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In early 2026, Michael Wiebe&#8217;s comment on Moretti (2021) was accepted at the <em>American Economic Review</em>. The comment showed that the causal results in one of the most cited recent papers on agglomeration and innovation were driven by coding errors. The paper had studied whether bigger technology clusters cause inventors to patent more, a question with direct implications for housing policy, urban planning, and billions of dollars in public subsidies to attract tech firms. Moretti found large positive effects. The event study showed a clear jump in patenting when inventors moved to bigger clusters. The instrumental variables strategy confirmed the causal interpretation.</p><p>Both results were wrong. The errors were plain bugs in the code, not subtle econometric judgments or debatable identification assumptions.</p><p>The event study specification was wrong. Among other problems, the treatment variable was not properly interacted with all event-year indicators, including year zero. The coefficient for the move year was therefore estimated using data from all years, not just the move year, inflating the apparent effect. The instrumental variable, based on the number of inventors in the same field working for firms in other cities, was constructed from data that had not been sorted by city. The code computed first-differences across different cities rather than within them. A single missing sort command meant the instrument was mixing up city A&#8217;s value this year with city B&#8217;s value last year. Correcting the sort order and rerunning the IV regressions produces first-stage F-statistics around 4.5 to 7, far below conventional thresholds for instrument strength, and null second-stage estimates.</p><p>Wiebe spent three years on this replication. He emailed Moretti in July 2023 about the event study problem. He never received a response. The comment went through a full round of review and revision at the AER before acceptance. Along the way, Wiebe documented ten distinct issues: the event study bug, the IV sorting error, a problematic log transformation using log(y + 0.00001) that reversed the citation quality results, an identifier based on inventor names that conflated different people in different cities, and cleaning code that used many-to-many merges with nonunique sort orders, producing different samples every time the code was run.</p><p>This story is worth telling not because Moretti made mistakes, since everyone makes mistakes, but because of what it reveals about the infrastructure of empirical economics. We have spent thirty years refining our identification strategies. We have developed sophisticated tools for causal inference: difference-in-differences with heterogeneous treatment effects, synthetic controls, regression discontinuity designs, shift-share instruments with proper inference. The econometric methodology has never been better. But the code that implements these methods often runs on the digital equivalent of duct tape and hope. Fragile Stata scripts with undocumented variable names, unsorted datasets, arbitrary choices buried deep in data-cleaning routines that no one else will ever read.</p><p>The credibility revolution taught economists to take identification seriously. We are now learning, more slowly and more painfully, that we also need to take implementation seriously.</p>
      <p>
          <a href="/__u/carloschavezp29.substack.com/p/the-missing-sort-command">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[LATE for the Party]]></title><description><![CDATA[On parallel discoveries in causal inference]]></description><link>https://carloschavezp29.substack.com/p/late-for-the-party</link><guid isPermaLink="false">https://carloschavezp29.substack.com/p/late-for-the-party</guid><dc:creator><![CDATA[Carlos Chavez]]></dc:creator><pubDate>Mon, 16 Mar 2026 04:25:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!vj6g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618b50e8-97c0-40eb-a32c-0ac7d5a09d7a_830x524.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In 1994, Guido Imbens and Joshua Angrist published a nine-page paper in <em>Econometrica</em> that formalized what instrumental variables actually estimate when treatment effects differ across people. They called it the local average treatment effect: the causal effect for compliers, the subgroup whose treatment status is changed by the instrument. The paper, together with the monotonicity assumption, the requirement that the instrument does not push anyone in the opposite direction, reshaped how economists interpret IV estimates. Imbens and Angrist shared the 2021 Nobel Prize in part for this work.</p><p>What is less well known is what was happening in other fields at the same time. And before.</p><p>At least eight groups of researchers, working independently across economics, biostatistics, and epidemiology, arrived at the same idea between 1983 and 1994.</p><p>I learned about this history a few weeks ago, when a reader named Stuart Baker sent me an email after reading one of my essays on instrumental variables. Baker is a retired mathematical statistician from the National Cancer Institute. He pointed me to a 2024 article he published in <em>CHANCE</em> magazine with Karen Lindeman, titled &#8220;Multiple Discoveries in Causal Inference: LATE for the Party.&#8221; The title borrows from Angrist&#8217;s own Nobel lecture, where Angrist acknowledged the parallel discoveries and described himself and Imbens as &#8220;late to the partial compliance party.&#8221;</p><p>Here is what Baker and Lindeman documented.</p><h2>The Eight Discoveries</h2><p>The earliest formulation belongs to Baker himself. In 1983, as a graduate student in biostatistics at the Harvard School of Public Health, he was working on Marvin Zelen&#8217;s &#8220;randomized consent&#8221; design, in which participants are randomized to either standard care or an offer of a new treatment. Some people offered the new treatment decline it. This creates noncompliance: not everyone in the treatment group actually gets the treatment.</p><p>Baker wanted to estimate the effect of actually receiving the treatment, not just being offered it. He worked out the problem from first principles. He defined what would later be called principal strata: never-takers who would refuse treatment regardless of assignment, and compliers who would accept when offered. He derived maximum likelihood estimates for binary outcomes under the assumption that randomization only affects outcomes through the treatment actually received, what economists would later call the exclusion restriction. On August 12, 1983, he completed a manuscript. He showed it to his mentors. They were not enthusiastic. The paper was not published.</p><p>One year later, Howard Bloom, at the Harvard School of Government, independently derived the same result for one-sided noncompliance and published in <em>Evaluation Review</em>. In 1989, Thomas Permutt and J. Richard Hebel, at the University of Maryland School of Medicine, formulated it for two-sided noncompliance using recursive equations and published in <em>Biometrics</em>. In 1991, Alfred Sommer and Scott Zeger at Johns Hopkins developed it using relative risk. The same year, Robert Connor, Philip Prorok, and Douglas Weed at the National Institutes of Health arrived at it independently for case-control designs. And around the same time, Imbens and Angrist were working it out at the Harvard Economics Department, initially for one-sided noncompliance, before extending it with monotonicity in their 1994 <em>Econometrica</em> paper.</p><p>Also in 1994, Baker and Karen Lindeman, an anesthesiologist at Johns Hopkins, published their version in <em>Statistics in Medicine</em>. They called it the &#8220;paired availability design.&#8221; The setting was different, before-and-after studies rather than randomized trials, but the algebra was the same. The principal strata were the same. The identifying assumptions were the same.</p><p>And the list does not end there. After this essay was first published, <a href="https://x.com/guido_imbens">Guido Imbens noted on Twitter</a> that there were even more independent formulations, including Cuzick, Edwards, and Segnan (1997), who developed the same approach for adjusting for noncompliance and contamination in clinical trials, published in <em>Statistics in Medicine</em>. Imbens also observed that part of the reason for the parallel discoveries was the lack of interaction between the disciplines: "The notion of instrumental variables was not well known in biostatistics, and so the context for the Permutt and Hebel and Baker papers was missing." The value of the Angrist-Imbens-Rubin (1996) JASA paper, he added, was in making the connection between the statistics and econometrics literatures explicit.</p><p>Eight groups across three disciplines, working on different applied problems, all arriving at the same algebra. The chart below lays out the timeline. The top row shows the early formulations, from Baker&#8217;s unpublished 1983 manuscript through Permutt and Hebel in 1989. The bottom row shows the cluster: four independent groups between 1991 and 1994, across biostatistics, epidemiology, and economics, publishing within three years of each other. The color marks the discipline. The density of the bottom row is the point: once the problem was ripe, the solution appeared everywhere at once.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!vj6g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618b50e8-97c0-40eb-a32c-0ac7d5a09d7a_830x524.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!vj6g!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618b50e8-97c0-40eb-a32c-0ac7d5a09d7a_830x524.png 424w, /__u/substackcdn.com/image/fetch/$s_!vj6g!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618b50e8-97c0-40eb-a32c-0ac7d5a09d7a_830x524.png 848w, /__u/substackcdn.com/image/fetch/$s_!vj6g!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618b50e8-97c0-40eb-a32c-0ac7d5a09d7a_830x524.png 1272w, /__u/substackcdn.com/image/fetch/$s_!vj6g!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618b50e8-97c0-40eb-a32c-0ac7d5a09d7a_830x524.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!vj6g!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618b50e8-97c0-40eb-a32c-0ac7d5a09d7a_830x524.png" width="830" height="524" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/618b50e8-97c0-40eb-a32c-0ac7d5a09d7a_830x524.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:524,&quot;width&quot;:830,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:68080,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/191093796?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618b50e8-97c0-40eb-a32c-0ac7d5a09d7a_830x524.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!vj6g!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618b50e8-97c0-40eb-a32c-0ac7d5a09d7a_830x524.png 424w, /__u/substackcdn.com/image/fetch/$s_!vj6g!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618b50e8-97c0-40eb-a32c-0ac7d5a09d7a_830x524.png 848w, /__u/substackcdn.com/image/fetch/$s_!vj6g!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618b50e8-97c0-40eb-a32c-0ac7d5a09d7a_830x524.png 1272w, /__u/substackcdn.com/image/fetch/$s_!vj6g!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F618b50e8-97c0-40eb-a32c-0ac7d5a09d7a_830x524.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Why the Same Idea?</h2><p>The reason is simple.</p><p>The problem that generates LATE arises whenever treatment assignment and treatment receipt diverge. In a clinical trial, some patients decline the offered drug. In an economic setting, some people are offered a program but do not participate. In a public health intervention, some communities adopt a new practice and others do not. The naive comparison of those who received treatment with those who did not is biased by self-selection. The randomized or quasi-randomized assignment provides a way to identify the causal effect for the subpopulation whose behavior was actually changed: the compliers.</p><p>The identification problem is the same in medicine, economics, and epidemiology. The notation differs. The estimation methods differ. Economists use instrumental variables and focus on mean differences. Biostatisticians use maximum likelihood and work with binary outcomes and relative risks. Epidemiologists think in terms of case-control designs. The underlying logic is the same: define the latent classes, impose exclusion and monotonicity, estimate the effect for compliers.</p><p>Farkas Bolyai, the mathematician, wrote about the invention of non-Euclidean geometry: &#8220;When the time is ripe for certain things, these things appear in different places in the manner of violets coming to light in early spring.&#8221; LATE is one of those things.</p><h2>What Travels and What Does Not</h2><p>The economics canon for LATE is Imbens-Angrist (1994) and Angrist-Imbens-Rubin (1996). Some economists know Bloom (1984). But Baker (1983), Permutt and Hebel (1989), and the public health formulations from the early 1990s have not crossed the disciplinary boundary. I had not encountered them before Baker&#8217;s email.</p><p>Why does this matter?</p><p>One reason is intellectual history. The standard narrative in economics is that LATE emerged from the credibility revolution: Angrist&#8217;s work on natural experiments in the late 1980s created the need for a framework to interpret what IV actually estimates when effects are heterogeneous, and Imbens and Angrist provided it. That narrative is true, but it is part of a larger story. The same intellectual need, distinguishing the effect of assignment from the effect of receipt, had already generated the same answer in biostatistics. The credibility revolution gave economics its own reason to formalize the problem, and Imbens and Angrist provided the formalization that the discipline adopted. But the problem, and its solution, were already there.</p><p>Another reason is that the variations across disciplines matter substantively. Economists default to risk differences and mean differences: the treatment raises earnings by $X, or reduces the probability of recidivism by Y percentage points. These are additive effects. The biostatistics tradition works naturally with relative risk: the treatment doubles the probability of survival, or cuts the hazard in half. These are multiplicative effects. When treatment effects are heterogeneous, the additive LATE and the multiplicative LATE weight the complier population differently and can lead to different conclusions about who benefits from treatment. Baker&#8217;s extensions to survival data and missing outcomes address complications that arise constantly in empirical economics but are rarely discussed in the economics LATE literature.</p><h2>Late to the Party</h2><p>Angrist knew about the parallel work. In his Nobel lecture, he cited Bloom&#8217;s 1984 paper and described himself and Imbens as &#8220;late to the partial compliance party.&#8221; Baker borrowed the phrase for his article&#8217;s title.</p><p>Baker and Lindeman, for their part, write: &#8220;Although Karen and I did not get a trip to Stockholm, we had the satisfaction of knowing that the idea we had nurtured for years had a substantial impact on a wide range of fields.&#8221;</p><p>Baker retired from the NIH but continues to publish. He remains interested in the generalizability of LATE, particularly in how plotting the estimated effect against the complier fraction across multiple studies can diagnose treatment effect heterogeneity without estimating the full MTE curve.</p><h2>What This Story Is About</h2><p>Ideas cross disciplinary boundaries slowly, when they cross at all.</p><p>Economics and biostatistics use much of the same technical machinery. Both fields use regression, maximum likelihood, instrumental variables, and potential outcomes. Both care about causal inference. But they publish in different journals, attend different conferences, use different notation, and cite different canonical papers. So the same idea gets discovered more than once.</p><p>The LATE story is one example among many. Propensity score methods were developed by Rosenbaum and Rubin in biostatistics before they were widely adopted in economics. The difference-in-differences design has roots in epidemiology as well as econometrics. The regression discontinuity design was proposed by Thistlethwaite and Campbell in education research in 1960, decades before economists formalized it.</p><p>The boundaries between fields are thinner than our citation practices suggest. The next important idea in causal inference may already exist in a biostatistics journal, or a computer science conference, or an unpublished manuscript from 1983 sitting in a filing cabinet at the Harvard School of Public Health.</p><p>The history of econometrics is full of these rediscoveries. Ideas travel slowly across disciplinary boundaries. Sometimes they arrive decades late. And sometimes, as in the case of LATE, the economists who made the idea famous were themselves late to the party.</p><h2>Where to Start</h2><p>Baker, S. G., and Lindeman, K. S. (2024). &#8220;Multiple Discoveries in Causal Inference: LATE for the Party.&#8221; <em>CHANCE</em> 37(2): 21-25. The article that documents the eight independent formulations. Short, readable, and freely available through PubMed Central. Start here.</p><p>Imbens, G. W., and Angrist, J. D. (1994). &#8220;Identification and Estimation of Local Average Treatment Effects.&#8221; <em>Econometrica</em> 62(2): 467-475. The paper that gave LATE its name and its place in economics. Nine pages that changed what IV means.</p><p>Angrist, J. D., Imbens, G. W., and Rubin, D. B. (1996). &#8220;Identification of Causal Effects Using Instrumental Variables.&#8221; <em>Journal of the American Statistical Association</em> 91(434): 444-472. The full framework: compliers, always-takers, never-takers, defiers. Connects IV to the potential outcomes model.</p><p>Bloom, H. S. (1984). &#8220;Accounting for No-Shows in Experimental Evaluation Designs.&#8221; <em>Evaluation Review</em> 8(2): 225-246. The earliest published formulation of LATE for one-sided noncompliance, from the Harvard School of Government. Derived from first principles, independently of the economics literature.</p><p>Baker, S. G., and Lindeman, K. S. (1994). &#8220;The Paired Availability Design: A Proposal for Evaluating Epidural Analgesia During Labor.&#8221; <em>Statistics in Medicine</em> 13: 2269-2278. Baker and Lindeman&#8217;s formulation for before-and-after studies with different treatment availabilities. The same principal strata, the same exclusion restriction, different notation.</p><p>Permutt, T., and Hebel, J. R. (1989). &#8220;Simultaneous-Equation Estimation in a Clinical Trial of the Effect of Smoking on Birth Weight.&#8221; <em>Biometrics</em> 45: 619-622. An early formulation using recursive equations. Less transparent than the later versions, but the identification is the same.</p><p>Baker, S. G., Kramer, B. S., and Lindeman, K. S. (2016). &#8220;Latent Class Instrumental Variables: A Clinical and Biostatistical Perspective.&#8221; <em>Statistics in Medicine</em> 35: 147-160. A review that bridges the economics and biostatistics LATE literatures. The supplementary material includes Baker&#8217;s unpublished 1983 manuscript.</p><p>Angrist, J. D. (2022). &#8220;Empirical Strategies in Economics: Illuminating the Path from Cause to Effect.&#8221; <em>Econometrica</em> 90: 2509-2539. Angrist&#8217;s Nobel lecture. Contains his acknowledgment of the parallel discoveries and the &#8220;late to the partial compliance party&#8221; line.</p><div><hr></div><p><em>This essay is part of a series on instrumental variables and econometric methods. The most recent essay, &#8220;A Family of Instruments,&#8221; maps the variety of IV designs and the identification questions that separate them. The next essay will examine the Roy model and what it means for how economists think about selection, sorting, and self-selection into treatment.</em></p>]]></content:encoded></item><item><title><![CDATA[The New Age of Social Science]]></title><description><![CDATA[Or: why the future belongs to people who can still think without it]]></description><link>https://carloschavezp29.substack.com/p/the-new-age-of-social-science</link><guid isPermaLink="false">https://carloschavezp29.substack.com/p/the-new-age-of-social-science</guid><dc:creator><![CDATA[Carlos Chavez]]></dc:creator><pubDate>Fri, 13 Mar 2026 13:19:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!bal8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F728140da-a7ca-46b9-a36e-1096981eb082_2000x1429.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In 1766, Joseph Wright of Derby painted <em>A Philosopher Lecturing on the Orrery</em>. A group of men, women, and children gathered in the dark around a mechanical model of the solar system. A lamp sits at the center, where the sun would be, and the light radiates outward through the model&#8217;s orbiting spheres. The people lean in, their faces lit not by the world outside but by the model they built to represent it. No one in the painting is looking at the sky. They are looking at the machine that makes the sky intelligible.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!bal8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F728140da-a7ca-46b9-a36e-1096981eb082_2000x1429.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!bal8!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F728140da-a7ca-46b9-a36e-1096981eb082_2000x1429.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!bal8!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F728140da-a7ca-46b9-a36e-1096981eb082_2000x1429.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!bal8!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F728140da-a7ca-46b9-a36e-1096981eb082_2000x1429.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!bal8!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F728140da-a7ca-46b9-a36e-1096981eb082_2000x1429.jpeg 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!bal8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F728140da-a7ca-46b9-a36e-1096981eb082_2000x1429.jpeg" width="1456" height="1040" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/728140da-a7ca-46b9-a36e-1096981eb082_2000x1429.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1040,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The Orrery&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The Orrery" title="The Orrery" srcset="/__u/substackcdn.com/image/fetch/$s_!bal8!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F728140da-a7ca-46b9-a36e-1096981eb082_2000x1429.jpeg 424w, /__u/substackcdn.com/image/fetch/$s_!bal8!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F728140da-a7ca-46b9-a36e-1096981eb082_2000x1429.jpeg 848w, /__u/substackcdn.com/image/fetch/$s_!bal8!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F728140da-a7ca-46b9-a36e-1096981eb082_2000x1429.jpeg 1272w, /__u/substackcdn.com/image/fetch/$s_!bal8!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F728140da-a7ca-46b9-a36e-1096981eb082_2000x1429.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">No one in the painting is looking at the sky. Joseph Wright of Derby, 1766.</figcaption></figure></div><p>That image captures something about how social science has always worked. We do not observe the economy directly. We build models, systems of equations, identification strategies, thought experiments, and the light comes from them. The quality of the illumination depends on the quality of the model. A bad model casts shadows in the wrong places. A good one lets you see things you could not see before.</p><p>The question this essay asks is what happens when the machine gets an upgrade.</p><p>Something is changing in the way we do research. A study published in <em>Economics Letters</em> this year found that the use of LLM-characteristic terminology in leading economics journals nearly doubled between 2023 and 2024 (Feyzollahi and Rafizadeh, 2025). Doubled. In one year. The profession has not had time to think about what this means.</p><p>This essay is not about whether AI is good or bad for social science. That framing is useless. AI is already inside the research process, in the coding, the literature reviews, the drafting, the data cleaning, the hypothesis generation. The better question is what kind of discipline social science becomes as these tools move from supplementary to constitutive. To answer that, it helps to know what social science was before.</p><h2>Two Revolutions</h2><p>The history of empirical social science in the twentieth century can be organized around two methodological revolutions. Each one redefined what it meant for research to be credible.</p><p>The first revolution began in the 1940s, with the work of Trygve Haavelmo, the Cowles Commission, and the generation of economists who believed that the path to knowledge ran through structural models. The ambition was enormous: write down a complete model of the economy, a system of simultaneous equations describing supply, demand, production, and investment, then use statistical methods to recover the parameters of that system from observed data. Credibility, in this world, meant theoretical coherence. A paper was persuasive if its model was internally consistent, if the identifying restrictions were derived from economic theory, and if the estimation procedure was appropriate for the model&#8217;s structure.</p><p>This tradition produced some of the most important ideas in social science. The concept of identification, the question of whether, given the data and the assumptions, we can uniquely recover the object of interest, was formalized in this era. So were the tools of instrumental variables, maximum likelihood estimation, and the entire apparatus of structural econometrics. Heckman&#8217;s selection model, the discrete choice models of McFadden, and the dynamic programming estimators of Rust all descend from this lineage.</p><p>But the structural era had a problem that became increasingly difficult to ignore. The models were elaborate, the assumptions were strong, and the results were often sensitive to specifications that the reader could not easily evaluate. By the 1980s, Edward Leamer had already raised the alarm. His 1983 paper, &#8220;Let&#8217;s Take the Con out of Econometrics,&#8221; argued that applied econometrics was a form of persuasion in which researchers had too many degrees of freedom and too little discipline about reporting how their choices affected results. The critique was devastating, and the profession heard it.</p><p>The second revolution was the response. Beginning in the late 1980s and accelerating through the 1990s and 2000s, a new generation of empirical researchers, led by figures like Joshua Angrist, David Card, Alan Krueger, and Guido Imbens, argued that credibility came not from the elegance of the model but from the strength of the research design. The question shifted from &#8220;Is the model correctly specified?&#8221; to &#8220;Is the identifying variation plausibly exogenous?&#8221; Natural experiments, regression discontinuities, difference-in-differences, and instrumental variables grounded in institutional features of the world rather than in theoretical restrictions became the gold standard.</p><p>This was the credibility revolution, and it worked. It produced a body of empirical knowledge in labor economics, public finance, development, and education that is among the most reliable in the social sciences. It also narrowed the scope of what economists were willing to study. If you could not find a clean source of exogenous variation, you could not write the paper. The set of answerable questions shrank, even as the quality of the answers improved.</p><p>These two revolutions were never fully reconciled. The structuralists worried that the design-based researchers were identifying parameters that were too local, too specific, and too disconnected from the models needed for policy evaluation. The experimentalists worried that the structuralists were recovering parameters that depended on assumptions no one could verify. Heckman (2010) made the case that the credibility revolution had thrown out too much. Angrist and Pischke (2010) made the case that it had not thrown out enough. The debate continues.</p><p>But both sides agreed on one thing: the fundamental constraint of social science was not computation. It was not data. It was the quality of the argument, the identification strategy, the thought experiment, the logic that connected observed variation to the causal question. Computation was a means. The argument was the end.</p><p>That consensus is now being tested.</p><h2>What AI Actually Changes</h2><p>The most visible thing AI changes is the production function of research. Tasks that used to take weeks, cleaning a large dataset, coding a complex estimator, reviewing a literature, formatting results, can now be done in hours or minutes. In a profession where the bottleneck to publishing was often the sheer labor of implementation, collapsing that bottleneck changes who can do research, how fast projects move, how many papers get written, and how quickly ideas diffuse.</p><p>But this is also the least interesting thing AI changes, because it is purely quantitative. Faster implementation of the same research process produces more of the same research. The question is whether AI changes the process itself, whether it creates possibilities that did not exist before, or destroys constraints that held the discipline together.</p><p>I think it does both. Three changes, taken together, amount to something qualitatively different from what came before.</p><h3>New Research Objects</h3><p>For most of its history, empirical social science studied human behavior by observing humans: in surveys, in experiments, in administrative data, in the field. AI introduces a different kind of object altogether. John Horton, at MIT, coined the term <em>homo silicus</em> to describe the possibility of using large language models as simulated economic agents (Horton, 2023). The idea is that because LLMs are trained on vast corpora of human-generated text, they internalize, in some statistical sense, the patterns of human reasoning, including economic reasoning. You can give an LLM an endowment, a set of preferences, and a choice problem, and observe its behavior. You can run experiments on it that would be expensive, slow, or ethically impossible to run on humans.</p><p>Horton and his co-authors replicated classic experiments from behavioral economics, Charness and Rabin (2002), Kahneman, Knetsch, and Thaler (1986), Samuelson and Zeckhauser (1988), using GPT as the subject, and found qualitatively similar results. Manning, Zhu, and Horton (2024) went further, building a system that automatically generates and tests social scientific hypotheses using LLMs as both the scientist and the subject.</p><p>Whether <em>homo silicus</em> is a reliable proxy for <em>homo sapiens</em> is an open and contested question. The LLM has no body, no survival instinct, no lived experience. Its &#8220;preferences&#8221; are a reflection of its training data, not of any actual utility function. But the methodological point stands: social science now has access to a class of subjects that can be studied at near-zero marginal cost, at arbitrary scale, under conditions of total experimental control. This changes the feasibility frontier of the discipline.</p>
      <p>
          <a href="/__u/carloschavezp29.substack.com/p/the-new-age-of-social-science">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[A Family of Instruments]]></title><description><![CDATA[What Economists Found, Constructed, and Argued About]]></description><link>https://carloschavezp29.substack.com/p/a-family-of-instruments</link><guid isPermaLink="false">https://carloschavezp29.substack.com/p/a-family-of-instruments</guid><dc:creator><![CDATA[Carlos Chavez]]></dc:creator><pubDate>Sat, 07 Mar 2026 22:56:34 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!chy6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3b80eac-8dc0-4788-8db1-2275078cc093_760x430.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><a href="/__u/substack.com/home/post/p-189502356">The previous essay in this series traced how the two-stage architecture</a>, estimate something in the first stage, use it in the second, became the dominant mode of empirical economics. That essay focused on the structure. This one focuses on the object that makes the structure work: the instrument.</p><p>Economists have found instruments in draft lotteries, birth dates, distances to colleges, colonial mortality rates, the leniency of judges, and the industrial composition of cities. They have derived instruments from structural models and discovered them in natural experiments. They have constructed instruments from the interaction of local exposure and national shocks. These instruments look nothing alike. They come from different fields, invoke different arguments, and identify different parameters. But they all do the same job: they provide variation in a treatment or endogenous variable that is, the researcher argues, unrelated to the outcome except through the treatment itself.</p><p>This essay maps the family. The organizing principle is where the exclusion restriction gets its credibility. Structural instruments derive their credibility from the economic model: the theory says a variable is excluded from one equation, and this exclusion identifies the parameter. Natural experiment instruments derive credibility from the design of the world: a lottery, a policy discontinuity, a historical accident generated as-if-random variation. Constructed instruments, like the Bartik shift-share, derive credibility from the exogeneity of their components, the shares or the shocks, and researchers can disagree about which component does the identifying work. Judge and examiner designs derive credibility from quasi-random assignment of cases to decision-makers, but their interpretation hinges on a behavioral assumption, monotonicity, that is harder to verify than in other designs. Chart 1 summarizes this taxonomy and its relationship to the two identification frameworks, LATE and MTE, that give the family its deepest structure.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!0dJp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf30732-0a2a-4769-9b92-46aeef02dee8_692x511.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!0dJp!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf30732-0a2a-4769-9b92-46aeef02dee8_692x511.png 424w, /__u/substackcdn.com/image/fetch/$s_!0dJp!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf30732-0a2a-4769-9b92-46aeef02dee8_692x511.png 848w, /__u/substackcdn.com/image/fetch/$s_!0dJp!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf30732-0a2a-4769-9b92-46aeef02dee8_692x511.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0dJp!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf30732-0a2a-4769-9b92-46aeef02dee8_692x511.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!0dJp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf30732-0a2a-4769-9b92-46aeef02dee8_692x511.png" width="692" height="511" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9bf30732-0a2a-4769-9b92-46aeef02dee8_692x511.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:511,&quot;width&quot;:692,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:59649,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/190232216?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf30732-0a2a-4769-9b92-46aeef02dee8_692x511.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!0dJp!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf30732-0a2a-4769-9b92-46aeef02dee8_692x511.png 424w, /__u/substackcdn.com/image/fetch/$s_!0dJp!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf30732-0a2a-4769-9b92-46aeef02dee8_692x511.png 848w, /__u/substackcdn.com/image/fetch/$s_!0dJp!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf30732-0a2a-4769-9b92-46aeef02dee8_692x511.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0dJp!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bf30732-0a2a-4769-9b92-46aeef02dee8_692x511.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The essay asks two questions. First, where do instruments come from? The answer turns out to be historically contingent, a story about what the profession valued at different points in time and what kinds of arguments it found persuasive. Second, what do instruments identify? The answer, formalized by Imbens and Angrist in 1994 and deepened by Heckman and Vytlacil in 2005, is that different instruments identify different things, even when applied to the same question. The variety of instruments is not just a catalog of clever research designs. It is a window into the deepest unresolved question in applied econometrics: what parameter are we estimating, and for whom?</p><p>Chart 2 traces this evolution through the seminal papers, from Wright&#8217;s first instrument in 1928 to the modern formalization of shift-share designs. The timeline is also a map of what the profession learned to worry about: first, where instruments come from; then, what they identify; then, how they can fail.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!chy6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3b80eac-8dc0-4788-8db1-2275078cc093_760x430.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!chy6!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3b80eac-8dc0-4788-8db1-2275078cc093_760x430.png 424w, /__u/substackcdn.com/image/fetch/$s_!chy6!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3b80eac-8dc0-4788-8db1-2275078cc093_760x430.png 848w, /__u/substackcdn.com/image/fetch/$s_!chy6!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3b80eac-8dc0-4788-8db1-2275078cc093_760x430.png 1272w, /__u/substackcdn.com/image/fetch/$s_!chy6!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3b80eac-8dc0-4788-8db1-2275078cc093_760x430.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!chy6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3b80eac-8dc0-4788-8db1-2275078cc093_760x430.png" width="760" height="430" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f3b80eac-8dc0-4788-8db1-2275078cc093_760x430.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:430,&quot;width&quot;:760,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:58385,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/190232216?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3b80eac-8dc0-4788-8db1-2275078cc093_760x430.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!chy6!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3b80eac-8dc0-4788-8db1-2275078cc093_760x430.png 424w, /__u/substackcdn.com/image/fetch/$s_!chy6!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3b80eac-8dc0-4788-8db1-2275078cc093_760x430.png 848w, /__u/substackcdn.com/image/fetch/$s_!chy6!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3b80eac-8dc0-4788-8db1-2275078cc093_760x430.png 1272w, /__u/substackcdn.com/image/fetch/$s_!chy6!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3b80eac-8dc0-4788-8db1-2275078cc093_760x430.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h2>The Structural Instruments</h2><p><a href="/__u/substack.com/@carloschavezp29/p-188895762">The first instruments were derived, not discovered. Philip Wright</a>, in a 1928 appendix to a book about tariffs on animal and vegetable oils, showed how to separate supply and demand curves using variables that shift one curve but not the other. Demand shifters, variables that affect how much consumers want to buy but do not directly affect production costs, can instrument for price in a supply equation. Supply shifters do the reverse. The logic was algebraic: if you can write down the system of equations and identify which variables are excluded from which equations, you can solve for the structural parameters.</p><p>The Cowles Commission formalized this in the 1940s and 1950s. Haavelmo, Koopmans, and their colleagues at the University of Chicago developed the theory of identification for systems of simultaneous equations. The key contribution was the rank and order conditions: a structural equation is identified if and only if the excluded instruments provide enough independent variation. Theil formalized two-stage least squares in 1953, giving applied researchers a practical way to implement these ideas. The instrument was whatever the economic model said was excluded.</p><p>These were structural instruments in a precise sense. They came from theory. You wrote down a model of the labor market, or the market for butter, or the macroeconomy, and the model told you what could serve as an instrument. The exclusion restriction was defended internally, as a maintained implication of the model, rather than externally via institutional design. If the model was right, the instrument was valid.</p><p>The strength of this approach was its coherence. The weakness was its dependence on the model being correct. If the theory excluded a variable from one equation but the world did not, the instrument was invalid, and nothing in the data would tell you.</p><h2>The Natural Experiments</h2><p>The credibility revolution changed the source of the argument. Instead of deriving instruments from models, researchers began looking for instruments in the world, in institutional rules, historical accidents, and randomized or quasi-randomized variation that nature or policy had produced. The shift was not in the econometrics. The algebra of IV estimation was the same. What changed was the basis for believing the exclusion restriction.</p><p>Three papers from the early 1990s illustrate the shift, and each reveals a different kind of natural experiment.</p><p><strong>The draft lottery.</strong> In 1990, Joshua Angrist published a paper in the <em>American Economic Review</em> that used the Vietnam-era draft lottery as an instrument for military service. The question was whether veterans earned less than nonveterans. The problem was selection: men who served in Vietnam differed from men who did not in ways that also affected their earnings. Some volunteered, some avoided service, and the decision to serve or not was correlated with unobservable characteristics like risk tolerance, health, and ability.</p><p>The draft lottery provided the instrument. Between 1970 and 1973, priority for military service was randomly assigned based on birth date. Each day of the year was drawn in a public lottery and assigned a rank. Men with low lottery numbers were called first and were far more likely to be drafted. In 1970, the highest random sequence number reached was 195: men with numbers below that threshold were draft-eligible. The assignment was random, a physical lottery, and this randomness is what makes it an instrument. A man&#8217;s lottery number was uncorrelated with his ability, ambition, health, or anything else that might affect his earnings, because the number was determined by the order in which dates were drawn from a container. The exclusion restriction holds not because an economic model says so but because of the physical mechanism of randomization.</p><p>Angrist used Social Security earnings records to estimate that white veterans earned approximately 15 percent less than comparable nonveterans in the early 1980s. The identifying variation is transparent: compare the earnings of men who, by the luck of the draw, faced a high probability of being drafted with the earnings of men who faced a low probability, and scale the earnings difference by the difference in service rates. This is the Wald estimator, the ratio of the reduced form to the first stage, and it is the simplest form of IV.</p><p>The paper was a template. It showed that a natural experiment, an event that generated as-if-random variation in a treatment, could substitute for the structural model as the source of instrument validity. The argument for the exclusion restriction shifted from &#8220;the model excludes this variable from this equation&#8221; to &#8220;the lottery was random.&#8221; The source of credibility moved from theory to design.</p><p><strong>Quarter of birth.</strong> In 1991, Angrist and Alan Krueger published a paper in the <em>Quarterly Journal of Economics</em> that used quarter of birth as an instrument for years of schooling. The logic was subtle. Compulsory schooling laws require students to stay in school until a certain age, typically sixteen. Because of school entry rules, children born earlier in the year start school at an older age and can therefore drop out legally after completing fewer years of schooling than children born later in the year. Quarter of birth affects schooling through this institutional channel but, the authors argued, does not directly affect earnings.</p><p>The paper produced IV estimates of the return to schooling that were close to the OLS estimates, suggesting little ability bias in the standard estimates. It was widely influential. But it also became a cautionary tale. In 1995, John Bound, David Jaeger, and Regina Baker published a paper in the <em>Journal of the American Statistical Association</em> showing that the quarter-of-birth instruments were extremely weak. With 180 instruments (quarters interacted with year and state of birth), the correlation between the instruments and schooling was tiny. In finite samples, 2SLS with many weak instruments is biased toward OLS, and the apparent agreement between IV and OLS may have been an artifact of that bias rather than evidence of no ability bias.</p><p>The Bound-Jaeger-Baker critique did not kill the quarter-of-birth design, but it forced the profession to take weak instruments seriously. The first-stage F-statistic became a standard diagnostic. Stock and Yogo (2005) provided critical values. The rule of thumb that F should exceed 10, crude but useful, entered every empirical researcher&#8217;s toolkit. The episode demonstrated that finding a natural experiment was not enough. The instrument also had to be strong.</p><p><strong>College proximity.</strong> In 1995, David Card published a chapter using proximity to a four-year college as an instrument for schooling. Men who grew up near a college completed more education and earned more than men who grew up farther away. Card argued that living near a college reduces the cost of attending, increasing the probability of enrollment, but does not directly affect earnings conditional on education.</p><p>The instrument produced IV estimates of the return to schooling that were 25 to 60 percent higher than OLS, the opposite of what you would expect if the main bias in OLS were upward (due to omitted ability). Card interpreted this as evidence that the returns to schooling are highest for men from disadvantaged backgrounds, who are the ones most affected by the cost of college. The compliers in this design, the men whose schooling was changed by proximity to a college, had higher returns than the average person in the sample.</p><p>This interpretation anticipated, by a year, the formal result that Imbens and Angrist would name the local average treatment effect.</p><p><strong>Colonial mortality.</strong> The other natural experiments described above operated through causal chains that were short and mechanical: a lottery assigns a number, the number determines draft eligibility, eligibility affects service. The most ambitious natural experiment instrument stretched the causal chain across centuries. In 2001, Daron Acemoglu, Simon Johnson, and James Robinson published a paper in the <em>American Economic Review</em> that used European settler mortality rates in colonies as an instrument for contemporary institutions. The causal chain was long: in colonies where European settlers faced high mortality (from malaria, yellow fever, and other tropical diseases), the colonizers established extractive institutions designed to transfer resources to the metropole. In colonies where settlers survived, they established inclusive institutions with property rights protections and checks on executive power. These institutional differences persisted for centuries. Settler mortality from two hundred years ago, the authors argued, affects income today only through its effect on institutional quality.</p><p>The paper demonstrated that institutions have a large causal effect on income per capita. It also demonstrated how far the natural experiment approach could stretch. The exclusion restriction requires that settler mortality rates from the eighteenth and nineteenth centuries have no direct effect on GDP per capita today except through institutions. This is a much harder argument to make than &#8220;the lottery was random.&#8221; The instrument has been intensely debated. Critics have questioned the mortality data, the exclusion restriction, the measurement of institutions, and the functional form. Albouy (2012) challenged the historical mortality estimates directly. The debate has not been fully resolved.</p><p>The Acemoglu-Johnson-Robinson instrument sits at the boundary of the natural experiment approach. The variation is historical and arguably exogenous to the individual country&#8217;s current economic performance, but the causal chain is long and mediated by many factors. It illustrates both the power and the vulnerability of the instrument: power because it addresses one of the largest questions in economics (why are some countries rich and others poor?), vulnerability because the exclusion restriction stretches across centuries and continents.</p><h2>The Constructed Instruments</h2><p>Not all instruments are found in nature. Some are constructed.</p><p>The Bartik instrument, named after Timothy Bartik&#8217;s 1991 book on state and local economic development, is the most prominent example. Bartik instruments are also called shift-share instruments, and the name describes the construction. Take a set of local areas. Each area has a pre-existing industrial composition, the &#8220;shares.&#8221; National employment growth varies by industry, the &#8220;shifts.&#8221; The Bartik instrument for a given area is the predicted employment growth obtained by weighting the national industry growth rates by the area&#8217;s initial industry shares. The idea is that national industry trends are exogenous to any individual local area, and the local area&#8217;s exposure to those trends depends on its pre-existing industrial structure.</p><p>Bartik instruments have been used hundreds of times, in labor economics, in trade, in public finance, in urban economics. The canonical modern application is Autor, Dorn, and Hanson&#8217;s (2013) study of the effects of Chinese import competition on U.S. labor markets, where the shift-share instrument combines local industry employment shares with changes in Chinese exports to other high-income countries.</p><p>For decades, the instrument was used without a formal econometric framework. That changed with three papers published between 2019 and 2022. Goldsmith-Pinkham, Sorkin, and Swift (2020) showed in the <em>American Economic Review</em> that the Bartik instrument is numerically equivalent to using the local industry shares as multiple instruments. This means identification comes from the exogeneity of the shares: the assumption that a local area&#8217;s initial industrial composition is uncorrelated with unobserved determinants of the outcome, conditional on controls. Borusyak, Hull, and Jaravel (2022) showed in the <em>Review of Economic Studies</em> that the same instrument can be analyzed from the opposite direction: identification comes from the exogeneity of the shocks, the national industry growth rates, while the shares are allowed to be endogenous. Ad&#227;o, Koles&#225;r, and Morales (2019) showed in the <em>Quarterly Journal of Economics</em> that standard errors in shift-share regressions are wrong because the residuals are correlated across regions that share similar industrial structures, and they proposed a correction.</p><p>The three frameworks agree on the mechanics of the instrument but disagree on what makes it valid. If you follow Goldsmith-Pinkham, Sorkin, and Swift, you need to defend the exogeneity of the shares. If you follow Borusyak, Hull, and Jaravel, you need to defend the exogeneity of the shocks. Two papers can use the same shift-share instrument and disagree about its validity because they are defending different primitives. In practice, these are different arguments with different empirical implications, and they lead to different diagnostic tests. The shares-based view asks: are the Rotemberg weights concentrated on a few industries whose shares might be endogenous? The shocks-based view asks: are the industry-level shocks uncorrelated with industry-level confounders?</p><p>The Bartik instrument is interesting for this essay because it is neither a structural instrument in the Cowles Commission sense (it does not come from writing down a system of equations) nor a natural experiment in the Angrist sense (there is no single randomized or quasi-randomized event). It is a hybrid: a constructed object that combines pre-existing structure with external variation. The exclusion restriction is argued differently depending on which framework you adopt, and reasonable researchers can disagree about which argument is more convincing. Borusyak, Hull, and Jaravel (2025) synthesized the entire debate into a practical guide in the <em>Journal of Economic Perspectives</em>, with checklists for both the shares-based and shocks-based approaches.</p><p>Other constructed instruments follow similar logic. Hausman instruments, introduced in the context of demand estimation, instrument for the price of a good in one market with the prices of the same good in other markets. The reasoning is that a common cost shock (say, a change in input prices) affects prices in all markets simultaneously, while demand shocks are local. If I want to estimate the demand curve for cereal in Chicago, I can use the price of the same cereal in Los Angeles as an instrument, because the L.A. price reflects cost variation that also drives the Chicago price but is unrelated to Chicago-specific demand conditions. The assumption is that after controlling for observables, the only thing that makes cereal prices move together across cities is the common cost component. Nevo (2001) used this logic in his influential study of the ready-to-eat cereal industry, and variants appear throughout the industrial organization literature on demand estimation.</p><p>Gravity-based instruments in trade use predicted bilateral trade flows, based on geographic distance and country size, to instrument for actual trade volumes. The identifying assumption is that geography affects trade through trade costs but does not directly affect the outcome of interest (such as income or growth) except through trade itself. Frankel and Romer (1999) used this approach to estimate the effect of trade on income, though the exclusion restriction, that geography does not affect income through any channel other than trade, has been questioned.</p><p>In each case, the instrument is not found in nature but assembled from components, and the argument for validity depends on assumptions about which components are exogenous. The construction gives the researcher flexibility but also introduces ambiguity: the more the instrument is constructed rather than discovered, the more the exclusion restriction depends on modeling choices rather than institutional features of the world.</p><h2>The Judge and Examiner Designs</h2><p>A particularly productive branch of the family uses variation in the assignment of decision-makers as instruments. The idea is simple: if cases (defendants, disability applicants, foster children) are randomly or quasi-randomly assigned to judges or examiners who differ in their tendencies, then the identity of the assigned decision-maker is an instrument for the decision itself.</p><p>The canonical application is incarceration. Jeffrey Kling published a paper in the <em>American Economic Review</em> in 2006 showing that longer incarceration reduces post-release employment and earnings. He exploited the random assignment of federal cases to judges who varied in their sentencing severity. A defendant assigned to a harsh judge receives a longer sentence than a defendant assigned to a lenient judge, and this variation is independent of the defendant&#8217;s characteristics (conditional on the randomization scheme). The judge&#8217;s average sentencing behavior, typically measured as a leave-out mean, serves as the instrument.</p><p>Joseph Doyle applied the same logic to child protection in 2007, using the random assignment of child abuse investigators who varied in their propensity to place children in foster care. Nicole Maestas, Kathleen Mullen, and Alexander Strand used it in 2013 for Social Security disability insurance, exploiting the random assignment of disability examiners who differed in their allowance rates. The design has since been applied to bail decisions, immigration courts, patent examiners, and many other settings where cases are quasi-randomly assigned to decision-makers.</p><p>The judge design is attractive because the source of the instrument, random assignment to decision-makers, is institutional and often well-documented. But it has a specific vulnerability. The monotonicity assumption requires that the instrument pushes all affected individuals in the same direction. In judge designs, this can fail if judges have crossing leniency rankings across case types: judge A is more lenient than judge B for drug offenses but harsher for violent offenses. If leniency rankings cross, the standard LATE interpretation breaks down, and the IV estimand is a weighted average of treatment effects with some negative weights, which need not correspond to any identifiable group. The recent literature has proposed tests for monotonicity and alternative assumptions (such as unordered monotonicity) that weaken the requirement. Goldsmith-Pinkham, Hull, and Koles&#225;r (2026) provide a comprehensive practical guide to these designs, covering identification, monotonicity testing, and implementation. But the judge design illustrates a general principle: the cleaner the source of randomization, the more demanding the behavioral assumptions needed to give the estimate a sharp interpretation.</p><p>The taxonomy so far has been about where instruments come from: from models, from nature, from construction, from institutional assignment. The next question is what they buy you: which causal object the IV estimand equals.</p><h2>What the Instrument Identifies</h2><p>For forty years after Wright, the answer seemed obvious. If the model is correctly specified and the instrument is valid, IV estimates the structural parameter in the equation. The return to schooling. The elasticity of labor supply. The effect of incarceration on recidivism. The parameter was assumed to be the same for everyone, or at least the same on average, and the instrument&#8217;s job was to provide consistent estimation.</p><p>In 1994, Guido Imbens and Joshua Angrist published a nine-page paper in <em>Econometrica</em> that changed this answer. They showed that when treatment effects are heterogeneous, meaning that different people benefit differently from treatment, IV does not estimate the average treatment effect for the population. Under an additional assumption they called monotonicity, IV estimates the local average treatment effect: the average treatment effect for compliers, the subpopulation whose treatment status is changed by the instrument.</p><p>The monotonicity assumption says that the instrument moves everyone in the same direction. If the instrument is a draft lottery, monotonicity says that no one who would have volunteered without a low lottery number refused to serve because of the lottery. If the instrument is college proximity, monotonicity says that no one who would have attended college far away was deterred from attending because a college was nearby. In most applications, this is plausible. But it is an assumption, and it matters because without it the IV estimand is not a well-defined average of treatment effects.</p><p>The LATE result was profound. It meant that different instruments, applied to the same question, estimate different parameters. The draft lottery identifies the effect of military service for men who were drafted and served but would not have volunteered. Quarter of birth identifies the effect of schooling for students who stayed in school only because compulsory schooling laws forced them to. College proximity identifies the effect of schooling for students who attended college only because there was one nearby. These are different subpopulations. They may have very different treatment effects.</p><p>Card&#8217;s college proximity instrument produced IV estimates of the return to schooling that were substantially higher than OLS. Angrist and Krueger&#8217;s quarter-of-birth instrument produced estimates that were roughly similar to OLS. Both are consistent with the LATE interpretation. The college proximity compliers are disadvantaged students with high marginal returns to schooling. The quarter-of-birth compliers are students at the low end of the schooling distribution, who may have more moderate returns.</p><p>The Imbens-Angrist result was part of the work for which they shared the 2021 Nobel Prize in Economics with David Card. It formalized something that the profession had not confronted directly: the instrument is not just a tool for estimation. It defines the estimand. The choice of instrument is a choice about whose treatment effect you learn.</p><p>This has practical implications that the profession is still absorbing. When two instruments give different IV estimates for the same treatment, the standard reaction used to be: one of them is wrong. The LATE framework says: both might be right. They are estimating effects for different complier populations. The Hausman test for overidentification, which tests whether instruments give statistically similar estimates, is often interpreted as a test of validity. But if treatment effects are heterogeneous, the test rejects even when all instruments are valid. The rejection tells you about heterogeneity, not invalidity.</p><p>Distinguishing between invalidity and heterogeneity requires economic reasoning, not statistical tests. It requires thinking about who the compliers are and why they might differ.</p><h2>The MTE Unification</h2><p><a href="/__u/substack.com/@carloschavezp29/p-181470966">James Heckman saw the LATE result as a symptom of a deeper problem and an invitation to solve it.</a></p><p>In a 1997 paper in the <em>Journal of Human Resources</em> titled &#8220;Instrumental Variables: A Study of Implicit Behavioral Assumptions Used in Making Program Evaluations,&#8221; Heckman argued that the LATE framework conceals important assumptions about economic behavior and limits the policy relevance of IV estimates. The compliers are defined by the instrument, not by the policy question. If a policymaker wants to know the effect of expanding college access to a new population, the relevant treatment effect is the effect for the marginal entrants under the proposed policy, who may be entirely different from the compliers identified by any particular instrument.</p><p>In 2005, Heckman and Edward Vytlacil published a seventy-page paper in *Econometrica* that provided the unifying framework. The marginal treatment effect, MTE(x,uD)MTE(x, u_D) MTE(x,uD&#8203;), is defined as the treatment effect for an individual with observable characteristics xx x and unobserved resistance to treatment uDu_D uD&#8203;. It is the effect for the person who is just indifferent between treatment and control at a given level of unobserved distaste for treatment.</p><p>The MTE curve is the fundamental object. Every treatment effect parameter, ATE, ATT, LATE, or any policy-relevant treatment effect, can be expressed as a weighted average of the MTE curve. The weights differ. The ATE puts equal weight on everyone. The ATT puts more weight on people who chose treatment. The LATE, for a particular instrument, puts weight on the part of the MTE curve that the instrument&#8217;s variation illuminates, the compliers for that instrument.</p><p>This means that different instruments do not estimate &#8220;different effects&#8221; in some vague sense. They weight different parts of the same underlying curve. The draft lottery traces out one segment. College proximity traces out another. A Bartik instrument traces out yet another. If the MTE curve is flat, all instruments agree and LATE equals ATE. If the curve slopes, they disagree, and the disagreement is informative about the shape of the underlying heterogeneity.</p><p>Return to the schooling example to see this concretely. Card&#8217;s college proximity instrument produces IV estimates of the return to schooling that are substantially higher than OLS. Angrist and Krueger&#8217;s quarter-of-birth instrument produces estimates close to OLS. In the MTE framework, this is not contradictory. College proximity moves students from disadvantaged backgrounds into college, students at the high end of the MTE curve who face high costs and have high returns. Quarter of birth moves students near the dropout margin, students at a different point on the curve. If the MTE curve slopes downward, students who are hardest to push into treatment (high uDu_D uD&#8203;, high resistance) have the highest returns, both results are consistent with the same underlying curve. The instruments are not giving &#8220;conflicting&#8221; estimates. They are illuminating different parts of the same object.</p><p>Chart 3 makes this visible. The horizontal axis is unobserved resistance to treatment: people on the left select into treatment easily, people on the right resist. If treatment gains vary systematically with that resistance, the MTE curve slopes, and each instrument illuminates a different segment. In canonical applications like schooling, the curve slopes downward: proximity-based instruments, which shift cost-sensitive students, weight the left side of the curve where returns are highest. Lottery-based instruments weight a broader middle range. Exposure-based instruments like Bartik shift-share weight yet another segment. Each instrument&#8217;s LATE is the average height of the MTE curve over its complier region. The location of these regions is application-specific, not universal, but the principle holds generally: if the curve has shape, instruments will disagree, and the disagreement is the information.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!U5TH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1f92049-9415-4a0d-b3c4-643f50527a56_704x675.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!U5TH!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1f92049-9415-4a0d-b3c4-643f50527a56_704x675.png 424w, /__u/substackcdn.com/image/fetch/$s_!U5TH!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1f92049-9415-4a0d-b3c4-643f50527a56_704x675.png 848w, /__u/substackcdn.com/image/fetch/$s_!U5TH!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1f92049-9415-4a0d-b3c4-643f50527a56_704x675.png 1272w, /__u/substackcdn.com/image/fetch/$s_!U5TH!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1f92049-9415-4a0d-b3c4-643f50527a56_704x675.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!U5TH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1f92049-9415-4a0d-b3c4-643f50527a56_704x675.png" width="704" height="675" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c1f92049-9415-4a0d-b3c4-643f50527a56_704x675.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:675,&quot;width&quot;:704,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:90772,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/190232216?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1f92049-9415-4a0d-b3c4-643f50527a56_704x675.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!U5TH!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1f92049-9415-4a0d-b3c4-643f50527a56_704x675.png 424w, /__u/substackcdn.com/image/fetch/$s_!U5TH!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1f92049-9415-4a0d-b3c4-643f50527a56_704x675.png 848w, /__u/substackcdn.com/image/fetch/$s_!U5TH!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1f92049-9415-4a0d-b3c4-643f50527a56_704x675.png 1272w, /__u/substackcdn.com/image/fetch/$s_!U5TH!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1f92049-9415-4a0d-b3c4-643f50527a56_704x675.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The MTE framework does not eliminate the instrument choice problem. It reframes it. Instead of asking &#8220;which instrument gives the right answer?&#8221; you ask &#8220;which part of the MTE curve does this instrument illuminate, and is that the part I care about?&#8221; For policy evaluation, the answer depends on the proposed policy. A policy that makes college free weights different margins than a policy that builds new colleges in rural areas. The instruments that are informative for one policy may be uninformative for the other.</p><p>Heckman and Vytlacil showed that under certain conditions, the full MTE curve can be nonparametrically identified from the data, provided you have a continuously distributed instrument with sufficient support. When this is possible, the researcher can compute any treatment effect parameter by integrating the MTE curve with the appropriate weights. When the instrument&#8217;s support is limited, identification is partial, and the MTE curve can be traced only over the range where the instrument has bite.</p><p>The MTE framework is more demanding than the LATE approach. It requires stronger assumptions, a continuous instrument with large support, and a monotone selection model. It requires specifying more about the model. But it answers a wider class of questions. The LATE approach is honest about what it identifies but limited in what it can say about policy. The MTE approach asks for more but delivers more. (I wrote at length about the MTE framework, its identification, and its policy implications in an earlier essay, &#8220;<a href="/__u/substack.com/@carloschavezp29/p-181470966">Why Don&#8217;t We Talk Enough About Marginal Treatment Effects?</a>&#8220;)</p><h2>What Makes an Instrument Good or Bad</h2><p>The family portrait is incomplete without a discussion of pathology. What goes wrong with instruments, and how do you tell?</p><p>Three conditions must hold for an instrument to be valid: relevance, exclusion, and (for the LATE interpretation) monotonicity. Of these, only relevance is directly testable.</p><p><strong>Relevance.</strong> The instrument must be correlated with the endogenous variable. The first-stage F-statistic is the standard diagnostic. Staiger and Stock (1997) showed that when the instruments are weak (low F), the 2SLS estimator is biased toward OLS, confidence intervals have incorrect coverage, and tests have wrong size. The rule of thumb that F should exceed 10 comes from Stock and Yogo&#8217;s (2005) tabulations, which give critical values for the degree of bias acceptable to the researcher. More recently, Montiel Olea and Pflueger (2013) provided a robust version of the effective F-statistic that is valid with heteroscedasticity and clustering, and Lee, McCrary, Moreira, and Porter developed weak-IV-robust inference methods that remain valid regardless of instrument strength.</p><p><strong>Exclusion.</strong> The instrument must affect the outcome only through the endogenous variable. This is the condition that cannot be tested. The overidentification test, the Hansen or Sargan J-test, tests whether the instruments are mutually consistent but cannot detect a common violation. If all your instruments are invalid in the same direction, the J-test will not pick it up.</p><p>The exclusion restriction is always an argument. For the draft lottery, the argument is that lottery numbers affect earnings only through military service. For quarter of birth, the argument is that being born in January rather than April affects earnings only through schooling. For college proximity, the argument is that growing up near a college affects earnings only through education. Each of these arguments has been challenged. The quality of an IV paper depends on how well the researcher makes the case, and reasonable people can disagree.</p><p>Conley, Hansen, and Rossi (2012) proposed a framework for &#8220;plausibly exogenous&#8221; instruments that relaxes the exclusion restriction. Instead of assuming that the instrument has zero direct effect on the outcome, they allow a small direct effect and derive bounds on the treatment effect that are robust to a specified degree of violation. This sensitivity analysis has become increasingly common in applied work. It does not rescue a bad instrument, but it allows the researcher to show how much the results depend on the exclusion restriction being exactly right.</p><p><strong>Monotonicity.</strong> The instrument must move all affected individuals in the same direction. This matters for the LATE interpretation. It is partially testable: you can check whether the instrument&#8217;s first-stage effect is positive for every observable subgroup, though this does not rule out violations within subgroups. In judge designs, where the instrument is a decision-maker&#8217;s average tendency, violations of monotonicity are a particular concern because different judges may rank cases differently.</p><h2>The Unresolved Tension</h2><p>The family of instruments is united by a logic and divided by a question.</p><p>The logic is this: find variation in treatment that is independent of confounders. Use that variation to identify a causal effect. The exclusion restriction is the lynchpin. If it holds, the instrument works. If it does not, nothing else can save the estimate.</p><p>The question is this: what does the instrument estimate, and is that the right thing to estimate?</p><p>The design-based tradition, associated with Angrist and Imbens, answers: the instrument estimates the LATE for compliers. Be honest about what you have identified. Do not extrapolate beyond what the data support. If you want the ATE, find an instrument that moves everyone, or acknowledge that you have identified something more local.</p><p>The structural tradition, associated with Heckman, answers: the instrument illuminates one part of the MTE curve. If you want to answer a policy question, you need the right part of the curve, or ideally the whole curve. Identify the MTE, integrate with the right weights, and compute the policy-relevant treatment effect. The LATE is a byproduct, not the target.</p><p>These positions are not logically incompatible. They are different responses to the same tradeoff between credibility and ambition. The design-based approach prioritizes transparency and defensibility: tell the reader exactly what you have identified and let them judge whether it is useful. The structural approach prioritizes policy relevance: build a framework that can answer the question the policymaker actually asks, even if this requires stronger assumptions.</p><p>The tension plays out in every applied paper that uses instruments. When Angrist uses the draft lottery and reports a LATE, he is being transparent about his estimand but silent about what would happen if the policy changed. When Heckman estimates an MTE curve and computes a policy-relevant treatment effect, he is answering the policy question but relying on assumptions that are harder to verify.</p><p>The 2021 Nobel Prize went to both sides of this debate, Card and Angrist and Imbens, for developing natural experiments and the LATE framework. Heckman had received the Nobel in 2000 for the selection correction and the structural approach. The prizes bracket the tension. They do not resolve it.</p><p>The profession has not converged on one position. In practice, most applied papers in the design-based tradition report LATEs and acknowledge the limitation. A growing number also discuss what their estimates imply about heterogeneity and policy relevance. The best papers in the structural tradition use MTE methods when the data support them, typically when there is a continuous instrument with substantial variation, and fall back on LATE when they do not. The two traditions are converging operationally even if they remain philosophically distinct.</p><p>No other social science uses instruments the way economics does. Political scientists and epidemiologists have adopted the language and some of the methods, but the instrument remains the economist&#8217;s most distinctive empirical tool. The family keeps growing: recent years have brought instruments based on peer effects, network structures, algorithmic recommendations, and machine-learning-predicted treatment propensities. Each new member inherits the same two burdens. The exclusion restriction is always a claim, never a proof. And the estimand depends on the instrument, because different instruments move different people into treatment, and people differ in their responses. The choice of instrument is a choice about whose experience counts.</p><p>Whether you work in the design-based or the structural tradition, the argument for the instrument&#8217;s validity is where credibility lives. Everything else, the choice of estimator, the treatment of heterogeneity, the inference procedure, is downstream. The family of instruments is the family of arguments that economists make about causation. Understanding its history, its internal disagreements, and the formal results that structure those disagreements is the fastest route to understanding what empirical economics can and cannot claim to know.</p><p>The next essay turns to the model that underlies much of this debate: the Roy model. It is the model of self-selection, the formal account of why people sort into treatment differently, and it connects directly to the MTE framework that gives the family of instruments its deepest structure.</p><h2>Where to Start</h2><p>Angrist, J. D. (1990). &#8220;Lifetime Earnings and the Vietnam Era Draft Lottery: Evidence from Social Security Administrative Records.&#8221; <em>American Economic Review</em> 80(3): 313-336. The paper that made the natural experiment the standard source of instruments. The draft lottery as an instrument for veteran status.</p><p>Imbens, G. W., and Angrist, J. D. (1994). &#8220;Identification and Estimation of Local Average Treatment Effects.&#8221; <em>Econometrica</em> 62(2): 467-475. Nine pages that changed what IV means. With heterogeneous effects and monotonicity, IV estimates the effect for compliers, not for everyone.</p><p>Heckman, J. J., and Vytlacil, E. (2005). &#8220;Structural Equations, Treatment Effects, and Econometric Policy Evaluation.&#8221; <em>Econometrica</em> 73(3): 669-738. The MTE unification. Every treatment parameter, ATE, ATT, LATE, is a weighted average of the same marginal treatment effect curve. Different instruments weight different parts.</p><p>Angrist, J. D., and Krueger, A. B. (1991). &#8220;Does Compulsory School Attendance Affect Schooling and Earnings?&#8221; <em>Quarterly Journal of Economics</em> 106(4): 979-1014. Quarter of birth as instrument for schooling. Influential and cautionary: the weak instruments problem emerged from this paper.</p><p>Bound, J., Jaeger, D. A., and Baker, R. M. (1995). &#8220;Problems with Instrumental Variables Estimation When the Correlation Between the Instruments and the Endogenous Explanatory Variable Is Weak.&#8221; <em>Journal of the American Statistical Association</em> 90(430): 443-450. The critique that forced the profession to take weak instruments seriously.</p><p>Goldsmith-Pinkham, P., Sorkin, I., and Swift, H. (2020). &#8220;Bartik Instruments: What, When, Why, and How.&#8221; <em>American Economic Review</em> 110(8): 2586-2624. Identification in Bartik designs comes from the exogeneity of the shares. A formal framework for the most widely used constructed instrument.</p><p>Borusyak, K., Hull, P., and Jaravel, X. (2022). &#8220;Quasi-Experimental Shift-Share Research Designs.&#8221; <em>Review of Economic Studies</em> 89(1): 181-213. The alternative: identification comes from the exogeneity of the shocks. Together with Goldsmith-Pinkham et al., this paper defines the modern understanding of shift-share instruments.</p><h2>Some Additional Readings</h2><p><strong>On natural experiments and the credibility revolution</strong></p><p>Card, D. (1995). &#8220;Using Geographic Variation in College Proximity to Estimate the Return to Schooling.&#8221; In <em>Aspects of Labour Market Behaviour: Essays in Honour of John Vanderkamp</em>, edited by L. N. Christofides, E. K. Grant, and R. Swidinsky, 201-222. Toronto: University of Toronto Press. College proximity as an instrument for schooling. The IV estimates are higher than OLS, suggesting that compliers, disadvantaged students, have above-average returns.</p><p>Acemoglu, D., Johnson, S., and Robinson, J. A. (2001). &#8220;The Colonial Origins of Comparative Development: An Empirical Investigation.&#8221; <em>American Economic Review</em> 91(5): 1369-1401. European settler mortality as an instrument for contemporary institutions. The most ambitious, and most debated, natural experiment instrument in development economics.</p><p>Angrist, J. D., Imbens, G. W., and Rubin, D. B. (1996). &#8220;Identification of Causal Effects Using Instrumental Variables.&#8221; <em>Journal of the American Statistical Association</em> 91(434): 444-472. The extended treatment of the LATE framework, with the taxonomy of compliers, always-takers, never-takers, and defiers. The paper that connected IV to the potential outcomes framework.</p><p><strong>On shift-share instruments</strong></p><p>Bartik, T. J. (1991). <em>Who Benefits from State and Local Economic Development Policies?</em> Kalamazoo, MI: W.E. Upjohn Institute for Employment Research. The book that introduced the instrument that bears his name, though the formal analysis came decades later.</p><p>Ad&#227;o, R., Koles&#225;r, M., and Morales, E. (2019). &#8220;Shift-Share Designs: Theory and Inference.&#8221; <em>Quarterly Journal of Economics</em> 134(4): 1949-2010. Standard errors in shift-share regressions are wrong. The correction accounts for cross-regional correlation induced by common industry shocks.</p><p>Borusyak, K., Hull, P., and Jaravel, X. (2025). &#8220;A Practical Guide to Shift-Share Instruments.&#8221; <em>Journal of Economic Perspectives</em> 39(1): 181-204. The state-of-the-art practitioner&#8217;s guide. Synthesizes the shares-vs-shocks debate into simple checklists for applied researchers. The paper to read first if you want to use a Bartik instrument correctly.</p><p><strong>On judge and examiner designs</strong></p><p>Kling, J. R. (2006). &#8220;Incarceration Length, Employment, and Earnings.&#8221; <em>American Economic Review</em> 96(3): 863-876. The canonical judge leniency design. Random assignment to federal judges instruments for sentence length.</p><p>Maestas, N., Mullen, K. J., and Strand, A. (2013). &#8220;Does Disability Insurance Receipt Discourage Work? Using Examiner Assignment to Estimate Causal Effects of SSDI Receipt.&#8221; <em>American Economic Review</em> 103(5): 1797-1829. Examiner leniency as an instrument for disability insurance receipt. Found that SSDI receipt reduces labor force participation by about 28 percentage points.</p><p>Goldsmith-Pinkham, P., Hull, P., and Koles&#225;r, M. (2026). &#8220;Leniency Designs: An Operator&#8217;s Manual.&#8221; Forthcoming, <em>Journal of Economic Perspectives</em>. The modern guide to judge and examiner designs. Covers identification, monotonicity testing, and practical implementation. The companion to the BHJ practical guide, but for assignment-based instruments.</p><p><strong>On weak instruments and diagnostics</strong></p><p>Staiger, D., and Stock, J. H. (1997). &#8220;Instrumental Variables Regression with Weak Instruments.&#8221; <em>Econometrica</em> 65(3): 557-586. The paper that formalized the weak instruments problem. When instruments are weak, IV is biased, and inference is unreliable.</p><p>Conley, T. G., Hansen, C. B., and Rossi, P. E. (2012). &#8220;Plausibly Exogenous.&#8221; <em>Review of Economics and Statistics</em> 94(1): 260-272. What if the exclusion restriction is only approximately true? A framework for sensitivity analysis that bounds the treatment effect under small violations.</p><p><strong>On the Heckman-Angrist debate</strong></p><p>Heckman, J. J. (1997). &#8220;Instrumental Variables: A Study of Implicit Behavioral Assumptions Used in Making Program Evaluations.&#8221; <em>Journal of Human Resources</em> 32(3): 441-462. Heckman&#8217;s critique of the LATE framework: the compliers are defined by the instrument, not by the policy question, and LATE may be uninformative about the effects of the policies that matter.</p><div><hr></div><p><em>This is the third essay in a series on instrumental variables and econometric methods. The first essay covered the history of instrumental variables from Wright (1928) to the credibility revolution. The second traced the two-stage architecture from Heckman&#8217;s selection correction to double machine learning. The next essay will examine the Roy model and what it means for how economists think about selection, sorting, and self-selection into treatment.</em></p>]]></content:encoded></item><item><title><![CDATA[Thought Experiments in the Age of AI]]></title><description><![CDATA[Economists and the Super Research Assistant]]></description><link>https://carloschavezp29.substack.com/p/thought-experiments-in-the-age-of</link><guid isPermaLink="false">https://carloschavezp29.substack.com/p/thought-experiments-in-the-age-of</guid><dc:creator><![CDATA[Carlos Chavez]]></dc:creator><pubDate>Wed, 04 Mar 2026 18:46:19 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!zh0V!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffff9b4a5-53ee-4548-85da-563173f13249_2666x2666.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Economists are adopting large language models at unprecedented speed.  LLM-characteristic terminology in leading economics journals has nearly doubled between 2023 and 2024 (Feyzollahi and Rafizadeh, 2025). Systems like ChatGPT and Claude are now used for coding, literature review, hypothesis generation, robustness checks, and drafting. Dell (2025) surveys the expanding toolkit. Korinek (2023) catalogs use cases.</p><p>This essay is about economists using LLMs as partners in the upstream work of causal research, posing questions, defining interventions, and choosing identification strategies. I set aside downstream uses like coding, text classification, and flexible estimation inside researcher-specified models. The subject is the economist who sits down with Claude or ChatGPT to do the intellectual work that precedes estimation.</p><p>Large language models drastically reduce the cost of implementation in empirical research. They do not reduce the cost of specifying the causal thought experiment that defines the parameter being estimated (yet). This asymmetry creates a substitution risk: researchers may adopt AI-generated causal narratives without constructing the institutional and structural understanding that makes those narratives scientifically meaningful. In that relationship, thought experiments, the foundation of causal inference since Haavelmo (1943, 1944) and recently formalized by Heckman and Pinto (2024), can be complementary to the AI, or they can be replaced by it.</p><p>The thesis of this essay is that substitution is the default. Complementarity is possible, but it requires discipline that the tools themselves do not impose.</p><h2>Haavelmo&#8217;s Distinction</h2><p>In 1943, Trygve Haavelmo published &#8220;The Statistical Implications of a System of Simultaneous Equations&#8221; in <em>Econometrica</em> 11(1): 1&#8211;12. The following year, he published &#8220;The Probability Approach in Econometrics&#8221; as a supplement to <em>Econometrica</em>, volume 12. Together, these papers drew a line between two things that look similar but are not the same: the data we observe and the hypothetical manipulations we imagine when we ask causal questions.</p><p>To ask &#8220;what would happen to output if we increased investment by 10%?&#8221; is to conduct a thought experiment. You fix some inputs, vary others, and trace the consequences through a model. The answer depends on the model, on what you chose to hold fixed, and on the structural relationships you posited. It does not fall out of the data by itself. The data can tell you what happened. It cannot tell you what would have happened under a different arrangement of the world, not without a model that specifies how the arrangement matters.</p><p>Ragnar Frisch, Haavelmo&#8217;s teacher at Oslo, had put the point bluntly: causality is in the mind. The causal parameter lives in the researcher&#8217;s hypothetical model, not in the statistical properties of the observed variables. The data can help you estimate the parameter, but the data does not define it. The definition comes from the intervention: the specification of what is fixed, what adjusts, and why. Here a &#8220;thought experiment&#8221; means an intervention embedded in a hypothetical model, what is fixed, what is allowed to adjust, and which structural relationships remain in force.</p><p>Heckman and Pinto (2024) formalize this distinction in modern notation. Their paper, &#8220;Econometric Causality: The Central Role of Thought Experiments&#8221; in the <em>Journal of Econometrics</em>, distinguishes two models. The empirical model is the data-generating process that produces the observations you collect. The hypothetical model is the mental construct you use to define what a causal effect means. The hypothetical model shares the same structural equations and error distributions as the empirical model, but it replaces the actual mechanism determining some variable with an external, researcher-controlled manipulation. This replacement is the &#8220;fix&#8221; operator, the formalization of Marshall&#8217;s <em>ceteris paribus</em>.</p><p>Take supply and demand. Prices and quantities are determined simultaneously. A regression of quantity on price recovers neither the demand curve nor the supply curve. To identify demand, you fix price externally, hold the demand shifters constant, and trace quantity demanded. That is the thought experiment.</p><p>Conditioning on P = p in the data and looking at Q is not the same operation as imagining that P is set to p by fiat, replacing the supply equation with an external rule, and tracing Q through the demand equation alone. The first is statistics. The second is the thought experiment that defines a causal effect. Heckman and Pinto&#8217;s fix operator formalizes the second: it replaces one variable&#8217;s natural determination with an external assignment while leaving the rest of the system intact.</p><p>The fix operator is not a statistical procedure. It is a claim about the economic system, about which equations govern behavior, which variables adjust, and what it means to intervene on one of them. Specifying the fix requires understanding the system. That requirement is the point.</p>
      <p>
          <a href="/__u/carloschavezp29.substack.com/p/thought-experiments-in-the-age-of">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[The Two-Stage Idea]]></title><description><![CDATA[How One Architecture Solved Half of Econometrics]]></description><link>https://carloschavezp29.substack.com/p/the-two-stage-idea</link><guid isPermaLink="false">https://carloschavezp29.substack.com/p/the-two-stage-idea</guid><dc:creator><![CDATA[Carlos Chavez]]></dc:creator><pubDate>Sun, 01 Mar 2026 20:58:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!zwv4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fa29a64-554b-45c3-96d6-80d39e51dc99_829x663.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><a href="/__u/substack.com/@carloschavezp29/p-188895762">The previous essay in this series traced the histor<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>y of instrumental variables from Philip Wright&#8217;s 1928 appendix to the credibility revolution of the 1990s.</a> That essay ended with a claim: the two-stage logic Henri Theil formalized in 1953: use a first stage to remove the blockage, then feed its output into a second stage to recover what you actually care about. That logic is far more general than instrumental variables alone. This essay makes good on the claim.</p><p>The same architecture underlies the Heckman correction for selection bias, the control function approach, propensity score methods, and a range of estimators that break hard problems into tractable pieces. These methods look different on the surface. They solve different problems, invoke different assumptions, and show up in different literatures. But they share a common structure, and understanding that structure is the fastest way to see what each method actually does, what it assumes, and where it can go wrong. If you&#8217;ve ever estimated something you don&#8217;t ultimately care about, just to make the main estimate credible, you were doing Theil&#8217;s move.</p><p>The structure is this. You face a problem you cannot solve in one step because something you need, a quantity, a distribution, a counterfactual, is not directly available. So you estimate it in a first stage. Then you take that estimate and use it in a second stage to recover the parameter you care about. The first stage handles the part of the problem that blocks direct estimation. The second stage does the rest.</p><p>Econometrics has a hundred names for this move. Underneath, it is always the same architecture:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!zwv4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fa29a64-554b-45c3-96d6-80d39e51dc99_829x663.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!zwv4!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fa29a64-554b-45c3-96d6-80d39e51dc99_829x663.png 424w, /__u/substackcdn.com/image/fetch/$s_!zwv4!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fa29a64-554b-45c3-96d6-80d39e51dc99_829x663.png 848w, /__u/substackcdn.com/image/fetch/$s_!zwv4!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fa29a64-554b-45c3-96d6-80d39e51dc99_829x663.png 1272w, /__u/substackcdn.com/image/fetch/$s_!zwv4!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fa29a64-554b-45c3-96d6-80d39e51dc99_829x663.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!zwv4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fa29a64-554b-45c3-96d6-80d39e51dc99_829x663.png" width="829" height="663" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8fa29a64-554b-45c3-96d6-80d39e51dc99_829x663.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:663,&quot;width&quot;:829,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:128485,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/189502356?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fa29a64-554b-45c3-96d6-80d39e51dc99_829x663.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!zwv4!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fa29a64-554b-45c3-96d6-80d39e51dc99_829x663.png 424w, /__u/substackcdn.com/image/fetch/$s_!zwv4!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fa29a64-554b-45c3-96d6-80d39e51dc99_829x663.png 848w, /__u/substackcdn.com/image/fetch/$s_!zwv4!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fa29a64-554b-45c3-96d6-80d39e51dc99_829x663.png 1272w, /__u/substackcdn.com/image/fetch/$s_!zwv4!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fa29a64-554b-45c3-96d6-80d39e51dc99_829x663.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The differences across methods are in what fills the <em>g</em> slot and why. The rest of this essay fills in the blanks.</p><h2>Heckman&#8217;s Problem</h2><p>In 1974, James Heckman was studying married women&#8217;s labor supply. The question was standard: how do wages affect hours worked? The data problem was not. You observe wages only for women who are employed. Women who choose not to work have no observed wage. A regression of hours on wages, estimated on the sample of workers, ignores everyone who opted out, and the women who opted out are not a random subset of the population. They are the women for whom the market wage fell below their reservation wage. The sample is selected on the dependent variable, and the selection is driven by unobservables.</p><p>Gronau (1974) and H. Gregg Lewis (1974) had recognized the problem. Both published papers in the same issue of the <em>Journal of Political Economy</em> pointing out that wage comparisons across groups are contaminated by selectivity: the composition of who works differs across groups, and this compositional difference biases estimated wage gaps. The theoretical roots ran even deeper. Roy (1951) had shown that when individuals select into occupations based on comparative advantage, the observed distribution of earnings in any sector is a truncated version of the population distribution. The Roy model&#8217;s influence extends far beyond selection corrections, into trade, migration, inequality, and the foundations of treatment effect heterogeneity, and it merits an essay of its own in this series. But recognizing a problem and solving it are different things. Gronau, Lewis, and Roy described the bias. Heckman built the fix.</p><p>The fix came in stages. In a 1974 paper in <em>Econometrica</em>, Heckman formalized the selection model for female labor supply and developed an econometric method for the problem. In a 1976 paper in the <em>Annals of Economic and Social Measurement</em>, he showed that truncation, sample selection, and limited dependent variables all share a common statistical structure. Then, in 1979, in a paper that would eventually earn him a share of the Nobel Prize, he published the result that made everything operational.</p><p>The paper was &#8220;Sample Selection Bias as a Specification Error,&#8221; published in <em>Econometrica</em> 47(1): 153-161. The NBER working paper version had circulated since 1977. The core insight is in the title: selection bias is an omitted variable problem. Selection changes E[&#949;&#7522; | D&#7522; = 1] from zero to a nonzero function of Z&#7522; through the selection rule. That function is the omitted variable. If you can estimate it and include it in your regression, the bias disappears.</p><h2>The Inverse Mills Ratio</h2><p>Here is how the argument works. Suppose you want to estimate a wage equation:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;% Eq 1: Wage equation\n \\displaystyle Y_i = X_i \\beta + \\varepsilon_i &quot;,&quot;id&quot;:&quot;YEZAZKZNBP&quot;}" data-component-name="LatexBlockToDOM"></div><p>but Y&#7522; is observed only when a woman works. Whether she works depends on a selection equation:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\displaystyle D_i^* = Z_i \\gamma + u_i&quot;,&quot;id&quot;:&quot;ZHREIZZTHH&quot;}" data-component-name="LatexBlockToDOM"></div><p>where D&#7522; = 1 if D&#7522;* &gt; 0 (she works) and D&#7522; = 0 otherwise. The error terms &#949;&#7522; and u&#7522; are jointly normal and potentially correlated. That correlation is the source of the problem. If u&#7522; and &#949;&#7522; are correlated, then conditioning on D&#7522; = 1 (selecting workers) distorts the distribution of &#949;&#7522; in the selected sample, and OLS on the wage equation is biased.</p><p>Heckman showed that the conditional expectation of the wage, given that the woman works, is:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\displaystyle E[Y_i \\mid X_i,\\; D_i = 1] = X_i \\beta + \\rho \\sigma_\\varepsilon \\,\\lambda(Z_i \\gamma)&quot;,&quot;id&quot;:&quot;UAUXTJXUPS&quot;}" data-component-name="LatexBlockToDOM"></div><p>where &#955;(&#183;) = &#966;(&#183;)/&#934;(&#183;) is the inverse Mills ratio, the ratio of the standard normal density to the cumulative distribution function evaluated at Z&#7522;&#947;. The term &#961;&#963;_&#949; captures the strength and direction of the correlation between the selection error and the outcome error. If &#961; = 0, there is no selection bias and &#955; drops out. If &#961; &#8800; 0, omitting &#955; is a specification error, exactly like omitting any other relevant variable.</p><p>The two-stage procedure follows immediately. In the first stage, estimate a probit of D&#7522; on Z&#7522; to get &#947;&#770;, and compute &#955;&#770;&#7522; = &#966;(Z&#7522;&#947;&#770;)/&#934;(Z&#7522;&#947;&#770;) for each observation. In the second stage, regress Y&#7522; on X&#7522; and &#955;&#770;&#7522; using OLS on the selected sample. The coefficient on &#955;&#770;&#7522; estimates &#961;&#963;_&#949;, and including it as a regressor purges the selection bias from &#946;&#770;.</p><p>This is the Heckman correction. Applied researchers know it as the Heckit model. It has been used thousands of times, in labor economics, in health economics, in development, in industrial organization, in political science. It was a central part of the work for which Heckman shared the Nobel Prize in 2000 with Daniel McFadden. And it is, at its core, a two-stage method: estimate the selection mechanism in the first stage, correct for it in the second.</p><h2>What Makes the Heckman Correction Work, and What Can Break It</h2><p>Three assumptions drive the result. First, joint normality of the errors. The inverse Mills ratio is the correct control function only under bivariate normality. If the errors are not normal, the functional form of the correction term changes, and the two-step estimator is inconsistent. This is a strong assumption, and the method&#8217;s sensitivity to it was recognized early. Goldberger (1983) showed in a paper titled &#8220;Abnormal Selection Bias&#8221; that misspecifying the distribution can make the Heckman correction perform worse than ignoring selection altogether.</p><p>Second, the method works best with an exclusion restriction in the selection equation. Formally, you need at least one variable in Z&#7522; that does not appear in X&#7522;: a variable that affects whether a woman works but does not directly affect her wage. Children, nonlabor income, and spousal earnings are commonly used. Without an exclusion restriction, identification comes entirely from the nonlinearity of the inverse Mills ratio, which is functional form identification, and applied researchers are right to distrust it. The distinction matters: with a good exclusion restriction, the method identifies the selection effect from genuine variation in the probability of working. Without one, it identifies it from the curvature of the normal distribution, which is much more fragile. When identification comes only from nonlinearity, small misspecification can impersonate &#8220;selection.&#8221;</p><p>Third, the standard errors from the second stage are wrong. This is the generated regressors problem, which Pagan (1984) analyzed in general and Murphy and Topel (1985) addressed with a correction for two-step estimators. The issue is simple. The second stage treats &#955;&#770;&#7522; as if it were known, ignoring the fact that it was estimated in the first stage. This understates the standard errors, sometimes by a lot. The fix is either to use the Murphy-Topel correction or to bootstrap the entire two-stage procedure. Applied papers that report naive second-stage standard errors without correction are overstating their precision, and reviewers have become increasingly strict about this.</p><p>Despite these limitations, the Heckman correction represented a genuine breakthrough. Before 1979, applied researchers faced a hard choice: ignore selection and produce biased estimates, or attempt maximum likelihood estimation of the full model, which was computationally expensive and often numerically unstable with the computing resources of the time. Heckman&#8217;s two-step estimator offered a third option: a simple, computationally cheap procedure that could be implemented with two OLS regressions and a probit. The price was efficiency. The two-step estimator is less efficient than full MLE, but the gain in accessibility was enormous. This tradeoff between efficiency and usability is the defining feature of two-stage methods, and it is the same tradeoff that made 2SLS dominant over LIML and full-information methods in the instrumental variables literature.</p><h2>The Control Function</h2><p>The Heckman correction is a special case of something more general. The inverse Mills ratio is a particular function, specific to the normal distribution, that you include as a regressor to absorb the correlation between the selection error and the outcome error. But the underlying logic does not require normality. It does not even require a selection problem. The logic is: if something in the error term of your equation is correlated with a regressor, find a way to estimate that problematic component and include it as a control. Once you condition on it, the remaining error is clean, and OLS recovers the parameter of interest.</p><p>This is the control function approach. The term was formalized by Heckman and Robb (1985) in a paper on methods for evaluating the impact of interventions, though the principle traces back to Heckman&#8217;s earlier work. The basic setup is familiar from instrumental variables. You have a structural equation Y = &#946;X + &#949; where X is endogenous, and a first-stage equation X = Z&#960; + V where Z is a valid instrument. The endogeneity of X comes from the fact that V and &#949; are correlated. If they were not correlated, X would be exogenous and OLS would be fine.</p><p>The control function approach exploits the fact that once you know V, the correlation between X and &#949; is captured by the correlation between V and &#949;. So if you include V as an additional regressor, you are conditioning on the source of the endogeneity, and the coefficient on X in the augmented regression is consistent:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot; \\displaystyle Y = \\beta X + \\delta V + \\eta \n&quot;,&quot;id&quot;:&quot;PVDLFUCKAX&quot;}" data-component-name="LatexBlockToDOM"></div><p>where &#951; is now uncorrelated with X by construction. The coefficient &#948; captures the degree of endogeneity. If &#948; = 0, X was exogenous all along and the correction was unnecessary, which also gives you a direct test for endogeneity: just test whether &#948;&#770; is significantly different from zero.</p><p>In practice, V is not observed, so you estimate it. Run OLS of X on Z to get residuals V&#770;&#7522; = X&#7522; &#8722; Z&#7522;&#960;&#770;, then include V&#770;&#7522; in the second-stage regression. This is sometimes called two-stage residual inclusion, to distinguish it from two-stage least squares, which is two-stage residual substitution (replacing X with X&#770;). In linear IV, the OLS residual is the right control. In nonlinear models the control is typically a function of the residual, or a generalized residual from a nonlinear first stage, and that is where Rivers-Vuong and Blundell-Smith live.</p><p>In the linear model, the two procedures give identical results. If you replace X with X&#770; (2SLS) or if you include V&#770; alongside X (control function), you get the same estimate of &#946;. The algebra guarantees this because X = X&#770; + V&#770;, so including one and substituting the other are just reparameterizations.</p><p>But in nonlinear models, they diverge, and the control function approach is usually preferred. Rivers and Vuong (1988) showed how to apply the control function to probit models with endogenous regressors. Blundell and Smith (1986) extended it to Tobit models. Wooldridge (2015) provided a comprehensive modern survey. The advantage of the control function in nonlinear settings is that it works within the structural model rather than replacing the endogenous variable with a linear projection. In a probit, for example, the nonlinearity of the link function means that plugging in X&#770; does not have the same interpretation as it does in the linear case. Including V&#770; as an additional regressor, by contrast, directly addresses the source of the endogeneity and preserves the structural interpretation of the model.</p><p>The connection to Heckman&#8217;s selection model is now visible. The inverse Mills ratio &#955;&#770;&#7522; is a control function. It captures the part of the outcome error that is correlated with the selection decision, and including it as a regressor removes that correlation. The specific functional form (the ratio &#966;/&#934;) comes from the assumption of normality, but the logic, estimate the problematic component and control for it, is the same logic that underlies the Rivers-Vuong correction, the Blundell-Smith extension, and every other control function estimator. What differs across applications is the form of the correction term and the distributional assumptions that justify it.</p><h2>The Propensity Score</h2><p>In 1983, Paul Rosenbaum and Donald Rubin published a paper in <em>Biometrika</em> that solved a different problem with the same two-stage architecture. The problem was confounding in observational studies. If treatment assignment is not random, treated and untreated units may differ in observable characteristics, and a naive comparison of their outcomes confounds the treatment effect with these pre-existing differences. The solution, in principle, is to compare treated and untreated units that are similar on all relevant covariates. The difficulty is dimensionality. If you have twenty covariates, matching on all of them simultaneously is practically impossible. The data are too sparse in twenty-dimensional space.</p><p>Rosenbaum and Rubin&#8217;s theorem states that you do not need to match on all twenty covariates. You only need to match on a single number: the propensity score, defined as e(x) = P(D = 1 | X = x), the conditional probability of treatment given the covariates. They proved that if treatment assignment is strongly ignorable given X, meaning that potential outcomes are independent of treatment conditional on X, then they are also independent of treatment conditional on the propensity score alone. The propensity score is a sufficient statistic for confounding. All the information in a high-dimensional covariate vector that matters for confounding is compressed into a single scalar.</p><p>The paper appeared in <em>Biometrika</em> 70(1): 41-55, under the title &#8220;The Central Role of the Propensity Score in Observational Studies for Causal Effects.&#8221; Rosenbaum had completed his PhD under Rubin at Harvard just three years earlier. Rubin was at Harvard. The paper proposed three applications: matched sampling on the propensity score, subclassification on the propensity score, and covariance adjustment using the propensity score. All three are implemented as two-stage procedures. In the first stage, estimate the propensity score, typically with a logit or probit. In the second stage, use the estimated propensity score to match, stratify, or reweight the data and estimate the treatment effect.</p><p>The identifying assumption is strong ignorability, which has two parts. First, unconfoundedness: conditional on the covariates, treatment assignment is independent of potential outcomes. Second, overlap: every unit has a positive probability of receiving either treatment status, so 0 &lt; e(x) &lt; 1 for all x. Unconfoundedness is the propensity score analog of the exclusion restriction in IV. Both are untestable. Both require the researcher to argue that the variation used for identification is clean. The difference is in the source of the argument. In IV, you need a variable that affects treatment but not the outcome directly. In propensity score methods, you need a set of covariates that, once conditioned on, make treatment assignment as good as random. The IV researcher needs an excluded instrument. The propensity score researcher needs a complete set of confounders.</p><p>Neither assumption is testable. The quality of the design depends on the argument, not the estimator.</p><h2>LaLonde, and What Propensity Scores Fixed</h2><p>The canonical demonstration of propensity score methods came from a destructive test. In 1986, Robert LaLonde published a paper in the <em>American Economic Review</em> that asked a devastating question: if we have the experimental answer, can econometric methods recover it from observational data? The experiment was the National Supported Work Demonstration, which randomly assigned applicants to a job training program. LaLonde took the experimental treatment effect, then replaced the randomized control group with non-experimental comparison groups drawn from the Panel Study of Income Dynamics and the Current Population Survey. He applied the standard econometric methods of the time, including regression adjustment and Heckman selection models, and found that they largely failed. Different comparison groups and different econometric specifications produced wildly different estimates, many of which were far from the experimental benchmark.</p><p>The paper was a warning shot. It suggested that non-experimental methods were unreliable, that the econometric corrections available in the mid-1980s could not substitute for randomization. The result was widely cited and deeply influential.</p><p>Then, in 1999, Rajeev Dehejia and Sadek Wahba revisited LaLonde&#8217;s data using propensity score matching. Their paper, published in the <em>Journal of the American Statistical Association</em>, showed that propensity score methods, applied carefully, came much closer to the experimental benchmark than the methods LaLonde had used. The key was that propensity score matching reduced the comparison to treated and control units with similar probabilities of treatment, which trimmed the most problematic comparison observations, the ones that were so different from the treated group that no amount of regression adjustment could make them comparable.</p><p>The Dehejia-Wahba result was not the final word. Smith and Todd (2005) raised questions about the sensitivity of propensity score estimates to specification choices. The debate continues. But the LaLonde-Dehejia-Wahba sequence established propensity score methods as a serious tool for observational causal inference and, just as critically, established the principle that non-experimental methods should be evaluated against experimental benchmarks whenever possible.</p><h2>Inverse Probability Weighting</h2><p>Propensity scores can be used for more than matching. One of the most powerful applications is inverse probability weighting (IPW). The idea is to reweight the observed sample so that it looks like a random sample. If a treated unit with covariates x has a probability e(x) of being treated, then it is overrepresented relative to a random sample by a factor of 1/e(x). Weighting treated observations by 1/e(x) and control observations by 1/(1 &#8722; e(x)) creates a pseudo-population in which treatment assignment is independent of the covariates in expectation, conditional on the propensity model and overlap.</p><p>The estimator of the average treatment effect under IPW is:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\displaystyle \\hat{\\tau}_{IPW} = \\frac{1}{N} \\sum_{i=1}^{N} \\left[ \\frac{D_i \\, Y_i}{\\hat{e}(X_i)} - \\frac{(1 - D_i) \\, Y_i}{1 - \\hat{e}(X_i)} \\right]&quot;,&quot;id&quot;:&quot;EPDTUNLITI&quot;}" data-component-name="LatexBlockToDOM"></div><p>This is again a two-stage method. The first stage estimates e(x). The second stage computes the weighted average. The method traces back to Horvitz and Thompson (1952) in the survey sampling literature, but its modern application in causal inference owes primarily to work by James Robins and collaborators in biostatistics.</p><p>IPW has a known vulnerability: when e(x) is close to zero or one, the weights explode and the estimator becomes unstable. Applied researchers recognize this pathology as variance blow-up. This is the practical expression of the overlap assumption failing. Units with extreme propensity scores contribute disproportionately to the estimate, and a single observation with a very small weight in the denominator can dominate everything. Trimming extreme propensity scores or using stabilized weights can help, but the instability is a genuine limitation.</p><h2>Doubly Robust Methods</h2><p>A natural question arises: what if the propensity score model is wrong? What if the outcome model is wrong? Doubly robust methods, developed by Robins, Rotnitzky, and Zhao (1994), address this by combining both models. The estimator includes an outcome regression (predicting Y from X for each treatment group) and a propensity score weighting. The remarkable property is that the estimator is consistent if either the outcome model or the propensity score model is correctly specified. You get two chances to be right.</p><p>The augmented inverse probability weighted (AIPW) estimator augments the IPW estimator with a regression adjustment:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\displaystyle \\hat{\\tau}_{AIPW} = \\frac{1}{N} \\sum_{i=1}^{N} \\left[ \\hat{\\mu}_1(X_i) - \\hat{\\mu}_0(X_i) + \\frac{D_i \\bigl(Y_i - \\hat{\\mu}_1(X_i)\\bigr)}{\\hat{e}(X_i)} - \\frac{(1 - D_i)\\bigl(Y_i - \\hat{\\mu}_0(X_i)\\bigr)}{1 - \\hat{e}(X_i)} \\right]&quot;,&quot;id&quot;:&quot;ESECXLXLGW&quot;}" data-component-name="LatexBlockToDOM"></div><p>where &#956;&#770;&#8321; and &#956;&#770;&#8320; are the estimated conditional means of the outcome under treatment and control. The regression adjustment handles the case where the propensity score is wrong, and the weighting handles the case where the outcome model is wrong. If both are right, the estimator is more efficient than either alone.</p><p>This two-stage-with-insurance architecture has become standard in modern program evaluation. It extends naturally to settings where machine learning is used in the first stage, which brings us to the current frontier.</p><h2>Machine Learning Enters the First Stage</h2><p>In 2018, Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins published a paper in <em>The Econometrics Journal</em> that brought the two-stage idea into the machine learning era. The paper, &#8220;Double/Debiased Machine Learning for Treatment and Structural Parameters,&#8221; showed how to use flexible machine learning methods (random forests, LASSO, neural networks) in the first stage while maintaining valid inference in the second stage.</p><p>The problem they solved is this. If you use a flexible first-stage estimator that overfits, the bias from overfitting contaminates your second-stage estimate. The solution involves two ingredients: sample splitting (estimate the first stage on one half of the data and the second stage on the other, then swap and average) and a Neyman-orthogonal score function that makes the second-stage estimator insensitive to small errors in the first stage. The result is that you can estimate nuisance parameters with whatever machine learning method you want, and as long as the nuisance estimates converge at a fast enough rate, your treatment effect estimate is &#8730;n-consistent and asymptotically normal.</p><p>This is the two-stage idea at its most general. The first stage has been democratized: you are no longer limited to probit or OLS or any particular parametric form. The second stage retains classical econometric properties. DML is still two-stage, just with orthogonality as the stitching technology. The architecture is the same one Theil formalized in 1953 and Heckman applied in 1979. What has changed is the toolkit available for the first stage and the theoretical machinery for making the two stages play well together.</p><h2>What Unifies Them</h2><p>Step back and look at the family through the template from the introduction. Every method fills in the same blanks:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Gscp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7bcf1e0-2212-4055-9ab3-88c54c4ba525_631x159.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Gscp!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7bcf1e0-2212-4055-9ab3-88c54c4ba525_631x159.png 424w, /__u/substackcdn.com/image/fetch/$s_!Gscp!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7bcf1e0-2212-4055-9ab3-88c54c4ba525_631x159.png 848w, /__u/substackcdn.com/image/fetch/$s_!Gscp!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7bcf1e0-2212-4055-9ab3-88c54c4ba525_631x159.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Gscp!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7bcf1e0-2212-4055-9ab3-88c54c4ba525_631x159.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Gscp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7bcf1e0-2212-4055-9ab3-88c54c4ba525_631x159.png" width="631" height="159" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d7bcf1e0-2212-4055-9ab3-88c54c4ba525_631x159.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:159,&quot;width&quot;:631,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:33034,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/189502356?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7bcf1e0-2212-4055-9ab3-88c54c4ba525_631x159.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Gscp!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7bcf1e0-2212-4055-9ab3-88c54c4ba525_631x159.png 424w, /__u/substackcdn.com/image/fetch/$s_!Gscp!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7bcf1e0-2212-4055-9ab3-88c54c4ba525_631x159.png 848w, /__u/substackcdn.com/image/fetch/$s_!Gscp!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7bcf1e0-2212-4055-9ab3-88c54c4ba525_631x159.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Gscp!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7bcf1e0-2212-4055-9ab3-88c54c4ba525_631x159.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>The pattern is the same in every row. The first stage estimates something that is not the parameter of interest but that is necessary to remove a bias, a confound, or an identification failure that prevents direct estimation. The second stage uses the first-stage output to recover the target parameter. And in every case, the method inherits the assumptions and the errors of its first stage.</p><p>This last point is the most important and the most underappreciated. Two-stage methods are modular, and modularity is their greatest strength. They let you decompose a hard problem into pieces. You do not need to estimate the entire system simultaneously. You do not need to get everything right at once. You can focus on one equation, one selection mechanism, one propensity score, and then plug the result into the next step. This is why they dominate applied work. The computational and conceptual burden is lower than joint estimation.</p><p>But modularity is also the vulnerability. If the first stage is wrong, the second stage is wrong. If the probit in the Heckman correction misspecifies the selection equation, &#955;&#770; is wrong and the selection correction is wrong. If the logit in the propensity score model omits an important confounder, the propensity scores are wrong and the treatment effect estimate is wrong. If the first-stage F-statistic in 2SLS is too low, the IV estimate is biased toward OLS. The specific failure mode varies, but the structure of the failure is always the same: garbage in the first stage, garbage in the second.</p><p>The generated regressors problem is the technical manifestation of this. The second stage treats its first-stage inputs as if they were known with certainty, but they are estimates. Pagan (1984) showed that ignoring the estimation uncertainty from the first stage produces biased standard errors in the second stage. Murphy and Topel (1985) provided a general correction. Modern practice increasingly uses bootstrap methods that re-estimate the entire two-stage procedure in each replication, which automatically accounts for first-stage uncertainty without requiring analytical corrections.</p><h2>Why the Profession Chose Modularity</h2><p>There is a question worth asking about the history. Every two-stage method has a joint estimation counterpart that is, in principle, more efficient. The Heckman two-step estimator is less efficient than full information maximum likelihood of the selection model. 2SLS is less efficient than full information maximum likelihood of the simultaneous system. Propensity score weighting is less efficient than correctly specified outcome regression.</p><p>So why did two-stage methods win?</p><p>The answer is the same answer that explains why 2SLS displaced full-information maximum likelihood in the simultaneous equations literature. Two-stage methods are resistant to certain kinds of misspecification. If you estimate the full system jointly and one equation is wrong, the contamination can spread to every parameter estimate. If you estimate one equation at a time with a two-stage approach, the misspecification in one equation does not automatically infect the others. <em>You sacrifice efficiency for containment.</em></p><p>And two-stage methods are simple. Heckman&#8217;s two-step procedure requires a probit and an OLS regression. 2SLS requires two OLS regressions. Propensity score matching requires a logit and a matching algorithm. None of these require numerical optimization of a complicated likelihood function. In the 1970s and 1980s, when computing was slow and expensive, this mattered enormously. In the 2020s, when computing is cheap, it still matters because simplicity reduces the scope for coding errors, makes results reproducible, and forces the researcher to be explicit about each step.</p><p>The profession chose modularity because it worked, because it scaled, and because it made the researcher&#8217;s assumptions transparent. Every two-stage method wears its assumptions on its sleeve. The first stage is visible. The second stage is visible. The connection between them is visible. When something goes wrong, you can usually tell which stage failed. This transparency is not a byproduct. It is the feature.</p><h2>The Two-Stage Logic, Beyond Econometrics</h2><p>The same architecture appears outside of econometrics. In machine learning, stacking and boosting use the output of one model as input to another. In Bayesian statistics, empirical Bayes methods estimate hyperparameters in a first stage and condition on them in a second. In survey sampling, raking and calibration adjust sample weights in a first stage to improve estimates in a second. The two-stage idea is not an invention of econometrics. It is a recurring solution to a recurring problem: when you cannot get what you want directly, estimate what you need first and use it to get the rest.</p><p>What econometrics contributed was the theoretical framework for understanding when this works and when it does not. The Cowles Commission established the conditions for identification. Heckman showed how to turn identification into a tractable estimator. Rosenbaum and Rubin showed how to reduce the dimensionality of confounding to a single score. Pagan and Murphy-Topel showed what happens to inference when the first stage introduces estimation error. Chernozhukov and collaborators showed how to make the architecture compatible with modern machine learning. Each step in this progression deepened our understanding of the same fundamental idea.</p><p>The two-stage method is not a single technique. It is an approach to problem-solving, one that says: if the whole problem is too hard, cut it in two, solve the pieces, and stitch them together. The stitching is where the subtlety lives. Getting it right requires understanding what the first stage is doing, what assumptions it carries, and how its errors propagate into the second stage. Getting it wrong can produce estimates that look precise and are completely misleading.</p><p>This is why the history matters. The researchers who built these methods were not just solving technical problems. They were building an architecture for empirical research, one that lets applied economists work with the data they have rather than the data they wish they had. That architecture is now so embedded in practice that most applied researchers use two-stage methods without thinking of them as two-stage methods. Every time you run a Heckman correction, a propensity score analysis, a 2SLS regression, or a double ML procedure, you are using the same idea, formalized seventy years ago in The Hague, extended to Chicago and Cambridge and London, and still expanding. Econometrics did not just invent estimators. It invented a way of breaking impossible problems into solvable pieces.</p><h2>Where to Start</h2><p>Heckman, J. J. (1979). &#8220;Sample Selection Bias as a Specification Error.&#8221; <em>Econometrica</em> 47(1): 153-161. The paper that reframed selection bias as an omitted variable problem and gave applied researchers a simple two-step fix. The paper that launched a thousand Heckit models.</p><p>Heckman, J. J. (1976). &#8220;The Common Structure of Statistical Models of Truncation, Sample Selection and Limited Dependent Variables and a Simple Estimator for Such Models.&#8221; <em>Annals of Economic and Social Measurement</em> 5(4): 475-492. The unifying paper that showed these three problems share a single structure.</p><p>Gronau, R. (1974). &#8220;Wage Comparisons: A Selectivity Bias.&#8221; <em>Journal of Political Economy</em> 82(6): 1119-1143. The precursor that identified the selection bias problem in wage comparisons but did not solve it.</p><p>Rosenbaum, P. R., and Rubin, D. B. (1983). &#8220;The Central Role of the Propensity Score in Observational Studies for Causal Effects.&#8221; <em>Biometrika</em> 70(1): 41-55. The theorem that collapses multidimensional confounding into a single scalar.</p><p>LaLonde, R. J. (1986). &#8220;Evaluating the Econometric Evaluations of Training Programs with Experimental Data.&#8221; <em>American Economic Review</em> 76(4): 604-620. The destructive test that asked whether econometric methods could replicate experimental results. The answer was mostly no.</p><p>Dehejia, R. H., and Wahba, S. (1999). &#8220;Causal Effects in Nonexperimental Studies: Reevaluating the Evaluation of Training Programs.&#8221; <em>Journal of the American Statistical Association</em> 94(448): 1053-1062. The propensity score response to LaLonde, showing that careful matching could get closer to the experimental benchmark.</p><p>Smith, J. A., and Todd, P. E. (2005). &#8220;Does Matching Overcome LaLonde&#8217;s Critique of Nonexperimental Estimators?&#8221; <em>Journal of Econometrics</em> 125(1-2): 305-353. The important follow-up that questioned the sensitivity of propensity score estimates to specification choices and comparison group construction.</p><h2>Some Additional Readings</h2><p><strong>On the control function approach</strong></p><p>Heckman, J. J., and Robb, R. (1985). &#8220;Alternative Methods for Evaluating the Impact of Interventions: An Overview.&#8221; <em>Journal of Econometrics</em> 30(1-2): 239-267. Where the control function idea was named and extended beyond the selection model.</p><p>Rivers, D., and Vuong, Q. H. (1988). &#8220;Limited Information Estimators and Exogeneity Tests for Simultaneous Probit Models.&#8221; <em>Journal of Econometrics</em> 39(3): 347-366. The control function extended to probit with endogenous regressors. Shows that including first-stage residuals as controls works in nonlinear models where 2SLS does not.</p><p>Wooldridge, J. M. (2015). &#8220;Control Function Methods in Applied Econometrics.&#8221; <em>Journal of Human Resources</em> 50(2): 420-445. The modern survey. Clear exposition of when control functions and 2SLS coincide and when they diverge.</p><p><strong>On the generated regressors problem</strong></p><p>Pagan, A. (1984). &#8220;Econometric Issues in the Analysis of Regressions with Generated Regressors.&#8221; <em>International Economic Review</em> 25(1): 221-247. The first systematic treatment of what goes wrong with standard errors when a regressor is estimated.</p><p>Murphy, K. M., and Topel, R. H. (1985). &#8220;Estimation and Inference in Two-Step Econometric Models.&#8221; <em>Journal of Business &amp; Economic Statistics</em> 3(4): 370-379. The correction for two-step standard errors that should be applied in any two-stage procedure.</p><p><strong>On doubly robust methods and modern developments</strong></p><p>Robins, J. M., Rotnitzky, A., and Zhao, L. P. (1994). &#8220;Estimation of Regression Coefficients When Some Regressors Are Not Always Observed.&#8221; <em>Journal of the American Statistical Association</em> 89(427): 846-866. The paper that introduced doubly robust estimation.</p><p>Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., and Robins, J. (2018). &#8220;Double/Debiased Machine Learning for Treatment and Structural Parameters.&#8221; <em>The Econometrics Journal</em> 21(1): C1-C68. The paper that married the two-stage idea with machine learning, showing how to use flexible first-stage estimators while preserving valid second-stage inference.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>This is the second essay in a series on instrumental variables and two-stage methods. The first essay covered the history of instrumental variables from Wright (1928) to the credibility revolution.  </p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[Ghost GDP and the Two Frontiers]]></title><description><![CDATA[Why the AI labor debate is asking the wrong question, and a prediction that will resolve it]]></description><link>https://carloschavezp29.substack.com/p/ghost-gdp-and-the-two-frontiers</link><guid isPermaLink="false">https://carloschavezp29.substack.com/p/ghost-gdp-and-the-two-frontiers</guid><dc:creator><![CDATA[Carlos Chavez]]></dc:creator><pubDate>Fri, 27 Feb 2026 10:00:45 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v11X!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbf5601-4d1b-4706-b816-2544066f3494_892x602.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>On Sunday, Citrini Research published a thought experiment on Substack titled &#8220;The 2028 Global Intelligence Crisis.&#8221; The piece is written as a dispatch from June 2028. In it, AI-driven automation has destroyed white-collar employment, unemployment breaks into double digits, and U.S. equities are deep into a drawdown. On Monday, IBM had its worst day in decades; Visa, Mastercard, DoorDash, and a range of software names sold off. Michael Burry shared the piece. Nassim Taleb added that markets are underpricing structural risks. A Substack post, containing no model and no estimation, moved billions of dollars in market capitalization in a single day.</p><p>The piece is well written. The scenario is vivid. Several of the mechanisms it describes are real. But the economics are underspecified.</p><p>Noah Smith responded today with a post arguing that Citrini is &#8220;just a scary bedtime story.&#8221; His central objection is correct: Citrini does not write down an explicit macroeconomic model, so the assumptions driving the scenario are invisible. You cannot tell what they think about general equilibrium, about adjustment mechanisms, about historical base rates. The narrative does all the work. Smith separates the microeconomic thesis, which industries and jobs AI will disrupt, from the macroeconomic thesis, whether that disruption causes a recession. He finds the micro thesis debatable and the macro thesis underdeveloped. The distinction is useful. A technology can destroy specific business models without crashing the economy, and confusing the two leads to exactly the kind of panic selling that followed the Citrini post.</p><p>Alex Imas, a professor at the University of Chicago Booth School of Business, asked a version of Citrini&#8217;s question two months before them, but with a model. In a January essay on his Substack, Imas worked through whether advanced AI could actually lead to negative economic growth. The mechanism is demand collapse: if automation shifts income from workers with high marginal propensities to consume to capital owners who are already satiated, aggregate demand falls even as productive capacity expands. Machines sit idle because there is no one to buy what they could produce. Imas derived the conditions under which this occurs and concluded that they are probably too extreme to hold in practice, but that with less extreme assumptions, automation can depress demand and push growth toward the lower end of current forecasts. The exercise illustrates exactly what Citrini lacks. Writing down the model makes the assumptions visible, and therefore contestable. Imas evaluated them and found them wanting. Citrini never gave their readers that opportunity.</p><p>Separately, S&#233;b Krier published a guest essay on Imas&#8217;s Substack arguing for what he calls &#8220;The Cyborg Era.&#8221; Krier, who leads AGI policy development at Google DeepMind, contends that full labor displacement requires that AI alone outperforms the combination of AI and humans. As long as complementarity holds, human labor retains comparative advantage. He argues that this complementarity is likely to persist for ten to twenty years, sustained by the jagged, uneven nature of AI capability gains and by intrinsic preferences for human involvement in social and relational goods.</p><p>Four positions, then. Citrini sees a doom spiral. Imas formalizes the demand collapse and finds it unlikely but not impossible. Krier sees durable complementarity. Smith says the micro disruption is real but the macro panic is overblown. All four share a common limitation: they treat automation as scalar. They argue over the sign of a single aggregate effect, instead of asking whether different frontiers move different parts of the wage distribution in opposite directions, and then cancel under aggregation.</p><p>This essay does four things. It checks Citrini&#8217;s implied productivity claims against historical benchmarks. It examines their central concept, Ghost GDP, and shows where it connects to real economics and where it departs. It explains why the demand and complementarity arguments, while correct in principle, miss a critical asymmetry. And it introduces a framework, drawn from my own research, that produces a different and more specific prediction about what AI will do to wages than anything in the current debate.</p><p>The core claim of this essay is simple: when automation frontiers push inequality in opposite directions, any single index can produce a precisely estimated zero, and a precisely wrong conclusion.</p><h2>The Solow Residual as a Sanity Check</h2><p>The part of economic growth that cannot be explained by increases in labor or capital is called total factor productivity, or TFP. Economists call it the Solow residual, after Robert Solow&#8217;s 1957 decomposition. It captures technological change, organizational improvement, and everything else the model cannot account for. Moses Abramovitz called it &#8220;a measure of our ignorance.&#8221; The description is accurate.</p><p>As a result, short-run TFP movements are noisy, but sustained accelerations are hard to fake. If AI is the transformative technology its advocates claim, the effect should eventually appear as accelerating TFP growth. The question is how much acceleration is plausible.</p><p>Electricity, during its peak diffusion in the 1920s, contributed roughly 0.5 to 0.7 percentage points per year to U.S. TFP growth. The information and communications technology revolution, personal computers, the internet, enterprise software, added about 0.5 to 1.0 percentage points during 1995 to 2004. These were technologies that reorganized production across most sectors of the economy. They are the upper bound of what general-purpose technologies have delivered in modern history.</p><p>Acemoglu (2024), in the most careful calibration to date, estimates that AI will contribute between 0.5 and 0.9 percentage points to TFP growth over the next decade. That is cumulative over the decade, not an annual growth rate. It amounts to roughly one-tenth of what the most optimistic AI valuations require. Figure 1 puts this in context. The solid bars show the peak annual TFP contributions from electricity and ICT, the two general-purpose technologies closest in scope to what AI promises. The small red bar is Acemoglu&#8217;s estimate, annualized. The dashed ghost bar on the right is what current AI equity valuations would require: a back-of-the-envelope calculation starting from the roughly $8&#8211;13 trillion in AI-attributable market capitalization premium, converting to implied excess earnings via a perpetuity model, and mapping to TFP through a standard Solow decomposition. The gap between what the most careful calibration estimates and what the market is pricing is an order of magnitude.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!v11X!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbf5601-4d1b-4706-b816-2544066f3494_892x602.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!v11X!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbf5601-4d1b-4706-b816-2544066f3494_892x602.png 424w, /__u/substackcdn.com/image/fetch/$s_!v11X!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbf5601-4d1b-4706-b816-2544066f3494_892x602.png 848w, /__u/substackcdn.com/image/fetch/$s_!v11X!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbf5601-4d1b-4706-b816-2544066f3494_892x602.png 1272w, /__u/substackcdn.com/image/fetch/$s_!v11X!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbf5601-4d1b-4706-b816-2544066f3494_892x602.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!v11X!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbf5601-4d1b-4706-b816-2544066f3494_892x602.png" width="892" height="602" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bfbf5601-4d1b-4706-b816-2544066f3494_892x602.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:602,&quot;width&quot;:892,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:150589,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/189323231?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbf5601-4d1b-4706-b816-2544066f3494_892x602.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!v11X!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbf5601-4d1b-4706-b816-2544066f3494_892x602.png 424w, /__u/substackcdn.com/image/fetch/$s_!v11X!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbf5601-4d1b-4706-b816-2544066f3494_892x602.png 848w, /__u/substackcdn.com/image/fetch/$s_!v11X!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbf5601-4d1b-4706-b816-2544066f3494_892x602.png 1272w, /__u/substackcdn.com/image/fetch/$s_!v11X!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbf5601-4d1b-4706-b816-2544066f3494_892x602.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p>
      <p>
          <a href="/__u/carloschavezp29.substack.com/p/ghost-gdp-and-the-two-frontiers">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[What Is an Instrument?]]></title><description><![CDATA[From Wright's Appendix to the Credibility Revolution]]></description><link>https://carloschavezp29.substack.com/p/what-is-an-instrument</link><guid isPermaLink="false">https://carloschavezp29.substack.com/p/what-is-an-instrument</guid><dc:creator><![CDATA[Carlos Chavez]]></dc:creator><pubDate>Mon, 23 Feb 2026 13:14:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!KymY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49d28510-1327-4551-accc-79d9382940b4_847x436.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In 1928, Philip Wright published <em>The Tariff on Animal and Vegetable Oils</em>, a study of the butter and flaxseed markets commissioned by the Brookings Institution. The book is largely forgotten. But buried in Appendix B is a method that changed how economists think about causation from observational data. Wright&#8217;s problem was that price and quantity are determined simultaneously in the market. A regression of quantity on price does not estimate demand. It estimates a composite object that mixes supply and demand shocks, with no structural interpretation. (The appendix is attributed to Philip Wright, though historians debate the role of his son Sewall Wright, who pioneered path analysis in genetics. Stock and Trebbi (2003) investigate the question in detail.)</p><p>Wright&#8217;s move was to find a variable that shifts supply without shifting demand. If a supply shifter moves the equilibrium price while demand stays put, then the observed price variation is entirely supply-driven, and the quantity response traces out the demand curve. That is the logic of instrumental variables: identify causal responses using variation in an endogenous regressor that is generated by a source plausibly orthogonal to the structural error. The instrument does not need to be interesting in itself. It needs only to move the variable of interest for reasons unrelated to the error term.</p><p>That idea, in various forms, has become one of the central tools of empirical economics. It underlies the natural experiments of Angrist and Card, the shift-share designs of Bartik, the judge leniency instruments of Kling, and thousands of applied papers across labor, development, trade, and public finance. Yet the history of how this idea was formalized, who developed it, and what problems it was originally designed to solve, is less widely known than it should be. Understanding that history matters, because the method carries assumptions that came out of the problems it was built to solve, and those assumptions still determine what the method can and cannot tell us.</p><p>Wright himself is probably the most underappreciated figure in the history of economics. He was not a young prodigy at a top department. He was a former college professor turned Brookings researcher, in his late sixties, writing about tariffs on oils, and he solved the identification problem in an appendix. The profession ignored the solution for two decades. His contribution was rediscovered only after others had reinvented it. There is something worth reflecting on in the fact that the most important methodological idea in empirical economics was first worked out by someone the discipline has largely forgotten, and I plan to write more about Wright in a future essay.</p><h2>The Identification Problem</h2><p>The problem that instruments solve is older than the solution. Economists in the early twentieth century understood that most economic variables are jointly determined. Prices affect quantities, but quantities also affect prices. Education affects wages, but ability affects both education and wages. Police deployment correlates with crime, but the direction of causation runs both ways.</p><p>In each case, ordinary least squares produces biased estimates because the regressor is correlated with the error term. The error term captures everything the researcher cannot observe, ability, expectations, political motivations, strategic responses, and in most applied settings these unobservables are almost always tangled up with the variables of interest. Running a regression as if the regressor were exogenous does not solve this. It produces a number, but the number does not have a causal interpretation.</p><p>The supply and demand case is the textbook illustration, but it is worth dwelling on because the structure of the problem recurs everywhere. A market generates equilibrium prices and quantities. We observe many markets, or the same market over time. We want to know the demand elasticity. But the data are generated by the intersection of supply and demand, and without additional information we cannot separate the two. Regressing quantity on price gives a coefficient that reflects the relative variability of supply and demand shocks, not the slope of either curve.</p><p>Alfred Marshall understood this. So did the economists who built the first macroeconometric models in the 1930s. The question was not whether simultaneity was a problem. The question was what to do about it.</p><h2>The Cowles Commission</h2><p>The answer came from the Cowles Commission for Research in Economics, housed at the University of Chicago from 1939 to 1955. This group, including Trygve Haavelmo, Tjalling Koopmans, Jacob Marschak, and Theodore Anderson, built the theoretical foundations of identification in simultaneous systems.</p><p>Haavelmo&#8217;s 1944 paper, &#8220;The Probability Approach in Econometrics,&#8221; was the starting point. He argued that the equations in an empirical model should be interpreted as structural relationships, descriptions of how the economy actually works, rather than statistical summaries. The coefficients in a demand equation are causal parameters. They tell you what would happen if you intervened to change the price. Estimating them correctly requires taking simultaneity seriously. This was a fundamental reorientation. Before Haavelmo, the field was largely about fitting curves to data. After Haavelmo, it became about estimating the parameters of a model assumed to describe how the world works. The distinction between a statistical summary and a structural relationship, between a correlation and a causal parameter, runs through everything that followed. I discussed Haavelmo briefly in my essay on causality, but his 1944 paper deserves a dedicated treatment, and I intend to give it one.</p><p>The Cowles Commission itself is a remarkable institution that keeps appearing in this series, sometimes directly, sometimes as background. It has surfaced in my essays on causality, on identification, and now on instrumental variables. This is not a coincidence. The Cowles program, for roughly fifteen years in Chicago, produced the theoretical infrastructure on which most of modern empirical work rests: identification theory, simultaneous equations estimation, the probability approach. Many of the ideas I write about in this Substack trace back to work done in that building on the South Side of Chicago between 1939 and 1955. A full account of the Cowles Commission, its intellectual ambitions, its internal debates, and its complicated relationship with the emerging Chicago school of economics, is something I plan to write as a standalone essay.</p><p>The Cowles program aimed to specify complete systems of simultaneous equations, supply and demand, consumption and investment, money demand and money supply, and estimate all the structural parameters at once. Koopmans and his collaborators worked out when this was possible. The answer involved exclusion restrictions. To identify the demand curve, you need at least one variable that appears in the supply equation but not in the demand equation. That variable shifts supply, which moves the equilibrium price, and the resulting variation in price is exogenous in the sense that it is uncorrelated with unobserved demand shifters.</p><p>This is the formal origin of the instrument concept. An instrument is just a variable. What makes it an instrument is the exclusion restriction: the claim that some variable affects one equation but not another. Rainfall affects agricultural supply but not consumer demand. Draft lottery numbers affect military service but not earnings directly. The restriction cannot be tested from the same data. It is an assumption about how the world works, justified by theory or institutional knowledge.</p><p>Anderson and Rubin, working within the Cowles framework, developed the limited information maximum likelihood (LIML) estimator in 1949 and 1950 for estimating the parameters of a single structural equation. What is less appreciated, and what Anderson himself pointed out in a 2005 retrospective, is that their 1950 paper implicitly contained the two-stage least squares estimator and its asymptotic distribution. The notation was difficult and the exposition obscure. The connection was not made explicit, and the method we now call 2SLS would be formalized independently by three other researchers.</p><p>The Cowles Commission was not the only path to instrumental variables. A parallel tradition, motivated by measurement error rather than simultaneity, was converging on the same idea. If the variable you want to study is measured with noise, OLS is biased toward zero (attenuation bias). Olav Reiers&#248;l in Norway in the 1940s and R.C. Geary in Ireland in 1949 worked out that you could fix this by finding a variable correlated with the true value of the mismeasured regressor but uncorrelated with the noise. That variable is an instrument, though nobody in the errors-in-variables tradition called it that at first. Abraham Wald&#8217;s 1940 paper on fitting straight lines when both variables contain errors is sometimes cited as an even earlier precursor. What is remarkable is that two completely different problems, simultaneity and measurement error, led to the same estimator. The math does not care why your regressor is correlated with the error term. It only cares that you have a variable that breaks the correlation.</p><h2>Three Countries, One Estimator</h2><p>Between 1953 and 1958, three researchers working independently in three countries converged on the same estimation technique. The convergence tells us something about the state of the field. The identification problem was well understood, LIML existed but was computationally demanding, and the need for a simpler equation-by-equation estimator was widely felt. The solution was, in a sense, waiting to be found.</p><p>Henri Theil, working at the Central Planning Bureau in The Hague, published the first description of two-stage least squares in 1953. The approach was simple. In the first stage, regress the endogenous variable on all the exogenous variables in the system to obtain predicted values purged of correlation with the structural error. In the second stage, use those predicted values in place of the original endogenous variable. The resulting estimates are consistent. The method required only two ordinary least squares regressions and could be applied one equation at a time, without specifying the entire system. This was a substantial practical advantage over the full-information methods of the Cowles Commission.</p><p>Robert Basmann arrived at the same estimator independently in 1957. Basmann had earned his PhD under Gerard Tintner at Iowa State, won a Fulbright to the University of Oslo, and taught for a year at Northwestern University. His writings during that period made deep contributions to simultaneous equation models and identification. His derivation of 2SLS took a slightly different but mathematically equivalent approach to Theil&#8217;s, emphasizing the generalized classical linear estimation framework rather than the two-stage procedure. Basmann eventually left academia for litigation consulting, and served as the statistical consultant on Oprah Winfrey&#8217;s defense team in the trial against the Texas cattle producers, which is not the career arc anyone would have predicted from his dissertation on simultaneous equations.</p><p>The third independent discoverer was J. Denis Sargan at the London School of Economics. Sargan published his version in 1958 in <em>Econometrica</em>, under the title &#8220;The Estimation of Economic Relationships Using Instrumental Variables.&#8221; Beyond deriving 2SLS, Sargan contributed something that would prove equally important: the overidentification test. When you have more instruments than endogenous regressors, the model is overidentified, and you can test whether the extra instruments are consistent with the maintained assumptions. The Sargan test, which Hansen generalized into the J-test for GMM in 1982, remains a standard diagnostic in applied work.</p><p>A note on names. Denis Sargan at the LSE, who developed 2SLS and the overidentification test in the 1950s, is not Thomas Sargent at NYU and Stanford, who won the Nobel Prize for rational expectations and dynamic macroeconomics. They are different people with different contributions, and the confusion between them is common even among economists.</p><p>By 1958, three researchers across the Netherlands, the United States, and the United Kingdom had independently formalized the same method. The Cowles Commission had framed the problem. LIML provided a solution, but it was computationally demanding. Theil, Basmann, and Sargan each saw that a simpler, limited-information approach was possible, and each found it.</p><h2>What the Estimator Does</h2><p>The formal requirements are straightforward. Given a structural equation</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;Y = \\beta X + \\varepsilon&quot;,&quot;id&quot;:&quot;WEOGZKKIKH&quot;}" data-component-name="LatexBlockToDOM"></div><p>where X is endogenous, meaning correlated with &#949;, an instrument Z must satisfy two conditions. First, relevance: Z must be correlated with X. The instrument must actually move the endogenous variable. This is testable. You run the first-stage regression and check the F-statistic. The rule-of-thumb threshold, following Stock and Yogo (2005), is F above ten, though more recent work by Montiel Olea and Pflueger (2013) has developed robust versions of this test.</p><p>Second, the exclusion restriction: Z must be uncorrelated with &#949;. The instrument must affect Y only through its effect on X, not through any other channel. This is not testable from the same data. You defend it with theory, with institutional knowledge, with a narrative about why the instrument is exogenous. The entire weight of the empirical strategy rests on this assumption.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!KymY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49d28510-1327-4551-accc-79d9382940b4_847x436.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!KymY!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49d28510-1327-4551-accc-79d9382940b4_847x436.png 424w, /__u/substackcdn.com/image/fetch/$s_!KymY!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49d28510-1327-4551-accc-79d9382940b4_847x436.png 848w, /__u/substackcdn.com/image/fetch/$s_!KymY!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49d28510-1327-4551-accc-79d9382940b4_847x436.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KymY!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49d28510-1327-4551-accc-79d9382940b4_847x436.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!KymY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49d28510-1327-4551-accc-79d9382940b4_847x436.png" width="847" height="436" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/49d28510-1327-4551-accc-79d9382940b4_847x436.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:436,&quot;width&quot;:847,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:69254,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/188895762?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49d28510-1327-4551-accc-79d9382940b4_847x436.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!KymY!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49d28510-1327-4551-accc-79d9382940b4_847x436.png 424w, /__u/substackcdn.com/image/fetch/$s_!KymY!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49d28510-1327-4551-accc-79d9382940b4_847x436.png 848w, /__u/substackcdn.com/image/fetch/$s_!KymY!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49d28510-1327-4551-accc-79d9382940b4_847x436.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KymY!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49d28510-1327-4551-accc-79d9382940b4_847x436.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>Figure 1. The instrument Z affects the outcome Y only through the endogenous variable X (solid arrows). The exclusion restriction (crossed-out curve) rules out any direct path from Z to Y. The unobserved error &#949; creates the endogeneity problem by affecting both X and Y (dashed arrows).</p><p>Under these conditions, the IV estimator is</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\hat{\\beta}_{IV} = \\frac{\\text{Cov}(Z, Y)}{\\text{Cov}(Z, X)}&quot;,&quot;id&quot;:&quot;MNPTGSPGXK&quot;}" data-component-name="LatexBlockToDOM"></div><p>This is a Wald ratio: reduced form over first stage. The numerator captures how much Y moves when Z moves. The denominator captures how much X moves when Z moves. Divide one by the other and you get how much Y moves per unit of X, using only the variation in X that Z generates. That is the causal parameter, or at least it is if the exclusion restriction holds.</p><p>Consider Miguel, Satyanath, and Sergenti (2004), who asked whether negative income shocks cause civil conflict in sub-Saharan Africa. Conflict and GDP growth are jointly determined: war destroys output, but bad harvests and collapsing incomes may also start wars. OLS on conflict against growth confounds both directions.</p><p>They used rainfall as an instrument. In countries where agriculture dominates the economy, rainfall shocks hit GDP hard, so the first stage is strong. The exclusion restriction requires that rainfall moves conflict only via income, not through other channels. If that holds, the Wald ratio gives you the causal effect of an income decline on the probability of war.</p><p>The estimates were large. A rainfall-driven growth decline of five percentage points raised the probability of civil conflict by roughly one half the following year. You could never get that number from OLS, because OLS mixes the effect of income on conflict with the reverse channel, conflict destroying income. The instrument pulls those apart.</p><p>But does rainfall really work only via GDP? Droughts also cause displacement, competition for water, migration, and psychological stress. If any of those channels operate independently of income, the exclusion restriction fails. Rainfall would then be an instrument for income plus all the other disruptions that droughts cause, not for income alone. The subsequent literature has fought over exactly this point. The instrument is only as good as the argument that it operates through the specified channel and no other.</p><p>Two-stage least squares mechanizes this. The first stage projects X onto the instrument space to obtain fitted values that capture only the instrument-driven variation. The second stage regresses Y on those fitted values. Because the fitted values are uncorrelated with &#949; by construction, OLS in the second stage is consistent.</p><p>The asymmetry between the two conditions matters for practice. Relevance is testable and routinely tested. The exclusion restriction is not testable from the same data and often asserted rather than defended. This is why the quality of an instrumental variables paper depends less on the estimation than on the argument for why the exclusion restriction holds.</p><p>Even when the exclusion restriction is right, relevance can kill you. A weak first stage means the Wald ratio is dividing two small numbers, and small perturbations in either one blow up the estimate. Staiger and Stock (1997) showed that weak instruments produce finite-sample bias toward OLS and size distortions that make your t-statistics unreliable. Stock and Yogo (2005) proposed the rule-of-thumb F &gt; 10 threshold to flag dangerous cases. And the problem is not hypothetical. Bound, Jaeger, and Baker (1995) went back to Angrist and Krueger&#8217;s famous quarter-of-birth instrument for schooling and found that the first-stage F-statistics were, in some specifications, uncomfortably low. A clever design with a plausible story can still give you garbage if the instrument barely moves the endogenous variable.</p><p>Applied papers also routinely report a Hausman test, which compares OLS and IV. If X is actually exogenous, both are consistent but OLS is more efficient, so they should agree. If they disagree, you reject exogeneity. But notice what the test assumes: it takes the instrument as valid and asks whether you needed it. It tests the disease (endogeneity), not the cure (instrument validity). A significant Hausman statistic means OLS and IV point in different directions. Whether that tells you OLS is biased or the instrument is broken depends entirely on whether you believe the exclusion restriction.</p><h2>What 2SLS Changed</h2><p>The Cowles approach required you to specify and estimate the whole system at once. Get one equation wrong and every other equation is contaminated. 2SLS broke that dependency. You could estimate one equation at a time, as long as you had an exclusion restriction for that equation. You did not need to know or estimate what was happening in the rest of the system. This made the method modular, which made it usable, which made it dominant. The price was that identification now depended entirely on the credibility of each instrument, case by case, with no structural model to discipline the choice.</p><p>This modularity had consequences that played out over decades. When the credibility revolution arrived in the 1990s, its toolkit was essentially the instrumental variables approach freed from the Cowles requirement to specify the entire system. Angrist (1990) used Vietnam draft lottery numbers to estimate the effect of military service on earnings. Angrist and Krueger (1991) used quarter of birth, which determines school entry age and therefore exposure to compulsory schooling laws, to estimate returns to education. Card (1995) used geographic proximity to colleges as an instrument for years of schooling. Each of these papers used a single instrument in a single equation, justified not by a full structural model but by an argument about why the instrument was as good as randomly assigned. The technical foundation had been laid by Theil, Basmann, and Sargan forty years earlier. What changed in the 1990s was the source of the exclusion restriction: it came from institutions, lotteries, and geographic accidents rather than from theory about how markets work.</p><p>Sargan&#8217;s overidentification test added a diagnostic dimension. When you have more instruments than endogenous regressors, you can test whether they all point to the same structural parameter. If they do not, either an instrument violates the exclusion restriction or the model is misspecified. The test does not tell you which instrument is invalid, but it signals that something is wrong. It is important to understand what the test does not do: it checks the mutual consistency of instruments, not the truth of the exclusion restriction. A set of instruments can all be wrong in the same way and still pass the test. Zellner and Theil (1962) extended the approach to three-stage least squares, which estimates entire systems while accounting for cross-equation error correlations. Hansen&#8217;s (1982) generalized method of moments provided a unifying framework that nests 2SLS as a special case. But for the vast majority of applied work, 2SLS remains the default. It is simple, it forces the researcher to think about identification, and it produces results that are easy to interpret and communicate.</p><h2>The Foundations, and What Comes Next</h2><p>The history of instrumental variables is a history of economists confronting the limits of observational data. We cannot run experiments on trade policy, or randomly assign education, or control for ability. Economic data are generated by choices, and choices are endogenous. Instruments provide a way to extract variation from this tangle, variation that moves for reasons the researcher can argue are orthogonal to the unobservables that create the endogeneity problem in the first place.</p><p>Wright saw this in 1928, when he realized that supply shifters could trace out demand curves. The Cowles Commission systematized it in the 1940s, embedding the idea in a theory of identification for simultaneous systems. Theil, Basmann, and Sargan operationalized it in the 1950s, giving the profession an estimator simple enough to use on the data and computing resources available at the time.</p><p>But the estimator was only the beginning. What an instrument estimates, what parameter it identifies, depends on assumptions that the pioneers did not fully articulate. Under the structural approach of the Cowles Commission, instruments recover the parameters of a model assumed to hold for everyone. Under the local average treatment effects framework of Imbens and Angrist (1994), the same estimator recovers something more limited: the causal effect for the subpopulation of compliers whose behavior is actually changed by the instrument. Different instruments, applied to the same question, can recover different parameters, not because one is wrong but because they weight the underlying heterogeneity differently. I have written about this distinction before. It is one of the most important open questions in modern econometrics.</p><p>And the two-stage logic that Theil formalized, use a first stage to solve one problem, then plug the result into a second stage, turns out to be far more general than instrumental variables alone. The same architecture underlies the Heckman correction for selection bias, the control function approach, propensity score methods, and a range of estimators that break an intractable problem into two tractable steps. That will be the subject of the next essay in this series.</p><p>The instrumental variable, for all its technical machinery, rests on a simple idea. Find variation that is clean. Use it to answer a question that dirty variation cannot answer. The hard part was never the estimation. The hard part was, and remains, the argument for why the variation is clean.</p><h2>Where to Start</h2><p>Wright, P. G. (1928). <em>The Tariff on Animal and Vegetable Oils</em>. Macmillan. The first instrumental variables application, buried in Appendix B of a tariff study. Stock and Trebbi (2003) provide a modern forensic analysis of whether it was Philip or Sewall Wright who wrote the econometric appendix.</p><p>Haavelmo, T. (1944). &#8220;The Probability Approach in Econometrics.&#8221; <em>Econometrica</em>. The manifesto that reframed econometric equations as structural, causal relationships rather than statistical summaries.</p><p>Theil, H. (1953). &#8220;Repeated Least Squares Applied to Complete Equation Systems&#8221; and &#8220;Estimation and Simultaneous Correlation in Complete Equation Systems.&#8221; Central Planning Bureau, The Hague. Two mimeographed papers that introduced 2SLS. Theil&#8217;s textbook, <em>Principles of Econometrics</em> (1971), gives a mature treatment.</p><p>Basmann, R. L. (1957). &#8220;A Generalized Classical Method of Linear Estimation of Coefficients in a Structural Equation.&#8221; <em>Econometrica</em>. The independent derivation of 2SLS from Northwestern, emphasizing the classical estimation framework.</p><p>Sargan, J. D. (1958). &#8220;The Estimation of Economic Relationships Using Instrumental Variables.&#8221; <em>Econometrica</em>. The third independent derivation, plus the overidentification test that bears his name.</p><p>Anderson, T. W. (2005). &#8220;Origins of the Limited Information Maximum Likelihood and Two-Stage Least Squares Estimators.&#8221; <em>Journal of Econometrics</em>. Anderson&#8217;s own account of how LIML implicitly contained 2SLS, and why nobody noticed at the time.</p><p>Hansen, L. P. (1982). &#8220;Large Sample Properties of Generalized Method of Moments Estimators.&#8221; <em>Econometrica</em>. The framework that nests 2SLS, provides the J-test generalization of Sargan&#8217;s overidentification test, and connects instrumental variables to a broader class of moment condition estimators.</p><h2>Some Additional Readings</h2><p><strong>On the history and philosophy of identification</strong></p><p>Angrist, J. D., &amp; Krueger, A. B. (1999). &#8220;Empirical Strategies in Labor Economics.&#8221; <em>Handbook of Labor Economics</em>. The discussion of what IV coefficients estimate under heterogeneous effects is essential and underappreciated.</p><p>Stock, J. H., &amp; Trebbi, F. (2003). &#8220;Retrospectives: Who Invented Instrumental Variable Regression?&#8221; <em>Journal of Economic Perspectives</em>. The detective work on Philip versus Sewall Wright.</p><p><strong>On what instruments identify under heterogeneity</strong></p><p>Imbens, G. W., &amp; Angrist, J. D. (1994). &#8220;Identification and Estimation of Local Average Treatment Effects.&#8221; <em>Econometrica</em>. The paper that showed IV recovers effects for compliers, not the population, and started the debate about what instruments actually estimate.</p><p>Heckman, J. J., &amp; Vytlacil, E. (2005). &#8220;Structural Equations, Treatment Effects, and Econometric Policy Evaluation.&#8221; <em>Econometrica</em>. Shows that different IV estimators using different instruments recover different weighted averages of the marginal treatment effect. The unifying framework that connects the structural and design-based traditions.</p><p><strong>On weak instruments and diagnostics</strong></p><p>Stock, J. H., &amp; Yogo, M. (2005). &#8220;Testing for Weak Instruments in Linear IV Regression.&#8221; In <em>Identification and Inference for Econometric Models</em>. The standard reference for first-stage F-statistic thresholds.</p><p>Andrews, I., Stock, J. H., &amp; Sun, L. (2019). &#8220;Weak Instruments in Instrumental Variables Regression: Theory and Practice.&#8221; <em>Annual Review of Economics</em>. A modern treatment of weak identification, with practical guidance.</p><p>Montiel Olea, J. L., &amp; Pflueger, C. E. (2013). &#8220;A Robust Test for Weak Instruments.&#8221; <em>Journal of Business &amp; Economic Statistics</em>. Extends Stock-Yogo to settings with heteroskedasticity and autocorrelation.</p><p></p>]]></content:encoded></item><item><title><![CDATA[What Trade Theory Misses About Developing Countries]]></title><description><![CDATA[How Ignoring Informality Fundamentally Distorts Our Understanding of Trade Liberalization.]]></description><link>https://carloschavezp29.substack.com/p/what-trade-theory-misses-about-developing</link><guid isPermaLink="false">https://carloschavezp29.substack.com/p/what-trade-theory-misses-about-developing</guid><dc:creator><![CDATA[Carlos Chavez]]></dc:creator><pubDate>Fri, 20 Feb 2026 11:53:50 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!yYR7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76c52f3-992e-4536-ad56-0266340e6e36_778x445.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Two billion workers. That is the ILO&#8217;s latest estimate of the global informal workforce, people employed without contracts, social protection, or legal recognition. In Sub-Saharan Africa and South Asia, informal employment rates exceed 85%. In Latin America, between 35% and 80% of workers operate outside the formal economy, depending on the country. The IMF estimates informality accounts for roughly 35% of GDP in low- and middle-income countries, compared to about 15% in advanced economies.</p><p>And yet, when economists evaluate trade liberalization in developing countries, they overwhelmingly rely on models that contain no informal sector. Heckscher-Ohlin, Melitz, and their descendants model a world where all firms operate under one set of rules. These frameworks were not wrong for the questions they were built to answer. But applying them to economies where half the workforce evades taxes, regulation, and labor laws means the model's economy and the actual economy are two different objects.</p><p><a href="https://www.nber.org/system/files/working_papers/w28391/revisions/w28391.rev0.pdf">A new paper forthcoming in </a><em><a href="https://www.nber.org/system/files/working_papers/w28391/revisions/w28391.rev0.pdf">Econometrica</a></em><a href="https://www.nber.org/system/files/working_papers/w28391/revisions/w28391.rev0.pdf">, &#8220;Trade and Domestic Distortions: The Case of Informality&#8221;, by Rafael Dix-Carneiro (Duke), Pinelopi Goldberg (Yale), Costas Meghir (Yale), and Gabriel Ulyssea (UCL) fills this gap directly. </a>And the results overturn several things we thought we knew.</p><h2>The Problem With Ignoring Informality</h2><p>Here is the basic tension. In developing countries, governments impose taxes, minimum wages, and labor regulations on firms. But enforcement is imperfect. Some firms register, comply, and bear the full regulatory burden. These are formal firms. Others fly under the radar, avoiding taxes and regulations but facing constraints on growth. These are informal firms.</p><p>This creates what the authors call a <em>size-dependent distortion</em>. Larger, more productive firms tend to be formal and therefore face heavier regulatory costs. They underproduce relative to the social optimum. Smaller, less productive firms tend to be informal and face fewer distortions. They overproduce relative to what is efficient. The result is systematic misallocation of resources: too much labor and output in low-productivity informal firms, too little in high-productivity formal ones.</p><p>Standard trade theory does not account for this. In a Melitz (2003) world, trade liberalization reallocates resources from less productive to more productive firms, generating well-known efficiency gains. But when informality creates an additional layer of distortion, the reallocation channel becomes far more consequential, and the standard models undercount the gains.</p>
      <p>
          <a href="/__u/carloschavezp29.substack.com/p/what-trade-theory-misses-about-developing">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[What Does "Identified" Mean in Macro?]]></title><description><![CDATA[The State of the Art and Why This Progress Is Underappreciated]]></description><link>https://carloschavezp29.substack.com/p/what-does-identified-mean-in-macro</link><guid isPermaLink="false">https://carloschavezp29.substack.com/p/what-does-identified-mean-in-macro</guid><dc:creator><![CDATA[Carlos Chavez]]></dc:creator><pubDate>Mon, 16 Feb 2026 12:59:17 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!923T!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f32769b-2ab8-4d90-a392-f6f42eac2108_895x884.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><a href="/__u/substack.com/@carloschavezp29/p-183084772">Last month I argued that macroeconomics never had a credibility revolution</a> &#8211; not because macroeconomists lacked rigor, but because the questions they inherited did not permit the research designs that made micro&#8217;s revolution possible. You cannot randomize monetary policy. There is no control group for the 2008 financial crisis. The economy is the unit of observation.</p><p>That piece ended with convergence. Applied micro, as its ambitions have grown, is rediscovering what macro and structural IO always knew: local estimates do not automatically aggregate, and structure is unavoidable when the question demands it.</p><p>This piece picks up where that one left off. If macro cannot import micro&#8217;s toolkit wholesale, what does rigorous identification <em>actually</em> look like in macroeconomics? The answer, I will argue, is more sophisticated and more coherent than the standard narrative suggests. Macro&#8217;s constraints &#8211; the impossibility of randomization, the unavoidability of general equilibrium, the necessity of forward-looking agents &#8211; forced a different kind of methodological creativity. And the problems macro was forced to confront decades ago &#8211; aggregation, counterfactual policy evaluation, the discipline of structure &#8211; are precisely the problems that applied micro is now encountering as its ambitions grow.</p><h2>Identification Is Always Relative to a Question</h2><p>The first confusion to clear away is that &#8220;identification&#8221; is not a single thing. When a macroeconomist says a model is identified or a shock is identified, they could mean any of several quite different claims &#8211; and the identifying assumptions required differ dramatically across them.</p><p>Consider the simplest representation of a macroeconomic data-generating process. Observable macro variables follow a structural moving average:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\ny_t = \\sum_{\\ell=0}^{\\infty} \\Theta_\\ell \\, \\varepsilon_{t-\\ell}\n&quot;,&quot;id&quot;:&quot;QLEYRRLLQV&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here y_t   is a vector of observables &#8211; output, inflation, interest rates &#8211; and  &#949;_t&#8203; is a vector of structural shocks &#8211; monetary, fiscal, technology, demand &#8211; that the econometrician does not observe. The matrices &#920;&#8203; capture how each shock propagates through the economy at each horizon.</p><p>Every empirical method in macroeconomics is trying to recover some feature of the &#920; matrices or the &#949; shocks. But different questions require recovering different features, and the assumptions needed escalate as you move from simpler to more ambitious objects.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!923T!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f32769b-2ab8-4d90-a392-f6f42eac2108_895x884.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!923T!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f32769b-2ab8-4d90-a392-f6f42eac2108_895x884.png 424w, /__u/substackcdn.com/image/fetch/$s_!923T!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f32769b-2ab8-4d90-a392-f6f42eac2108_895x884.png 848w, /__u/substackcdn.com/image/fetch/$s_!923T!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f32769b-2ab8-4d90-a392-f6f42eac2108_895x884.png 1272w, /__u/substackcdn.com/image/fetch/$s_!923T!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f32769b-2ab8-4d90-a392-f6f42eac2108_895x884.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!923T!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f32769b-2ab8-4d90-a392-f6f42eac2108_895x884.png" width="727" height="718.0648044692738" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1f32769b-2ab8-4d90-a392-f6f42eac2108_895x884.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:884,&quot;width&quot;:895,&quot;resizeWidth&quot;:727,&quot;bytes&quot;:174522,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/188107069?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f32769b-2ab8-4d90-a392-f6f42eac2108_895x884.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!923T!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f32769b-2ab8-4d90-a392-f6f42eac2108_895x884.png 424w, /__u/substackcdn.com/image/fetch/$s_!923T!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f32769b-2ab8-4d90-a392-f6f42eac2108_895x884.png 848w, /__u/substackcdn.com/image/fetch/$s_!923T!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f32769b-2ab8-4d90-a392-f6f42eac2108_895x884.png 1272w, /__u/substackcdn.com/image/fetch/$s_!923T!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f32769b-2ab8-4d90-a392-f6f42eac2108_895x884.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p><em> </em>At Stage I, you estimate the reduced-form dynamics &#8211; a VAR or local projection. This requires no structural assumptions at all, only that you have enough data and the right lag structure. At Stage II, you try to isolate a particular structural shock &#8211; a monetary policy shock, say. This requires identifying assumptions: timing restrictions, narrative classification, external instruments. At Stage III, you trace out impulse responses &#8211; how variables respond to that shock over time. At Stage IV, you ask how much of the variation in the data is driven by that shock, which requires stronger assumptions still. And at Stage V, you construct counterfactuals: what would have happened under a different policy <em>rule</em>? This is the question policymakers actually care about, and it requires the most structure &#8211; but, as we will see, far less structure than was previously believed.</p><p>Much confusion in empirical macro arises from failing to distinguish between these stages. A researcher who has cleanly identified a monetary policy shock (Stage II) and estimated its impulse response (Stage III) has accomplished something valuable. But they have not yet answered the question of how much monetary policy matters for inflation (Stage IV), let alone what alternative policy would have been optimal (Stage V). Conversely, a fully specified DSGE model that purports to answer Stage V questions is only as credible as its identifying assumptions at Stage II.</p><p>The field&#8217;s progress over the past decade has been to clarify these distinctions and to develop methods appropriate to each stage &#8211; methods that are explicit about what they assume and honest about what they cannot deliver.</p><h2>The Ground Is Clearer Than It Looks</h2><p>Before turning to the frontier problems, two results have simplified the methodological landscape considerably.</p><p>The first is Plagborg-M&#248;ller and Wolf&#8217;s 2021 proof in <em>Econometrica</em> that local projections and vector autoregressions estimate the same impulse responses. This is not a technical footnote. For two decades, applied macroeconomists debated whether LPs or VARs were the right approach, with Jord&#224;&#8217;s LP advocates emphasizing robustness to misspecification and VAR practitioners emphasizing efficiency. The debate sometimes felt like a paradigm choice.</p><p>It was not. LPs and VARs are different finite-sample estimators of the same population object. The proof requires only unrestricted lag structures &#8211; no parametric assumptions. What seemed like a methodological divide is a bias-variance tradeoff. VARs impose more structure and gain efficiency; LPs impose less and gain robustness. In their simulation study across thousands of data-generating processes, Li, Plagborg-M&#248;ller, and Wolf (2024) found that shrinkage methods &#8211; Bayesian VARs, penalized LPs &#8211; tend to dominate at the intermediate and long horizons that matter most for policy.</p><p>The LP vs. VAR debate is over. The debate that matters is upstream: <em>what shock are you identifying, and are your identifying assumptions defensible?</em></p><p>The second clarifying result concerns the external instruments &#8211; or proxy SVARs &#8211; that have become the leading identification approach in empirical macro. Introduced by Mertens and Ravn (2013) and systematized by Stock and Watson (2018), the method was further generalized by Miranda-Agrippino and Ricco (2023), who relaxed the identification conditions. The standard proxy SVAR identifies the effects of a structural shock &#949;_1,t&#8203; using an external instrument z_t  &#8203; that satisfies: </p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;E[z_t \\, \\varepsilon_{1,t}] \\neq 0 \\quad \\text{and} \\quad E[z_t \\, \\varepsilon_{j,t}] = 0 \\quad \\text{for } j \\neq 1\n&quot;,&quot;id&quot;:&quot;YHZTKYXRDJ&quot;}" data-component-name="LatexBlockToDOM"></div><p>Relevance and exogeneity &#8211; the same conditions as instrumental variables in micro. This is not a coincidence; it <em>is</em> instrumental variables, applied to time series. The contribution of Miranda-Agrippino and Ricco was to show that identification requires conditions only on the shock of interest and its relationship to the instrument, not on all the shocks in the system. This is a useful generalization because in practice, the macroeconomist has a candidate instrument for one shock &#8211; a monetary policy surprise, a narrative fiscal shock &#8211; and wants to learn about that shock's effects without taking a stand on everything else.</p><h2>The Information Problem</h2><p>Now we turn to where the frontier is genuinely unsettled. The leading application of high-frequency identification in macro uses interest rate movements in narrow windows around FOMC announcements as instruments for monetary policy shocks. In a 30-minute window, output and employment do not change, so whatever moves interest rates in that window is plausibly monetary.</p><p>This logic, pioneered by Kuttner (2001) and extended by G&#252;rkaynak, Sack, and Swanson (2005), and refined by Gertler and Karadi (2015), seemed to provide macro with something close to an experiment. But three separate literatures have identified a problem &#8211; and they disagree about what it is.</p><p>The problem shows up in a simple fact: monetary policy surprises are predictable. Bauer and Swanson (2023) documented that standard surprises around FOMC announcements can be predicted from macroeconomic and financial data available <em>before</em> the announcement, with R-squared values of 10 to 40 percent. If the surprises are predictable, they are not fully exogenous.</p><p>Why are they predictable? Here the three camps diverge.</p><p><strong>The Fed information effect.</strong> Nakamura and Steinsson (2018) documented that when the Fed raises rates unexpectedly, markets sometimes revise <em>upward</em> their expectations for growth. If the rate hike conveyed bad news about the economy, expectations should fall. The fact that they rise suggests the announcement conveys good news about fundamentals &#8211; the Fed knows something the market does not, and by raising rates, it reveals that news. The monetary policy surprise is contaminated by the information the action conveys.</p><p><strong>Informational rigidities.</strong> Miranda-Agrippino and Ricco (2021) proposed that the problem arises because the Fed and the private sector have different information sets, and the private sector updates sluggishly. They showed that projecting monetary surprises onto the Fed&#8217;s internal Greenbook forecasts removes the contamination. Their instrument produces textbook-clean impulse responses: a contractionary monetary shock lowers output and inflation without any price puzzle.</p><p><strong>The Fed response to news.</strong> Bauer and Swanson (2023) offered a third explanation. The predictability of monetary surprises does not reflect the Fed revealing private information; it reflects markets systematically underestimating how strongly the Fed responds to publicly available economic news. When strong employment numbers arrive, the market adjusts its rate expectations, but not by enough &#8211; so the next FOMC announcement delivers a hawkish surprise that is correlated with the earlier employment data. The solution: orthogonalize the surprises against pre-announcement macroeconomic and financial data.</p><p>Bauer and Swanson&#8217;s solution is notable for two additional innovations. They expand the set of monetary policy events to include speeches by the Fed Chair &#8211; which, as Swanson and Jayawickrema (2024) showed, are actually more important for financial markets than FOMC announcements themselves. And their orthogonalized, expanded surprise series produces macroeconomic effects that are <em>larger and more statistically significant</em> than previous high-frequency estimates &#8211; resolving the persistent concern that monetary policy shocks seemed too small to matter.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Dyzs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19ba5a7e-c654-444a-b39f-aaf8a23efcfe_820x758.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Dyzs!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19ba5a7e-c654-444a-b39f-aaf8a23efcfe_820x758.png 424w, /__u/substackcdn.com/image/fetch/$s_!Dyzs!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19ba5a7e-c654-444a-b39f-aaf8a23efcfe_820x758.png 848w, /__u/substackcdn.com/image/fetch/$s_!Dyzs!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19ba5a7e-c654-444a-b39f-aaf8a23efcfe_820x758.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Dyzs!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19ba5a7e-c654-444a-b39f-aaf8a23efcfe_820x758.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Dyzs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19ba5a7e-c654-444a-b39f-aaf8a23efcfe_820x758.png" width="820" height="758" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/19ba5a7e-c654-444a-b39f-aaf8a23efcfe_820x758.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:758,&quot;width&quot;:820,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:153356,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/188107069?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19ba5a7e-c654-444a-b39f-aaf8a23efcfe_820x758.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Dyzs!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19ba5a7e-c654-444a-b39f-aaf8a23efcfe_820x758.png 424w, /__u/substackcdn.com/image/fetch/$s_!Dyzs!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19ba5a7e-c654-444a-b39f-aaf8a23efcfe_820x758.png 848w, /__u/substackcdn.com/image/fetch/$s_!Dyzs!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19ba5a7e-c654-444a-b39f-aaf8a23efcfe_820x758.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Dyzs!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19ba5a7e-c654-444a-b39f-aaf8a23efcfe_820x758.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The figure above illustrates the stakes. The same conceptual exercise &#8211; estimate the effect of a contractionary monetary policy shock on output and inflation &#8211; yields qualitatively different answers depending on which identification approach you use. Gertler and Karadi&#8217;s raw surprises produce a price puzzle in inflation (prices <em>rise</em> initially after a tightening). Miranda-Agrippino and Ricco&#8217;s information-purged instrument eliminates the puzzle. Bauer and Swanson&#8217;s orthogonalized series produces the largest effects.</p><p>The choice of how you separate the policy shock from the information it conveys <em>changes your estimate of how monetary policy works.</em> The three competing stories &#8211; information effects, informational rigidities, Fed response to news &#8211; imply different things about what central bank communication does and how transparent the Fed should be. The identification question and the economic question are the same question.</p><h2>What the Econometrician Cannot See</h2><p>A deeper problem lurks behind the information debate. Standard structural VAR analysis assumes <em>invertibility</em>: that the econometrician can recover the structural shocks from the observed data. Formally, it assumes that the history of reduced-form residuals spans the history of structural shocks &#8211; that nothing important is hidden.</p><p>This is a strong assumption. If agents in the economy observe signals about future policy, future productivity, or future demand that the econometrician does not observe &#8211; news, announcements, private information &#8211; then the structural shocks are <em>not</em> recoverable from the data. The VAR is non-invertible.</p><p>Plagborg-M&#248;ller and Wolf (2022) showed what happens to identification under non-invertibility. The good news: impulse responses remain identified when you have a valid external instrument, even if the shock is non-invertible. The LP-IV estimator is robust to this problem. The bad news: variance decompositions &#8211; <em>how much of the fluctuation in a variable is due to a particular shock</em> &#8211; become only interval-identified. You get bounds, not point estimates.</p><p>But the bounds can be informative, and the paper&#8217;s application to monetary policy produced a striking result.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Jkze!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e8cfdf9-ff41-4f35-a112-4f0352a88c5e_835x825.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Jkze!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e8cfdf9-ff41-4f35-a112-4f0352a88c5e_835x825.png 424w, /__u/substackcdn.com/image/fetch/$s_!Jkze!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e8cfdf9-ff41-4f35-a112-4f0352a88c5e_835x825.png 848w, /__u/substackcdn.com/image/fetch/$s_!Jkze!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e8cfdf9-ff41-4f35-a112-4f0352a88c5e_835x825.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Jkze!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e8cfdf9-ff41-4f35-a112-4f0352a88c5e_835x825.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Jkze!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e8cfdf9-ff41-4f35-a112-4f0352a88c5e_835x825.png" width="835" height="825" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6e8cfdf9-ff41-4f35-a112-4f0352a88c5e_835x825.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:825,&quot;width&quot;:835,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:149046,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/188107069?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e8cfdf9-ff41-4f35-a112-4f0352a88c5e_835x825.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Jkze!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e8cfdf9-ff41-4f35-a112-4f0352a88c5e_835x825.png 424w, /__u/substackcdn.com/image/fetch/$s_!Jkze!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e8cfdf9-ff41-4f35-a112-4f0352a88c5e_835x825.png 848w, /__u/substackcdn.com/image/fetch/$s_!Jkze!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e8cfdf9-ff41-4f35-a112-4f0352a88c5e_835x825.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Jkze!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e8cfdf9-ff41-4f35-a112-4f0352a88c5e_835x825.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In U.S. data from 1990 onward, the robust confidence intervals for the forecast variance contribution of monetary policy shocks to inflation are remarkably tight &#8211; ruling out contributions above single-digit percentages at essentially all horizons. Meanwhile, traditional SVAR analysis that assumes invertibility estimates substantially higher contributions.</p><p>The implication is worth sitting with. If monetary <em>shocks</em> &#8211; the erratic, surprise component of policy &#8211; account for only a small fraction of inflation variation, then the action is somewhere else. It is in the <em>systematic</em> component: the policy rule itself. The Taylor rule, or whatever framework the central bank follows, matters far more than the deviations from it. The interest rate path that the Fed lays down quarter after quarter, responding to inflation and output in predictable ways &#8211; that is what shapes inflation dynamics. The 30-minute surprises around FOMC announcements, which have absorbed so much of the profession&#8217;s identification energy, are capturing something real but something small.</p><p>This reorients the entire research agenda. If the rule is the object, then the question is not &#8220;what is the effect of a monetary shock?&#8221; but &#8220;what would have happened under a different rule?&#8221; The high-frequency identification literature &#8211; Stages II and III in our pipeline &#8211; is valuable as a diagnostic tool, as a source of identified moments that discipline models. But the policy question lives at Stage V. And Stage V requires counterfactual analysis of rules, not shocks.</p><h2>Just Enough Structure: The Sufficient Statistics Revolution</h2><p>The Lucas critique &#8211; the observation that private behavior changes when policy rules change, so that reduced-form relationships estimated under one regime do not hold under another &#8211; has haunted macroeconomic policy analysis since 1976. The standard response has been: estimate a fully specified structural model with &#8220;deep&#8221; parameters that remain invariant to policy changes. Hence the DSGE industry.</p><p>But fully specified structural models are expensive. They require functional form assumptions, calibration choices, and an enormous number of parameters, many of which are weakly identified. And the results are often fragile to specification choices that papers do not emphasize.</p><p>The breakthrough of the past few years has been to show that you can do policy counterfactuals &#8211; counterfactuals that are robust to the Lucas critique &#8211; with far less structure than a full DSGE model. Three papers, published almost simultaneously, establish this from different angles.</p><p><strong>Beraja (2023, JPE): Counterfactual equivalence.</strong> Consider three different New Keynesian models &#8211; a textbook version, a behavioral variant with cognitive discounting (&#224; la Gabaix), and a working capital version where firms borrow to produce. All three can be calibrated to match the same data under the current monetary policy rule. They are observationally equivalent. Now change the rule. The textbook and behavioral models generate <em>identical</em> counterfactual predictions, despite their different microfoundations. The working capital model generates a different prediction.</p><p>Beraja&#8217;s key theorem formalizes this. To identify a policy counterfactual, you do not need to identify all parameters of a complete structural model. You need enough restrictions to pin down the objects that <em>matter</em> for the counterfactual &#8211; but no more. The set of models you must distinguish between is smaller than the Lucas critique suggests.</p><p><strong>McKay and Wolf (2023, Econometrica): Impulse responses as sufficient statistics.</strong> In any linearized macroeconomic model where agents care about the expected path of the policy instrument &#8211; not whether that path comes from the rule or from shocks &#8211; knowledge of the causal effects of policy shocks (both contemporaneous and anticipated) is sufficient to construct counterfactuals under alternative policy rules. No structural model is required. The counterfactuals are, by construction, robust to the Lucas critique.</p><p>The linearity requirement is not trivial &#8211; it rules out certain forms of state dependence and nonlinearity at the zero lower bound, though the policy rule itself can be nonlinear. But the class of models nested in this framework is broad: representative-agent and heterogeneous-agent New Keynesian models, behavioral variants with cognitive discounting, and many of the workhorse specifications in current use.</p><p>The intuition is powerful. If you know how the economy responds to a 25-basis-point surprise rate hike, and how it responds to an anticipated rate hike six months ahead, you can reconstruct the economy&#8217;s response to <em>any</em> alternative path of interest rates &#8211; and therefore to any alternative policy rule that implies a different path.</p><p><strong>Barnichon and Mesters (2023, AER): Sufficient statistics for macro policy.</strong> Two statistics are sufficient to detect and correct suboptimal macroeconomic policy: (i) forecasts for the policy objectives conditional on the current policy choice, and (ii) impulse responses of those objectives to policy shocks. If these two objects are &#8220;orthogonal&#8221; &#8211; if you cannot use the impulse responses to adjust the forecasts and lower the loss function &#8211; then policy is optimal. If they are not orthogonal, the gap tells you <em>how</em> and <em>by how much</em> to adjust.</p><p>These are the macro analogues of Chetty&#8217;s (2009) sufficient statistics for welfare analysis. Just as you do not need to estimate the entire demand system to evaluate a tax change &#8211; only the elasticities that enter the welfare formula &#8211; you do not need a complete structural model to evaluate monetary policy. You need forecasts and impulse responses. Both can be estimated from the data.</p><p>Taken together, these three papers constitute what I would call a sufficient statistics revolution in macroeconomics. The parallel to micro is exact:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!WeH4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd9e3665-87ed-4f1b-8414-b00f9007f46b_680x261.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!WeH4!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd9e3665-87ed-4f1b-8414-b00f9007f46b_680x261.png 424w, /__u/substackcdn.com/image/fetch/$s_!WeH4!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd9e3665-87ed-4f1b-8414-b00f9007f46b_680x261.png 848w, /__u/substackcdn.com/image/fetch/$s_!WeH4!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd9e3665-87ed-4f1b-8414-b00f9007f46b_680x261.png 1272w, /__u/substackcdn.com/image/fetch/$s_!WeH4!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd9e3665-87ed-4f1b-8414-b00f9007f46b_680x261.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!WeH4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd9e3665-87ed-4f1b-8414-b00f9007f46b_680x261.png" width="680" height="261" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dd9e3665-87ed-4f1b-8414-b00f9007f46b_680x261.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:261,&quot;width&quot;:680,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:38140,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/188107069?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd9e3665-87ed-4f1b-8414-b00f9007f46b_680x261.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!WeH4!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd9e3665-87ed-4f1b-8414-b00f9007f46b_680x261.png 424w, /__u/substackcdn.com/image/fetch/$s_!WeH4!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd9e3665-87ed-4f1b-8414-b00f9007f46b_680x261.png 848w, /__u/substackcdn.com/image/fetch/$s_!WeH4!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd9e3665-87ed-4f1b-8414-b00f9007f46b_680x261.png 1272w, /__u/substackcdn.com/image/fetch/$s_!WeH4!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd9e3665-87ed-4f1b-8414-b00f9007f46b_680x261.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The difference is that in micro, this idea has been around since 2009 and is now widely internalized. In macro, the formalization is happening now &#8211; and it is, in some respects, more ambitious, because it must handle forward-looking agents, general equilibrium feedback, and the Lucas critique.</p><h2>What Has Changed</h2><p>Step back and consider the shape of the field now versus a decade ago.</p><p>The estimation debate has largely been resolved. LPs and VARs estimate the same object. The choice between them is a finite-sample question about bias and variance, not a philosophical one. Shrinkage helps. This is settled.</p><p>The identification debate has not been resolved &#8211; but it has been <em>clarified</em>. The information problem in monetary policy is now understood as a substantive question about what central bank communication conveys, not a technical nuisance to be corrected and forgotten. The three competing diagnoses &#8211; information effects, informational rigidities, Fed response to news &#8211; imply different things about central bank transparency, about the nature of expectations formation, and about the size of monetary transmission. A paper that reports results under only one shock measure is implicitly taking a position on all of this, whether the authors realize it or not.</p><p>The invertibility problem has moved from a theoretical curiosity to an empirical constraint. The data are consistent with substantial non-invertibility in the monetary policy setting. This does not invalidate impulse response analysis &#8211; LP-IV is robust &#8211; but it means that variance decompositions from traditional SVARs carry an assumption that the data may not support. The gap between the invertibility-assuming and invertibility-robust estimates of monetary shock importance is not small.</p><p>And the counterfactual problem &#8211; the Lucas critique, the reason macroeconomists built DSGE models in the first place &#8211; turns out to require less structure than four decades of practice suggested. The sufficient statistics program shows that impulse responses to policy shocks, combined with forecasts, can do the work that was thought to require a fully specified model. The structure is still there, but it lives in the identifying assumptions at Stages II and III rather than in the parametric choices of a DSGE at Stage V. The discipline has shifted from &#8220;specify the right model&#8221; to &#8220;identify the right moments and use them carefully.&#8221;</p><p>This is a genuine intellectual achievement, and it has happened quietly. No Nobel lecture announced it. No JEP symposium declared a revolution. But the distance between the macro identification toolkit of 2015 and the one available today is substantial &#8211; and the direction of travel is toward more credibility, not less.</p><h2>What Remains</h2><p>The standard narrative about macroeconomics &#8211; that it failed to participate in the credibility revolution and remains stuck in a world of untestable structural assumptions &#8211; is not just incomplete. It is increasingly outdated.</p><p>Over the past decade, macroeconomists have developed a coherent identification program that moves from clean shock identification (high-frequency, narrative, proxy SVARs) through robustly estimated impulse responses (LP-IV, Bayesian methods) to policy counterfactuals that require only the structure needed to answer the question at hand (semistructural methods, sufficient statistics).</p><p>This program owes a debt to micro. The credibility revolution&#8217;s insistence on transparency &#8211; show your assumptions, defend your exclusion restriction, be honest about what you cannot identify &#8211; disciplined the entire profession, macro included. The narrative and high-frequency identification literatures are micro&#8217;s logic applied to macro&#8217;s questions. That debt is real and should be acknowledged.</p><p>But the traffic now runs in both directions. On counterfactual policy evaluation &#8211; the question of what would happen under a rule that has never been tried &#8211; macro currently has a more coherent toolkit than micro needs or possesses. The sufficient statistics approach to policy evaluation, the semistructural counterfactual methods, the partial identification results under non-invertibility: these address a class of problems that micro has not faced in the same form. As micro takes on more ambitious questions &#8211; national policies, long-run effects, general equilibrium responses &#8211; it is encountering precisely the aggregation and counterfactual challenges that macro has been wrestling with for decades. The DiD crisis was a preview. The need for structure will not go away.</p><p>None of this means macro has solved its identification problems. The information problem in high-frequency identification remains unresolved. Non-invertibility may be pervasive. Weak identification in structural models is widespread. The frontier is genuinely difficult.</p><p>But the trajectory is unmistakable. The field has moved from the false confidence of the 1970s, through the necessary nihilism of the Sims critique, to a mature synthesis that is explicit about assumptions, honest about limits, and increasingly powerful in what it can deliver.</p><p>Macro did not have a credibility revolution because it never had the luxury of avoiding structure. It had to learn how to discipline it. That discipline &#8211; born of constraint, not choice &#8211; is what the rest of the profession is now converging toward.</p><h2>A Question for Part III</h2><p>This piece has described what identification in macro looks like now &#8211; the toolkit, the unresolved tensions, the emerging synthesis. But it has not asked the deeper question: <em>What would it mean for macro to actually get identification right?</em></p><p>Suppose we could identify the aggregate fiscal multiplier &#8211; not the regional estimate that absorbs general equilibrium through time fixed effects, but the national multiplier under a specific monetary regime, varying with the state of the business cycle. Suppose we could trace the full causal chain from a change in the Fed&#8217;s policy rule to its effect on output, inflation, employment, and inequality, at every horizon, with known heterogeneity across sectors and households.</p><p>The model wars would end. The fiscal multiplier pins down price stickiness, the importance of demand shocks, the role of liquidity constraints. One well-identified number collapses a large parameter space. The Phillips curve debate &#8211; sixty years and counting &#8211; would be resolved not philosophically but empirically. Central banks could optimize in a meaningful sense, computing rate paths the way engineers compute trajectories, rather than averaging across dozens of structural models that disagree with each other.</p><p>But here is the thing. Macro probably <em>cannot</em> achieve full identification in the micro sense. The economy is one interconnected system where every intervention propagates through expectations, prices, wages, asset markets simultaneously. General equilibrium cannot be held constant the way a control group can. You can credibly identify local effects. But aggregate effects &#8211; the objects that actually matter for policy &#8211; would require running the same economy under different policies at the same time. We do not have parallel universes.</p><p>This is precisely why the sufficient statistics program matters. It represents the field accepting that full identification is unattainable for aggregate questions and asking instead: what is the <em>minimum</em> we need to identify to answer the policy question? The answer &#8211; impulse responses plus forecasts &#8211; is achievable. Not everything. But possibly enough.</p><p>The next piece in this series will take up that question directly. What would a fully identified macroeconomics look like? What would it deliver for policy? And why is the gap between what we have and what we would need not a counsel of despair but a map of the work that remains?</p><h2>Some Worth Readings</h2><p><strong>Methodological Unification</strong></p><p>Plagborg-M&#248;ller, M. &amp; Wolf, C. (2021). &#8220;Local Projections and VARs Estimate the Same Impulse Responses.&#8221; <em>Econometrica</em> 89(2): 955&#8211;980.</p><p>Li, D., Plagborg-M&#248;ller, M. &amp; Wolf, C. (2024). &#8220;Local Projections vs. VARs: Lessons From Thousands of DGPs.&#8221; <em>Journal of Econometrics</em> 244(2).</p><p><strong>External Instruments and Proxy SVARs</strong></p><p>Mertens, K. &amp; Ravn, M. (2013). &#8220;The Dynamic Effects of Personal and Corporate Income Tax Changes in the United States.&#8221; <em>American Economic Review</em> 103(4): 1212&#8211;1247.</p><p>Stock, J. &amp; Watson, M. (2018). &#8220;Identification and Estimation of Dynamic Causal Effects in Macroeconomics Using External Instruments.&#8221; <em>Economic Journal</em> 128: 917&#8211;948.</p><p>Miranda-Agrippino, S. &amp; Ricco, G. (2023). &#8220;Identification with External Instruments in Structural VARs.&#8221; <em>Journal of Monetary Economics</em> 135: 1&#8211;19.</p><p><strong>The Information Problem</strong></p><p>Nakamura, E. &amp; Steinsson, J. (2018). &#8220;High-Frequency Identification of Monetary Non-Neutrality: The Information Effect.&#8221; <em>Quarterly Journal of Economics</em> 133(3): 1283&#8211;1330.</p><p>Miranda-Agrippino, S. &amp; Ricco, G. (2021). &#8220;The Transmission of Monetary Policy Shocks.&#8221; <em>American Economic Journal: Macroeconomics</em> 13(3): 74&#8211;107.</p><p>Bauer, M. &amp; Swanson, E. (2023). &#8220;A Reassessment of Monetary Policy Surprises and High-Frequency Identification.&#8221; <em>NBER Macroeconomics Annual</em> 37: 87&#8211;155.</p><p>Swanson, E. &amp; Jayawickrema, V. (2024). &#8220;Speeches by the Fed Chair Are More Important Than FOMC Announcements: An Improved High-Frequency Measure of U.S. Monetary Policy Shocks.&#8221; Working Paper, University of California, Irvine.</p><p><strong>Non-Invertibility and Partial Identification</strong></p><p>Plagborg-M&#248;ller, M. &amp; Wolf, C. (2022). &#8220;Instrumental Variable Identification of Dynamic Variance Decompositions.&#8221; <em>Journal of Political Economy</em> 130(8): 2164&#8211;2202.</p><p><strong>The Sufficient Statistics Program</strong></p><p>Beraja, M. (2023). &#8220;A Semistructural Methodology for Policy Counterfactuals.&#8221; <em>Journal of Political Economy</em> 131(1): 190&#8211;201.</p><p>McKay, A. &amp; Wolf, C. (2023). &#8220;What Can Time-Series Regressions Tell Us About Policy Counterfactuals?&#8221; <em>Econometrica</em> 91(5): 1695&#8211;1725.</p><p>Barnichon, R. &amp; Mesters, G. (2023). &#8220;A Sufficient Statistics Approach for Macro Policy.&#8221; <em>American Economic Review</em> 113(11): 2809&#8211;2845.</p><p>Chetty, R. (2009). &#8220;Sufficient Statistics for Welfare Analysis: A Bridge Between Structural and Reduced-Form Methods.&#8221; <em>Annual Review of Economics</em> 1: 451&#8211;488.</p><p><strong>Background</strong></p><p>Kuttner, K. (2001). &#8220;Monetary Policy Surprises and Interest Rates: Evidence from the Fed Funds Futures Market.&#8221; <em>Journal of Monetary Economics</em> 47(3): 523&#8211;544.</p><p>G&#252;rkaynak, R., Sack, B. &amp; Swanson, E. (2005). &#8220;Do Actions Speak Louder Than Words? The Response of Asset Prices to Monetary Policy Actions and Statements.&#8221; <em>International Journal of Central Banking</em> 1(1): 55&#8211;93.</p><p>Gertler, M. &amp; Karadi, P. (2015). &#8220;Monetary Policy Surprises, Credit Costs, and Economic Activity.&#8221; <em>American Economic Journal: Macroeconomics</em> 7(1): 44&#8211;76.</p><p>Nakamura, E. &amp; Steinsson, J. (2018). &#8220;Identification in Macroeconomics.&#8221; <em>Journal of Economic Perspectives</em> 32(3): 59&#8211;86.</p><p>Ramey, V. (2016). &#8220;Macroeconomic Shocks and Their Propagation.&#8221; <em>Handbook of Macroeconomics</em> Vol. 2.</p><p>Fern&#225;ndez-Villaverde, J., Rubio-Ram&#237;rez, J. &amp; Schorfheide, F. (2016). &#8220;Solution and Estimation Methods for DSGE Models.&#8221; <em>Handbook of Macroeconomics</em> Vol. 2.</p>]]></content:encoded></item><item><title><![CDATA[AI Cannot Cook Your Dinner]]></title><description><![CDATA[On the Limits of Artificial Intelligence in Social Science Research]]></description><link>https://carloschavezp29.substack.com/p/ai-cannot-cook-your-dinner</link><guid isPermaLink="false">https://carloschavezp29.substack.com/p/ai-cannot-cook-your-dinner</guid><dc:creator><![CDATA[Carlos Chavez]]></dc:creator><pubDate>Fri, 13 Feb 2026 00:14:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!zh0V!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffff9b4a5-53ee-4548-85da-563173f13249_2666x2666.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>There is a robot in Germany that prepares 120 meals per hour inside a glass box at a supermarket. There is a humanoid in Shenzhen that can pick a cherry by its stem with two fingers, plate it beside toast, and pour milk without spilling a drop. There is a kitchen in London, built by Moley Robotics, that claims to replicate any recipe with sub-millisecond precision.</p><p>And yet none of these machines can walk into your kitchen after a long day, open your half-empty refrigerator, look at what is there, and decide what to make for dinner. Not because the technology is not advanced enough. Because the problem is fundamentally different from what the technology solves.</p><p>I have been thinking about this distinction, between executing a recipe and deciding what to cook, because the same confusion runs through the current conversation about artificial intelligence and social science. There is enormous excitement about what AI can do for economics, political science, and sociology. Some of it is justified. But much of it rests on the same error: mistaking execution for judgment, pattern recognition for understanding, and prediction for explanation.</p><p>This essay argues that AI in social science faces two ceilings. The first is soft, a set of tasks where AI is useful but less revolutionary than advertised. The second is hard, a set of problems where AI cannot help, not because the technology is immature, but because the problems require something that computation cannot provide. Understanding where each ceiling lies matters, because confusing the two leads to either irrational exuberance or unnecessary fear. Neither serves us well.</p><h2>The soft ceiling: What AI actually does for us</h2><p>The honest case for AI in social science is more modest than the headlines suggest, but it is real. AI accelerates work that used to be slow. It automates tasks that used to be tedious. It handles scale that used to be impossible.</p><p>Consider text classification. A political scientist studying legislative responses to social movements once had to read thousands of congressional bills, manually coding each one by topic, framing, and policy domain. This was valuable work, but it took months. A fine-tuned language model can now do the first pass in hours, with accuracy that rivals, and in some cases exceeds, trained human coders.</p><p>Consider literature synthesis. An economist exploring a new research question used to spend weeks reading papers, tracing citations, mapping the intellectual landscape. An LLM can compress this process. It can summarize findings, identify methodological patterns, and flag contradictions across hundreds of papers. The output is not perfect, but it is a starting point that previously required significant labor.</p><p>Consider data cleaning. Anyone who has worked with household survey data, ENAHO, LSMS, DHS, knows that harmonizing variables across years, handling missing values, and detecting coding errors consumes a disproportionate share of research time. AI tools can now assist with these tasks, identifying anomalies, suggesting recoding schemes, and flagging inconsistencies that a human researcher might miss on the first pass.</p><p>These are genuine productivity gains. They free researchers to spend more time thinking and less time processing. They lower the barrier to entry for empirical work. They make it possible for a single researcher to do what once required a team.</p><p>But notice what all these tasks have in common. They are tasks where the researcher already knows what needs to be done and is using AI to do it faster. The classification scheme exists; AI applies it. The literature exists; AI summarizes it. The data structure exists; AI cleans it. In each case, the intellectual work, defining the categories, choosing the question, designing the data architecture, has already been done by a human.</p><p>This is the soft ceiling. AI is a remarkably powerful assistant. It is not, in any meaningful sense, a researcher.</p><h2>The hard ceiling: Where AI cannot go</h2><p>The hard ceiling is not about computational limitations that will be overcome with better models. It is about the nature of social science itself, about what it means to ask a good question, to identify a causal effect, and to know why an answer matters.</p><h3>The question problem</h3><p>Every important piece of research begins with a question. Not a topic, not a dataset, not a method, a question. And the quality of the question determines everything that follows.</p><p>What would happen to caloric poverty in Peru if you lowered tariffs on food imports? Does exposure to political violence during childhood affect trust in democratic institutions thirty years later? Do cognitive skills and personality traits jointly determine economic preferences, or do they operate through independent channels?</p><p>These are not questions that emerge from patterns in data. They emerge from knowing a country, understanding a literature, recognizing a gap, and having the institutional knowledge to see that a particular source of variation might answer something important. They emerge from living in Barranca and reading trade theory and thinking about what happened to food prices when Peru liberalized. They emerge from conversations, years of conversations, about what matters and what is possible.</p><p>An LLM trained on the entire corpus of economics papers can tell you what questions have been asked. It cannot tell you what questions should be asked next. It can identify patterns in existing research. It cannot identify the absence,  the question that no one has thought to pose because it requires a combination of local knowledge, theoretical intuition, and disciplinary taste that no training corpus contains.</p><p>Susan Athey, in her October 2025 Distinguished Lecture at Georgetown&#8217;s Massive Data Institute, put this precisely: the hardest question is not how to predict, but how to measure. Deciding what to measure, what outcome matters, for whom, and why, is a judgment that requires understanding the world in a way that pattern recognition does not provide.</p><h3>The identification problem</h3><p>Suppose you have a question. Now you need to answer it. In economics, this means designing an identification strategy, finding a source of variation that allows you to isolate a causal effect from the tangle of confounding, reverse causation, and selection that characterizes observational data.</p><p>This is where the gap between AI and social science becomes sharpest. Identification requires what I would call institutional imagination, the ability to look at the world and see a natural experiment hiding in plain sight.</p><p>Joshua Angrist saw that quarter of birth, combined with compulsory schooling laws, created exogenous variation in years of education. David Card saw that the Mariel boatlift created a sudden, unexpected increase in labor supply in Miami but not in comparable cities. Melissa Dell saw that the historical boundary of the Mita forced labor system in colonial Peru created a discontinuity in development outcomes that persists to this day.</p><p>None of these identification strategies came from data. They came from understanding institutions, history, and geography, and from the creative leap of recognizing that these features of the world could be exploited to answer a causal question.</p><p>Can AI help here? Marginally. Given a treatment and an outcome, an LLM can propose candidate instruments by drawing on its training data. It might suggest &#8220;compulsory schooling laws&#8221; for education, or &#8220;distance to the coast&#8221; for trade exposure. But these suggestions are recombinations of what already exists in the literature. The instrument that Dell used, the boundary of a colonial labor institution, required something more precise than knowledge. An LLM trained on every history book ever written about Peru would contain the relevant facts: that the Mita system existed, that it had a geographic boundary, that colonial institutions shaped development. The bottleneck was not knowledge. It was the creative leap of seeing that two apparently unrelated facts, a boundary drawn by the Spanish crown in the sixteenth century and a pattern of underdevelopment in the twenty-first, could be connected through a regression discontinuity design. This is not retrieval. It is imaginative synthesis, the ability to see that an institutional artifact can serve as a source of exogenous variation. LLMs are trained to predict the next token in a sequence. Imaginative synthesis requires generating a connection that does not exist in any sequence the model has seen.</p><p>Judea Pearl, the computer scientist who developed the graphical approach to causal inference, has argued that the impressive achievements of deep learning amount to curve fitting. The key, he insists, is to move from reasoning by association to causal reasoning, the ability to infer causes from observed phenomena, to distinguish seeing from doing, to reason about counterfactuals. Pearl is right about the diagnosis. But his proposed solution, equipping machines with causal graphs, still requires someone to draw the graph. And drawing the graph correctly requires knowing which arrows exist and which do not, which is exactly the identification problem restated in different notation.</p><h3>The judgment problem</h3><p>Even after you have identified a causal effect, you face a question that no algorithm can answer: does this matter?</p><p>Statistical significance tells you whether an effect is distinguishable from zero. It does not tell you whether the effect is economically meaningful, policy-relevant, or important for human welfare. A coefficient of 0.03 on a variable measuring the impact of trade liberalization on caloric intake could be trivial or transformative, depending on the baseline, the population, and what else is happening in people&#8217;s lives.</p><p>This judgment: Is the effect large enough to care about? For whom? Compared to what?  Requires understanding the context in which the number lives. It requires knowing that a 3% increase in caloric intake means something different in a population near the starvation threshold than in one with adequate nutrition. It requires understanding the policy alternatives. It requires, in a word, wisdom, the kind that accumulates through years of engagement with a subject, not through training on text.</p>
      <p>
          <a href="/__u/carloschavezp29.substack.com/p/ai-cannot-cook-your-dinner">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Why Economists Linearize Everything]]></title><description><![CDATA[And What We Win and Lose When We Do]]></description><link>https://carloschavezp29.substack.com/p/why-economists-linearize-everything</link><guid isPermaLink="false">https://carloschavezp29.substack.com/p/why-economists-linearize-everything</guid><dc:creator><![CDATA[Carlos Chavez]]></dc:creator><pubDate>Mon, 09 Feb 2026 09:59:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Fk8o!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8315da6-a901-4a2c-91a9-cc91e68e4f70_895x776.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In 2012, Melissa Dell, Benjamin Jones, and Benjamin Olken published an influential paper on temperature and economic growth. It was a landmark contribution: the first rigorous panel study to use within-country temperature fluctuations to identify causal effects on aggregate output, and it opened entirely new channels, showing that temperature affects not just agriculture but industrial output, political stability, and capital accumulation. They ran a linear regression, GDP growth on temperature shocks, with country fixed effects and the standard controls. They even tested for nonlinear effects, but found little evidence at the macro level. Crucially, though, their tests were conducted within a linear framework: they checked whether the linear slope differed across subgroups, hot versus cold countries, rich versus poor, but never asked whether the relationship itself was curved. For rich countries clustered near the peak of what would turn out to be an inverted-U, the positive effects just below the optimum and the negative effects just above it cancelled within the group, producing a flat slope that looked like no effect. The result was clear: higher temperatures substantially reduce growth in poor countries, but have little effect in rich ones. The coefficient for wealthy nations was small, statistically indistinguishable from zero. Temperature, it seemed, was a problem for the developing world.</p><p>Three years later, Marshall Burke, Solomon Hsiang, and Edward Miguel looked at the same type of data and asked a different question. Instead of forcing a straight line through the relationship, they allowed it to curve. They fit a quadratic. What emerged was a striking inverted-U: economic productivity peaks at around 13&#176;C and declines sharply at both extremes. The effect was large, statistically powerful, and universal across rich and poor countries alike. The linear result had not been evidence of &#8220;no effect&#8221; in rich countries. The positive effects of warming at one end of the distribution and the negative effects at the other had cancelled in the average. The linear model had not summarized the relationship. It had erased it. This is not a criticism of Dell, Jones, and Olken, whose paper remains essential reading. It is a demonstration of how much functional form can matter, even when the identification strategy is sound and the researchers are careful.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Fk8o!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8315da6-a901-4a2c-91a9-cc91e68e4f70_895x776.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Fk8o!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8315da6-a901-4a2c-91a9-cc91e68e4f70_895x776.png 424w, /__u/substackcdn.com/image/fetch/$s_!Fk8o!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8315da6-a901-4a2c-91a9-cc91e68e4f70_895x776.png 848w, /__u/substackcdn.com/image/fetch/$s_!Fk8o!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8315da6-a901-4a2c-91a9-cc91e68e4f70_895x776.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Fk8o!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8315da6-a901-4a2c-91a9-cc91e68e4f70_895x776.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Fk8o!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8315da6-a901-4a2c-91a9-cc91e68e4f70_895x776.png" width="895" height="776" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f8315da6-a901-4a2c-91a9-cc91e68e4f70_895x776.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:776,&quot;width&quot;:895,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:269919,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/187371804?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8315da6-a901-4a2c-91a9-cc91e68e4f70_895x776.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Fk8o!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8315da6-a901-4a2c-91a9-cc91e68e4f70_895x776.png 424w, /__u/substackcdn.com/image/fetch/$s_!Fk8o!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8315da6-a901-4a2c-91a9-cc91e68e4f70_895x776.png 848w, /__u/substackcdn.com/image/fetch/$s_!Fk8o!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8315da6-a901-4a2c-91a9-cc91e68e4f70_895x776.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Fk8o!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8315da6-a901-4a2c-91a9-cc91e68e4f70_895x776.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The story of temperature and GDP is dramatic, but it is not special. It is what happens, in subtler ways, every time an economist assumes that the relationship between X and Y is a straight line. And that assumption is made almost every time. Before the identification strategy, before the robustness checks, before the standard errors are clustered, linearity is already baked in. The returns to an additional year of education? A constant percentage, whether you&#8217;re going from year 4 to year 5 or from year 15 to year 16. The effect of a minimum wage increase on employment? The same whether the increase is from $7 to $8 or from $14 to $15.</p><p>None of these assumptions are likely to be true. We know this. We have known this for a long time. And yet linearity remains the default specification in virtually every field of economics - labor, development, trade, public finance, macroeconomics. The question is why.</p><p>The easy answer is convenience. Linear models are tractable. They have closed-form solutions. They satisfy the Gauss-Markov conditions. Software runs them instantly. But convenience is a description, not an explanation. Economists are not lazy people. We spend years learning sophisticated mathematics. We develop elaborate identification strategies. We agonize over the exclusion restriction in our instrumental variable. And then we assume the relationship is linear without a second thought.</p><p>The real answer is deeper, more interesting, and more consequential than most economists realize. It involves the history of how our discipline borrowed from physics, the mathematical properties that make linear models uniquely tractable, and a set of habits that survived long after the conditions that created them disappeared. Understanding why we linearize everything reveals something important about how economics works, and about where its methods break down</p><h2>The Physics Inheritance</h2><p>Economics did not develop its mathematical methods in isolation. The founders of mathematical economics drew heavily from classical mechanics, and linearity came with the package.</p><p>L&#233;on Walras, writing in the 1870s, modeled general equilibrium as a system of simultaneous equations, essentially the same mathematics that physicists used to describe systems of forces in balance. Irving Fisher, whose 1892 dissertation laid the groundwork for modern utility theory, earned Yale&#8217;s first PhD in economics, but wrote it under the supervision of the physicist Josiah Willard Gibbs, and the work was saturated with mechanical analogies between thermodynamics and economic equilibrium. Jan Tinbergen, the first Nobel laureate in economics and founder of econometric modeling, had studied physics under Paul Ehrenfest in Leiden. Ragnar Frisch, who shared that first Nobel with Tinbergen, explicitly modeled economic dynamics on mechanical oscillations.</p><p>The borrowing was not superficial. These scholars imported not just techniques but a worldview: that economic systems, like physical systems, could be described by deterministic equations relating observable quantities. And in classical mechanics, linear approximation is everywhere. Small oscillations around equilibrium are linear. Hooke&#8217;s law is linear. Superposition, the principle that effects from multiple causes simply add up, is a property of linear systems.</p><p>Philip Mirowski, in <em>More Heat than Light</em> (1989), documented this transfer in exhaustive detail. The neoclassical revolution in economics was, in his account, a wholesale adoption of the mathematics of mid-nineteenth-century physics, specifically, the formalism of constrained optimization borrowed from Lagrangian mechanics. Utility maximization is formally identical to energy minimization. The first-order conditions of the consumer&#8217;s problem are structurally the same as the equilibrium conditions of a physical system. The mathematics came with an implicit assumption: that the systems being modeled were well-behaved, smooth, and amenable to local linear analysis.</p><p>Not everyone followed the physicists. Alfred Marshall argued that &#8220;the Mecca of the economist lies in economic biology rather than in economic dynamics&#8221;, an economics built on organic metaphors of evolution and adaptation rather than mechanical equilibrium. A biological economics would have been far more hospitable to nonlinearity, heterogeneity, and path dependence. But biology offered no clean formalism, and the physics-based approach won precisely because its equations could be written down, and solved by linearizing them.</p><p>This inheritance shaped how economists thought about relationships in data. If the world is composed of additive, separable effects, then linear regression is not an approximation - it is the correct model. The econometrician&#8217;s job is to estimate the coefficients, not to question the functional form.</p><p>When Tinbergen built the first macroeconometric models in the 1930s, systems of linear equations meant to describe the Dutch and then the American economy, he was doing what a physicist would do: writing down equations of motion for an aggregate system. Keynes, in his famous 1939 review of Tinbergen&#8217;s work, objected to almost everything about the enterprise, including linearity. He called the assumption that all relationships are linear &#8220;a very drastic and usually improbable postulate... indeed, it is ridiculous.&#8221; But the deeper point is that Keynes&#8217;s objection was ignored. The assumption was so embedded in the method that the profession moved forward as if it were not an assumption at all.</p><p>The problem is that economies are not physical systems. Prices are not forces. Agents are not particles. The superposition principle does not hold when people respond strategically to each other&#8217;s behavior. Economies exhibit increasing returns, network effects, tipping points, and regime changes, all fundamentally nonlinear phenomena. But by the time these differences became clear, the linear framework was already entrenched in how economists were trained and how journals evaluated their work.</p><h2>Why Linearity Is Mathematically Seductive</h2><p>Put aside history for a moment. There are genuine mathematical reasons why linear models dominate, and understanding them helps explain why the habit persists even among economists who know better.</p><p>The first reason is the Gauss-Markov theorem. In its simplest form: among all unbiased linear estimators, OLS has the smallest variance. This is an extraordinary guarantee. You do not need to know the distribution of the errors. You do not need to worry about choosing among different estimation methods. If your model is linear and your errors satisfy a few conditions - zero mean, constant variance, uncorrelated with each other - then OLS is the best you can do.</p><p>The second reason is the algebra of expectations. The conditional expectation function E[Y|X] decomposes the relationship between any two variables into signal and noise. If this conditional expectation happens to be linear, then the OLS coefficient recovers it exactly. And even when the conditional expectation is nonlinear, OLS gives you the best linear approximation, the line that minimizes the mean squared prediction error. This is the best linear predictor (BLP), and it exists regardless of the true functional form.</p><p>This is why many econometricians are not troubled by linearity. They argue: I am not claiming the world is linear. I am estimating a weighted average derivative, the slope of the best linear approximation. If the true relationship is curved, my coefficient tells me the average slope. This is not wrong. But it conceals something important about <em>whose</em> average it is.</p><p>The linear coefficient &#946; in a regression of Y on X is not an unweighted average of the marginal effect of X on Y across the distribution. It places more weight on observations with extreme values of X, observations far from the mean, because those observations contribute more to the covariance between X and Y. If the marginal effect varies across the distribution, the linear coefficient overrepresents the effects at the tails and underrepresents the effects in the middle. This can be problematic if you care about the effect for the typical person rather than the person with an extreme value of the treatment.</p><p>This weighting problem is not a curiosity - it is the thread that connects most of what goes wrong when we impose linearity in causal inference. When we move from OLS to instrumental variables, the weights change, but the fundamental issue remains. Angrist and Krueger (1999) showed that the IV estimator does not recover the average effect of X on Y in the population. It recovers a weighted average of marginal effects, with weights determined by the distribution of the instrument&#8217;s effect on the endogenous variable. Change the instrument, change the weights, and you can get a different &#8220;effect&#8221;, not because the underlying relationship changed, but because the linear summary selected a different slice of it. The marginal treatment effect (MTE) framework, developed by Heckman and Vytlacil (2005), makes this fully explicit: different linear IV estimators using different instruments are all estimating different weighted averages of the same underlying <em>distribution</em> of effects. They look like they disagree about &#8220;the effect,&#8221; but they are answering different questions because they weight the heterogeneous effects differently. The linearity assumption, one coefficient summarizing the entire relationship, is what creates the illusion of a single answer. In all three cases - OLS, IV, and MTE - the problem is the same: a linear specification compresses a potentially rich, heterogeneous relationship into a single number, and the weights that determine that number are chosen by the mathematics, not by the researcher.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!20Rv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcec81c-d4af-42d9-919f-eaca96b5a84b_740x804.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!20Rv!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcec81c-d4af-42d9-919f-eaca96b5a84b_740x804.png 424w, /__u/substackcdn.com/image/fetch/$s_!20Rv!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcec81c-d4af-42d9-919f-eaca96b5a84b_740x804.png 848w, /__u/substackcdn.com/image/fetch/$s_!20Rv!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcec81c-d4af-42d9-919f-eaca96b5a84b_740x804.png 1272w, /__u/substackcdn.com/image/fetch/$s_!20Rv!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcec81c-d4af-42d9-919f-eaca96b5a84b_740x804.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!20Rv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcec81c-d4af-42d9-919f-eaca96b5a84b_740x804.png" width="740" height="804" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9bcec81c-d4af-42d9-919f-eaca96b5a84b_740x804.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:804,&quot;width&quot;:740,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:193764,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/187371804?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcec81c-d4af-42d9-919f-eaca96b5a84b_740x804.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!20Rv!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcec81c-d4af-42d9-919f-eaca96b5a84b_740x804.png 424w, /__u/substackcdn.com/image/fetch/$s_!20Rv!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcec81c-d4af-42d9-919f-eaca96b5a84b_740x804.png 848w, /__u/substackcdn.com/image/fetch/$s_!20Rv!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcec81c-d4af-42d9-919f-eaca96b5a84b_740x804.png 1272w, /__u/substackcdn.com/image/fetch/$s_!20Rv!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcec81c-d4af-42d9-919f-eaca96b5a84b_740x804.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>.The third reason linearity persists is computational. Before the 1990s, estimating nonlinear models on large datasets was either impossible or prohibitively expensive. Maximum likelihood for nonlinear specifications required iterative algorithms that sometimes failed to converge. Nonparametric methods required bandwidth choices that were poorly understood. Even something as simple as adding a quadratic term created collinearity problems that made standard errors unstable.</p><p>Linear regression, by contrast, requires nothing more than matrix inversion. For decades, this was the only thing computers could do reliably with economic data. The software was designed for linearity, the textbooks taught linearity, and the referees expected linearity. Researchers who wanted to estimate nonlinear models faced not just technical challenges but institutional resistance.</p><h2>The Taylor Expansion Defense</h2><p>The most sophisticated defense of linearity comes from approximation theory. Any smooth function can be approximated locally by a linear function, the first-order Taylor expansion, the tangent line at a point. If we are studying small perturbations around an equilibrium, the linear approximation is not just convenient; it is mathematically justified.</p><p>This argument has been enormously productive. In macroeconomics, log-linearization around a steady state is the standard technique for solving dynamic stochastic general equilibrium (DSGE) models. Take a complex, nonlinear system of equations - household optimization, firm pricing, monetary policy rules - replace each equation with its first-order approximation around the long-run equilibrium, and the resulting linear system can be solved analytically. The Taylor expansion is also the implicit justification for using logarithms. When we estimate log-linear models, log wages on years of education, log GDP on trade openness, we are implicitly assuming that the relationship between the levels is multiplicative, and the log transformation makes it additive. The coefficient &#946; in a log-linear model represents a constant percentage change, which is a particular form of nonlinearity in levels that becomes linear after transformation.</p><p>But the Taylor expansion defense has a critical limitation: it only works locally. Move far from the point of expansion, and the linear approximation can be arbitrarily bad. This matters in economics more than in physics, because the perturbations we study are often not small.</p><p>Consider the 2008 financial crisis. DSGE models, log-linearized around a steady state with functioning credit markets, had no mechanism to generate the kind of collapse that occurred. The economy moved so far from the steady state that the linear approximation was meaningless. As Fern&#225;ndez-Villaverde and Rubio-Ram&#237;rez have argued, the difference between first-order and higher-order perturbation methods matters precisely when the economy is in the region where interesting things happen: financial crises, liquidity traps, regime changes.</p><p>The zero lower bound on interest rates provides another illustration. In a linearized New Keynesian model, the central bank follows a Taylor rule: raise rates when inflation is high, lower them when it is low. But when the nominal interest rate hits zero, the Taylor rule ceases to operate. The constraint is a hard nonlinearity, a kink, that linearized models cannot capture. Fern&#225;ndez-Villaverde, Gordon, Guerr&#243;n-Quintana, and Rubio-Ram&#237;rez (2015) showed that solving the same model with global methods, which respect the nonlinearity of the zero lower bound, produces dramatically different predictions for the duration of zero-rate episodes, the effectiveness of forward guidance, and the dynamics of recovery.</p><p>Or consider asset pricing. Log-linearization of the consumption Euler equation eliminates precisely the higher-order moments - variance, skewness, kurtosis of consumption growth - that drive risk premia. The equity premium puzzle exists in part because the linearized model cannot generate enough risk aversion from the observed smoothness of aggregate consumption.</p><p>This is the fundamental irony. We linearize to make models tractable for policy analysis. But the situations where policy analysis matters most - recessions, crises, structural transformations - are exactly the situations where linearization fails.</p><h2>What Linearity Hides in Empirical Work</h2><p>The consequences of linearity in applied microeconomics are different from macro, but equally important. Here the issue is not approximation around a steady state - it is what happens when you estimate a single coefficient for a relationship that varies across the distribution.</p><p>Start with the most famous linear model in labor economics: the Mincer equation. Log wages are regressed on years of schooling and years of experience (plus experience squared, the one nonlinearity everyone allows). The coefficient on schooling is interpreted as the &#8220;return to education&#8221;: each additional year of school increases wages by, say, 8 percent.</p><p>But does anyone believe that going from 0 to 1 year of schooling has the same proportional effect as going from 15 to 16? That the return to the first year of a PhD program is the same as the return to the last? The Mincer equation imposes this by construction.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!cLH4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2286c1ce-5a61-4af2-b255-74a58f2ad5f4_875x687.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!cLH4!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2286c1ce-5a61-4af2-b255-74a58f2ad5f4_875x687.png 424w, /__u/substackcdn.com/image/fetch/$s_!cLH4!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2286c1ce-5a61-4af2-b255-74a58f2ad5f4_875x687.png 848w, /__u/substackcdn.com/image/fetch/$s_!cLH4!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2286c1ce-5a61-4af2-b255-74a58f2ad5f4_875x687.png 1272w, /__u/substackcdn.com/image/fetch/$s_!cLH4!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2286c1ce-5a61-4af2-b255-74a58f2ad5f4_875x687.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!cLH4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2286c1ce-5a61-4af2-b255-74a58f2ad5f4_875x687.png" width="875" height="687" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2286c1ce-5a61-4af2-b255-74a58f2ad5f4_875x687.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:687,&quot;width&quot;:875,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:295761,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/187371804?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2286c1ce-5a61-4af2-b255-74a58f2ad5f4_875x687.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!cLH4!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2286c1ce-5a61-4af2-b255-74a58f2ad5f4_875x687.png 424w, /__u/substackcdn.com/image/fetch/$s_!cLH4!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2286c1ce-5a61-4af2-b255-74a58f2ad5f4_875x687.png 848w, /__u/substackcdn.com/image/fetch/$s_!cLH4!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2286c1ce-5a61-4af2-b255-74a58f2ad5f4_875x687.png 1272w, /__u/substackcdn.com/image/fetch/$s_!cLH4!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2286c1ce-5a61-4af2-b255-74a58f2ad5f4_875x687.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>The figure above illustrates the point with simulated data. The dashed red line shows what OLS sees: a constant return of roughly 8 percent per additional year of schooling. The solid blue line, a nonparametric LOWESS fit on the same data, tells a different story. Returns are modest at low levels of education, accelerate sharply around the transition to college, and then flatten at the graduate level. The linear fit captures none of this. It overstates returns at the bottom, understates them in the middle, and overstates them again at the top. The &#8220;8 percent return to education&#8221; is an average that describes nobody&#8217;s actual experience. The true returns to education almost certainly vary by level - primary, secondary, tertiary - and by field, by country, by cohort, and by individual characteristics. Heckman, Lochner, and Todd (2006) showed that the classic Mincer specification actually fits the data quite poorly once you allow for more flexible functional forms, and that the &#8220;returns to education&#8221; depend heavily on how you model the relationship.</p><p>The same problem appears in treatment effect estimation, and it connects directly to the weighting issue I described in the previous section. When we estimate the effect of a program using instrumental variables, we get a Local Average Treatment Effect (LATE), the effect for compliers, the people whose treatment status was changed by the instrument. The standard IV estimator assumes this effect is constant across compliers. If the effect varies, the LATE is a weighted average that may not correspond to any individual&#8217;s actual experience, and as the MTE framework shows, changing the instrument changes the weights, producing a different &#8220;effect&#8221; from the same underlying heterogeneity.</p><p>Or consider difference-in-differences, the workhorse of causal inference in applied economics. The standard DiD estimator assumes that treatment effects are additive: the treatment shifts the outcome by a constant amount. The parallel trends assumption is about levels (or logs). But what if the treatment effect is proportional, larger in absolute terms for units with higher baseline outcomes? Then the parallel trends assumption might hold in logs but not in levels, or vice versa. The choice between linear and log specifications is not innocuous; it determines what &#8220;parallel trends&#8221; means and what the estimator recovers.</p><p>Roth and Sant&#8217;Anna (2023) have shown that many DiD applications are sensitive to functional form in ways that researchers rarely investigate. The same data can produce significant effects or null results depending on whether you model the outcome in levels, logs, or ranks. Linearity is doing real work in these papers, and it is work that is almost never scrutinized.</p><p>The two-way fixed effects estimator, until recently the default for DiD with staggered treatment adoption, makes this worse. Goodman-Bacon (2021) decomposed the TWFE estimator into a weighted average of all possible two-by-two DiD comparisons in the data. Some of these comparisons use already-treated units as controls, which is only valid under strict linearity assumptions about how treatment effects evolve over time. When treatment effects are heterogeneous or dynamic, the TWFE estimator can produce a negative coefficient even when the treatment effect is positive for every single unit. The correction methods proposed by Callaway and Sant&#8217;Anna (2021), Sun and Abraham (2021), and others address the weighting problem, but they still rely on functional form assumptions about how outcomes are generated. The linearity question runs deeper than most of the recent DiD literature acknowledges.</p><h2>The Nonlinear Effects That Disappear</h2><p>Perhaps the most troubling consequence of linearity is that it can make real effects invisible. I touched on this in <a href="/__u/substack.com/home/post/p-186694857">my essay on causality</a>, but it is worth developing further here.</p><p>If the true relationship between X and Y is U-shaped, or inverted U-shaped, then a linear regression will estimate a slope near zero. The positive effects at one end and the negative effects at the other end cancel out in the average. The researcher concludes that X has no effect on Y, when in fact X has strong effects that depend on where you are in the distribution.</p><p>The temperature-growth story from the opening is the starkest example, but the pattern is everywhere. The relationship between stress and performance (the Yerkes-Dodson curve) is inverted-U. The relationship between firm size and innovation is often found to be nonlinear. The relationship between income inequality and growth has produced decades of contradictory linear regressions, probably because the true relationship depends on the type of inequality, the level of development, and the institutional context, none of which a linear specification can capture.</p><p>In development economics, the relationship between rainfall and conflict appears to be nonlinear: both droughts and floods increase conflict, but normal rainfall does not. Estimate a linear model and you might find nothing, because the effects of too-little and too-much rain cancel in the average. The same data, analyzed with a more flexible specification, reveals a strong relationship.</p><p>Trade economists have encountered similar issues. The gravity model of trade, one of the most successful empirical regularities in economics, is log-linear by convention. But Santos Silva and Tenreyro (2006) showed that the standard practice of estimating gravity equations in logs with OLS produces inconsistent estimates when the error term is heteroskedastic. The nonlinearity of the logarithmic transformation interacts with the error structure in ways that bias the coefficients. Their proposed alternative, Poisson pseudo-maximum likelihood, respects the nonlinear structure of the model and produces meaningfully different trade elasticities.</p><p>In health economics, the dose-response relationship between alcohol consumption and mortality is J-shaped: moderate drinkers have lower mortality than both abstainers and heavy drinkers. A linear regression of mortality on alcohol consumption would estimate a small positive or negative slope, depending on the sample, and miss the entire structure of the relationship. The policy implications are completely different: a linear model might suggest that reducing alcohol consumption uniformly would improve health, while the nonlinear relationship suggests that the effects depend entirely on where individuals start.</p><p>Even in the core of labor economics, Card and Krueger&#8217;s famous minimum wage study, and the decades of debate that followed, is partly a debate about functional form. If the employment response to minimum wage increases is nonlinear, small at low wage levels, large at high ones, then the conflicting results across studies might reflect different parts of the same nonlinear relationship rather than genuine disagreement about the sign of the effect. Cengiz, Dube, Lindner, and Zipperer (2019) used bunching estimators precisely to avoid imposing linearity on this relationship.</p><p>The lesson is stark. When we impose linearity, we are not just simplifying - we are censoring. We are making a decision about what effects we are willing to find. A linear model cannot discover a U-shaped relationship. It cannot detect threshold effects. It cannot identify effects that change sign depending on the context. It can only find monotonic, constant-slope relationships, and when reality is more complicated, it returns a null result that we mistake for evidence of no effect.</p><h2>Why the Profession Has Been Slow to Change</h2><p>If nonlinearity matters so much, why hasn&#8217;t the profession moved away from linear models more aggressively?</p><p>Part of the answer is institutional. The credibility revolution, which transformed empirical economics starting in the 1990s, was focused on identification, not functional form. The question became: do you have a credible source of exogenous variation? If you had a good instrument, a clean natural experiment, a sharp regression discontinuity, the functional form was secondary. Angrist and Pischke&#8217;s <em>Mostly Harmless Econometrics</em>, the manifesto of the credibility revolution, is explicit about this: focus on the linear model, interpret the coefficient as a weighted average derivative, and do not get distracted by functional form.</p><p>This was liberating. It meant that researchers could focus on finding clever research designs instead of agonizing over whether the relationship was quadratic or cubic or exponential. The linear model became a vehicle for the identification strategy, and the identification strategy was what mattered.</p><p>But this created a blind spot. The credibility revolution gave us excellent tools for answering &#8220;does X affect Y?&#8221; but poor tools for answering &#8220;how does the effect of X on Y vary across the distribution of X, Y, and everything else?&#8221; The linear coefficient, even when causally identified, is an average, and averages can mislead when the underlying distribution has interesting structure.</p><h3>The Flexible Controls Defense</h3><p>A common response from applied researchers is that they do not really assume linearity, not in any meaningful sense. The argument goes: with enough fixed effects, interactions, and flexible controls, the linear model can approximate any function. A researcher who includes region-by-year fixed effects, bins the treatment variable into deciles, or adds polynomial terms is not imposing a straight line on the data. The variation that identifies the coefficient is local, and locally, any smooth function is approximately linear.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!SVu-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4677eb2-6452-4687-98bc-90a5aa30fd13_818x725.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!SVu-!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4677eb2-6452-4687-98bc-90a5aa30fd13_818x725.png 424w, /__u/substackcdn.com/image/fetch/$s_!SVu-!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4677eb2-6452-4687-98bc-90a5aa30fd13_818x725.png 848w, /__u/substackcdn.com/image/fetch/$s_!SVu-!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4677eb2-6452-4687-98bc-90a5aa30fd13_818x725.png 1272w, /__u/substackcdn.com/image/fetch/$s_!SVu-!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_webp, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4677eb2-6452-4687-98bc-90a5aa30fd13_818x725.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!SVu-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4677eb2-6452-4687-98bc-90a5aa30fd13_818x725.png" width="818" height="725" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e4677eb2-6452-4687-98bc-90a5aa30fd13_818x725.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:725,&quot;width&quot;:818,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:170387,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://carloschavezp29.substack.com/i/187371804?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4677eb2-6452-4687-98bc-90a5aa30fd13_818x725.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!SVu-!, /__u/carloschavezp29.substack.com/w_424, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4677eb2-6452-4687-98bc-90a5aa30fd13_818x725.png 424w, /__u/substackcdn.com/image/fetch/$s_!SVu-!, /__u/carloschavezp29.substack.com/w_848, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4677eb2-6452-4687-98bc-90a5aa30fd13_818x725.png 848w, /__u/substackcdn.com/image/fetch/$s_!SVu-!, /__u/carloschavezp29.substack.com/w_1272, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4677eb2-6452-4687-98bc-90a5aa30fd13_818x725.png 1272w, /__u/substackcdn.com/image/fetch/$s_!SVu-!, /__u/carloschavezp29.substack.com/w_1456, /__u/carloschavezp29.substack.com/c_limit, /__u/carloschavezp29.substack.com/f_auto, /__u/carloschavezp29.substack.com/q_auto:good, /__u/carloschavezp29.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4677eb2-6452-4687-98bc-90a5aa30fd13_818x725.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This defense has real merit. A regression with sufficiently rich controls is performing something like local linear approximation within cells defined by the controls. The identification comes from within-cell variation, and if the cells are small enough, the linearity assumption is nearly innocuous for each cell.</p><p>But the defense works less well than it appears. First, the headline coefficient is still a single number, a weighted average across all those cells. Even if the relationship is approximately linear within each cell, the marginal effects can vary dramatically across cells, and the linear summary still compresses that variation into one &#946;. Second, the &#8220;enough controls&#8221; argument does not address the functional form of the outcome equation itself. Fixed effects absorb level differences across groups, but they do not allow the <em>slope</em> of the treatment effect to vary across groups unless the researcher explicitly interacts the treatment with group indicators, which most do not. Third, the argument proves too much: if you need region-by-year fixed effects, decile bins, and polynomial terms to make linearity approximately valid, then linearity is not innocent. You are doing a great deal of work to rescue the assumption, and the question becomes whether a more flexible specification would have answered the question more directly.</p><p>The deeper issue is that &#8220;flexible controls&#8221; address confounding, they help ensure that the coefficient is causal, but they do not address functional form. You can have a perfectly identified, causally interpretable coefficient that is nonetheless a misleading summary of the underlying relationship because the relationship is not constant across the support of the treatment variable. Identification and functional form are separate problems, and solving one does not solve the other.</p><h3>Machine Learning and the Path Forward</h3><p>Machine learning has begun to change this landscape, and it represents the most promising path beyond the constant-effect assumption, though the profession is still in the early stages of integrating these tools into standard practice.</p><p>Methods like causal forests, developed by Athey and Imbens (2019), allow researchers to estimate heterogeneous treatment effects without specifying the functional form in advance. The algorithm partitions the data to find subgroups with different treatment effects, letting the data reveal the structure of heterogeneity rather than imposing it through pre-specified interactions. The sorted effects method of Chernozhukov, Fern&#225;ndez-Val, and Luo (2018) provides a complementary approach: a way to visualize the full distribution of treatment effects across individuals, making transparent what the average hides. Double/debiased machine learning (Chernozhukov et al., 2018) offers a framework for using flexible, high-dimensional methods to model nuisance functions, the control variables and the first stage, while still recovering valid inference on the causal parameter of interest. The key insight is that you can let machine learning handle the parts of the model where functional form matters most (the controls) while maintaining the interpretability and inferential properties that economists value.</p><p>These tools are becoming more common, but they are still far from standard practice. I plan to write a dedicated essay on how machine learning methods are reshaping causal inference, what they can and cannot do, and where they fit relative to the structural and reduced-form traditions. For now, the point is that the tools to move beyond linearity exist. The question is whether the profession&#8217;s incentives will catch up to its capabilities.</p><h3>The Sociology of Linearity</h3><p>There is also a sociological dimension. Linearity makes results easy to communicate. &#8220;An additional year of education increases earnings by 8 percent&#8221; is a sentence anyone can understand. &#8220;The returns to education range from 3 to 15 percent depending on the individual&#8217;s unobserved ability, the type of schooling, and the local labor market, with the relationship being concave at higher levels and potentially non-monotonic for certain subgroups&#8221; is accurate but unpublishable in the abstract of a top journal. The discipline rewards clean, memorable results. Nonlinear findings are messy, conditional, and harder to summarize in a tweet.</p><p>Referees reinforce this. A paper with a quadratic or spline specification will face questions that a linear specification will not. Why that number of knots? Why a quadratic and not a cubic? Is the nonlinearity robust to the choice of control variables? These questions are legitimate, but they impose a higher burden on nonlinear specifications than on linear ones. The linear model benefits from being the default: it does not need to justify itself. Everything else does.</p><p>The result is a kind of methodological conservatism. Researchers know that nonlinearity matters. They sometimes include robustness checks with quadratic terms or run binned scatter plots. But the headline result, the one in the abstract and the introduction, is almost always a linear coefficient. The nonlinearity, if acknowledged at all, is relegated to the appendix.</p><h3>The Structural Alternative</h3><p>The structural econometrics tradition, the one associated with the Cowles Commission, with Heckman, with the work I see every day, never had this problem in the same way. Structural models specify the economic mechanism, and the mechanism can be nonlinear. A utility function can have decreasing marginal returns. A production function can have complementarities. A selection model can generate nonlinear relationships between observed and unobserved variables. The functional form is derived from theory, not imposed for convenience.</p><p>This is one of the underappreciated virtues of structural work. When you write down a model of behavior and estimate its parameters, you can capture nonlinearities that reduced-form methods average away. The cost is real: you need to commit to a specific model, and if the model is wrong, your nonlinearities are wrong too. Structural models trade one set of functional form assumptions for another - they are derived from economic theory rather than imposed for algebraic convenience, but they are assumptions nonetheless, and they can be just as consequential. A misspecified utility function will generate misspecified nonlinearities with great precision. The benefit is that you can ask questions that linear methods cannot answer: What happens at the margin? How do effects vary across the distribution? What does the policy response look like for individuals who are far from the average?</p><p>The tension between structural and reduced-form approaches, which has defined much of the methodological debate in economics over the past two decades, is partly a debate about linearity. Reduced-form methods buy credibility by minimizing assumptions, but one of the assumptions they retain is linearity. Structural methods are more demanding about specification, but they can accommodate the nonlinear features of economic behavior that reduced-form methods must ignore. Neither tradition has fully solved the functional form problem. But recognizing that each makes different functional form compromises, rather than treating one as assumption-free, is the first step toward being honest about what our methods can and cannot tell us.</p><h2>The World Is Nonlinear</h2><p>So where does this leave us? I do not think economists should stop using linear models. That would be absurd. Linear regression remains the right tool for many questions, and the discipline of focusing on identification rather than functional form has been enormously productive. Some of the best papers in economics are simple linear regressions with clever identification strategies.</p><p>But I think we have been insufficiently honest about what linearity costs us.</p><p>When we report a linear coefficient, we should be clear that it is a weighted average, and that the weights may not correspond to any policy-relevant quantity. When we find a null result with a linear specification, we should ask whether a more flexible specification would tell a different story. This matters even more because journals are notoriously reluctant to publish null results. A real, nonlinear relationship that a linear specification reduces to a near-zero coefficient does not just mislead the researcher, it never reaches the literature at all. The effect disappears twice: once in the estimation, and once in the publication process. When we log-linearize a macro model, we should acknowledge that the approximation fails exactly when the model is most needed. And when we teach econometrics, we should spend as much time on functional form as we do on endogeneity.</p><p>The irony of modern empirical economics is that we have become extraordinarily sophisticated about one source of bias - endogeneity, selection, confounding - while remaining remarkably casual about another: the assumption that the world arranges itself into straight lines. We have elaborate tools for ensuring that our coefficient is causal. We have almost no routine for checking whether a single coefficient is the right summary of the relationship.</p><p>The next time you read an empirical paper, or write one, ask yourself: what would this result look like if the relationship were not a straight line? If the answer is &#8220;I don&#8217;t know,&#8221; that is worth investigating. If the answer is &#8220;it would probably look different,&#8221; then the linear coefficient is not summarizing the relationship. It is hiding it.</p><p>This is not a call to abandon linearity. It is a call to treat it as what it is: an assumption, as consequential as any exclusion restriction, and one that deserves the same scrutiny.</p><h2>Where to Start</h2><p><strong>Heckman, Lochner, &amp; Todd (2006).</strong> &#8220;Earnings Functions, Rates of Return and Treatment Effects: The Mincer Equation and Beyond.&#8221; <em>Handbook of the Economics of Education.</em> The systematic case that the standard Mincer specification fits the data poorly and that &#8220;returns to education&#8221; depend on functional form.</p><p><strong>Heckman &amp; Vytlacil (2005).</strong> &#8220;Structural Equations, Treatment Effects, and Econometric Policy Evaluation.&#8221; <em>Econometrica.</em> Shows that different IV estimators recover different weighted averages of the marginal treatment effect function, the paper that unifies the weighting discussion.</p><p><strong>Roth &amp; Sant&#8217;Anna (2023).</strong> &#8220;When Is Parallel Trends Sensitive to Functional Form?&#8221; <em>Econometrica.</em> Should make every DiD practitioner uncomfortable. The same data can support or reject parallel trends depending on the transformation.</p><p><strong>Burke, Hsiang, &amp; Miguel (2015).</strong> &#8220;Global Non-linear Effect of Temperature on Economic Production.&#8221; <em>Nature.</em> Resolved the contradiction between micro and macro climate estimates by allowing for nonlinearity. The linear specification had been hiding a universal relationship.</p><h2>Some Additional Worth Readings`</h2><h3>On the history of linearity in economics</h3><p>Mirowski, P. (1989). <em>More Heat than Light: Economics as Social Physics, Physics as Nature&#8217;s Economics.</em> The definitive account of how economics borrowed its mathematics from physics, and what was lost in translation. Controversial, polemical, and impossible to ignore.</p><p>Leamer, E. E. (1983). &#8220;Let&#8217;s Take the Con out of Econometrics.&#8221; <em>American Economic Review.</em> Still devastating. If specification choices determine your results, and functional form is a specification choice, then linearity is not innocent.</p><h3>On what linearity hides in labor economics and treatment effects</h3><p>Angrist, J. D., &amp; Krueger, A. B. (1999). &#8220;Empirical Strategies in Labor Economics.&#8221; <em>Handbook of Labor Economics.</em> The discussion of what IV coefficients actually estimate under heterogeneous effects is buried in a handbook chapter that every applied economist should read twice.</p><p>Imbens, G. W., &amp; Angrist, J. D. (1994). &#8220;Identification and Estimation of Local Average Treatment Effects.&#8221; <em>Econometrica.</em> The paper that formalized the LATE framework. Understanding that IV recovers effects for compliers, not the population, was the first step toward recognizing how linearity and instrument choice shape what we find.</p><p>Angrist, J. D., &amp; Pischke, J.-S. (2009). <em>Mostly Harmless Econometrics.</em> The canonical defense of the linear model as a vehicle for identification. Essential reading, even, especially, when you disagree with its conclusions about functional form.</p><h3>On linearity in difference-in-differences and causal inference design</h3><p>Athey, S., &amp; Imbens, G. W. (2006). &#8220;Identification and Inference in Nonlinear Difference-in-Differences Models.&#8221; <em>Econometrica.</em> The theoretical framework for what happens when parallel trends holds for a nonlinear transformation of outcomes rather than levels. Published nearly two decades before the recent DiD revolution caught up to its insights.</p><p>Goodman-Bacon, A. (2021). &#8220;Difference-in-Differences with Variation in Treatment Timing.&#8221; <em>Econometrica.</em> The decomposition that launched a thousand DiD papers. Shows how the TWFE estimator weights different comparisons, and how linearity assumptions drive the weights.</p><p>de Chaisemartin, C., &amp; D&#8217;Haultf&#339;uille, X. (2020). &#8220;Two-Way Fixed Effects Estimators with Heterogeneous Treatment Effects.&#8221; <em>American Economic Review.</em> Shows the TWFE estimator is a weighted sum of average treatment effects with potentially negative weights.</p><p>Cengiz, D., Dube, A., Lindner, A., &amp; Zipperer, B. (2019). &#8220;The Effect of Minimum Wages on Low-Wage Jobs.&#8221; <em>Quarterly Journal of Economics.</em> Uses bunching estimators to avoid imposing linearity on the employment response to minimum wages.</p><p>Callaway, B., &amp; Sant&#8217;Anna, P. H. C. (2021). &#8220;Difference-in-Differences with Multiple Time Periods.&#8221; <em>Journal of Econometrics.</em> Proposes estimators that avoid the problematic weighting of TWFE under treatment effect heterogeneity.</p><p>Sun, L., &amp; Abraham, S. (2021). &#8220;Estimating Dynamic Treatment Effects in Event Studies with Heterogeneous Treatment Effects.&#8221; <em>Journal of Econometrics.</em> Shows how to construct robust event-study estimates under staggered adoption, still within a linear framework, but with honest weights.</p><h3>On linearization in macroeconomics</h3><p>Fern&#225;ndez-Villaverde, J., &amp; Rubio-Ram&#237;rez, J. F. (2006). &#8220;Solving DSGE Models with Perturbation Methods and a Change of Variables.&#8221; <em>Journal of Economic Dynamics and Control.</em> The case for higher-order perturbation methods.</p><p>Fern&#225;ndez-Villaverde, J., Gordon, G., Guerr&#243;n-Quintana, P., &amp; Rubio-Ram&#237;rez, J. F. (2015). &#8220;Nonlinear Adventures at the Zero Lower Bound.&#8221; <em>Journal of Economic Dynamics and Control.</em> Shows that global nonlinear methods produce dramatically different predictions for zero-rate episodes.</p><h3>On nonlinear alternatives and heterogeneous effects</h3><p>Santos Silva, J. M. C., &amp; Tenreyro, S. (2006). &#8220;The Log of Gravity.&#8221; <em>Review of Economics and Statistics.</em> A paper that should have ended the practice of estimating gravity equations in log-linear OLS. It hasn&#8217;t, which tells you something about how slowly functional form assumptions change.</p><p>Dell, M., Jones, B. F., &amp; Olken, B. A. (2012). &#8220;Temperature Shocks and Economic Growth.&#8221; <em>American Economic Journal: Macroeconomics.</em> The landmark panel study that opened the climate-economy literature. BHM's nonlinear approach later showed that the linear specification masked a universal relationship affecting rich and poor countries alike.</p><p>Athey, S., &amp; Imbens, G. W. (2019). &#8220;Machine Learning Methods That Economists Should Know About.&#8221; <em>Annual Review of Economics.</em> The bridge between the linear tradition and data-driven methods.</p><p>Chernozhukov, V., Fern&#225;ndez-Val, I., &amp; Luo, Y. (2018). &#8220;The Sorted Effects Method: Discovering Heterogeneous Effects Beyond Their Averages.&#8221; <em>Econometrica.</em> A method for seeing what the average hides.</p><p>Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., &amp; Robins, J. (2018). &#8220;Double/Debiased Machine Learning for Treatment and Structural Parameters.&#8221; <em>The Econometrics Journal.</em> The framework for using machine learning to handle nuisance functions while maintaining valid causal inference.</p>]]></content:encoded></item><item><title><![CDATA[The Variable We Left Out]]></title><description><![CDATA[On class, the credibility revolution, and the walls you can't see]]></description><link>https://carloschavezp29.substack.com/p/the-variable-we-left-out</link><guid isPermaLink="false">https://carloschavezp29.substack.com/p/the-variable-we-left-out</guid><dc:creator><![CDATA[Carlos Chavez]]></dc:creator><pubDate>Fri, 06 Feb 2026 13:31:29 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!noHI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce1865ae-bf7b-4b5c-ba65-ee05988f782a_726x521.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This is the first paid Friday edition. Monday posts will always be free. Fridays are where I get personal: opinions I've been sitting on, papers that changed how I think, and writing that comes from experience as much as from evidence. Let's start with a paper that hit close to home.</em></p><p>There&#8217;s a question I&#8217;ve been asked at every stage of my career, always phrased politely, always carrying the same quiet surprise: <em>&#8220;Where are you from?&#8221;</em></p><p>Not which university. Not which department. Where. As in: how did you end up here?</p><p>I grew up in Barranca, a small city on the coast of Peru. My parents didn&#8217;t go to college. Nobody in my family had set foot in a research university. When I first found myself sitting in a seminar room full of people who seemed to know exactly how to perform the role of &#8220;academic,&#8221; I had no script. I didn&#8217;t know you were supposed to interrupt. I didn&#8217;t know that asking a question at a talk was partly about showing you could ask a good question. I didn&#8217;t know that &#8220;networking&#8221; at conferences meant something different from being friendly. I learned these things the way you learn a second language as an adult: by watching, imitating, and getting things wrong enough times that eventually you stop getting them wrong.</p><p>I knew I was not the only one navigating this. I&#8217;m glad that there&#8217;s a paper about it now.</p><h2>The Paper</h2><p>Anna Stansbury (MIT) and Kyra Rodriguez (Berkeley) just released what I think is one of the most important labor papers of the year: <em><a href="https://annastansbury.github.io/website/StansburyRodriguez_The_Class_Gap_in_Career_Progression__Evidence_from_academia.pdf">&#8220;The Class Gap in Career Progression: Evidence from US Academia.&#8221;</a></em></p><p>The setting is US tenure-track academia. The data comes from the NSF&#8217;s Survey of Doctorate Recipients (1993&#8211;2021), linked with Web of Science bibliometric records and NSF award data. The identification strategy is elegant in its simplicity: compare people who got their PhD from the same institution, in the same field, in roughly the same period, and ask whether their socioeconomic background predicts where they end up.</p><p>It does. By a lot.</p><p>First-generation college graduates, compared to PhD classmates whose parents hold non-PhD graduate degrees (JDs, MDs, MBAs), are 10% less likely to be tenured at an R1 university, end up at institutions ranked 11% lower, earn 3% less, and report 5% lower job satisfaction. These are not small numbers. And they show up at both major career junctures: the tenure-track job market and the tenure decision itself.</p><p>Here&#8217;s the part that should make empirical researchers sit up: this gap exists entirely on the <em>intensive</em> margin. There is no class gap in whether you stay in academia; only in <em>where</em> you end up within it. Lower-SEB academics don&#8217;t leave. They get sorted downward.</p>
      <p>
          <a href="/__u/carloschavezp29.substack.com/p/the-variable-we-left-out">
              Read more
          </a>
      </p>
   ]]></content:encoded></item></channel></rss>