<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Paperclips and Other Alignment Problems]]></title><description><![CDATA[Paperclips and Other Alignment Problems is a Substack about what it takes to aim machines at the right targets when we are not even fully sure what the targets are. ]]></description><link>https://carlolc.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!9Skm!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fcarlolc.substack.com%2Fimg%2Fsubstack.png</url><title>Paperclips and Other Alignment Problems</title><link>https://carlolc.substack.com</link></image><generator>Substack</generator><lastBuildDate>Sat, 05 Sep 2026 02:37:36 GMT</lastBuildDate><atom:link href="/__u/carlolc.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Carlo Ludovico Cordasco]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[carlolc@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[carlolc@substack.com]]></itunes:email><itunes:name><![CDATA[Carlo Ludovico Cordasco]]></itunes:name></itunes:owner><itunes:author><![CDATA[Carlo Ludovico Cordasco]]></itunes:author><googleplay:owner><![CDATA[carlolc@substack.com]]></googleplay:owner><googleplay:email><![CDATA[carlolc@substack.com]]></googleplay:email><googleplay:author><![CDATA[Carlo Ludovico Cordasco]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[A follow up on Acemoglu et al. (2026)’s garbling result]]></title><description><![CDATA[I fear they might be wrong once again!]]></description><link>https://carlolc.substack.com/p/a-follow-up-on-acemoglu-et-al-2026s</link><guid isPermaLink="false">https://carlolc.substack.com/p/a-follow-up-on-acemoglu-et-al-2026s</guid><dc:creator><![CDATA[Carlo Ludovico Cordasco]]></dc:creator><pubDate>Sun, 30 Aug 2026 21:29:19 GMT</pubDate><content:encoded><![CDATA[<p>In the past few months, I&#8217;ve been thinking about what <a href="/__u/carlolc.substack.com/p/acemoglu-et-al-2026-are-wrong-about">I wrote about Acemoglu et al</a>. This is probably what one ought to do after publishing a post entitled &#8220;Acemoglu et al. are wrong,&#8221; although I&#8217;m told that the more efficient strategy is to publish another post explaining why everyone misunderstood the first one.</p><p>I still think the original argument was right. Acemoglu, Kong and Ozdaglar model a world in which AI can cause people to lose existing forms of knowledge, but the forms of knowledge that matter, the way they enter production, and the competences people can acquire are all fixed before the model begins. The target moves. The game stays the same.</p><p>I now think there is another problem.</p><p>Suppose we grant them the fixed game. Suppose AI really does reduce people&#8217;s incentive to acquire a socially valuable kind of knowledge. Suppose this knowledge will deteriorate unless something changes. Does it follow that we should deliberately make AI worse?</p><p>I don&#8217;t think it does.</p><p>The reason is that they treat two things as though they must remain bundled together: producing useful output and reproducing the knowledge that society will need in the future. Historically, they often were bundled. AI may allow us to separate them.</p><p>Once that possibility appears, the policy problem changes. We can restore the old constraint so that people continue learning for the same reason they learned before. Or we can preserve the gain from lifting the constraint and rebuild the valuable functions that the old production process happened to perform.</p><p>The second option is harder than it sounds. It requires us to explain who will validate AI-mediated work, how the next generation of experts will be trained, and why the market might fail to produce enough of them. But it is also the option that preserves what AI has made possible.</p><h2>The argument in short</h2><p>Acemoglu et al.&#8217;s mechanism is elegant.</p><p>People acquire general knowledge partly because it helps them solve their own problems. Their learning also contributes a small amount to a shared stock of knowledge from which other people, future people and AI systems can benefit. This contribution is a public good. Individuals do not capture its full value.</p><p>Now give each person a sufficiently accurate AI system. The AI supplies the context-specific information they previously had to acquire for themselves. People rationally learn less. Unfortunately, the public contribution that accompanied their private learning disappears as well. Individually sensible reliance on AI gradually depletes the shared knowledge base.</p><p>The proposed solution is to reduce the precision of the AI. Acemoglu et al. call this garbling. If the system is less reliable, people must continue learning, and the public stock of knowledge survives.</p><p>There is a genuine insight here. The activity through which we produce something is often also the activity through which we learn how to produce it. Remove the activity and you may remove the training system along with it.</p><p>But notice what garbling does. It preserves the valuable by-product of an old production constraint by partially restoring the constraint.</p><p>Before deciding whether this is sensible, we need to understand what that constraint was doing.</p><h2>The constraint was doing three jobs</h2><p>Suppose that, before AI, producing a competent output in domain X required competence in X. To write a good philosophical paper, you needed to know some philosophy. To design a protein, you needed to know some chemistry. To produce reliable software, you generally needed to know how to program.</p><p>That requirement did at least three things.</p><p>To begin with, it limited production. Acquiring X was costly, so only people who had spent time developing the relevant competence could produce much useful X-output.</p><p>It also placed much of the capacity for evaluation inside the producer. The person constructing the argument, writing the software, or designing the experiment would not be infallible, obviously, but they normally possessed some ability to recognize when the result was nonsense. Production and epistemic accountability were imperfectly joined together.</p><p>Most importantly, the same production process helped to create the next generation of competent people. Juniors began with easier tasks, made mistakes, observed experts, received feedback and gradually took responsibility for harder work. The constraint on production supplied the work through which X was learned.</p><p><strong>AI may break all three connections.</strong></p><p>A person can now acquire a different competence, call it Y, which consists in getting AI systems to do useful work. Y might include decomposing a problem, supplying context, constructing a workflow, choosing between models, combining outputs, running AI-assisted experiments, or knowing how to elicit several competing approaches before selecting one.</p><p>The claim does not require X and Y to be completely independent. Often they will reinforce each other. A good programmer may use a coding model better than someone who has never programmed. A chemist may know what information a protein-design system needs. A philosopher may know which objection is worth developing, and so forth.</p><p>The important point is comparative. Acquiring Y can produce a large improvement in someone&#8217;s ability to generate X-output through AI without producing a comparable improvement in their unaided ability to produce X or to evaluate it independently.</p><p>Someone may become much better at producing working-looking software through AI while remaining unable to debug it without AI. Someone may generate promising biological candidates without fully understanding the mechanism that makes them promising. Someone may produce a polished philosophical argument without being able to determine whether the argument works.</p><p>AI has then made two substantially different human competencies into alternative routes into the same production process.</p><h2>The non-expert is now on the production side</h2><p>Knowledge has always been divided across people, including inside production. A laboratory technician may run an assay they could not design or fully interpret. A paralegal may prepare material whose legal significance must be assessed by someone else. A machine operator may produce a component without understanding the engineering theory behind it.</p><p>So the unusual feature of AI cannot simply be that nonexperts now contribute to expert production. They already did.</p><p>The difference is that older divisions of labour often made the limits of each contribution reasonably visible. The technician produces a result for the scientist to interpret. The paralegal prepares a document for a lawyer to review. Responsibility for validation sits somewhere identifiable in the organization.</p><p>Generative AI allows someone with little X to produce something that has the surface form of a complete X-output. It can produce the whole memorandum, the whole program, the whole diagnosis, the whole argument. The boundary of production moves much faster than the boundary of epistemic authority.</p><p>One can generate the artefact without acquiring the competence needed to know whether the artefact is good.</p><p>That is the real problem. The arrival of AI does not merely deepen an existing division of labour. It can hide the division inside an apparently finished product.</p><p>Sometimes this does not matter very much. If the quality of the output can be tested cheaply and independently, the competence of the person who generated it may be largely irrelevant. Software can sometimes be checked against a strong test suite. A design can be run through a reliable simulator. Biological candidates can be subjected to experimental assays. These tests may be costly and incomplete, and experts were usually required to construct them, but they can still allow someone without deep X to participate productively in generation.</p><p>Elsewhere, the test is itself an exercise of X.</p><h2>Generation, validation and renewal</h2><p>It helps to divide the process into three stages.</p><p><strong>Generation</strong> produces candidate answers, designs, arguments, diagnoses, molecules or programs.</p><p><strong>Validation</strong> determines which candidates are correct, useful or worth pursuing.</p><p><strong>Renewal</strong> improves the methods, identifies new problems, develops better tests and produces the people who will carry the domain forward.</p><p>AI may make Y a strong substitute for X at the first stage. Someone with little domain competence can generate an impressive candidate.</p><p>It does not follow that Y substitutes for X at the other two stages.</p><p>Nolan Lovett makes a related distinction between <strong>Internalized Mastery</strong>, deep competence developed through sustained practice, and <strong>Distributed Mastery</strong>, the ability to orchestrate human&#8211;AI systems. He calls the continuing dependence of distributed mastery on internalized expertise the <strong>Validation Tether</strong>. AI-mediated production can become widespread while remaining dependent on a pool of people capable of judging what it produces.</p><p>Philosophy is a useful example because the validation problem is unusually visible there. I also happen to have spent enough time around philosophy to recognize polished rubbish in its natural habitat.</p><p>An LLM can generate objections, distinctions, examples and replies almost without limit. It can produce fifty candidate arguments before a philosophy department has finished deciding where to go for lunch.</p><p>The scarce capacity lies elsewhere.</p><p>Does the objection actually work? Is the distinction substantial, or has the language model merely assigned two names to the same thing? Does the example isolate the relevant feature? Has the argument uncovered a problem, or manufactured the appearance of one? Is it new? Is it an old point with the serial numbers removed? Does the reply answer the objection, or simply repeat the original claim at greater length?</p><p>These questions do not have immediate independent tests. Answering them is part of doing philosophy.</p><p>A person with little philosophy may therefore become quite good at generating philosophical material while remaining unable to select among it. The AI has lifted the constraint on generation. Validation becomes the bottleneck.</p><p>Once we see this, the relation between AI and expertise becomes more complicated than either &#8220;substitution&#8221; or &#8220;complementarity.&#8221;</p><p>For an individual trying to produce a candidate, AI may substitute for X. For the system trying to turn candidates into reliable output, AI-generated abundance may increase the value of X.</p><p>Suppose one person using AI can generate a thousand candidate designs. Someone must still decide which are worth testing. If generation expands faster than validation, expert attention becomes more scarce, not less. Better AI can reduce the value of X for the person producing the first draft while increasing the value of X elsewhere in the production system.</p><p>The direction is not fixed. If AI also becomes excellent at testing, diagnosis and evaluation, the need for human validators may fall. Where reliable external tests exist, extensive specialization in Y may be efficient. Where quality is difficult to assess without X, AI can make X-validation more important precisely because X-generation has become cheap.</p><p>The relevant question is which constraint AI lifts.</p><h2>Why garbling is such a peculiar response</h2><p>Seen from this perspective, garbling is a particular way of preserving knowledge.</p><p>Suppose AI allows someone to produce the same output in half the time. Society has gained something. We can produce more, work less, explore more possibilities, or redirect some of the saved time toward validation, training and research.</p><p>The gain may be larger than a simple saving in labour. AI can allow people with different backgrounds to participate in a domain, bringing problems, intuitions and competences that would previously have remained outside it. It can make previously infeasible projects possible. It can enlarge the candidate set from which useful discoveries might emerge.</p><p>Garbling gives some of this back. It makes the system less informative so that people must continue performing more of the old cognitive task.</p><p>That may preserve their competence. It may also preserve the public signals that their learning produced. But we should be clear about the exchange. We protect a valuable by-product of the old production process by making the productive technology less useful.</p><p>Another response would preserve the full technology and reproduce the by-product separately.</p><p>Suppose the problem is that too few people will invest in X. Society could fund X-research, create well-paid validation roles, subsidize specialist training, support professional apprenticeships, require expert review in high-stakes settings, and build institutions through which AI users report anomalies to people capable of learning from them.</p><p>The general idea has appeared elsewhere in the recent literature. Afrouzi, Blanco, Drenik and Hurst model a technology that automates work through which human capital was previously accumulated, while also helping maintain and expand the productive frontier. Their preferred policy acts differently on the two margins, combining a tax on automation profits with a subsidy for frontier maintenance, rather than simply suppressing the technology throughout the economy.</p><p>The principle is straightforward. Where a technology creates a valuable output and damages a separate socially valuable activity, we should at least consider supporting the damaged activity directly.</p><p>This gives us two possible strategies.</p><p>The first restores the constraint. It makes AI less capable, less available or less reliable so that ordinary production continues to generate expertise incidentally.</p><p>The second preserves the gain from lifting the constraint and rebuilds validation and knowledge renewal as deliberate social functions.</p><p>The second sounds better. Unfortunately, saying &#8220;let us fund the experts&#8221; is not yet an answer.</p><p>The experts have to come from somewhere.</p><h2>Who trains the validators?</h2><p>This is the hardest problem for the specialization story.</p><p>Imagine that firms discover they can employ many people who are good at Y, supported by a small number of experienced X-experts. Current production might work very well. The experts handle difficult cases, validate outputs and intervene when the AI fails.</p><p>Ten years later, where do the experienced experts come from?</p><p>They did not begin as experts. They became experts by performing junior work, initially badly and then less badly, under the supervision of people who already knew what they were doing.</p><p>The old production system helped finance this process. Junior workers were not merely students. They also performed useful work. Their labour gave organizations a reason to employ them while they learned. Training was expensive, but some of its cost was recovered through production.</p><p>AI may remove that arrangement. If a senior worker can use AI to complete the junior tasks quickly, the organization loses the immediate reason to employ the novice. The work disappears before the need for future experts does.</p><p>Enrique Ide develops this problem in a model where novices acquire tacit knowledge by working alongside more knowledgeable experts. Because tacit knowledge is embodied and difficult to verify in a contract, the market cannot simply price and sell the future expertise in advance. Automation of entry-level work can raise current output while weakening the transmission of knowledge between generations.</p><p>Matt Beane&#8217;s research on robotic surgery shows what a disrupted learning pipeline can look like in practice. Traditional surgical training allowed trainees to develop competence through gradually increasing participation. Robotic surgery greatly reduced their role in the work. The minority who reached competence often did so through what Beane calls &#8220;shadow learning,&#8221; informal and sometimes norm-challenging practices developed because the approved route no longer supplied enough meaningful participation.</p><p>This is not a minor qualification. It may defeat the simple proposal that society should concentrate X in a smaller specialist sector. A specialist sector cannot survive unless it has a way to produce specialists.</p><p>The distinction between production and training still helps, but it does not solve everything.</p><p>We do not make autopilot unreliable during ordinary passenger flights in order to preserve pilots&#8217; manual competence. We use simulators, supervised training and deliberately constructed exercises. The production environment contains the best available technology. The training environment sometimes withholds it because the point of the exercise is to develop a particular human capacity.</p><p>Something similar may work elsewhere. Trainee programmers could be required to solve some problems without coding assistants. Junior lawyers could analyse cases before seeing AI-generated answers. Philosophy students could construct arguments before using a model to criticize them.</p><p>But simulations and classroom exercises cannot reproduce every form of expertise. Some judgment develops only through real cases, real responsibility and exposure to the consequences of error. The trainee must sometimes participate in production, even when giving the task to the AI or the senior expert would be more efficient.</p><p>Rebuilding the training pipeline may therefore require organizations to tolerate deliberate inefficiency. They may need to reserve tasks for trainees, subsidize supervised roles, give junior workers meaningful responsibility, and pay experts to teach rather than using every minute of their time for immediate production.</p><p>Who will pay for this?</p><p>The old system did not provide training for free, but training was partly bundled with commercially useful work. Once AI removes that work, the cost of reproducing expertise becomes more visible. Universities, firms, professional bodies and governments may have to pay explicitly for something that the production process previously financed indirectly.</p><p>The constraint-lifting dividend gives us resources with which to do this. Whether those resources will actually be used for training is a separate question.</p><h2>Won&#8217;t higher expert wages solve the problem?</h2><p>Suppose AI makes validation the bottleneck. Existing X-experts become more valuable. Their wages rise. People see the higher return and invest in X. Perhaps the market solves the problem without garbling or public intervention.</p><p>It might.</p><p>We should not infer a market failure simply because society continues to need experts. Scarcity creates a price signal. Firms also have reasons to retain expertise. They care about liability, reputation, continuity, model failures and difficult cases. Where they bear the cost of bad outputs, they may invest heavily in validation and training.</p><p>But the value of today&#8217;s experts and the production of tomorrow&#8217;s experts are different things.</p><p>AI can increase the rent earned by the existing stock of X while weakening the process through which the next stock is created. Training takes years. The current wage responds to current scarcity, while the people who begin training today may not become useful validators until much later.</p><p>Firms may also prefer to hire experts trained elsewhere rather than bear the cost of training their own. A company that finances years of apprenticeship cannot necessarily prevent the newly trained worker from leaving. Better tests, professional methods and standards spill across organizations. The firm that develops them bears the cost while competitors share the benefit.</p><p>Some of the value of X also lies in rare failures. Deep expertise may matter most when the system encounters a novel case, drifts outside its normal distribution or produces an error that ordinary monitoring cannot detect. Firms will sometimes pay for that option value. They may invest too little when the failure is unlikely during the tenure of the decision-maker, when losses fall on clients or the public, or when the resulting knowledge would protect the wider sector rather than only the firm financing it.</p><p>Higher scarcity wages can therefore coexist with an eroding training pipeline. The market can bid intensely for the remaining experts without reproducing enough of them.</p><p>This does not establish that underinvestment will occur everywhere. It tells us what must be shown before the policy argument works. We need some reason why the people who finance X capture less than the full benefit produced by validation, training and renewal.</p><p>The externality cannot simply be assumed.</p><h2>Specialization will not always be clean</h2><p>Even where an X-sector survives, the sensible division of labour will rarely consist of pure X-experts on one side and Y-operators who know nothing about the domain on the other.</p><p>A person using AI may need enough X to recognize that the system has misunderstood the task. They may need to identify which contextual details matter, know when to escalate a case, or understand why a validator has rejected the output.</p><p>The validator faces the corresponding problem. An expert who did not observe how the output was generated may lack relevant context. They may not know what instructions were given, which alternatives were discarded, or which parts of the answer are model-generated rather than supplied by the operator.</p><p>Communication can fail precisely because the person who has the information lacks the competence to recognize its importance.</p><p>Some knowledge also arises through widely dispersed experience. A small research group cannot necessarily replace thousands of practitioners encountering unusual cases in different settings. Doctors, engineers, teachers and managers learn partly because the world keeps surprising them locally. A worker with too little X may fail to notice which surprise deserves to be reported.</p><p>A thin expert layer also creates fragility. Validators can become overwhelmed. They may share the same blind spots. Their judgments may be ignored by organizations whose revenue depends on rapid production. Concentrating expertise can make it easier to capture.</p><p>The likely social structure therefore contains several groups. Some people will acquire deep X. Some will specialize mainly in Y. Many will need a working combination of both. How much X must remain distributed depends on the testability of the output, the cost of communicating with experts, the importance of local knowledge, and the consequences of error.</p><p>This is where Acemoglu et al.&#8217;s concern may be strongest. If collective knowledge can be maintained only through the widespread engagement of ordinary users, a small specialist sector will not replace what accurate AI removes. Some preservation of direct practice may then be justified.</p><p>But this is one possible structure of the problem. It should be established rather than assumed.</p><h2>What counts as X will change</h2><p>There is a final complication, and it brings me back to my earlier post.</p><p>The valuable domain competence we need to preserve is unlikely to remain fixed as AI develops.</p><p>Suppose AI becomes excellent at generating grammatically competent philosophical prose and reconstructing standard arguments. Those activities become less scarce. Philosophical expertise may shift toward choosing worthwhile questions, identifying the precise source of a disagreement, constructing examples that isolate one feature at a time, and recognizing when the existing literature has framed the problem badly.</p><p>Those capacities become the new X.</p><p>The same movement will occur elsewhere. If AI becomes good at generating candidate proteins, expertise may shift toward target selection, experimental design, anomaly interpretation and causal explanation. If AI becomes good at producing software, expertise may shift toward specification, architecture, security, integration and the design of tests capable of revealing unfamiliar failures.</p><p>Y changes as well. General competence in directing AI becomes increasingly domain-shaped. The person who knows which evidence to provide, which constraints to impose and which output deserves suspicion may need substantial X after all.</p><p>The boundary between X, Y and the hybrid competence needed to connect them will move as the technology develops.</p><p>Garbling may preserve yesterday&#8217;s X by keeping yesterday&#8217;s tasks artificially difficult. Meanwhile, the genuinely scarce competence has moved somewhere else.</p><p>This is why experimentation with the capable technology matters. We need to discover where the new bottlenecks are, which parts of validation can be automated, which forms of judgment remain difficult, and what kinds of training produce people capable of handling the remaining problems.</p><p>Making the technology worse can interfere with that discovery.</p><h2>Restoring the constraint or rebuilding the function?</h2><p>Several recent papers now converge on parts of this problem.</p><p>Lovett distinguishes internalized from AI-mediated mastery and worries about the regeneration of professional expertise. Ide focuses on the intergenerational transmission of tacit knowledge. Ide and Talam&#224;s model AI within knowledge hierarchies where less knowledgeable workers handle routine work and more knowledgeable people solve exceptions. Afrouzi and colleagues examine automation, learning and career dynamics. Beane shows how a new technology can weaken the normal route through which people learn expert work.</p><p>The convergence is useful. It suggests that the underlying problem is real.</p><p>The question I want to isolate concerns the appropriate response.</p><p>Before AI, producing X-output, validating it and training future X-experts were often bundled inside the same production process. AI can separate them. It allows people with a different competence to generate X-candidates, but it does not necessarily validate those candidates or reproduce the experts on whom validation depends.</p><p>Garbling restores the old bundle. It makes AI less useful so that ordinary production continues to generate expertise incidentally.</p><p>Sometimes this may be necessary. Where expert knowledge can develop only through real participation, where the relevant learning is widely dispersed, and where training cannot be separately financed, preserving parts of the old activity may be the only workable option.</p><p>Even then, the restriction need not apply everywhere. We might preserve unaided work within training, reserve real tasks for novices, require human engagement in high-stakes settings, or design workflows that force users to inspect and justify outputs. These interventions protect the route into expertise without deliberately degrading every productive use of the technology.</p><p>Where validation and training can be organized separately, the case for general garbling becomes much weaker. We can preserve the gain from lifting the production constraint and use part of that gain to support the epistemic functions on which the new system depends.</p><p>The aim should not be to preserve the pre-AI distribution of competence across individuals. A fall in average unaided X does not by itself amount to knowledge collapse.</p><p>The relevant questions are whether society retains enough capacity to evaluate what it produces, respond when the systems fail, improve the methods it uses, and train the people who will be able to do these things in the future.</p><p>The old constraint helped to produce that capacity accidentally. People acquired X because they could not produce X-output without it.</p><p>AI may lift the constraint.</p><p>The hard task is to reproduce the valuable functions deliberately, without throwing away the gain that made the reorganization possible.</p><p>Making AI worse is one way to preserve human knowledge. It should not be the first one we try.</p><h3>Some relevant work</h3><p>Acemoglu, D., Kong, D., and Ozdaglar, A. (2026). &#8220;AI, Human Cognition and Knowledge Collapse.&#8221; NBER Working Paper No. 34910.</p><p>Lovett, N. (2026). &#8220;The Tragedy of the Cognitive Commons: How AI Could Disrupt the Regeneration of Professional Expertise.&#8221; <em>Human Resource Development Review</em>, OnlineFirst.</p><p>Ide, E. (2026). &#8220;Automation, AI, and the Intergenerational Transmission of Knowledge.&#8221; arXiv:2507.16078.</p><p>Afrouzi, H., Blanco, A., Drenik, A., and Hurst, E. (2026). &#8220;Automation, Learning, and Career Dynamics.&#8221; NBER Working Paper No. 35157; Federal Reserve Bank of Atlanta Working Paper 2026-6.</p><p>Ide, E., and Talam&#224;s, E. (2025). &#8220;Artificial Intelligence in the Knowledge Economy.&#8221; <em>Journal of Political Economy</em>, 133(12), 3762&#8211;3800.</p><p>Beane, M. (2019). Shadow Learning: Building Robotic Surgical Skill When Approved Means Fail. <em>Administrative Science Quarterly</em>, <em>64</em>(1), 87-123.</p>]]></content:encoded></item><item><title><![CDATA[Bad AI-assisted research is the PR we don’t deserve]]></title><description><![CDATA[On research as commitment]]></description><link>https://carlolc.substack.com/p/bad-ai-assisted-research-is-the-pr</link><guid isPermaLink="false">https://carlolc.substack.com/p/bad-ai-assisted-research-is-the-pr</guid><dc:creator><![CDATA[Carlo Ludovico Cordasco]]></dc:creator><pubDate>Sat, 29 Aug 2026 11:24:51 GMT</pubDate><content:encoded><![CDATA[<p>Today I&#8217;m writing without AI assistance in <strong>protest against AI slop</strong>. Yes, you read this correctly: against AI slop. ChatGPT would say that I&#8217;m &#8220;right to push back&#8221; against slop. Or perhaps it would say that protesting AI slop &#8220;strengthens my argument&#8221;. And, for once, it would be right.</p><p>When LLMs arrived, I did get overexcited about how much I could do with them. I remember that my first use case was reformatting a bunch of references in BibTeX format, except that I had not noticed it had got all the DOIs wrong. Little did I know about hallucinations back then, and a tweet with my <em>incredible</em> finding (&#8220;Hey you all, ChatGPT can translate references into BibTeX format&#8221;) went viral, only to attract a bunch of angry luddites shortly after, rightly pointing out that the DOIs were indeed incorrect.</p><p>I know I get overexcited about innovation. This is just my nature. I try to contain it sometimes, but I can&#8217;t. Some people get excited by Walkmans in the Spotify era; some others would rather spend 15 million euros on a 1967 Rolex 6239 than &#163;500 on an Apple Watch; I like new stuff. Stuff that most of the time lifts a constraint or solves a problem, though at other times it might create new ones.</p><p>AI for me was a revolution. I thought I was lazy, and AI taught me I wasn&#8217;t. LLMs showed me the light at the end of any tunnel (aka project) I would embark on. I could have an argument developed in the space of an afternoon. No more demoralisation when I couldn&#8217;t quite make something work; no need to let it sink for weeks waiting for a generous colleague or for a talk at a conference to solve a problem. I could do it in a few hours, and I could even get a draft in a few minutes once the argument was fully developed. I started working 14 hours a day because I saw immediate results, and when you are already a short-term satisfaction monkey, and it turns out you can get immediate satisfaction from tasks and outputs that once required months, if not years, you just cannot stop.</p><p>Admittedly, it was not the greatest prose ever, but at the beginning, I had not formed my current distaste for default AI prose. We had not been inundated with endless colon constructions, staccato, &#8220;not x, it&#8217;s y&#8221;, empty announcers at the beginning of every damn paragraph, and so forth. At the beginning, the polished prose did indeed seem polished. And the tasks I would run through LLMs mostly stemmed from a backlog of old ideas developed during my PhD that I didn&#8217;t have time/incentives to publish, given that I transitioned rather abruptly from political philosophy to business schools.</p><p>My initial use case was thus not linked to developing new ideas. It was about refining and formalising intuitions I had already formed and putting them into paper form.</p><p>Then the backlog of ideas expired. I had to start thinking about new things, and that&#8217;s when the frustration started to pick up. Not immediately, likely because the old backlog gave rise to a bunch of derivative ideas that led me to write and submit a bunch of other papers. But after a while, I realised that if LLMs could help me explore an argument in a few hours and have a first draft in a few minutes, they could not help me come up with research ideas from scratch. I mean, they surely can. You can prompt an LLM to suggest cool ideas to pursue, but that&#8217;s not how research projects are usually born, at least in my book.</p><p>Research projects are commitments, like relationships, I like to say (and, yes, people do make fun of me for saying that). They don&#8217;t grow out of nowhere. You need to fall in love, get obnubilated by your ideas as if nothing else mattered in the world, then get past the honeymoon phase and check whether they are viable or something else indeed matters more in the world. Commitment is meaningful only after these things have happened, and writing without commitment is soulless. And it shows. Whether the prose is human or AI, you can tell whether it stems from a researcher&#8217;s commitment.</p><p>In the past few weeks, I&#8217;ve reviewed several papers. Most of them AI-written. And I don&#8217;t mind: time is not endless, and young untenured scholars are under real pressure. I use AI prose in papers all the time, and AI prose can be good if you show it some love. You can prompt your way into writing better stuff: it takes sentence-by-sentence care and a rather specific knowledge of syntax, rhythm and pacing, though you can learn the relevant terms from the LLM itself as you go. It is painfully slow at first. Eventually, you build a set of instructions, learn to recognise what isn&#8217;t working properly, and end up writing both faster and better. That, at least, has been my experience.</p><p>Pangram will likely still detect your prose as AI-generated, but, honestly, who cares? We can&#8217;t write for Pangram, nor for colleagues who might reproach you for a high Pangram score. Our goal is to write clearly, articulate ideas, and solve problems. When older, tenured purists treat any use of AI as evidence of laziness, their contempt reeks of privilege and is sometimes merely rent protection.</p><p><strong>The problem is that the papers I reviewed were bad.</strong> And it was not just the prose. They had been written with the commitment of someone putting a coin into a slot machine and waiting to see whether a paper would come out of it. Now, this cannot be the way we do research. Lower exploration and drafting costs cannot take away the one thing that truly makes research enjoyable and meaningful: falling in love with a question, seeing where it takes you, trying to break down the problem into as many parts as you can, and being ready to challenge the initial intuitions you had. AI cannot take away our commitment to a question, because commitment is the one and only thing that matters in this profession, both intrinsically and instrumentally.</p><p>The brainy smurfs might tell you that letting AI write for you will make you dumb, that you&#8217;re cheating, that serious scholarship requires the same friction they had to endure. They are wrong in principle, but there is also a sense in which they are right. When exploration costs are high, when drafting takes ages, you need to commit, and good things can only come out of sincere and meaningful commitments. Human academic writing wasn&#8217;t necessarily proof of thought (we all know how much pre-AI nonsense is out there), but in many cases it was evidence of commitment, and that is what I fear AI is taking away from many of us.</p><p>The anti-AI crowd doesn&#8217;t need help from us. They can already leverage our innate status quo bias, the fear of change, and that human and disheartening sense of losing one&#8217;s professional identity once a technological change steps in. They don&#8217;t need endless uninspiring papers, or text written just for the sake of writing &#8216;submitted&#8217; on a CV, to make a case against AI in research and lobby for restrictions. More than that, we, qua committed researchers and AI enthusiasts, don&#8217;t need this PR. We should fight this lack of commitment more than they do. And that&#8217;s why yesterday I wrote my first ranty review report. I might have overdone it, and if so, I apologise. But I care too much about this profession, and about what AI can do for us, to leave it at a simple rejection.</p><p>The technology is great. Our use determines how it will be received. This is a collective action problem; make sure you do your part.</p>]]></content:encoded></item><item><title><![CDATA[The Pangram Tax]]></title><description><![CDATA[How AI detectors can waste effort, hide good collaboration, and slow down what humans and AI learn to do together]]></description><link>https://carlolc.substack.com/p/the-pangram-tax</link><guid isPermaLink="false">https://carlolc.substack.com/p/the-pangram-tax</guid><dc:creator><![CDATA[Carlo Ludovico Cordasco]]></dc:creator><pubDate>Wed, 26 Aug 2026 15:54:24 GMT</pubDate><content:encoded><![CDATA[<p>Over the past few days, I have changed my view of the new AI policy at <em>Philosophy &amp; Public Affairs</em> more than once.</p><p>My <a href="/__u/carlolc.substack.com/p/on-ppas-new-ai-policy">first post</a> was perhaps too critical. In my <a href="/__u/carlolc.substack.com/p/a-reply-to-seth-lazar-on-ppas-ai">second</a>, replying to Seth Lazar, I conceded lots of ground. I accepted that requiring authors to write the final prose themselves does keep some bad papers away. Rewriting a whole manuscript takes time. Someone who hoped to generate something plausible, skim it, and submit it may decide that the additional work is not worthwhile.</p><p>I still think that concession was correct. I remained concerned, however, about what happens once authors can run the detector themselves.</p><p>You can use a model while you work, but the finished text must be yours. You declare that it is yours, and someone runs the wording through a detector to assess that declaration. Universities are adopting rules of this kind. So are journals, employers, and clients. I will use journal submissions as the clearest example, although the argument applies more broadly.</p><p>What happens when the score becomes part of the writing process?</p><p>My friend Niclas Berggren <a href="https://nonicoclolasos.com/2026/07/22/what-ai-detection-tools-miss/">made this point before I did</a>, and I want to begin with his argument. A detector analyses the wording. It cannot determine who had the ideas, who made the important decisions, who checked the analysis, or who could defend the work under questioning. Any use of AI that does not appear in the prose is unavailable to it.</p><p>As Niclas observes, one consequence is that people will rewrite perfectly good sentences solely to reduce the score.</p><p>I want to develop that point.</p><h2>Who runs Pangram first?</h2><p>Consider a simplified case. Anything scoring above a particular threshold is rejected. Writers know the threshold and can run the detector themselves before submitting. Actual rules are less precise, but the simplification makes the mechanism easier to see.</p><p>It matters who runs the detector first. If only the institution ran it, after the paper had been submitted, authors would face a risk of rejection. Some would submit and fail. But authors can run Pangram themselves, so that risk can be removed in the simplified case. If the score is too high, they can rewrite the paper and check again, or decide not to submit it. Model-written prose that remains above the threshold is therefore rewritten or abandoned before anyone at the institution reads it.</p><p>By the time the institution runs its test, it is not examining what people originally wrote. It is examining what they decided to submit after checking. The abandoned drafts do not appear. The rewritten drafts arrive in their revised form, and the institution cannot see what they previously looked like.</p><p>Now consider a paper that arrives with a low score. It could have reached that score in at least three ways. The author may not have used AI. The author may have conducted careful AI-assisted research and written the prose independently. Or the author may have produced a poor paper with a model and altered the wording until its score fell below the threshold.</p><p>The score does not distinguish among these cases. It does not tell you whether the paper is good. Nor does it tell you how many other papers were rewritten or abandoned before this one was submitted.</p><h2>The Pangram tax</h2><p>Some rewriting occurs only because of the score. I want to distinguish this from ordinary revision, much of which improves a paper regardless of what a detector says.</p><p>Consider two ways of producing a good paper with a model.</p><p>One person asks for a draft and then carries out the substantive revision directly in the final prose. They check the claims, return to the sources, reconstruct the argument, and replace vague passages with more precise ones. This work would have been useful in any case. It also lowers the score because, by the end, the sentences on the page are theirs. The detector requires no additional work from them.</p><p>Now imagine a different way of working. Instead of revising the draft sentence by sentence in your own prose, you return to the model. You ask for alternative ways of developing the argument, present objections, and compare the responses. This process helps you determine what you think. When you want the wording changed, you ask the model to change it. The resulting text is one you have directed throughout but typed very little of. Every sentence was ultimately produced by the model.</p><p>The detector requires no further work from the first writer. You must rewrite the whole manuscript.</p><p>Before submitting, you must reproduce in your own words a text you already have. You must do so in a way that changes the score, even though the argument was completed before the rewriting began. The rule requires the prose to be yours, while the detector assesses whether it resembles human-written prose. You therefore do the work twice: first by developing the argument and then by expressing it again in different words.</p><p>This additional work is what I call the tax. It does not measure laziness or the use of a shortcut. Both writers did serious work. Its size depends on where the substantive revision occurred and how much of the final prose the model produced.</p><p>In some cases, the required rewriting will correspond to the amount of intellectual work still needed. A barely read generated draft may require substantial and useful revision before it falls below the threshold. But the relationship is unreliable. A paper that involved a great deal of thought may still require extensive rewriting because the thinking occurred through dialogue with the model and the final wording remained model-generated.</p><p>The amount of rewriting a paper requires does not, by itself, tell us how much thought went into it.</p><p>The relevant test is counterfactual: would you have made this change without the rule or the score? If not, and if the change does not improve the paper, the time spent making it is part of the tax.</p><p>A clear sentence is replaced with another clear sentence. A paragraph is reordered without becoming easier to follow. The quality of the paper remains the same, but the detector is less likely to identify it.</p><p>Pangram reports a percentage, but under a threshold rule only one distinction matters: whether the score is above or below the threshold. A lower score provides no benefit until it crosses that threshold. Once it has done so, there is no reason to continue. The threshold determines when the rewriting can stop.</p><p>A simple equation makes the cost explicit.</p><p>Suppose you have T hours for the paper and spend d of them on changes made solely for the detector. This leaves T &#8722; d hours for improving the paper. Let m be the paper&#8217;s quality before you spend those hours, and let &#945; represent how much each useful hour adds. The final quality is:</p><p>Q = m + &#945;(T &#8722; d),&#8195;&#945; &gt; 0</p><p>One might argue that the changes are harmless because one clear sentence is no worse for the reader than another. Even if the changes are harmless, however, they take time. Compared with spending all T hours on improvement, you have given up &#945;d. Those hours could have been used to check evidence, develop the argument, correct mistakes, or find a better solution to the problem.</p><h2>What the threshold selects for</h2><p>The prospect of having to spend those d hours also affects whether you submit the paper.</p><p>Let k represent the total cost of crossing the threshold. This includes the improvement forgone, the effort involved, and the other activities displaced. Let s represent the stakes of the submission, and let V(s) represent the expected benefit of submitting below the threshold rather than withdrawing, before paying the cost.</p><p>You pay when:</p><p>V(s) &gt; k</p><p>Suppose V(s) increases with the stakes and eventually exceeds k. Let s* be the point at which:</p><p>V(s*) = k</p><p>Below s*, you withdraw the paper. Above it, you pay the cost and submit.</p><p>Both sides vary. Crossing costs differ across papers, people, and methods of working. But if the cost is held constant, an author who expects a greater benefit from submitting has a stronger reason to pay it.</p><p>Some abandoned papers benefit a journal. Models make it easy to produce many weak but plausible papers. Someone who intended to submit such a paper at little cost may decide that the required rewriting is not worthwhile. The paper is not submitted, and no editor or reviewer has to read it. The anticipated cost has therefore prevented a weak submission without anyone actually incurring that cost.</p><p>The problem concerns the basis of this selection.</p><p>The threshold effectively selects according to the cost of crossing it and the value the author places on submitting. Neither is a reliable measure of quality. Someone may write a strong paper out of curiosity, decide that the additional work is not worthwhile, and withdraw it. Someone whose promotion depends on a publication may rewrite and submit a weak paper.</p><p>The cost can therefore discourage some low-effort submissions without reliably separating good work from bad.</p><p>As the amount of detector-specific rewriting increases, d becomes larger and k rises. More papers are deterred, while those who continue give up more time that might have improved their work. If crossing the threshold is easy, the deterrent effect and the opportunity cost are both small. The same mechanism produces both the benefit and the cost. Whether the trade-off is worthwhile is a genuine institutional question.</p><h2>What disappears from view</h2><p>Up to this point, the tax is an individual cost. It consumes an author&#8217;s time or prevents a paper from being submitted. Strategic adaptation may also impose a broader cost.</p><p>Once writers can test and rewrite their work before submitting, the final prose provides less usable evidence about the process that produced it. The rule already gives people a reason to conceal prohibited uses. Running the detector beforehand also tells them whether their wording is likely to reveal those uses. Once the score is below the threshold, the detector provides the institution with no further information about them.</p><p>Some rules ask authors to explain how they used AI. This may appear to solve the problem. An honest account can tell an institution a great deal about permitted uses. But if the disclosure remains private, nobody else learns from it. It will not include anything the author has decided to conceal. And if describing close collaboration is expected to bring greater scrutiny or informal sanctions, people will provide only the minimum information the form requires.</p><h2>What nobody gets to learn</h2><p>A broader consequence concerns experimentation.</p><p>People still need to determine how to divide a task with a model, when to trust it, which claims require checking, when it genuinely saves effort, and when it creates a different problem that is less immediately obvious. In the process, someone may develop a better way to test its output, build a useful tool, or identify a task that a person and a model can perform better together than either can perform alone.</p><p>Time used to replace equivalent sentences cannot be used for those experiments. Not every saved hour would have produced a discovery. Most would not. But experiments that are not conducted cannot produce new knowledge.</p><p>Another consequence concerns communication.</p><p>Some practical knowledge would remain private under any policy. But practical knowledge is sometimes shared. People describe what they can do. Consultants demonstrate their methods. Educators sometimes turn good practices into courses. Within a profession, people share enough information about what succeeded and failed to give others a basis from which to begin.</p><p>If admitting to successful AI collaboration carries professional costs, fewer accurate examples will circulate. Others will have to rediscover the method or copy what remains visible. Restrictions may reduce AI use, but they are unlikely to eliminate it. Adoption can therefore continue while more effective methods remain concealed.</p><p>Secrecy can consequently affect both how much AI is used and how well it is used.</p><p>This returns me to <a href="/__u/gus1365199.substack.com/p/will-ai-take-your-job">Gus Skorburg&#8217;s concern</a> about mediocre AI being adopted at scale. Detection rules do not improve model performance. Detector-backed restrictions may also make better methods less visible and more difficult to transfer.</p><p>Model providers have records of the interactions. Institutions observe the completed work. Neither source of information, by itself, connects a method to its result.</p><p>We could therefore end up with powerful models and accurate detectors while knowing very little about what people and models can accomplish together.</p><h2>What institutions actually need</h2><p>This is not an argument for applying lower standards to AI-assisted work. Model output is often shallow, repetitive, incorrect, and unsupported. It should be checked, challenged, and rewritten for as long as those changes improve the work.</p><p>The problem arises when Pangram is treated as the solution.</p><p>Institutions have legitimate reasons to care about how work was produced. A university needs to know whether a student learned anything. A journal may want to know whether the named authors did the thinking and can defend the claims. A client may need to know how confidential material was handled. When the process matters, institutions need evidence that tracks it as directly as reasonably possible.</p><p>One genuine benefit of an AI-restrictive policy is that it may reduce submission volume. Nobody can carefully read an unlimited number of plausible documents, and models can produce such documents quickly. A threshold may reduce the number by making some low-cost submissions no longer worthwhile.</p><p>But the threshold filters according to prose score, crossing cost, and author incentives. The journal wanted to know whether the work deserved its attention. Limits on how frequently one person can submit address volume more directly, although they will not solve every version of the problem.</p><p>When a rule permits some uses of AI, authors must also be able to report those uses safely. Someone who honestly reports a permitted use should not find that admission treated as evidence against them. Private disclosure is insufficient if we also want useful methods to spread. People need to be able to explain what worked without placing their own work under suspicion.</p><p>The general approach should therefore be to judge the finished work when the result is what matters, examine how it was produced when the process matters, and make honest disclosure of permitted uses genuinely safe.</p><p>Authors should not be required to make changes that improve only the detector score.</p><p>The Pangram tax consists first of effort spent lowering a score rather than improving the work. Its broader effects arise when effective methods remain hidden and potentially informative experiments are not conducted.</p><p>The detector may classify the wording accurately enough. The error occurs when its score is interpreted as a measure of the work behind it.</p><p>The result may be an expensive measurement error.</p><p></p><p>P.S. One point that was implicit in my concession to Seth Lazar needs to be made explicit. The cost of rewriting may select on quality as well as stakes. If an author thinks a paper is not good enough to justify rewriting, they may drop it. That is part of the deterrent effect I accepted, and it is a form of selection a journal might reasonably want.</p><p>Quality should therefore appear explicitly in the calculation. Let q be the author&#8217;s estimate of the paper&#8217;s quality. The expected value of submitting should be written as V(q,s), because a paper the author considers better will normally seem more likely to be accepted. The decision is:</p><p>V(q,s) &gt; k</p><p>The threshold is therefore not selecting only according to how much the author has riding on the submission. Holding the stakes and rewriting cost constant, an author should be more willing to rewrite a paper they consider good.</p><p>This gives the policy a genuine quality-screening effect. But the effect has three limitations.</p><p>First, it operates on the author&#8217;s estimate of quality, not directly on quality itself. That estimate may be informative: authors usually know something about the strengths and weaknesses of their own work. But it is also imperfect. The filter removes papers that authors consider unworthy of further effort, which is not necessarily the same as removing the weakest papers.</p><p>Second, quality competes with stakes. A selective journal rejects most submissions, including many good ones. Even a paper the author considers strong may therefore remain unlikely to succeed. When the value of success is high enough (e.g., a promotion case, a probation deadline, or a grant application), even a small chance may justify the rewriting. An author with less at stake may abandon a better paper because the same chance is not worth the cost.</p><p>Third, the screen applies unevenly. If your ordinary process of checking and revising a generated draft has already produced prose that falls below the threshold, you may face little or no additional detector-specific cost. The extra quality screen applies mainly to papers that still require rewriting for the detector. It therefore selects within a particular set of workflows rather than across the journal&#8217;s submissions as a whole.</p><p>This does not withdraw my concession to Seth. It states more precisely what that concession amounts to. The rewriting requirement may discourage some weak papers because their authors do not consider them worth the additional effort. It may therefore improve the quality of the submitted pool. But quality is only one element in the decision, alongside the author&#8217;s stakes and the cost created by their way of working. The result is a real quality screen, but a noisy, uneven, and indirect one.</p>]]></content:encoded></item><item><title><![CDATA[A reply to Seth Lazar on PPA’s AI policy]]></title><description><![CDATA[I was once again out jogging yesterday when I discovered that Seth Lazar had tweeted about my post on PPA&#8217;s AI policy. As my four hardcore followers know, I usually write these rants under the impression that nobody besides those four excellent friends will actually read them. Perhaps I should revisit that assumption :)]]></description><link>https://carlolc.substack.com/p/a-reply-to-seth-lazar-on-ppas-ai</link><guid isPermaLink="false">https://carlolc.substack.com/p/a-reply-to-seth-lazar-on-ppas-ai</guid><dc:creator><![CDATA[Carlo Ludovico Cordasco]]></dc:creator><pubDate>Sun, 23 Aug 2026 10:34:08 GMT</pubDate><content:encoded><![CDATA[<p>I was once again out jogging yesterday when I discovered that <a href="https://x.com/sethlazar/status/2091176895857463758?s=46">Seth Lazar had tweeted about my post on PPA&#8217;s AI policy</a>. As my four hardcore followers know, I usually write these rants under the impression that nobody besides those four excellent friends will actually read them. Perhaps I should revisit that assumption :)</p><p>Seth&#8217;s thread is excellent, and it has helped me better understand PPA&#8217;s position. In my original post, I already conceded that its policy made strategic sense. A prestigious journal bears the visible reputational cost of publishing a bad paper while bearing almost none of the cost when an excellent AI-assisted paper is never submitted or is rejected. It is perfectly rational for PPA to care much more about the first kind of mistake.</p><p>Seth has persuaded me to go somewhat further. <strong>I still think it is perfectly legitimate to prefer a professional equilibrium in which experimentation with AI is widespread and open. But I also think that PPA is under no obligation to bear the costs and risks of bringing that equilibrium about</strong>, especially when it already has an editorial model that works well. Seth is also right that those of us who are AI-friendly (and I take him to be one of us) should be prepared to put our money where our mouths are.</p><p>I am nevertheless left with two issues. The first is what the requirement to rewrite model-generated prose adds once upstream uses of AI are already permitted, and what it costs. The second is how reliable the resulting signal will remain once authors begin adapting to the rule.</p><h2>What does rewriting add?</h2><p>Seth&#8217;s strongest point, I think, concerns the relationship between a paper and the person whose name appears above it.</p><p>When a journal publishes a paper, it does more than identify a good collection of sentences. Publication also gives readers, hiring committees, and promotion panels some evidence about the author. It suggests that this person understands the argument, can defend and revise it, and will be able to use what they have learned when they encounter another problem. Even when credentialling is not the editors&#8217; immediate aim, it is one of the predictable consequences of their decisions.</p><p>While universities are ultimately responsible for their hiring decisions, PPA still has an institutional interest in preserving the credentialling value of publication. A PPA publication is professionally valuable not only because the paper is good, but because it is taken to say something about the person who wrote it. That signal attracts strong submissions, reviewing labour, and professional attention, all of which reinforce the journal&#8217;s prestige. If excellent papers stopped providing much evidence about their authors, part of what makes publication in PPA valuable would disappear.</p><p>AI makes that inference from paper to author less reliable, even when the paper itself is good.</p><p>Suppose that two AI-assisted papers are equally strong. One researcher may have spent weeks challenging a model, rejecting most of its suggestions, checking its claims, and gradually developing an argument that they understand completely. Another, the <em>one-shotter</em>, may have generated an equally impressive paper with a single prompt, inspected it rather casually, and submitted it without fully understanding what they had produced. An editor may be able to judge that both papers are good while having little basis for knowing whether they provide the same evidence about their authors.</p><p>There is an obvious direct rule available: <strong>PPA could require authors to disclose substantial AI assistance, to understand and endorse every substantive claim in the paper, and to be able to defend, revise, and develop the argument.</strong> It could also require them to have made a meaningful intellectual contribution to selecting the problem, constructing the argument, or testing it against objections.</p><p><em>If everyone complied honestly</em>, this would track the underlying concern more closely than a rule about who generated the final words. A researcher who had seriously directed the collaboration and intellectually owned the result could submit the paper. Someone who had merely accepted a fortunate output they barely understood could not honestly do so.</p><p>The issue is that intellectual ownership cannot be observed directly. Editors cannot look inside an author&#8217;s head to determine whether they understand an argument, nor can they easily reconstruct how much of the intellectual work consisted in guiding and evaluating a model rather than accepting whatever it produced. An author can falsely declare that they understand and contributed to a paper at almost no cost.</p><p>Requiring authors to write the final paper themselves is an indirect way of getting at the same property. It remains possible for me to ask an LLM to produce an argument that I barely understand and then rewrite the whole thing in my own prose. The result might comply with PPA&#8217;s policy while still providing poor evidence of my philosophical abilities. Conversely, I might understand and be able to defend every part of an argument while violating the policy because I retained model-generated language.</p><p>The writing requirement becomes attractive because it imposes a cost. Rewriting an entire paper takes time and effort, so it should deter at least some people who would otherwise submit one-shot outputs. It may also force authors to notice that they cannot explain a step, that two apparently distinct claims are doing the same work, or that a crucial inference depends on an ambiguity. Someone who rewrites a paper will probably understand it better than someone who copied and pasted the output five minutes before submitting it.</p><p>This is the point at which the assumptions about compliance begin to pull apart. Under full compliance, a direct requirement of understanding and intellectual ownership tracks the relevant abilities more closely. The case for the writing requirement arises because we expect at least some authors not to comply honestly with that direct rule. Author-written prose is therefore not the talent PPA ultimately wants to identify. It is a comparatively observable and costly proxy for it.</p><p><strong>This is a narrower claim than what I thought was Seth&#8217;s original appeal to talent. It is a comparative claim about which institutional mechanism works best under imperfect compliance.</strong></p><p>That strikes me as a good argument. <strong>But the additional assurance has a price.</strong></p><p>Thankfully, PPA permits upstream uses of AI. An author may use a model to generate possible arguments, formulate objections, compare alternative structures, identify counterexamples, or work through a distinction that initially seemed clear. The policy intervenes at the end of that process by requiring the author to replace any model-generated prose that survives into the submitted paper.</p><p>For authors who have accepted the model&#8217;s reasoning too quickly, rewriting may add a great deal. It may improve both their understanding and the paper. The same requirement, however, applies to authors who already understand the argument, have tested it carefully, and can answer for every claim. Rewriting may still improve the prose, as rewriting often does, but the rule continues to impose work after those improvements have run out. Any remaining model-generated passage must be replaced because of how it was produced, even when the author understands its content and would otherwise retain its wording.</p><p>That work consumes time.</p><p>Until recently, delegating the production of publication-ready prose to a machine was not a serious option. Authors could receive extensive help from co-authors, editors, referees, and unusually generous colleagues, but automated sentence production did not form part of the ordinary technological environment. LLMs now make it possible to relax some of the labour involved in producing final prose. PPA deliberately preserves that labour because performing it provides evidence about the author.</p><p>Once a technological limitation becomes an institutional requirement, however, we must count its opportunity cost. Time spent reconstructing an argument in permitted prose cannot also be spent reading more widely, testing the argument against further objections, examining additional cases, developing another idea, or deciding which problem is worth pursuing next. The cost may be modest when we consider a paragraph or a paper. Across a career, and then across an entire profession, it becomes more substantial.</p><p><strong>Constraints also influence which abilities people develop</strong>. Philosophers, like anybody else, become better at the activities to which they repeatedly devote their limited time and attention, particularly when those activities determine whether their work can appear in leading journals.</p><p>If LLMs become genuinely useful philosophical collaborators, competence may increasingly involve knowing which model is appropriate for which task, how to frame a problem and provide the relevant context, how to decompose a complicated argument into questions a model can usefully address, and how to move iteratively between prompting, criticism, and revision. It may involve knowing when to consult another model, when to distrust all of them, how to distinguish a genuine insight from extremely plausible bullshit, and how to verify that an apparently original objection was not borrowed from somewhere else.</p><p>These are not merely technical tricks appended to the real philosophical work. They might partly determine what intellectual work a human&#8211;AI system can perform. Competence in them is compatible with writing every sentence of the final paper oneself. But every hour can be used only once. Time spent reconstructing language that the author already understands leaves less time for developing these or other abilities.</p><p>Lifting the constraint would not determine how philosophers used the released capacity. Some would simply produce more papers. Others might test more versions of an argument, investigate more cases, or abandon weak ideas earlier. Some might spend more time choosing questions and less time executing solutions once the central philosophical work had already been done. Others would waste most of the time they saved.</p><p>We do not yet know what the overall effect would be, partly because institutional rules help determine which practices researchers have reason to develop.</p><p>A costly signal derives its evidential value from the cost it imposes. The same cost is also intellectual time that cannot be used elsewhere. We should therefore ask whether the additional evidence supplied by mandatory rewriting is worth the philosophical work and competence that the requirement may crowd out.</p><p>PPA may reasonably judge that the additional assurance is worth the cost. Editors often cannot distinguish serious human&#8211;AI collaboration from shallow prompting and thus have good reason to prefer a rough rule that strengthens the connection between a paper and its author.</p><p>My concern is about what happens when the same rule is adopted across leading journals. Philosophers will continue to invest heavily in unaided composition because their professional success depends upon it. They will have correspondingly weaker incentives to develop forms of philosophical competence built around more integrated collaboration with AI.</p><p>This can become self-reinforcing. Unaided writing remains central to professional philosophy partly because our institutions continue to require it. Its continued centrality can then be treated as evidence that the requirement was necessary. Publication policies thus do not merely respond to an existing conception of philosophical competence. They help reproduce it.</p><h2>The policy changes what it measures</h2><p>My second question concerns enforcement.</p><p>The attraction of the writing requirement is that it appears easier to enforce than a direct requirement of understanding and intellectual ownership. Seth is confident that AI-generated prose has a detectable signature. Even granting that present models produce patterns that experienced editors or specialised systems can recognise, detectability is not independent of the policy built around it.</p><p>Once recognisable AI prose becomes grounds for rejection, retraction, and professional sanctions, the policy creates demand for AI prose that cannot be recognised (yes, yet another Lucas-style objection).</p><p>The most obvious response is to rewrite the output. Models will be prompted or developed specifically to avoid the habits that editors currently associate with AI, while commercial services will test manuscripts for suspicious patterns before they reach a journal. Since publication carries large professional rewards, authors who intend to violate the rule will have strong reasons to invest time and money in concealment.</p><p>There is an important difference between the available responses. If I personally rewrite a model&#8217;s argument in my own prose, I may comply with the policy, and the exercise may force me to understand the argument more fully. If I use AI to produce an argument I do not understand, pay someone else to turn it into convincingly human prose, and conceal that process from the journal, I am cheating. PPA is entitled to sanction me.</p><p>The moral distinction is clear. The evidential problem remains. Once publication depends on displaying a particular observable signal, people have reason to manufacture that signal. The prose may appear to show that the author has worked through the argument when, in fact, someone else has borne the cost on their behalf.</p><p>Seth suggests that if authors know that future advances may expose them, leading to retraction and a professional ban, some will decide that cheating is not worth the risk. The threat should deter violations.</p><p>It also raises the value of successful concealment. Anyone who still intends to cheat now has a stronger reason to remove every trace, since avoiding detection protects them against a much larger future loss. The policy therefore produces two effects that pull in opposite directions: it discourages some low-effort submissions while pushing the remaining violators towards more sophisticated methods.</p><p>Which effect dominates is an empirical question. If the additional cost deters enough one-shot submissions, the writing requirement may improve the connection between publication and genuine philosophical competence even as the submissions that remain become harder to identify. But any improvement will arise because the policy changes behaviour, rather than because AI-generated prose carries a permanent and readily detectable signature.</p><p>The two issues therefore pull in opposite directions. Requiring authors to replace AI-generated prose may strengthen the connection between a paper and its author and deter some forms of abuse. At the same time, it preserves a constraint that consumes scarce intellectual time, influences which abilities philosophers develop, and gives those who intend to violate the rule stronger incentives to conceal what they have done.</p><p>Author-written prose plainly provides some evidence of philosophical competence. But that alone does not justify the rule. Under imperfect compliance, it must provide a more reliable signal than direct requirements of disclosure and intellectual ownership, and the improvement must be substantial enough to justify the time consumed and the incentives created.</p><p>For the strategic reasons already given, PPA may reasonably judge that this trade-off is worth accepting. It may have no general responsibility to perform the work of hiring committees, but it has a clear interest in protecting the credentialling value of publication. That value helps sustain its prestige, attract strong submissions, and make publication in the journal professionally desirable.</p><p><strong>The real problem emerges when the same locally rational choice is repeated across leading journals.</strong> The profession may preserve author-written prose as a central marker of philosophical talent, give researchers strong reasons to keep developing that ability, and receive increasingly selective evidence about the alternatives. An equilibrium can be defensible for each institution that participates in it while remaining technologically conservative and self-reinforcing as a whole.</p><p>And Seth is right about the practical conclusion. Those of us who want the experiment cannot simply demand that PPA conduct it for us. We should be prepared to put our money, time, and reputations behind an alternative.</p><p>We should create the journal.</p><p>Annoyingly, that sounds like work.</p>]]></content:encoded></item><item><title><![CDATA[On PPA’s new AI Policy]]></title><description><![CDATA[I was in the middle of my daily jog when a friend sent me Seth Lazar&#8217;s tweet about PPA&#8217;s new AI policy.]]></description><link>https://carlolc.substack.com/p/on-ppas-new-ai-policy</link><guid isPermaLink="false">https://carlolc.substack.com/p/on-ppas-new-ai-policy</guid><dc:creator><![CDATA[Carlo Ludovico Cordasco]]></dc:creator><pubDate>Sat, 22 Aug 2026 13:21:36 GMT</pubDate><content:encoded><![CDATA[<p>I was in the middle of my daily jog when a friend sent me Seth Lazar&#8217;s tweet about PPA&#8217;s new AI policy. I&#8217;m sad, of course. Let me try to explain why.</p><p>I should start by saying that PPA was the best editorial experience I&#8217;ve had so far (close second was an excellent, yet likely AI-written, report from Phil Review earlier this year, although that paper was ultimately rejected).</p><p>The editorial comments from PPA, in particular, dramatically improved my paper in several different ways. I also like the model in which the editorial team takes a central role in handling submissions. It aligns reputational incentives nicely: the people making the decisions are also the people whose reputations are most closely connected to the quality of what the journal publishes. It is probably not an accident that journals organised roughly along these lines (e.g., Ethics, Phil Review, etc.) consistently publish better stuff.</p><p>I should also say upfront that I have reviewed a lot this year, and I have declined to review an equally large number of papers. Some of the papers I read were good; some were bad. Some appeared to be substantially AI-written; others plainly were not. And, as many of you will know, bad AI-written papers are often much more time-consuming to review. I take this to be at the heart of PPA&#8217;s decision.</p><p>Before getting to that, however, let&#8217;s look at Seth&#8217;s explanation of the new policy:</p><blockquote><p>Academic journals serve at least two functions: the promulgation of new knowledge, and identification and credentialing of talented researchers. There are many other means available to share new knowledge, but few for credentialing talented researchers, where &#8220;talented&#8221; means roughly: able, on their own or together with other researchers, to make and communicate significant progress on fundamental problems. Submitting AI-authored essays makes the task of identifying talented researchers harder.</p><p>In addition, while it is at this time possible for a researcher to use AI to author a paper that is itself high quality, it is disproportionately unlikely that they will do so, and much more likely that the paper will have all the surface appearances of sophisticated work, but ultimately be insubstantial or incoherent. This makes it harder to peer review, and makes the task of peer review in general more onerous, at a time when it is already under significant strain.</p></blockquote><p>On the identification and credentialling of talented researchers, I think Seth is straight-up wrong, for reasons that are obvious but apparently worth repeating: talent does not exist in a vacuum. It is expressed through, and partly constituted by, the technological capabilities available to us.</p><p>One could have been a very talented cardiologist before ECGs and cardiac MRI were invented. Once those technologies exist, however, I would rather be diagnosed by a cardiologist who is competent in using them. That person&#8217;s contribution to the profession is likely to be much, much more valuable. It would be strange to say that their use of these tools makes it harder to identify whether they are talented cardiologists. Competence in using the available tools is constitutive of being a talented cardiologist.</p><p>The obvious reply is that AI is not merely an instrument of measurement. It can generate prose, objections, arguments, and sometimes even ideas. Fair enough. So the issue isn&#8217;t about researchers who use AI and researchers who do not. It is between researchers who intellectually own their papers and researchers who do not.</p><p>Does the author understand the argument? Can they identify its weaknesses, explain its significance, answer objections, and revise it when something goes wrong? If so, the fact that a machine contributed some (or even much) of the language does not establish that the researcher lacks talent. If not, the paper should not be published. But it should not be published because the author cannot answer for its intellectual content.</p><p>I don&#8217;t want to spend too many more words on this, because I think Lazar knows perfectly well that talent is technologically mediated. That is why I find the credentialing argument so strange.</p><p>The second argument is considerably stronger. Bad AI-written papers really are unusually annoying to review because they disable many of the heuristics on which reviewers have traditionally relied. Chief among those is bad prose.</p><p>Bad prose is not an infallible indication of bad thinking, but it is a very convenient one. When the writing is visibly confused, you can often identify the underlying problem quickly. AI can remove the visible confusion while leaving the underlying confusion exactly where it was. You then have to do the intellectual archaeology yourself. You read three pages feeling that an argument is about to appear. You reconstruct the claims charitably, test interpretation after interpretation, and eventually discover that there was nothing beneath the polished sentences. The prose merely created the surface appearance of a thought.</p><p>This is genuinely exhausting. I am indeed exhausted as a reviewer.</p><p>But we should not romanticise the old heuristics simply because AI has made them less reliable. Those heuristics were never innocent. Treating poor prose as evidence of poor thinking has always disproportionately penalised people writing in a second or third language, people outside elite academic networks, and people without access to expensive editing or unusually generous colleagues. Of course clear writing matters in philosophy. However, bad prose and bad philosophy are not coextensive, and the costs of pretending otherwise have always been distributed unequally.</p><p>AI exposes this because it separates linguistic fluency from intellectual quality more than we could before. It can make an empty paper look serious. But it can also allow someone with a serious argument to present it without being punished for entirely preventable linguistic deficiencies. A blanket prohibition sacrifices the second possibility in order to preserve the convenience of our old screening mechanisms.</p><p>This is why I think PPA&#8217;s decision risks locking us into a bad equilibrium. Journals face a growing number of submissions and rising screening costs. They respond by prohibiting substantial AI assistance. Conscientious authors comply and lose whatever legitimate benefits the technology might have offered them. Less conscientious authors conceal their use, because the rule will be extremely difficult to enforce consistently. Reviewers continue to receive polished bullshit, except that now no one is honest about how it was produced.</p><p>Everyone&#8217;s local response is understandable. The overall result is worse.</p><p>There are better possible equilibria. If AI has made it much cheaper to produce a superficially plausible submission, journals can regulate the number of submissions directly. PPA could, for example, allow each author to submit only one paper per year, as some other journals already do. This would give authors a reason to send their best work rather than treating the journal as a slot machine. It would address the volume problem directly, apply equally to everyone, and avoid requiring editors to become detectives trying to determine which paragraphs were touched by a machine.</p><p>AI might also help on the reviewing side. I do not mean uploading confidential manuscripts to ChatGPT, asking whether they should be accepted, and then going for lunch. I mean using secure, confidentiality-preserving systems to help reconstruct argument structures, identify unsupported transitions, check whether the conclusion follows from the stated premises, and locate points at which polished prose may be concealing an absence of content. The reviewer would remain responsible for the judgment. But some of the labour involved could be reduced. I see this is happening already in econ, where tools like refine.ink are now being deployed at scale.</p><p>Journals could also provide actual incentives for thorough reviewing. They could compensate reviewers for especially demanding reports, ask less frequently, or recognise sustained reviewing work properly. The precise mechanism is less important than abandoning the fiction that meticulous intellectual labour is an infinitely renewable free resource.</p><p><strong>I nevertheless understand why a journal with PPA&#8217;s prestige would choose the policy it has.</strong> Prestigious journals are much more worried about false positives than false negatives. A bullshit paper that survives review and gets published creates a visible reputational cost. An outstanding paper that is rejected at the desk simply disappears. The journal bears the cost of the first error and almost none of the cost of the second.</p><p>So it is rational, from the journal&#8217;s perspective, to optimise against false positives. But that does not make it the best policy for philosophy. By Seth&#8217;s own account, journals do not merely distribute credentials; they also promulgate knowledge. A system designed primarily to protect a journal from reputational risks will reject potentially valuable departures whenever those departures make evaluation more difficult. It will preserve the brand, but it may do so by sacrificing exactly the kind of work we should want journals to discover.</p><p>Perhaps collaboration with AI will produce very little outstanding philosophy. Perhaps it will produce a great deal. At present, none of us knows. A blanket ban guarantees only that PPA will not be where we find out.</p><p>None of this requires denying the obvious problems. AI-generated papers can be hollow. Using AI to impersonate an understanding one does not possess is dishonest. Flooding journals with barely inspected manuscripts is parasitic. A paper whose author cannot explain or defend its central argument should not be published. But these are failures of quality, intellectual ownership, and submission norms. They are not identical to the fact that AI played a substantial role in producing the prose.</p><p>Indeed, PPA&#8217;s own editorial model demonstrates that the value of a paper is not exhausted by what the author managed to produce unaided. Its editors substantially improved my paper. They helped restructure and clarify arguments in ways I had not managed by myself. <br>I am not equating a human editor with a language model, but I want to highlight that scholarship already permits extensive transformation between an author&#8217;s first formulation and the work that eventually appears in print. What matters is whether the author remains intellectually responsible for the result.</p><p>AI has made bullshit cheaper. That is a real problem. But it has also made competent expression cheaper, potentially widened access to philosophical discussion, and created new ways of testing and developing ideas. A serious institutional response should try to preserve those benefits while controlling the corresponding abuses.</p><p></p>]]></content:encoded></item><item><title><![CDATA[On why Governance by Shame in universities is a dead end]]></title><description><![CDATA[When I was an undergraduate, a story used to circulate about a dinner attended by professors and graduate students from the economics department.]]></description><link>https://carlolc.substack.com/p/on-why-governance-by-shame-in-universities</link><guid isPermaLink="false">https://carlolc.substack.com/p/on-why-governance-by-shame-in-universities</guid><dc:creator><![CDATA[Carlo Ludovico Cordasco]]></dc:creator><pubDate>Sun, 16 Aug 2026 15:08:09 GMT</pubDate><content:encoded><![CDATA[<p>When I was an undergraduate, a story used to circulate about a dinner attended by professors and graduate students from the economics department. I heard it often enough that I can no longer remember who first told it, which is usually a sign that either everyone was there or nobody was.</p><p>I will call the two protagonists Bob and Betty.</p><p>By the time the story begins, dinner had been going on for a while. People had eaten, several bottles had been opened, and the senior professors had reached that stage of the evening when saying exactly what you think begins to look like a service to the profession.</p><p>Bob mentioned an economist whom he wanted the department to hire. Betty was not impressed.</p><p>&#8220;Do you know what George Stigler&#8217;s greatest ambition was?&#8221; she asked.</p><p>Bob did not.</p><p>&#8220;It was to work in a department where he was the dumbest person.&#8221;</p><p>Then came the pause. At least, there was always a pause when the story was told.</p><p>&#8220;Bob, you have already achieved Stigler&#8217;s greatest ambition. Why would you want to hire someone even dumber than you?&#8221;</p><p>Did this exchange happen exactly as reported? Who knows. Stories told by graduate students about tipsy professors are not always models of historical accuracy. This one may even have improved over the years, as academic stories sometimes do when the people telling them realise that reality has failed to provide a satisfactory punchline.</p><p>At the time, I thought it was a story about the pleasures and dangers of drinking with economists. I now think it was also a story about how universities govern themselves.</p><p>Betty did not explain why the candidate was bad. She did not discuss their work, compare them with the other applicants or offer reasons that Bob might answer. She did something much quicker. She made it embarrassing for Bob to continue supporting them.</p><p>Bob could still argue for the appointment, of course. But Betty had changed the social meaning of doing so. Any further defence of the candidate could now be treated not as an argument to be answered, but as further evidence that Bob was a poor judge.</p><p>The other people at the dinner had also learned something. Supporting a candidate whom Betty regarded as weak might come at a personal cost. Even people who had not formed a view about the candidate now knew which view was safer to express.</p><p>Perhaps Betty helped the department. Formal appointment procedures can tell people who votes, what evidence must be considered and how candidates should be compared. They cannot remove judgment. Someone still has to decide whether a paper is genuinely important, whether a research programme is promising and whether a candidate is likely to make the department better. Informal pressure can stop people from using that discretion carelessly. Shame is one form of pressure, and academics have always been unusually skilled at administering it.</p><p>But whether Betty helped depended on facts that her &#8216;joke&#8217; did nothing to establish.</p><p>The people around the table needed reason to trust her judgment. Perhaps she knew the field extremely well and had an excellent record of spotting both talent and mediocrity. They also needed reason to trust how she used her standing. Perhaps she was willing to humiliate Bob because the appointment really would have been disastrous. Or perhaps she disliked the candidate, resented Bob or simply enjoyed reminding the table who had the highest status.</p><p>Her insult could draw upon trust that had already been earned. It could not create that trust.</p><p>Today, a great deal of academic politics looks like that dinner after several more bottles, except that the conversation now takes place online and nobody ever goes home.</p><p>One group discovers that a university has hired another white man, someone who once signed an objectionable letter, or a scholar whose work is said to make members of some group unsafe. Another discovers that a university has hired someone with fewer publications than a rejected candidate, someone who wrote a diversity statement with suspicious enthusiasm, or a scholar whose work appears to consist largely of identity politics.</p><p>Screenshots appear. Publication records are compared. Old comments are dug up. Within a few hours, people who have never read the candidate&#8217;s work have reached firm conclusions about its quality. Strangers who could not identify the department on a map begin demanding resignations.</p><p>The two sides disagree about nearly everything, but they have come to agree about the method. Everyone in the culture war is now an economist, although usually a rather cynical one. Raise the cost of a bad appointment, and universities will make fewer bad appointments. If internal procedures cannot be trusted, outside pressure must do the work. Administrators who expect public embarrassment, donor anger, student protest or government investigation will think twice before allowing ideology to displace merit.</p><p>Sometimes this is right. Universities do make indefensible decisions. Confidentiality can protect applicants, but it can also protect favouritism, discrimination and ordinary dishonesty. An institution should not be able to end every legitimate question by announcing that the process was confidential. There are occasions when people outside the room are right to demand an explanation.</p><p>The problem begins when public shaming becomes the normal way of correcting academic judgment.</p><p>A public campaign does not impose a cost on bad judgment itself. It imposes a cost on whatever can be presented to an audience as bad judgment. Those are not the same thing. The result depends on which account spreads most easily, which allegations the audience is already inclined to believe and which group has enough influence to turn attention into consequences.</p><p>A crowd can assess a screenshot much more quickly than it can assess a book. It can count publications more easily than it can judge whether they are any good. It can understand that a candidate belongs to the wrong political camp long before it could understand the candidate&#8217;s argument.</p><p>Once this kind of pressure becomes common, universities begin to anticipate it. Job ads are rewritten. Certain words disappear from official documents while others suddenly appear everywhere. Committees learn which reasons they may safely state in public and which judgments are better left unspoken. People do not necessarily change their minds. They simply become more careful about admitting what they think.</p><p>After a while, the result starts to look natural. A settlement produced by political pressure comes to be described as ordinary professional good sense. The people who benefit from it tell themselves that standards have improved and that only extremists could object.</p><p>Then power changes hands.</p><p>Anyone who thought that victories in the academic culture war were permanent should consider how quickly policy on diversity, equity and inclusion changed after Donald Trump returned to office in January 2025. His administration used executive action against DEI programmes, opened civil-rights investigations into universities and moved to change the way the federal government deals with accreditors. The right began speaking about merit, accountability and institutional reform with much the same confidence that the left had recently brought to equity, inclusion and safety.</p><p>The people cheering this reversal should pay attention to how easily it happened. What had looked like a deep change in academic values turned out to depend heavily on who had won an election. There is no reason to think that the new arrangement will be more permanent simply because a different group is now enjoying it.</p><p>Each side mistakes its present ability to punish people for a lasting agreement about who deserves punishment.</p><h2>Mutually assured cancellation</h2><p>There is an obvious response to all this. Perhaps trust is too much to ask. Perhaps the best we can do is make sure that both sides have weapons.</p><p>You can ruin our candidate, but we can ruin yours. You can investigate our hiring programme, but the next government can investigate yours. You can pressure a publisher to drop one of our books, but we can pressure a university to cancel one of your conferences. Once everyone understands that an attack will bring retaliation, perhaps everyone will decide to behave.</p><p>Call it mutually assured cancellation.</p><p>The comparison with nuclear deterrence is initially attractive. The United States and the Soviet Union did not need to trust one another, admire one another or reach an agreement about political morality. Each side needed only to believe that an attack would lead to devastating retaliation. Hostility remained, but fear gave both sides a reason not to take the final step.</p><p>Academic politics lacks most of the features that make even this grim arrangement possible.</p><p>There are no presidents in charge of the arsenals. A dean cannot order every anonymous social-media account to stand down. No one agrees about where the red lines are. No one can even agree about which side attacked first.</p><p>A campaign begins against Professor A. The people behind it say they are responding to the campaign against Professor B. That campaign, according to its defenders, was itself a response to something involving Professor C five years earlier. Everyone believes that they are retaliating. Nobody believes that the score has been evened. By the time someone has reconstructed the history, another six academics have been denounced.</p><p>The incentives are different too. Launching a nuclear weapon is generally understood to have certain disadvantages. Launching an academic pile-on can bring attention, followers and the pleasant feeling of having defended civilisation before breakfast. Restraint, by contrast, is almost invisible. Nobody goes viral for looking at an accusation, thinking carefully about it and deciding not to help destroy a stranger&#8217;s career.</p><p>But suppose we could solve all these problems. Imagine two perfectly organised academic factions. Each can punish any attack on its members. Both understand the other side&#8217;s red lines, and neither has an incentive to escalate by mistake. We have finally created a stable balance of terror.</p><p>It would still be a dreadful way to run a university.</p><h2>The land of the reasonable</h2><p>Hostile states can keep their distance. They can remain behind their borders, communicate when necessary and otherwise leave one another alone.</p><p>Academics cannot do this. We keep putting our work, our reputations and sometimes our careers in one another&#8217;s hands.</p><p>We review one another&#8217;s papers and grant applications. We decide who gets hired, promoted and invited. We write references, edit journals, organise conferences and teach one another&#8217;s arguments to students. Every one of these activities requires someone to make a judgment that cannot simply be read off a rule.</p><p>No procedure can decide whether an argument is genuinely original. No table can tell us whether an objection is fatal or merely annoying. A publication count cannot reveal whether a scholar has opened a promising line of research or spent ten years producing minor variations on the same article. At some point, another person has to read the work and decide what she thinks of it.</p><p>Competent people will often disagree. One referee may think that a paper contains an important argument. Another may think that the argument rests on a confusion. One department may regard a candidate as unusually exciting. Another may find the same work thin or fashionable. Neither disagreement, by itself, shows that someone is corrupt or stupid.</p><p>Still, decisions have to be made. The author gets rejected. One candidate gets the job and the others do not. Someone receives the grant, the promotion or the invitation.</p><p>This leaves us vulnerable to one another. That vulnerability is uncomfortable, but it cannot be designed away completely. We could remove it only by removing much of the judgment that makes intellectual work intellectual. A university in which every decision followed automatically from a spreadsheet would be easier to administer. It would not necessarily be worth having.</p><p>Rules and procedures still matter. They can require people to state reasons, declare conflicts of interest and apply the same basic process to each candidate. They can make some forms of favouritism harder to hide. Appeals and outside review can expose decisions that no one could plausibly defend.</p><p>What procedures cannot do is make people use them honestly.</p><p>&#8220;Fit&#8221; can be a real consideration, or it can be a polite name for wanting someone who shares the department&#8217;s politics. &#8220;Excellence&#8221; can describe intellectual quality, or it can be defined around the achievements of the candidate whom a committee already prefers. Publication numbers can be treated as decisive when they help one candidate and dismissed as crude when they help another. External referees can be selected because their judgment is respected or because their answer is predictable.</p><p>A longer rulebook may make this harder, but it cannot make it impossible. People who already know which result they want can enforce one provision strictly, interpret another generously and then describe the outcome as procedurally neutral.</p><p>Rules work best when the people applying them are trying, however imperfectly, to reach a fair judgment. Once that effort disappears, the rules give the quarrel more language and more paperwork. They do not end it.</p><p>I know this sounds banal, but what academic life needs is reasonableness.</p><p>By this I do not mean political moderation. I do not mean that everyone should occupy the centre, avoid strong conclusions or become pleasant company at dinner (though that would be great!). Betty may have been entirely reasonable to think that Bob&#8217;s candidate was terrible, even if her way of expressing the view left something to be desired.</p><p>Reasonable academics understand that many professional judgments remain open to disagreement even when the people making them are competent and acting in good faith. They allow for the possibility that a colleague can reach a different conclusion without being corrupt, stupid or secretly committed to destroying the university.</p><p>They also place limits on how they use power. Before taking advantage of an ambiguity in a rule, they ask whether they could defend this as a normal way for universities to operate when the same ambiguity is used by people they oppose. Before turning a disputed appointment into a public scandal, they ask whether they would accept that response whenever someone objected to an appointment they supported.</p><p>It&#8217;s not just about whether I would tolerate retaliation. A determined partisan might be perfectly happy with a system in which everyone behaves badly. The question is whether I could defend the practice as a fair way for people who disagree to work together.</p><p>Reasonableness leaves plenty of room for blunt criticism. Some appointments really are indefensible. Cronyism, discrimination and dishonesty should be exposed. Academic freedom does not require us to admire every piece of scholarship or treat every decision as equally respectable.</p><p>But we still need to distinguish between a judgment we think is wrong and an abuse of power that demands public punishment. We need to leave open the possibility that someone who reached a bad conclusion remains a colleague rather than becoming an enemy.</p><p>Deterrence can make restraint sensible. I leave your candidates alone because I am afraid of what you will do to mine.</p><p>Reasonableness gives us a different reason. I do not use every available weapon against you because disagreement does not place you outside the academic practice to which we both belong.</p><p>Academia should be the land of the reasonable. This is not because academics are unusually kind or emotionally mature. Anyone who has attended a departmental meeting will know better. It is because academic work depends unusually heavily on people exercising judgment over others with whom they may strongly disagree.</p><p>A balance of terror protects factions that can retaliate. It does much less for people who belong to neither camp, whose views do not form one of the approved political packages, or who simply lack influential friends. Indeed, the people most in need of fair treatment may be those with no faction capable of taking revenge on their behalf.</p><p>If fear of my allies is the only thing stopping you from abusing your power over me, anyone without allies remains at your mercy.</p><h2>Trusting first</h2><p>I do not have a procedure for restoring trust. A stern memorandum will not do it. Nor will adding the word &#8220;trust&#8221; to a university&#8217;s values page or asking everyone to attend a workshop about collegial behaviour.</p><p>Trust cannot simply be declared. It grows when people repeatedly behave in ways that make trusting them less foolish.</p><p>We see this when someone declines to use a weapon that would help her own faction. We see it when an academic acknowledges the quality of a scholar whose politics she dislikes, criticises indefensible behaviour by an ally or admits that an appointment she would not have made was still one that reasonable people could support.</p><p>We also see it when people allow an opponent to be wrong without trying to make that mistake the most expensive event of her professional life.</p><p>These acts are difficult because someone has to begin before trust is secure. Someone must exercise restraint without knowing whether others will return it. From within the logic of deterrence, this looks like unilateral disarmament. Within a university, it looks like behaving as though the other person might still be a colleague.</p><p>The alternative is to keep raising the cost of every alleged transgression. That may prevent a few bad appointments. It will also teach academics to hide what they think, flatter whichever faction is strongest and treat every hiring dispute as a possible opportunity for political escalation. Each victory will give the losing side another reason to seek revenge when power changes hands.</p><p>An academia held together by fear may achieve an armed peace. It will not become a community of inquiry. At best, as ChatGPT keeps suggesting to me, &#8220;it will be a cold war with footnotes&#8221;. Boy, if I hate LLMs&#8217; metaphors.</p><p><strong>P.S.</strong> Dear friends,</p><p>If you managed to read this post to the end, I&#8217;m extremely grateful. I also wanted to let you know that I&#8217;ll be taking a break from the blog.</p><p>I&#8217;m no longer sure that writing public-facing essays is for me. After my Aeon piece, I discovered that I&#8217;m not very good at dealing with comments that are partisan, prejudiced or simply uncharitable. Online discussions tend to bring out the worst in me, and life offline is far too good to spend it arguing with strangers.</p><p>There is too much great wine to drink, too many running paths I haven&#8217;t explored and too many papers I still want to write. I&#8217;m not a gifted analytic philosopher, but I do love analytic philosophy. I would rather spend my time trying to become a slightly better philosopher than becoming a slightly worse person online.</p><p>Perhaps I&#8217;ll return when I have something I genuinely want to say. For now, thank you for reading the blog, sharing the posts and occasionally letting me know that something I wrote was worth your time. That has meant a great deal to me.</p>]]></content:encoded></item><item><title><![CDATA[Rest In Peace Jason Arday]]></title><description><![CDATA[I am devastated by the death of Jason Arday.]]></description><link>https://carlolc.substack.com/p/rest-in-peace-jason-arday</link><guid isPermaLink="false">https://carlolc.substack.com/p/rest-in-peace-jason-arday</guid><dc:creator><![CDATA[Carlo Ludovico Cordasco]]></dc:creator><pubDate>Sat, 15 Aug 2026 12:33:13 GMT</pubDate><content:encoded><![CDATA[<p>I am devastated by the death of Jason Arday. I knew nothing about him before this whole media campaign began, and I had no desire whatsoever to learn more about the alleged plagiarism. My immediate reaction was that such accusations should be handled by the relevant universities and journals, through procedures designed to establish what happened, rather than weaponised for whichever bloody culture war currently needs ammunition.</p><p>What I find equally disheartening is that the reactions to Arday&#8217;s death have immediately become another round of the same culture war. It has absorbed everyone so thoroughly that it seems almost impossible to find anyone weighing what actually matters: the unnecessary suffering this whole story has generated. On one side, we have dumb takes about academia and meritocracy. On the other, we have people naming and shaming white professors who have also been accused of plagiarism but were not subjected to the same degree of public scrutiny. Apparently, the answer to disproportionate public punishment is to identify more people to punish.</p><p>This has to end. I am so sick of it.</p><p>And, in a way, I am glad that AI is dismantling some of the self-importance academics have accumulated around their profession. I do not mean that I am glad people (me included) might lose their jobs. I mean that I will be glad if academic employment ceases to be treated as a sacred calling that confers some special status on the person performing it.</p><p>Academia is a job. It can be useful and important, but so can countless other jobs. There is nothing morally special about being an academic, just as there is nothing morally special about being a banker or a plumber. An academic appointment is not a certificate of superior human worth. Failing to obtain one does not make someone a lesser person. Obtaining one does not make someone a more deserving human being.</p><p>Once we remember this, the argument about meritocracy becomes very simple.</p><p>We need competent institutions because we want them to solve problems. Universities should produce knowledge, educate people, help society understand the world, and ultimately contribute to reducing unnecessary suffering. If they allocate positions badly, they will do these things less effectively. In that perfectly ordinary sense, universities should be meritocratic: they should appoint the people most capable of performing the relevant work.</p><p>But this has nothing to do with moral desert. I do not believe that the person with the strongest publication record morally deserves an academic job more than everybody else. The point of appointing that person is not to reward their superior virtue. It is to place people where their abilities are most likely to produce valuable results.</p><p>Culture-war understanding of meritocracy turns every disputed appointment into a moral scandal. If someone is thought to have been appointed without sufficient merit, the successful candidate is treated as though they had stolen something belonging to a more deserving person. Their career becomes an insult to every supposedly better scholar who failed to obtain the same position. It then seems legitimate to examine every part of their life, amplify every error and expose them to unlimited public contempt.</p><p>That does not follow.</p><p>A university can make a bad appointment without the person appointed having committed a moral wrong. Candidates do not run hiring committees. They do not determine the standards by which applications are assessed, and they are not required to possess an impartial, God-like understanding of how they compare with every other possible candidate. Responsibility for an appointment lies principally with the institution that made it.</p><p>Plagiarism is, of course, a different matter. It is a violation of academic standards because it damages the practices through which knowledge is produced and trusted. Credible allegations should be investigated. Where universities or journals refuse to investigate properly, careful public scrutiny might be justified. <strong>But the seriousness of an allegation does not make every possible response to it legitimate.</strong> There is a difference between presenting evidence and setting up a show; between demanding an investigation and conducting a campaign of personal destruction; between criticising someone&#8217;s work and turning that person into a symbol of everything one hates about contemporary academia.</p><p>What useful purpose does the public show serve? Is the information accurate and relevant? Is the response proportionate? Could the institutional problem be addressed without imposing the same cost on an individual? And, once the evidence has been presented and an investigation demanded, what does the hundredth denunciation contribute?</p><p>Usually, very little. Each person joining a pile-on can tell themselves that their own contribution is insignificant. But the suffering is cumulative. The thousandth person is adding to something being experienced by an actual human being.</p><p>If the defence of these campaigns is meritocracy, then their defenders must explain how the personal suffering they produce is necessary for achieving more meritocratic institutions. It is not enough to invoke the importance of academic standards in the abstract. The means must make some plausible contribution to the end, and that contribution must be sufficient to justify the foreseeable harm.</p><p>In this case, the legitimate institutional questions could have been raised without making Jason Arday the object of an obsessive public campaign. The evidence could have been submitted to the relevant bodies. The adequacy of their investigations could have been questioned. Cambridge&#8217;s appointment procedures and public statements could have been scrutinised. None of that required treating one person&#8217;s entire life as public property.</p><p>The response from the other side is hardly better. Naming white academics accused of similar conduct may expose inconsistent enforcement, and inconsistent enforcement is a real institutional problem. But trying to correct that inconsistency by subjecting more individuals to the same treatment simply reproduces the original wrong. Equal cruelty is not justice. </p><p>There is no contradiction in saying that preventing plagiarism is important and that Jason Arday was treated cruelly. There is no contradiction in wanting universities to appoint competent people while denying that academic jobs measure human worth. There is no contradiction in demanding an investigation while refusing to participate in a public pile-on.</p><p>Our culture-war machinery cannot hold these thoughts together. It requires innocence or guilt, hero or fraud, vindication or destruction. The person disappears and is replaced by a symbol that each side can use against the other.</p><p>Jason Arday was a person before he became anyone&#8217;s symbol. Whatever the eventual judgment about his academic work might have been, he was capable of suffering. He had people who loved him and who now have to live with his absence. That fact should carry more moral weight than the satisfaction of winning another argument about the state of academia.</p><p>If, as I believe, the ultimate justification for better institutions is that they help us reduce unnecessary suffering, then we cannot treat the unnecessary suffering inflicted in their defence as irrelevant. Nothing about academic standards required this. </p>]]></content:encoded></item><item><title><![CDATA[No, letting AI write for us won't end civilisation]]></title><description><![CDATA[Saying dumb things about AI might, though]]></description><link>https://carlolc.substack.com/p/no-letting-ai-write-for-us-wont-end</link><guid isPermaLink="false">https://carlolc.substack.com/p/no-letting-ai-write-for-us-wont-end</guid><dc:creator><![CDATA[Carlo Ludovico Cordasco]]></dc:creator><pubDate>Fri, 07 Aug 2026 23:16:18 GMT</pubDate><content:encoded><![CDATA[<p>I&#8217;ve just finished reading Bret Stephens&#8217;s <a href="https://www.nytimes.com/2026/08/04/opinion/artificial-intelligence-ai-writing.html?unlocked_article_code=1.3FA.O0cB.j5hat-2QyTNd&amp;smid=url-share">piece</a> on the NYT and I am bored. And exhausted. And frustrated. So bored, exhausted, and frustrated that I don&#8217;t feel like delegating this blog post to Sol Extra High (yes, I&#8217;ve switched back to ChatGPT as my daily driver. Fable 5 had recently started writing like Deleuze). I will go back to writing as therapy, and the therapy I have in mind consists of vomiting all my frustration towards those brainy smurfs telling us how to use AI, and warning us that letting AI write for us is the end of the world, or civilisation, or whatever thing is supposed to end.</p><p>Yes, writing is thinking. What an incredible intuition you&#8217;ve got there, Bret! Never heard of it! Of course, a lot of times (before LLMs), I started drafting a paper with an idea in mind. And then writing helped me realise that the idea was dumb, or that it needed some refinement, or that &#8216;things are not as simple as they look&#8217; (a concept that brainy smurfs should start to take seriously). So, yes, writing helped me think around things. Own the argument, so to speak. Verify its soundness and robustness. Of course, I would still get things wrong most times, but there is a sense in which writing helped shape concretely the idea I had in mind.</p><p>But, hey, hear me out cause I have another breakthrough intuition: reading is also thinking! Didn&#8217;t perhaps Elena Ferrante make you deeply reflect on the relationship with your childhood best friend? Or on why we are irreparably attracted to narcissists? Didn&#8217;t Melville and Captain Ahab make you reflect on the danger of turning an injury into the sole purpose of your life? Didn&#8217;t Proust make you wonder which apparently forgotten part of your life might still be waiting inside a taste, a smell, or a room? </p><p>Writing is thinking. Reading is thinking. Writing is often just reading something that did not exist thirty seconds ago. You write a sentence, look at it, feel that something is wrong, and try again. The sentence has made your thought visible enough for you to disagree with it.</p><p>But hey, and here is the groundbreaking intuition: <strong>notice that the thinking does not reside magically in the act of typing the first sentence. It happens in the confrontation between what is in front of you and your sense that it is inadequate.</strong> Perhaps the sentence is false. Perhaps it is true but boring. Perhaps it says what you meant while revealing that what you meant was kinda meh. Perhaps it contains an implication you had not noticed. You become the reader of your own words, and that peculiar form of reading sends you back into thought.</p><p>Now suppose I have an idea for an argument. It&#8217;s still half-baked, but I know more or less what I want to say. I ask ChatGPT to have a go at writing it. It does, and I immediately see a problem. So I push back. ChatGPT tries again. That won&#8217;t do either. And while trying to explain why it won&#8217;t do, I finally understand what I was trying to argue in the first place. ChatGPT wrote the first version. Fine. But when, exactly, did I stop thinking?</p><p>The brainy smurf answer appears to be that none of this really counts. Apparently cognition resides in the fingertips: unless every word has passed individually through my keyboard, the resulting thought is counterfeit. Socrates could think through conversation; Proust&#8217;s could recover an entire world through the taste of a madeleine; my colleagues can make me rethink a paper by asking a question at a seminar. But if a machine writes down an argument, civilisation is over.</p><p>Of course, you can use AI without thinking. You can type &#8220;write me 1,000 words on Public Reason&#8221;, copy the answer, and remain gloriously untouched by any cognitive event. But people have always written without thinking (or at least that&#8217;s my interpretation of most French philosophy). They have padded essays, recycled arguments, paraphrased other people&#8217;s sentences, and written entire books whose principal achievement was to leave every relevant concept exactly where they found it. The fact that writing <em>can</em> be thinking never meant that every manually typed sentence was evidence of thought. Nor does the fact that AI <em>can</em> replace thinking mean that every AI-written sentence must have done so.</p><p>I don&#8217;t want to indulge in amateur sociology or psychology, but I am being forced to. We have a problem, folks, and it isn&#8217;t simply vested interests. Yes, writing well and quickly used to be an extraordinary comparative advantage, one that brought status and a beefy income premium. But something sillier is going on. People cannot shake the thought that if a particular skill made them smart (or at least made other people regard them as smart), then anyone who no longer needs to exercise that skill must be becoming dumb. The possibility that new tools might simply change where intelligence is exercised never quite bothers them.</p><p>As though, at some blessed moment in history, we discovered the one true path to civilisation, harmony, and human talent, and, miraculously, it consisted in doing things exactly as these people learned to do them, without letting AI write for us. I find this idea almost unbearably pretentious and astonishingly na&#239;ve.</p><p>When Del Piero left Juventus in 2012, I thought football didn&#8217;t make sense anymore. I had a brief crush on Allegri a few years later but, broadly speaking, football stopped making sense to me the day everyone seemed perfectly happy to carry on as though Juventus without Del Piero was still Juventus.</p><p>This sounds absurd because it is absurd. Football had not stopped making sense. It had stopped making sense <strong>to me</strong>. Somewhere, children were falling in love with players whose names I could barely remember and whose haircuts I found offensive. They weren&#8217;t experiencing football incorrectly. They had simply not built their entire relationship with the sport around Alessandro Del Piero.</p><p>I could have turned my grief into a grand theory. I could have argued that football had lost its soul, that the new generation lacked loyalty, that no one understood the beauty of a proper number ten anymore. I could probably have found a graph showing that attention spans began to collapse around the time Del Piero left Turin. But it would still have been my nostalgia dressed up as social science.</p><p>This is roughly what the brainy smurfs are doing with &#8216;pure&#8217; writing. Their Del Piero is the ability to turn a blank page into polished prose without any AI assistance. They (me included) built identities, careers, and self-esteem around that ability. Now a machine can do an indecent amount of it in a few seconds. This is painful. We are entitled to mourn. God knows I did when Del Piero left.</p><p>What they are not entitled to do is turn that mourning into a theory of human cognition. </p><p>Perhaps AI really will make us dumber. It is perfectly possible. If we stop checking/understanding what it says, lose the patience required to follow an argument, or become unable to express a thought without asking a machine to guess what it is, then yes, something important will have been lost. But that case has to be made. You cannot establish it simply by pointing out that people will no longer need to struggle in precisely the ways you struggled. Some of that struggle trained the mind. Some of it was just friction. </p><p>Del Piero left. Football continued. The bastards.</p>]]></content:encoded></item><item><title><![CDATA[On Leopold Aschenbrenner, Max Allegri, Citadel]]></title><description><![CDATA[and the Hard Law of Risk Management]]></description><link>https://carlolc.substack.com/p/on-leopold-aschenbrenner-max-allegri</link><guid isPermaLink="false">https://carlolc.substack.com/p/on-leopold-aschenbrenner-max-allegri</guid><dc:creator><![CDATA[Carlo Ludovico Cordasco]]></dc:creator><pubDate>Fri, 31 Jul 2026 10:27:22 GMT</pubDate><content:encoded><![CDATA[<p>For those of you who follow football, I have a confession to make: I am a big, big fan of Massimiliano Allegri.</p><p>I simply like the guy. He is funny, approachable and appears to know how to enjoy life. More importantly, I loved him as Juventus manager.</p><p>Nobody would describe Allegri as football&#8217;s most proactive coach. He does not begin with an elaborate vision of how every match must unfold and then demand that reality conform to it. He is excellent at managing risk, reading the game and exploiting the vulnerabilities of almost any opponent. He is perfectly happy to let the other team have the ball, appear more ambitious and collect compliments for playing the better football. What matters is retaining control of the match until the opportunity arrives.</p><p>To Italian readers, Allegri is almost the embodiment of <em>La dura legge del gol</em>, a 1997 song by the band 883. For English readers, the title means something like &#8220;the cruel law of goals.&#8221; Its central lesson is contained in six words: &#8220;<em>se non hai difesa, gli altri segnano</em>&#8221;&#8212;without a defence, the opposition scores. You may play the more beautiful football, but if you expose yourself and the other team takes its chance, it wins. The scoreboard is unmoved by your superior intentions.</p><p>So here is the question that has been bothering me. If I admire Allegri so much, why did I instinctively cast Ken Griffin as the villain in the fall of Leopold Aschenbrenner?</p><p>I have nothing personally against Griffin. Until yesterday, I had barely heard of him. I am sure he is a lovely guy, even if he appears to have some rather dumb ideas about AI. You can watch the relevant clip <a href="https://www.youtube.com/watch?v=5Ez0iC9x3ds">here</a>. More importantly, there is currently no evidence that Griffin or Citadel engineered Aschenbrenner&#8217;s collapse.</p><p>Nevertheless, the moment I learned that Citadel had acquired most of Aschenbrenner&#8217;s public-equity portfolio, I wanted Griffin to be the villain.</p><p>Why?</p><h2>A very twenty-first-century bank run</h2><p>Aschenbrenner is a former OpenAI researcher who, while still in his early twenties, founded a hedge fund called Situational Awareness. The name was not accidental. His central thesis was that artificial intelligence was advancing much faster than most people and institutions understood, and that the physical infrastructure required to support it&#8212;chips, memory, electricity, land and data centres&#8212;would become extraordinarily valuable.</p><p>The thesis was spectacularly successful. Situational Awareness reportedly grew to approximately $20 billion in assets under management in about two years. Its public portfolio contained large positions in precisely the companies that stood to benefit from the AI infrastructure buildout: memory manufacturers, power providers, data-centre operators and the so-called neoclouds.</p><p>Then, over the course of July, the trade turned violently against him.</p><p>AI-infrastructure stocks began falling together. A peculiar narrative took hold: perhaps the extraordinary amount of compute being built would soon prove excessive; perhaps hyperscalers would begin renting their unwanted capacity to others; perhaps neoclouds would be squeezed between falling compute prices and competition from companies with much larger balance sheets. Compute, we were repeatedly told, might be on the verge of becoming abundant.</p><p>This story never made much sense to me. Microsoft, Meta, Amazon and the frontier AI laboratories continued to describe themselves as compute-constrained. Capital-expenditure plans remained enormous. Demand for advanced chips and high-bandwidth memory remained extremely strong. Yet, for several weeks, the market behaved as though the AI-infrastructure thesis had been falsified.</p><p>I should disclose that I was not a neutral observer. I own shares in IREN, one of the companies caught in the sell-off. I also find Aschenbrenner&#8217;s account of AI progress substantially more convincing than Griffin&#8217;s. The decline was therefore not an abstract case study for me. I watched a thesis I believed to be broadly correct get demolished in real time by a narrative I believed to be largely nonsensical.</p><p>We now know more about what was happening underneath the prices.</p><p>According to the contents of an investor letter reported by the <em>Wall Street Journal</em>, Situational Awareness lost approximately 67 per cent in July. The losses generated margin calls from its lenders and eventually forced the firm to sell most of its public-stock portfolio to Citadel. Aschenbrenner told his investors that the fund had now removed all leverage.</p><p>The scale of the July loss is astonishing. Yet &#8220;collapse&#8221; requires qualification. Situational Awareness&#8217;s gains earlier in the year had been so enormous that, even after losing roughly two-thirds of its value in one month, the fund reportedly remained approximately 80 per cent up in 2026. It also retained a portfolio of private investments, including its stake in Anthropic.</p><p>Aschenbrenner had not lost everything. What he had lost was control of his public portfolio.</p><p>Most importantly, the investor letter partly blamed short sellers who had targeted the fund&#8217;s positions. It compared what happened to a bank run.</p><p>This is more than a theory invented on social media. It is the fund&#8217;s own contemporaneous account of the episode, albeit reported through a person familiar with the letter rather than independently demonstrated through trading records. It gives us evidence that Aschenbrenner believed other traders had identified his positions and were deliberately trading against his increasingly constrained ability to hold them.</p><p>There is a name for this mechanism in the academic literature: <a href="https://www.princeton.edu/~markus/research/papers/predatory_trading.pdf">predatory trading</a>.</p><p>Suppose a leveraged investor owns a large position in a relatively illiquid stock. Other traders discover, or correctly infer, that the investor may soon need to sell. They begin shorting the same stock. The resulting decline reduces the value of the investor&#8217;s collateral, prompting lenders to demand additional cash. To produce that cash, the investor must sell some of the position, pushing the price down further and validating the short sellers&#8217; original bet.</p><p>The short sellers can then cover their positions&#8212;or become buyers themselves&#8212;during the eventual forced liquidation.</p><p>Falling prices create forced sales. Forced sales create falling prices. The investor&#8217;s fundamental thesis may not have changed at all, but the market has shortened the period for which he is allowed to hold it.</p><p>This kind of trading need not involve illegal manipulation or explicit coordination. A number of funds can independently reach the same conclusion: the positions are visible, the investor appears overextended, and whoever sells first will be able to buy back later at a lower price. Each trader&#8217;s individually rational action helps bring about the liquidation everyone anticipates.</p><p>The short sellers do not need to believe that Aschenbrenner is wrong about compute. They need only believe that he cannot remain solvent long enough to be proven right.</p><h2>Reconstructing the spiral</h2><p>Some parts of the sequence remain speculative, but we can now construct a reasonably coherent account of how the crisis may have developed.</p><p>Situational Awareness held very large positions in a cluster of correlated AI-infrastructure stocks. Many of those holdings were publicly visible through regulatory filings. The exact amount of leverage remains unknown; social-media claims of fourfold leverage have not been verified. What the investor letter does establish is that leverage existed, that lenders made margin calls and that the firm has now removed it.</p><p>At approximately the same time, Aschenbrenner was attempting to participate in the US offering of SK Hynix, the leading producer of the high-bandwidth memory used in AI accelerators. Situational Awareness was one of three cornerstone investors&#8212;alongside Baillie Gifford and Coatue&#8212;to indicate interest in purchasing as much as a combined $7 billion in the offering.</p><p>The <a href="https://www.sec.gov/Archives/edgar/data/2120882/000119312526299963/d32785d424b4.htm">$7 billion indication</a> was shared across all three investors, was non-binding, and does not tell us the size of Situational Awareness&#8217;s eventual allocation. We cannot say that Aschenbrenner personally attempted to purchase $7 billion of Hynix.</p><p>It is nevertheless plausible that he was reducing some existing positions to make room for a substantial new investment. If so, his own selling may have contributed to the initial weakness in the same stocks whose falling prices would later place him under pressure. Once short sellers recognised the vulnerability, they could accelerate the process. Lower prices generated margin calls; margin calls required more liquidity; attempts to obtain liquidity confirmed that the fund was in difficulty.</p><p>The fund reportedly attempted to raise fresh capital. When it could not obtain enough money quickly enough, an orderly open-market exit was no longer realistic. A multibillion-dollar position cannot be sold at the price displayed on a trading screen. Selling such a portfolio over several days would advertise the liquidation to every sophisticated participant in the market. Buyers would withdraw their bids, short sellers would anticipate the remaining orders, and each successive sale would produce a worse price.</p><p>A private portfolio transaction offered something the public market no longer could: immediacy and certainty.</p><p>Citadel could agree a price for a large collection of positions, absorb them at once and hedge the risk through options, futures, shorts and related securities. It did not need to share Aschenbrenner&#8217;s long-term confidence in every company. It needed only to calculate the discount at which the entire package became attractive.</p><p>Reporting still differs slightly on the transaction&#8217;s precise scope. <a href="https://www.axios.com/2026/07/30/ai-hedge-fund-situational-awareness-citadel">Axios described</a> Citadel as purchasing the entire public-equity portfolio, while the <a href="https://www.ft.com/content/5fb44089-ecdf-4b48-bc14-1e8b4682b142">Financial Times</a> described it as purchasing a large portion of approximately $16 billion in public-equity holdings. The price has not been disclosed.</p><p>Then came the almost unbearable irony. As the hyperscalers reported earnings, they continued to describe strong AI demand, enormous capital expenditure and insufficient compute. Many of the same infrastructure stocks that had been collapsing rebounded violently. The compute-abundance narrative began to weaken almost immediately after Aschenbrenner had surrendered control of his public portfolio.</p><p>This sequence makes the darker interpretation psychologically irresistible. Aschenbrenner&#8217;s positions were known. His investor letter says short sellers targeted them. Griffin had publicly expressed scepticism about AI capex. Citadel then emerged as the buyer after the short sellers and the margin calls had done their work.</p><p>But we must stop where the evidence stops.</p><p>We now have evidence of a hunt. We do not know the identities of the hunters, and we certainly do not know that the eventual buyer was one of them. There is no evidence that Griffin or Citadel created the compute-abundance narrative, coordinated the short selling or deliberately triggered the margin calls. Citadel may simply have remained liquid until Aschenbrenner desperately needed someone with its balance sheet.</p><p>That may be precisely what I find unsettling.</p><h2>The Size of the Pie</h2><p>Allegri and the traders who targeted Aschenbrenner possess a similar kind of intelligence.</p><p>They minimise unnecessary exposure. They allow the other side to advance. They wait for a vulnerability and attack when the opponent is least capable of responding. They do not need to predict the whole game correctly. They need only to preserve their own agency until the other side loses its.</p><p>Citadel&#8217;s intervention follows the same basic logic. Aschenbrenner had the long-term thesis. Citadel had the liquidity. Aschenbrenner needed the future to arrive; Citadel needed only to survive the present.</p><p>Why, then, do I see prudence in Allegri and predation in finance?</p><p>The first answer is embarrassingly personal. Allegri used to set his traps on behalf of my football team. In this case, I identified with the person caught in the trap. I shared Aschenbrenner&#8217;s broad thesis and owned one of the affected stocks. Perhaps I admire opportunism when my side practises it and discover moral objections when it is practised against me.</p><p>But I think there is a more defensible distinction.</p><p>Football is a zero-sum game. Juventus can win only if its opponent does not. The purpose of the activity is supplied in advance and accepted by everyone: score more goals than the other team. Allegri has no obligation to enlarge the number of victories available to both sides. There is no larger pie to create.</p><p>Finance is different, or at least it is supposed to be.</p><p>Some financial activity is effectively redistributive. A trader recognises that somebody else will be forced to sell, gets ahead of the sale and later purchases the same asset at a discount. One participant&#8217;s superior execution becomes another participant&#8217;s loss.</p><p>But capital allocation can also be positive-sum. Capital can finance new power generation, semiconductor fabrication, data centres and scientific research. It can create productive capacity that did not previously exist. When investors correctly identify an important technological bottleneck and direct resources towards overcoming it, they can increase the size of the economic pie.</p><p>This is what I find attractive about people like Aschenbrenner. He was not merely predicting which line on a screen would move next. He had a substantive thesis about the future: intelligence would become enormously valuable; compute would remain scarce; and the companies capable of supplying chips, memory, electricity and data-centre capacity would enable a vast expansion in what the economy could accomplish.</p><p>That does not mean that every trade he made directly financed new productive capacity. Buying existing shares from another investor is not the same thing as paying for a new data centre. Yet valuations affect companies&#8217; ability to raise equity, borrow money and build. His participation in the Hynix offering was more direct still: it placed capital into a primary issuance intended partly to support expanded production.</p><p>Citadel&#8217;s expertise is of another kind. Its advantage lies in understanding positioning, liquidity, correlations, financing and constraints. It can profit from the relationship between other people&#8217;s beliefs and their ability to continue holding those beliefs.</p><p>Aschenbrenner was asking: <em>What will the world need?</em></p><p>The traders targeting his portfolio were asking: <em>How long can he afford to wait?</em></p><p>Citadel&#8217;s final question was simpler still: <em>At what price does his loss of choice become our opportunity?</em></p><p>Those questions can be answered brilliantly without having any substantive account of which future is worth building.</p><h2>The complication</h2><p>It would nevertheless be too easy to make Aschenbrenner the visionary hero and Griffin the parasitic villain.</p><p>Citadel supplied something real: immediate liquidity. By purchasing the portfolio privately, it may have prevented a chaotic open-market liquidation from pushing the same stocks much lower. The capacity to value, absorb, and hedge risk is economically useful. A distressed seller needs a buyer, and the buyer willing to act when nobody else will is entitled to demand a discount.</p><p>The same transaction can therefore look predatory from one side and stabilising from the other. Citadel benefited from Aschenbrenner&#8217;s disappearance as a voluntary economic agent, but its purchase may also have stopped that disappearance from inflicting greater damage on everyone else.</p><p>Nor does possessing a substantive vision excuse catastrophic risk management. If Aschenbrenner believed that the AI buildout would unfold over a decade, he should not have financed that belief in a way that allowed several weeks of adverse prices to determine whether he could remain invested.</p><p>A long-term thesis financed by short-term obligations is not truly a long-term position. Someone else owns the clock.</p><p>This is where Allegri returns to the story. Perhaps Aschenbrenner&#8217;s problem was not that finance rewards people like Allegri. Perhaps it was that he did not have enough Allegri in him.</p><p>Allegri&#8217;s risk management is not the absence of ambition. It preserves his team&#8217;s agency. He keeps the match alive until the moment at which his players can decide it. Aschenbrenner&#8217;s leverage did the opposite. It amplified the rewards from his vision, but it eventually transferred control over his time horizon to lenders, market prices and short sellers.</p><p>His opponents did not have to disprove his view of the future. They had only to make his route to that future financially impassable.</p><p>That may be the cruelest law of finance: it is not enough to be right. You must remain sufficiently liquid, solvent and independent to reach the period in which you will be proven right.</p><p>Perhaps this is why I still cannot admire Griffin in quite the way I admire Allegri. Allegri&#8217;s discipline exists in the service of an identifiable objective: a team winning the match. Aschenbrenner&#8217;s risk-taking existed in the service of a thesis about which productive capacity the future would require. Citadel&#8217;s discipline need not be attached to any first-order project beyond obtaining an attractive return.</p><p>A financial system populated entirely by Aschenbrenners might repeatedly blow itself up while trying to build the future. A system populated entirely by Citadels might be exceptionally liquid, exceptionally well-hedged and entirely dependent on somebody else to decide which future deserves to be built.</p><p>We need direction and discipline. Yet something remains troubling when the people with the strongest ideas about how to enlarge the pie repeatedly lose ownership to those whose greatest talent is recognising when they can no longer wait.</p><p>I still do not know whether Ken Griffin deserves to be the villain of this story. On the available evidence, probably not. The actual villain may be the financing structure that transformed a claim about the next decade into a position that had to survive every week along the way.</p><p>But I now understand why my mind cast Griffin in that role. He did not need to prove Aschenbrenner wrong. He needed only to survive him.</p><p>And, as Allegri could have told Leopold, sometimes that is enough to win.</p>]]></content:encoded></item><item><title><![CDATA[How much compute do we have? And how much will we need?]]></title><description><![CDATA[A simple exercise]]></description><link>https://carlolc.substack.com/p/how-much-compute-do-we-have-and-how</link><guid isPermaLink="false">https://carlolc.substack.com/p/how-much-compute-do-we-have-and-how</guid><dc:creator><![CDATA[Carlo Ludovico Cordasco]]></dc:creator><pubDate>Thu, 30 Jul 2026 12:00:10 GMT</pubDate><content:encoded><![CDATA[<p>This is a bit of a detour from my usual pieces of writing. Yesterday, I was reading Dwarkesh&#8217;s excellent <a href="https://www.dwarkesh.com/p/why-compute-might-get-10x-more-expensive">post</a> on the rising cost of compute, and I realised I have thoughts.</p><p>A big disclaimer first: I am much less persuaded than Dwarkesh that the long-run price of compute will reduce (assuming long-term isn&#8217;t 150 years as I don&#8217;t have meaningful predictions to make for such a time-span). I do think that we will remain compute-constrained for a long time. Compute may therefore be the principal constraint both on further development of frontier models and on the mass adoption of open-weight models, even at their current capabilities.</p><p>But long-term compute prices are not the subject of today&#8217;s post. I want to run a much simpler exercise. Assume, absurdly conservatively, that AI capabilities never substantially improve from where they are today. Where would that leave us on the adoption curve? How much economically productive adoption remains available using only capabilities that already exist?</p><p>The answer depends on what we mean by adoption. A tenfold increase in the number of users would be difficult. But productive inference depends on several variables: the number of users, the number of workflows in which AI is embedded, how frequently those workflows run, and how much computation each task consumes.</p><p>Today, only <a href="https://www.census.gov/library/working-papers/2026/adrm/CES-WP-26-25.html">around 18 per cent of US firms report using AI in any business function,</a> while only 12 per cent of workers use it daily (and I suspect the numbers are <em>much</em> lower in Europe). And &#8220;using AI&#8221; may mean asking ChatGPT to rewrite an email once a week. It generally does not mean embedding AI throughout customer support, coding, document processing, internal search, sales, compliance, analytics and administrative work.</p><p>Suppose the number of regular users eventually rises fourfold. Suppose each user runs AI across three to five times as many workflows. And suppose richer context, tool use and repeated checking make each workflow consume one to three times as much inference. That already gives us a plausible range of roughly 12 to 60 times today&#8217;s useful inference activity.</p><p>So I would put the remaining productive-adoption opportunity somewhere between 10x and 50x. The low end looks very conservative to me; the high end requires AI to become deeply embedded in routine work. This is a claim about useful inference performed, not about electricity consumption. Better chips, quantisation, smaller models and more efficient serving will reduce the power required for each unit of work.</p><p>Now consider the physical base from which this adoption must be served.</p><p>Estimates differ, but global AI-oriented data-centre capacity is probably around <a href="https://epoch.ai/data-insights/ai-datacenter-power/">30&#8211;40 GW today</a>. The <a href="https://epoch.ai/data-insights/ai-supercomputers-performance-share-by-country">US and Europe host roughly 80 per cent</a> of measured frontier-compute performance, overwhelmingly in the US. That suggests perhaps 22&#8211;30 GW of Western AI facility capacity.</p><p>But &#8220;available&#8221; would be the wrong word. Almost all of that capacity is already owned, leased or committed. <a href="https://www.jll.com/en-us/newsroom/global-data-center-sector-to-nearly-double-to-200gw-amid-ai-infrastructure-boom">Global data-centre occupancy is around 97 per cent</a>, and much of the small amount of vacant space cannot support modern high-density GPU racks. There may be less than 1 GW of genuinely AI-ready capacity immediately available to new customers.</p><p>What happens over the next five years?</p><p><a href="https://www.jll.com/en-us/newsroom/global-data-center-sector-to-nearly-double-to-200gw-amid-ai-infrastructure-boom">JLL expects total global data-centre capacity to rise from approximately 103 GW to 200 GW by 2030</a>, with AI accounting for around half of the final total. That would give us roughly 100 GW of AI capacity.</p><p>McKinsey is more aggressive. <a href="https://www.mckinsey.com/featured-insights/week-in-charts/the-future-of-ai-workloads">Its base case projects 219 GW of total data-centre demand in 2030, including:</a></p><ul><li><p>93 GW for AI inference;</p></li><li><p>62 GW for AI training;</p></li><li><p>64 GW for conventional cloud, storage and enterprise workloads.</p></li></ul><p>But our thought experiment assumes no further increase in capabilities. Under that assumption, there would be little reason for frontier-training demand to rise from approximately 23 GW today to 62 GW. Keep training and fine-tuning roughly flat, while allowing inference adoption to proceed as McKinsey expects, and the calculation becomes:</p><p>93 GW of inference + 23 GW of training = 116 GW of AI demand.</p><p>Add conventional workloads and total data-centre demand reaches approximately 180 GW.</p><p>That means mass adoption of capabilities that already exist could plausibly require 110&#8211;130 GW of AI capacity within five years. The industry&#8217;s planned supply, approximately 100 GW in JLL&#8217;s forecast, would be barely sufficient and perhaps moderately inadequate.</p><p>More importantly, <strong>planned capacity is not secured capacity</strong>. Only around 23&#8211;32 GW of all data-centre capacity is currently under construction globally. Perhaps 15&#8211;25 GW of that is AI-oriented. Adding it to existing capacity gives us only around 50&#8211;65 GW of AI infrastructure that is either operational or in a reasonably hard construction pipeline.</p><p>To reach 116 GW, another 50&#8211;70 GW must still obtain land, firm power, grid connections, financing, transformers, cooling equipment, and then actually be built.</p><p>The conclusion looks fairly obvious to me: we do not need AGI, recursive self-improvement or another dramatic capability jump to justify an enormous infrastructure buildout. Even if intelligence stopped improving today, widespread adoption of what already works would consume essentially the entire credible five-year pipeline.</p><p>Further capability growth is therefore upside to this calculation, not one of its assumptions. So are persistent agents, large-scale test-time reasoning, AI-generated video, scientific workloads, robotics and applications that have not yet been invented.</p><p>All of this suggests that the claim &#8220;too many data centres are being built&#8221; requires a surprisingly strong premise. It requires not merely that progress toward more capable models slows, but also that organisations fail to find productive uses for capabilities that already exist.</p><p>Now, none of this reflects my actual view of where capabilities are heading. I think we are only at the beginning of substantial further gains in intelligence.  And, partly for commercial reasons, and partly because frontier AI has become an object of geopolitical competition, frontier training will remain a priority and absorb a large share of the newest chips. That may also extend the economic life of older chips: as the latest hardware is directed towards training, older generations can continue serving inference workloads for models whose capabilities are already sufficient.</p><p>So, yes, I do think that we&#8217;re just at the beginning of a very large and dramatically scarce infrastructure buildout.</p>]]></content:encoded></item><item><title><![CDATA[On Pretty Woman, the Experience Machine, and Chalmer's VR]]></title><description><![CDATA[Pretty Woman is a film intellectuals are more or less required to treat with contempt.]]></description><link>https://carlolc.substack.com/p/on-pretty-woman-the-experience-machine</link><guid isPermaLink="false">https://carlolc.substack.com/p/on-pretty-woman-the-experience-machine</guid><dc:creator><![CDATA[Carlo Ludovico Cordasco]]></dc:creator><pubDate>Sat, 25 Jul 2026 13:00:34 GMT</pubDate><content:encoded><![CDATA[<p><em>Pretty Woman</em> is a film intellectuals are more or less required to treat with contempt. A Cinderella story about a corporate raider and a sex worker, scored to Roy Orbison, it is not where one expects to find a problem in moral philosophy. The weird thing is that it contains one.</p><p>Edward flies Vivian to San Francisco to see <em>La Traviata</em>. Before the opera begins, he offers her a theory: people&#8217;s first reactions to opera, he says, are decisive, and those who love it immediately will love it for life, while those who do not may learn to appreciate it without its ever becoming part of their soul. Then the opera starts, and Vivian, who a week earlier did not know she had feelings about nineteenth-century Italian opera, is undone by it.</p><p>Edward&#8217;s theory is wrong, and most of us are counterexamples to it, since people regularly learn to love things they disliked at first: black coffee, late Henry James, free jazz, running in the rain. But the scene gets something right all the same. Whatever Vivian comes to feel about opera, none of it could have happened before opera entered her life. Edward does not argue her into opera. He takes her to the opera.</p><p>I have been thinking about that step, because some of the best recent philosophy concerns the ways in which we become different people. L. A. Paul writes about transformative choices, choices that are difficult to assess beforehand because we cannot know what the experience will be like or how it will change us. Whether to have a child is a famous case, and you can read every parenting book ever written and still not know what parenthood will be like <em>for you</em>. Agnes Callard writes about aspiration, the slow work of coming to appreciate a value one can initially grasp only imperfectly. Her aspirant signs up for a music appreciation class and pinches herself to stay awake at the symphony, because she can tell that people she respects hear something in the music that she cannot yet hear, and she is trying to become someone who hears it too.</p><p>Paul&#8217;s chooser already knows that parenthood is an option; her problem is that she cannot see it from the inside. Callard is more alert to the earlier stage. In her work on liberal education, she argues that a university can make aspiration possible before a student knows what she wants to aspire to. The student encounters subjects, practices, teachers, friends, sports, politics, and art, and only later does one of them emerge as a possible direction for her life.</p><p>So philosophy has not ignored exposure, and what interests me is a consequence of Callard&#8217;s thought. If exposure helps to determine which aspirations become possible, then the environments we inhabit partly determine what we can become, and they do this before deliberation begins.</p><p>Before you can choose an experience or aspire to a value, something must bring it into view in a form that lets it register as potentially significant. Exposure comes in grades. Hearing jazz as background music in a caf&#233; gives you almost nothing beyond the knowledge that jazz exists, in roughly the way you know that Lampedusa exists. Watching a friend lean toward the speakers and distinguish things you cannot hear gives you something more, since it changes the meaning of your boredom. Your failure to enjoy the music becomes, possibly, a limitation of your ears, and that &#8220;possibly&#8221; is enough to get aspiration started.</p><p>I like to put this in terms of a menu, partly because I have argued elsewhere that moral theory has a menu problem too. Deliberation feels sovereign from the inside, as you survey the options, weigh them, and choose, but the options were printed before you sat down. You deliberate from a menu you mostly did not write, and the menu does not list what has been left off it.</p><p>Robert Nozick&#8217;s experience machine is the classic device for thinking about a life whose contents have been selected in advance. You plug in, and a computer supplies a lifetime of experiences. The debate that followed has mostly concerned whether pleasant experience is all that matters, and I suspect the machine takes away several things at once, which may be why the intuition against it is so difficult to pin on any single theory of well-being.</p><p>One loss deserves more attention than it has received. A machine can contain surprise, difficulty, apparent disagreement, and experiences designed to test the person inside it. Yet if its contents are limited to what that person commissioned, everything it contains must fall within a space she could specify before entering. Such a life may challenge her within that space, but it cannot be relied on to introduce what she had no way of asking for.</p><p>David Chalmers has done more than anyone to separate Nozick&#8217;s machine from virtual worlds in general. In <em>Reality+</em>, he observes that ordinary virtual reality can differ from the experience machine in several important ways. You can know you are in it, other people can enter it with you, and, most importantly, it can be interactive rather than preprogrammed, so that you make choices and alter what happens instead of merely undergoing a script. On Chalmers&#8217;s view, a sufficiently rich virtual world is a genuine reality, and a life lived there can be a good life.</p><p>I think he is right about this, and the concession can be pushed a step further, because virtuality turns out not to be where the problem lies, and on reflection neither is design as such.</p><p>Imagine an interactive simulation built from a comprehensive model of your present outlook. You know that you are entering it, no sequence of events has been fixed in advance, and you make consequential choices, undertake difficult projects, fail at some of them, and affect the people around you. The system has even been instructed to challenge you, supplying intelligent critics, serious arguments against your beliefs, and experiences likely to strain your commitments.</p><p>But everything that reaches you passes through the model. The system selects people, practices, works, and challenges by asking what someone with your history might find intelligible or significant. It can show you arguments you reject and experiences you fear, and it can change your mind about matters it already knows how to present to you. What it cannot reliably disclose is whatever falls outside its model of your relevance.</p><p>This world is not scripted in any ordinary sense, since your decisions are real and nobody knows in advance what you will do. Yet your present outlook governs the materials from which your future outlook will be formed. The menu can be revised, but only after the menu has decided what may count as a candidate for inclusion.</p><p>There are two levers here, and it is easy to run them together. One controls the source itself, as when an artificial critic generated to suit your profile says only what the system judges useful for you to hear. The other controls access, since a real person may form her judgment entirely independently while a personalized system decides whether her judgment will ever reach you.</p><p>The second case matters because selection does not turn other people into puppets. The writers, musicians, activists, and eccentrics excluded from my feed remain perfectly independent people, and what reflects my profile is their appearance in my life. A system can leave every voice untouched while arranging that I hear only the voices my past has prepared it to recognize.</p><p>The more I think about this, the less I trust the contrast between <em>designed</em> and <em>organic</em> worlds. A university curriculum is designed and may expose a student to something she never knew to want. A dinner party has a guest list and may still change a person&#8217;s life. An open virtual world can contain independent participants whose projects develop in ways no entrant or programmer anticipated, while a spontaneous community can be suffocatingly homogeneous. The important question is <strong>how much authority my present outlook has over what will reach me</strong>, and design can preserve openness or eliminate it, as can my own choices.</p><p>Other people remain central because they have their own sense of what matters, and they notice things for reasons I may not possess. When someone whose judgment I respect loves something I cannot yet parse, that love carries information, because it is evidence that there may be something there to get. Aspiration often runs on borrowed credit, and we invest time in Verdi, jazz, mathematics, or analytic philosophy on the strength of another person&#8217;s appreciation, before we have much appreciation of our own.</p><p>If a friend&#8217;s enthusiasm were manufactured to motivate me, it would tell me less about the object of that enthusiasm; part of the value of her response comes from the fact that it is hers. She may also put something before me that I could not have requested, because requesting it would have required a sense of its importance that I did not yet have.</p><p>None of this means that other people are the only source of new possibilities. An illness, a landscape, a work of art, or an accident can reorganize a life, and books and recordings carry the attention of people we may never meet. A designed process can expose us to material we did not select, provided its contents are allowed to exceed our present profile. Still, other people are among the most reliable sources of this kind of exposure, precisely because they do not begin from our map of what matters.</p><p>This reframes, at least slightly, what is wrong with echo chambers. The standard complaints concern misinformation and insufficient disagreement. Those are serious complaints, but a well-curated environment could correct both and remain closed in the sense that interests me, since it could keep supplying intelligent opponents whom I already know how to classify, so that every critic became another instance of an error I understand, and I could be surrounded by disagreement without ever encountering a possibility that unsettles my account of what the disagreement is about.</p><p>The same problem appears in ordinary self-curation. We choose friends, neighborhoods, institutions, news sources, and reading lists partly because they reflect things we already care about, and there is nothing inherently wrong with that; a life without selection would be impossible, and a life without commitment would be thin. The issue arises when our present values become a general admissions committee for our future, allowing through only what already has the right credentials.</p><p>Which brings this, inevitably, to the feeds. A recommender system optimized on my history mostly shows me extrapolations from my past. It does not create the people whose work appears there, but it increasingly governs which of them I meet. An AI companion tuned to my preferences has the outward form of another person, while its sense of what may interest me begins from a model of who I already am.</p><p>These systems can surprise us, since they draw on material produced by millions of other people, and engineers can give them room to explore beyond their predictions. Personalization itself is not the enemy; a good librarian personalizes, and so does a good friend. The question is whether a system permits anything to exceed the profile it has constructed: whether something can reach me that neither I nor a competent model of me would have selected. If we want that to happen, serendipity may have to become a design requirement rather than a pleasant accident.</p><p>Mill saw the social version of this long ago. His experiments in living are more than a plea for tolerance, because other people&#8217;s unfamiliar lives show us that our own practical menu may be incomplete, and their differences can disclose ways of living and caring that we had not considered. Hannah Arendt saw another version of the same point. Each person, she thought, is a new beginning, capable through action of introducing something others did not foresee, and a world remains open when it leaves room for such beginnings to matter.</p><p>Vivian&#8217;s tears at the opera were never quite the point; the point is that she got to have them. Edward chose the evening, but he did not choose what the music would become for her. He put her in contact with something that had not figured in her plans, and then the matter was hers.</p><p>We choose among futures only after some futures have become visible to us. What other people make possible is that futures we did not order can still get in.</p><p></p>]]></content:encoded></item><item><title><![CDATA[What do we owe to each other when we *really* don't know what we owe to each other?]]></title><description><![CDATA[Moral uncertainty, unawareness, and the dilemma of the missing menu]]></description><link>https://carlolc.substack.com/p/what-do-we-owe-to-each-other-when</link><guid isPermaLink="false">https://carlolc.substack.com/p/what-do-we-owe-to-each-other-when</guid><dc:creator><![CDATA[Carlo Ludovico Cordasco]]></dc:creator><pubDate>Fri, 17 Jul 2026 11:21:10 GMT</pubDate><content:encoded><![CDATA[<p>As some of you know, I&#8217;ve recently published a paper, stemming from my PhD thesis, in Economics &amp; Philosophy on how uncertainty about one&#8217;s future evaluative structure calls for abstracting from salient current preferences in contractarian bargaining procedures. The intuition is simple and comes from Kreps (1979): if I don&#8217;t know what I want to do tomorrow, I better keep my options open within reason. And what is interesting is that this flexibility-driven abstraction helps convergence among agents with widely diverse sets of preferences.</p><p>Notice that this paper is somewhat orthogonal to the moral uncertainty literature. Moral uncertainty folks talk about a different sort of uncertainty that is not prudential but rather normative. In many choice scenarios, I don&#8217;t know which moral theory is correct; hence I better perform some expected value procedure.</p><p>Let me reiterate, once again, that I am yet not deep enough in this literature to know everything that has been written. More importantly, I&#8217;m learning Bayesian epistemology as I go, thanks to LLMs. Yet, at first sight, I&#8217;m noticing that no one seems to have picked up on a problem that moral uncertainty frameworks seem to exhibit.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> Let&#8217;s call it <strong>the dilemma of the missing menu</strong>.</p><p>The basic concept of moral uncertainty is familiar enough. Suppose I assign some credence to utilitarianism, some to a Kantian view, some to Rossian pluralism, and perhaps smaller amounts to a range of other theories. Each theory tells me, conditional on its being correct, how choiceworthy my options are. I then weight those assessments by my credence in the relevant theory, add everything up, and choose the option with the highest expected choiceworthiness.</p><p>There is something genuinely attractive about this picture. It takes seriously the possibility that I may be wrong about morality, and it insists that this possibility should affect what I do. I find that idea compelling as I take my own moral uncertainty seriously, and I would rather not act as though whichever moral view currently seems most plausible to me must therefore be correct.</p><p>Still, the procedure needs something before it can get started. It needs a menu: a list of moral theories, each with enough substantive content to evaluate the available options. The problem is that our uncertainty may extend beyond the question of which theory on the menu is true. We may also have reason to think that the menu itself is incomplete. I do not mean theories that we have considered and rejected, or eccentric views sitting at the bottom of the credence distribution. I mean theories that are not represented at all because nobody has yet formulated them. </p><p>Imagine a thoughtful moral philosopher in 1950 deliberating about a population policy with effects extending far into the future. Her menu contains the best theories available to her, and she distributes her credences among them with admirable care. What it does not contain, and perhaps could not yet contain, is a person-affecting principle of the kind that philosophers would later develop in response to the non-identity problem.</p><p>The relevant distinctions have not yet been drawn. She may not possess a clear conceptual separation between harming a person and bringing into existence a person whose life is difficult but worth living. She may not yet distinguish making someone worse off from making a choice without which that very person would never have existed. These are not simply conclusions she has failed to reach. They are ways of structuring the moral problem that are not yet available to her.</p><p>No redistribution of credence among her existing theories captures what she is missing. She has not assigned the future person-affecting theory a very small probability. She has assigned it nothing because it is not among the possibilities her model can represent. Her uncertainty about the adequacy of the menu is not itself an item on the menu.</p><p>This does not seem like a merely hypothetical concern. The history of moral philosophy contains many episodes in which the space of available moral thought was enlarged. Consider the emergence of sustained philosophical attention to supererogation, the moral standing of animals, structural injustice, future generations, disability, etc. In each case, the change did not consist only in moving credence from one familiar theory to another. Often the change involved learning to notice a distinction, relation, or kind of claim that earlier theories had not adequately represented.</p><p>Someone attentive to this history therefore has inductive evidence that our current repertoire of moral theories is unlikely to be final. More importantly, the gaps are not randomly distributed. They seem especially likely to matter in domains involving long time horizons, new technologies, unfamiliar forms of agency, or effects on future people. Unfortunately, these are also the cases in which we most want moral guidance.</p><p>Suppose, then, that our agent wants her expected-choiceworthiness calculation to reflect this higher-order uncertainty. She has two obvious options.</p><p>She can calculate over the menu she currently has. The result will be perfectly well defined, but it will ignore a kind of uncertainty she has good reason to take seriously. Alternatively, she can add a residual hypothesis: some true moral theory that I have not yet conceived.</p><p>Assigning probability to that residual hypothesis is not the hard part. Bayesians use catch-all hypotheses all the time. The problem appears when we ask what the hypothesis says about the options. Conditional on the truth of some unconceived moral theory, how choiceworthy is policy A? How choiceworthy is policy B?</p><p>The residual hypothesis gives us no answer. It has no principles, no conclusions, no ranking, and no evaluative scale. It tells us only that the true theory lies outside the current menu. That is enough to receive probability mass, but not enough to generate a value that can be entered into an expectation.</p><p><strong>This is the dilemma of the missing menu. If we exclude yet unarticulated theories, expected choiceworthiness remains defined but fails to represent our uncertainty. If we include them, the expectation itself becomes undefined.</strong></p><p>Those familiar with the literature on <a href="https://www.jstor.org/stable/2672830">cluelessness</a> may notice a family resemblance, but the two problems are different. Cluelessness, as I understand it, is usually a problem about empirical consequences. Even if the evaluative standard is fixed, the long-run effects of our actions may be largely invisible to us. A consequentialist may then be tempted to impose convenient assumptions about unknown consequences in order to recover a determinate answer.</p><p>The missing-menu problem lies elsewhere. Give the philosopher of 1950 perfect knowledge of every empirical consequence of her policy, down to the identity and welfare of every future person, and the problem remains. What she lacks is not information about how the world will unfold. She lacks some of the evaluative concepts under which those facts may become morally significant. Cluelessness is ignorance within a moral map, whereas the missing-menu problem is reason to believe that part of the map has not yet been drawn.</p><p>At this point, one might object that I have moved too quickly. The fact that the residual hypothesis does not itself provide an evaluative score does not yet show that no reasonable score can be assigned to it. There are at least two ways of trying to complete the calculation.</p><p>The first is to treat the catch-all hypothesis as an ordinary residual category and assign it some value. The crudest version would simply give every option a neutral or middling score&#8212;five out of ten, say&#8212;and continue with the calculation. A more sophisticated version would point out that residual categories are common in probabilistic reasoning. We often assign some probability to &#8220;none of the possibilities explicitly listed&#8221; without knowing exactly what that residual possibility contains. Why should moral uncertainty be any different?</p><p>The answer is that a residual category is useful only when the quantity we are trying to estimate remains defined across all the possibilities it contains. Suppose a forecaster is comparing several possible weather patterns because she wants to estimate next year&#8217;s harvest. She may distinguish the most likely weather scenarios and place everything else into a residual category. That residual category is poorly specified, but the quantity of interest is not. Every possible weather pattern will result in some level of crop yield. The forecaster may not know that yield with much precision, but she knows what kind of quantity she is trying to estimate.</p><p>This is why she can assign an approximate value to the residual category. She is estimating the average value of a variable that is already defined across the possibilities grouped within it. Greater detail would improve the estimate, but it would not create the underlying variable.</p><p>The catch-all over unconceived moral theories is different. There is no evaluative variable that remains independently defined across the theories it contains. A choiceworthiness score acquires its meaning from a moral theory. Utilitarianism evaluates options by reference to welfare. A Kantian theory evaluates them through duties, constraints, or requirements of respect. Other theories may understand moral reasons and their relative importance in still other ways. Even when the theories are fully articulated, placing their assessments on a common scale is already a serious problem.</p><p>The difficulty with the catch-all is therefore not merely that we do not know its score. We do not yet know what would make something count as a score under the theories it contains. What would it mean to say that policy A receives five out of ten according to a theory nobody has articulated? There is no identified standard against which the five can be interpreted.</p><p>The comparison with empirical forecasting can make this easy to miss. In the weather case, the hypotheses describe different ways the world might be, while crop yield remains the same kind of quantity across them. In the moral case, the hypotheses concern different accounts of what makes an option choiceworthy in the first place. The theories do not merely assign different values to a shared and independently specified moral variable. They partly constitute the standards by which the options are evaluated.</p><p>The empirical residual category is therefore incomplete in its description of the possibilities it contains, but not in its specification of the quantity to be estimated. The moral residual category is incomplete in both respects. It tells us neither which theory is true nor how the relevant theories would evaluate the options. Assigning it a neutral or estimated score does not approximate a value that is already there. It supplies by stipulation the very evaluative structure that the catch-all lacks.</p><p>The second response could be more ambitious. Rather than trying to assign a value to the catch-all, it abandons the menu of moral theories altogether. We might instead take the space of all possible value functions over outcomes, place a probability measure over that space, and integrate. Every possible ranking or scoring of the options would then appear somewhere in the space, including those associated with theories nobody has yet conceived.</p><p>This seems cool, but I think it fails for two reasons.</p><p>The first is that a value function is not a moral theory. Moral theories have grounds, explanations, commitments, and internal structure. These features make them answerable to argument.</p><p>Suppose I read Parfit on non-identity and revise my credences. I do so because his cases bear on claims that the theories I understand actually make. A theory may fail to explain a distinction, generate an implausible implication, or conflict with a judgment in which I have confidence. Moral reasoning works through these points of contact.</p><p>Now compare a bare function that assigns 0.41 to policy A and 0.38 to policy B. There is little to interpret and almost nothing to argue with. I can eliminate functions that conflict with sufficiently secure conclusions. A function that systematically ranks gratuitous cruelty above kindness can receive very little weight. But most moral evidence does not arrive as a sequence of isolated rankings. It works through reasons, explanations, analogies, and principles. Bare functions contain none of these.</p><p>A probability measure over anonymous functions may therefore define an expectation, but it is unclear what that measure is supposed to answer to. It is especially unclear how future moral argument is meant to update it.</p><p>The second problem seems even more insurmountable. The function space is complete only relative to the outcomes over which the functions are defined. Yet the history of moral thought suggests that revision often changes the way outcomes themselves are represented.</p><p>Think of a spreadsheet. You may write every possible formula over the columns you currently have. In that sense, your formula space is complete. But if the quantity that later turns out to matter requires a column you never created, the completeness was not the kind you needed.</p><p>Something like this happens in moral philosophy. The empirical facts may already be available: who exists under each policy, how well each person lives, and which causal relations obtain. What changes is our understanding of which relations among those facts matter morally. The non-identity problem, for instance, forced philosophers to ask whether a complaint requires that someone be made worse off, whether the fact that a person owes her existence to a choice defeats a claim of harm, and how person-affecting and impersonal reasons interact.</p><p>These developments do not merely assign new numbers to a fixed set of morally transparent outcomes. They alter the descriptions under which those outcomes become normatively intelligible. A formally exhaustive space of functions can therefore remain incomplete because the representation over which those functions are defined is itself incomplete.</p><p>One may consider some subtler responses as well. For instance, one might follow one&#8217;s single favourite theory, abstain when the probability assigned to the residual hypothesis becomes too large, or replace theories with more minimal value hypotheses. </p><p>However, my impression is that these proposals would move the difficulty around rather than remove it. At some point, the decision procedure still needs substantive evaluative content connecting possibilities to rankings.</p><p>Does the dilemma imply paralysis? I don&#8217;t think so. In fact, the historical evidence that generates the problem may also support a modest practical principle.</p><p>Suppose I have reasons to think that the moral map of a domain is likely to be revised. Suppose also that my options differ substantially in how far they preserve my ability to recognize and respond to such a revision. In that case, I have a pro tanto reason to prefer the more reversible option.</p><p>The justification is not that flexibility is intrinsically valuable, as that would simply be another first-order moral claim, and therefore another entry on the menu. The reason is instead that I am acting through a model I have specific grounds to distrust. When the model may be incomplete in ways I cannot yet represent, it is sensible to avoid locking in its practical consequences unnecessarily.</p><p>Decision theory contains several related ideas. Robust control asks how a policy performs across imperfect models. Ambiguity-sensitive approaches resist treating poorly understood uncertainty as though it were ordinary risk. Quasi-option value captures the benefit of preserving future choice when additional information may arrive.</p><p>The missing-menu case belongs to the same broad family, but it is more radical. Here we may be unable to parameterize what we are missing. We do not merely lack confidence about which value a known variable will take. We may lack the variable.</p><p>This matters especially for AI systems. Much of the current discussion assumes that the central problem around AI ethics is how to encode the correct values, or at least how to make systems act under uncertainty about which currently available values are correct. But if the missing-menu problem is real (and it is!), then the issue is that some of the morally relevant concepts may not yet be available to encode. A system can be perfectly calibrated over the theories we know and still be systematically closed to forms of moral criticism that have not yet been articulated.</p><p>That possibility changes what responsible design should aim at. The goal cannot simply be to produce systems that implement today&#8217;s best moral judgments with increasing consistency and scale. A system that does this too rigidly may amplify not only our moral insight but also the limitations of our present moral vocabulary. The more widely such a system is deployed, the more difficult it may become to revise the practices and institutions built around it.</p><p>This is one reason why the language of alignment can be misleading. Alignment suggests a target that already exists and that merely needs to be specified accurately. Moral progress often has a different structure. We do not simply become better at applying a fixed set of values. We learn that some harms were misdescribed, some interests were ignored, some distinctions were morally spurious, and some categories were missing altogether. The target itself changes because our understanding of what counts as morally salient changes.</p><p>An AI system designed under these conditions should therefore be judged partly by its capacity for evaluative revision. It should preserve traces of the reasons behind its decisions, remain open to challenge from perspectives not represented in its original training or specification, and avoid making morally contestable judgments irreversible merely because they can be automated. The relevant virtue is <strong>revisability</strong>.</p><p>There is a broader political point here. Institutions have often obstructed moral progress by turning a partial moral understanding into an entrenched administrative category. AI systems may intensify this tendency because they can apply such categories with unprecedented speed, consistency, and reach. A mistake embedded in a human bureaucracy may be local, uneven, and open to discretion. The same mistake embedded in a widely deployed model can become standardized infrastructure.</p><p>The missing-menu problem therefore gives us a reason to resist the fantasy of final moral encoding. The responsible ambition is more modest and, in another sense, more demanding: to build systems that can act while remaining corrigible by forms of moral understanding that neither their designers nor their users yet possess. If moral progress partly consists in learning to see what our existing theories leave out, then a morally serious AI system must leave room for us to discover that it has been looking at the world in the wrong way.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>While no one has picked up on the specific dilemma of this post, there is excellent foundational work in philosophy on unawareness, which provides the basic concepts for the present argument. E.g., check out this excellent <a href="https://philpapers.org/archive/STEBUR.pdf">book</a>.</p></div></div>]]></content:encoded></item><item><title><![CDATA[What Do We Owe Each Other When We Don’t Know What We Owe Each Other?]]></title><description><![CDATA[On deciding well under moral uncertainty, and on what deciding well does not do.]]></description><link>https://carlolc.substack.com/p/what-do-we-owe-each-other-when-we</link><guid isPermaLink="false">https://carlolc.substack.com/p/what-do-we-owe-each-other-when-we</guid><dc:creator><![CDATA[Carlo Ludovico Cordasco]]></dc:creator><pubDate>Tue, 14 Jul 2026 10:58:53 GMT</pubDate><content:encoded><![CDATA[<p>I&#8217;m travelling back after four lovely days in Oxford, catching up with old friends and attending my favourite workshop of the year, R:ETRO, which Rita Mota and Alan Morrison have made into the best workshop in business ethics. I presented a paper on moral uncertainty, and I was lucky to have Lewis Williams as my discussant. He read the paper more charitably than it deserved and then asked the two questions I expect to be thinking about for months. A good part of what follows exists because he asked them.</p><p>As some of you know, I&#8217;m new to the moral uncertainty literature, and I&#8217;m making sense of it as I go. This post is longer and more careful than my usual fare, because the point I want to make only becomes visible once the literature has been set out properly. At the end I&#8217;ll compress everything into five claims and ask you, sincerely, what I&#8217;m getting wrong.</p><p>Start with the case from the paper. A department chair has to write a promotion recommendation for a colleague whose ageing father needs daily care. The university&#8217;s criteria are the usual three: research, teaching, service, and they say nothing about personal circumstances. By those criteria, the case is borderline. The colleague&#8217;s publications have slowed, his teaching is strong, and the service he declined, he declined because he was caring for his father.</p><p>The chair finds that she can understand fairness in two ways. On the first understanding, fairness is consistency. The criteria were announced in advance, everyone else has been judged by them, and an evaluator who starts adjusting for personal circumstances replaces a public standard with her private sympathies, treats this candidate differently from every previous one, and misleads the committee about what the letter in front of them is reporting. On this understanding, the caregiving deserves sympathy, perhaps some institutional support, but no place in the evaluation. On the second understanding, a fair evaluation measures what a person achieved against the circumstances he faced. A body of work produced while caring for a dying parent is not the same achievement as an identical one produced in ordinary circumstances, and so an evaluation that ignores the caregiving ignores the one fact that explains the file on her desk.</p><p>She thinks about this properly, for longer than the committee&#8217;s timetable really allows. And here is the difficult part. She cannot rule either understanding out. Each rests on considerations she cannot answer. Notice, because this matters for everything that follows, that her problem is not missing information. She could know every fact about his publications, his teaching, and his father, and the question of which understanding of fairness is correct would remain exactly where it is. Her uncertainty is about morality itself. And the letter is due Thursday.</p><p>Notice also that neither option is guaranteed to be morally acceptable. If the consistency understanding is correct, then the adjusted letter treats the other candidates unfairly, since they were measured by the stated criteria. If the contextual understanding is correct, then the criteria-based letter fails a colleague who had a claim to have his circumstances considered. Whichever letter she writes, she may be wronging someone, and she does not know which letter is the wrong one.</p><p>The chair&#8217;s situation is exactly what the moral uncertainty literature exists for, and that literature has become quite sophisticated. As far as I&#8217;m aware, the modern discussion largely begins with Ted Lockhart&#8217;s <em>Moral Uncertainty and Its Consequences</em> (2000), whose central idea was that we should treat uncertainty about morality the way decision theory treats uncertainty about ordinary facts. A rational person who is unsure whether a bridge will hold weighs the probability of collapse against the cost of going the long way round. In the same way, the thought goes, a rational person who is unsure which moral view is correct should weigh her confidence in each view against how much each view says is at stake. The best developed treatment is the book <em>Moral Uncertainty</em> (2020) by William MacAskill, Krister Bykvist and Toby Ord, building on earlier work by Andrew Sepielli among others. Their account depends on how much information the competing moral views provide. Where the views assign values that can be measured on a common scale and compared with one another, which is the situation they treat as central, they argue that one should maximise expected choiceworthiness. In plainer words, assign each option a value according to each moral view you take seriously, weight those values by how confident you are in each view, and choose the option with the highest total. Where the views provide less information than that, they recommend other methods, borrowed from the theory of voting, but those refinements will not matter here.</p><p>The literature is not naive about the difficulties. There is a serious problem about whether the stakes assigned by different moral theories can be compared on a common scale at all. There is a regress problem, since one can be uncertain about the theory of uncertainty too. And there is an opposing camp, the normative externalists, Brian Weatherson and Elizabeth Harman most prominently, who reject the whole project. On their view, what you should do is settled by the moral facts, and your uncertainty about those facts changes nothing. I will come back to them, because my disagreement with that camp turns out to be located somewhere unexpected. But set the objections aside for now, and let the chair have the best the literature can give her.</p><p>So suppose she does everything the leading theory asks. Her two understandings of fairness are similar enough in what they measure, the seriousness of a possible unfairness to a particular person, that she can reasonably treat their stakes as comparable, which puts her in the situation where the theory recommends the expected choiceworthiness calculation. She examines her confidence in each understanding honestly and finds it close to even, with perhaps a slight edge to the consistency view, whose institutional basis she finds harder to dismiss. She thinks carefully about the stakes each view assigns, how seriously the contextual view treats the failing of a caregiver, how seriously the consistency view treats the unequal treatment of past candidates. She compares them as honestly as the hard problems of comparison allow. And suppose the calculation comes out, given her confidence levels and the stakes she can defend, in favour of the criteria-based letter. Nothing in what follows depends on it coming out that way rather than the other. She writes the letter that applies the criteria as stated, submits it on Thursday, and she has, by the standards of the best available theory of choice under moral uncertainty, decided impeccably.</p><p>Now I want to ask a question that, as far as I can tell, the literature does not ask. What, exactly, did the procedure give her? It gave her a defensible way to choose. Her confidence levels were what they were, the stakes were what they were, a recommendation was due, and the procedure turned all of that into an action in a way she could justify to anyone who asked. If choosing under moral uncertainty is something one can do well or badly, she did it well. Call this procedural warrant. She has it, and she is entitled to it.</p><p>Here is what the procedure did not give her. It did not tell her which understanding of fairness is correct. It could not have. Its inputs were her confidence levels and the stakes the theories assign, and no calculation performed on those inputs produces information about fairness that was not already in them. Maximising expected choiceworthiness is a method for acting despite an open question. It is not a method for answering one. The question the chair could not answer on Wednesday, which understanding of fairness governs an evaluation like this, is exactly as open on Friday as it was before she ran the numbers. Her acting changed many things in the world. It changed nothing in her evidence.</p><p>So on Friday morning the chair is in a situation that I think deserves more attention than it gets. There is a completed act, her letter, sitting in the committee&#8217;s inbox. And there is a question (did that letter wrong the colleague?) which remains open, and which is now a question about something she has actually done. On Wednesday it was a question about a choice she had yet to make. On Friday it is a question about an act that exists. The deciding did not answer the question. It gave the question something to be about.</p><p>The moral uncertainty literature, without ever saying so, encourages a confusion at exactly this point. It offers procedures which, when you have carried them out, produce the feeling of being finished. You had a hard case, you did what the best theory recommends, and the matter seems dealt with. What that feeling hides is that two different questions were on the table, and the procedure addressed only one of them. Whether she chose well, given her uncertainty, is settled, and settled in her favour. Whether the act she chose wronged her colleague is not settled, and could not be settled by anything the procedure touched. Deciding how to act while a question is open and answering the question are different achievements, and only the first one has occurred. To have deliberated impeccably about what to do, given that you have not settled which principles apply, is not to have deliberated about which principles apply. The literature never asserts otherwise. It simply produces the feeling of completion at the exact place where this confusion arises, and then says nothing about anything after Thursday.</p><p>Philosophy is not silent about moral aftermaths in general. There is a rich literature on what philosophers call moral remainder, and it is worth seeing why none of it covers the chair.</p><p>Bernard Williams gave us the lorry driver who, through no fault of his own, runs over a child, and argued that the driver appropriately feels something a bystander should not, a burdened form of regret that comes with having been the one who did it. Michael Walzer gave us the politician who authorises torture to prevent a catastrophe and emerges with what he called dirty hands. Ruth Barcan Marcus argued that in genuine dilemmas, where two obligations apply and only one can be met, the unmet obligation leaves a remainder, which shows itself in the appropriateness of guilt and in the demand to make amends. These are all cases of something surviving a choice, and they are the natural place to look for the chair.</p><p>But consider what each remainder is a response to. The driver&#8217;s regret responds to an established harm. The child is dead, and the death was his doing. The politician&#8217;s dirty hands respond to an established violation. She knows exactly which principle she broke, because it is a principle she herself holds. Marcus&#8217;s dilemmas involve established obligations. The agent knows both requirements applied to her, and knows which one went unmet. In every case the leftover has a definite object, some wrong or harm or defeated claim whose existence the agent has established. The remainder literature is a literature about living with what you know you did.</p><p>The chair has established nothing of the kind. She does not know that her letter harmed anyone or violated anything. What she knows is that it might have, in a specific way she can state, resting on considerations she weighed and could not dismiss, and which nothing has answered. Put the two literatures side by side and a gap appears. The uncertainty literature deals with the time before the act, when the question is open. The remainder literature deals with the time after the act, when the wrong is established. The time after the act, when the question is still open, is dealt with by no one. I find this remarkable, because that combination, having acted and still not knowing, is not a rare or exotic situation. It may be the most common moral situation there is. Most hard choices are made without the question being settled, and the question does not settle itself once the deadline passes.</p><p>So what should we say about the chair on Friday? The view I have arrived at comes in two claims, one about who the relevant duties are owed to, and one about when they apply.</p><p>The first claim. The question the chair faced, what does fairness require in evaluating the colleague, is not a question addressed to morality in general. It concerns what he is owed. It is, in a sense I want to take seriously, his question. And if it is his question, then the duty to handle it well is a duty owed to him. This is why, I think, we would judge two chairs differently even when their letters are word for word identical. One wrote the criteria-based letter after genuinely weighing both understandings. The other wrote it because weighing them was tiring and the criteria were more convenient. The colleague has a complaint against the second chair, and the complaint concerns how she decided, whatever we end up saying about what she decided. It is worth noticing how rarely the moral uncertainty literature says this. Its norms are, almost always, addressed to no one in particular. They tell the agent how to be rational, or how to respond to her evidence, and they rarely mention the person the decision concerns. The nearest exception I have found is Chelsea Rosenthal, who has recently argued that reckless deciding under moral uncertainty disrespects the people it might wrong, a view I share. What I have not seen anyone do is explain what makes these duties belong to the person, who exactly holds them and why, or follow them past the moment of decision. My suspicion, which I&#8217;m developing into a paper, is that the norms of decision under moral uncertainty are owed to the very people the decisions are about, in something like the way that duties not to impose risks are owed to the people put at risk.</p><p>The second claim. If the deliberation was owed to him before the act, something is still owed to him after it, and for the same reason. The question is still his, and it is still open. What is owed, I think, is modest. She does not have to agonise, investigate, reopen the case, or avoid him in the corridor. She can put the entire business out of her mind, which is what deadlines and sanity require, because to stop thinking about a question is not to declare it answered. The one thing she may not do is quietly move from <em>I decided this</em> to <em>this was fine</em>, when the only thing supporting the move is that she decided it. If someone asks her in March whether the caregiver was treated fairly, the honest answer is still <em>I genuinely don&#8217;t know, I did what I could defend</em>. There is a small but real failure in the chair who instead says <em>well, it&#8217;s done</em>, as though having done it were a reason to think it was permissible. Her having acted bears on the question of whether she wronged him exactly as much as the Thursday deadline bore on the nature of fairness, which is to say not at all. The question is exactly as serious as it ever was. The failing chair has simply stopped treating it as a question. These two claims are, I think, one requirement meeting the agent at two times, before the act and after it.</p><p>And notice that the after-act version is where my disagreement with the normative externalists actually sits, if it turns out to be a disagreement at all. They hold that what the chair objectively ought to have done was determined by the true understanding of fairness all along, and that her confidence levels never changed it. On this, as far as I can tell, they are right, and nothing I have said requires denying it. My claim was never that her uncertainty changed what treatment he was owed. My claim is the more modest one that her uncertainty, being real and being reasonable, leaves the question of whether she failed him open after she acts, and that there is a better and a worse way to treat an open question about what another person was owed. Whether an externalist must reject even this modest claim, I am genuinely unsure. Perhaps they can accept it as a requirement of honesty about one&#8217;s own epistemic situation rather than as anything distinctively moral, in which case our disagreement would shrink to the question of how to classify the requirement rather than whether it exists. That would be a smaller disagreement than I expected to have, and I would be glad to have it instead of a larger one.</p><h2>Summarising</h2><p>Let me try and compress all of the above. First, people are owed treatment according to the moral standards that actually apply to them, whatever those turn out to be. Second, what those standards require is often genuinely unclear, and the uncertainty often survives careful, honest deliberation. Third, because of the first two claims, we owe the people affected by our decisions a genuine deliberation, and they, specifically, are the ones with a complaint if we treat the question as settled when nothing has settled it. Fourth, deliberating and deciding, however well, resolves the practical question without resolving the moral one, because no procedure operating on our confidence levels produces evidence about the principles themselves. Fifth, after acting we owe the person one further thing, which is not to pretend the question got answered. This last requirement is purely negative. It forbids a single move, treating the completed act as if it were a reason to believe the act was permissible.</p><h2>What I might be getting wrong</h2><p>This is an actual request rather than the blogger&#8217;s customary false modesty, so let me make it easy by listing the places where I think the argument is most vulnerable.</p><p>The moral uncertainty theorists could fairly reply that their frameworks never claimed to settle first-order moral questions, so I am objecting to a claim that no one makes. My answer is that I am pointing at something the literature fails to say rather than something it wrongly says, and that such omissions can matter, because a theory that produces the feeling of settledness at exactly the point where a confusion arises, and says nothing about the aftermath, encourages in practice a view it never states in print. But whether an omission is a fair target is itself a fair question.</p><p>The externalists could reply that once they grant that the objective wrong is what matters, my after-act requirement is either trivial, a demand to keep accurate track of what one knows, or mislabelled, an epistemic norm described in the language of what is owed to a person. I think the requirement is genuinely moral and genuinely owed to the colleague, because the question being mishandled is a question about what he was owed. But I concede this is the most serious worry I face. Is <em>owed to him</em> doing real work, or am I attaching his name to the ordinary duty of intellectual honesty?</p><p>Third, is the difference between honest deadlock and simple poor judgment as clear as my chair makes it look? Her case was built to be favourable, two understandings with real support, weighed by a conscientious person. Much apparent moral uncertainty is laziness, or a reluctance to admit what one already knows, and an account that dignified every hesitation would prove far too much. I rely on a condition, roughly, that the unanswered considerations be ones a responsible deliberator could not simply set aside, and I admit that conditions of this kind are easier to state roughly than precisely.</p><p>Fourth, does the fifth claim ever let anyone finish anything? If no decision settles anything, it can sound as though we must carry every hard choice for the rest of our lives, which cannot be right. I have tried to prevent this in two ways, by making the requirement purely negative, so that it forbids one move rather than demanding any ongoing activity, and by insisting that questions genuinely settled by reasons stay settled. Reasons can close a question. Acting cannot. You may think this is not enough.</p><p>And fifth, the whole argument assumes that moral questions have answers we can be right or wrong about, at least often enough for an open question to be a genuine status. Readers with less confident metaethics should tell me how much of this survives translation into their terms, because I am unsure.</p><p>Those are the five places I&#8217;d look first. If you see a sixth, the comments are open, and I promise to treat the question as a question.</p>]]></content:encoded></item><item><title><![CDATA[On Anaesthesia and AI]]></title><description><![CDATA[Recently, I wrote a piece for Aeon on the time asymmetry between the costs and benefits of LLMs.]]></description><link>https://carlolc.substack.com/p/on-anaesthesia-and-ai</link><guid isPermaLink="false">https://carlolc.substack.com/p/on-anaesthesia-and-ai</guid><dc:creator><![CDATA[Carlo Ludovico Cordasco]]></dc:creator><pubDate>Fri, 03 Jul 2026 08:13:28 GMT</pubDate><content:encoded><![CDATA[<p>Recently, I wrote a <a href="https://aeon.co/essays/what-we-cant-measure-about-ai-yet">piece</a> for Aeon on the time asymmetry between the costs and benefits of LLMs. Unsurprisingly, the essay has attracted many negative comments from the anti-AI crowd, including accusations that it was paid for by Anthropic (please, Dario, pay me or at least give me a free subscription to Claude Max!). It would be dumb to complain about it: I knew the audience and expected to be inundated with such comments, sometimes orthogonal to the piece's content. Foolishly, I&#8217;ve chosen this path of trying to provide what I take to be a nuanced analysis of the potential of a new technology and its application to extremely conservative sectors. </p><p>The essay leads with an analogy that I found useful: the introduction of anaesthesia. Here&#8217;s an excerpt from the Aeon piece:</p><blockquote><p>Before 1846, the scope of what a surgeon could attempt was bounded by what a conscious patient could endure in minutes. Patients were held down by assistants or strapped to the operating table (my mother&#8217;s tonsillectomy in early 1960s Sicily was still performed more or less this way), and the surgeon&#8217;s reputation depended on the economy of their movements, because every second of the procedure was a second of conscious agony. Robert Liston could amputate a leg in under <span>30 seconds,</span> because he had to. Operations were confined to the body&#8217;s surface, to amputations and the drainage of abscesses and the excision of superficial tumours, because no patient could withstand sustained work inside the thoracic, abdominal or cranial cavities.</p><p>When ether and chloroform arrived, some of the most respected figures in American and British medicine argued that rendering patients unconscious was a grave clinical error. John Pollard Harrison of the Medical College of Ohio, then vice-president of the American Medical Association, wrote in 1849 that &#8216;pain is curative &#8211; the actions of life are maintained by it &#8211; were it not for the stimulation induced by pain, surgical operations would more frequently be followed by dissolution.&#8217; Charles Meigs, who held the chair of obstetrics at Jefferson Medical College in Philadelphia, treated labour pains as a desirable, salutary and conservative manifestation of the life force. Other surgeons pointed out that conscious patients confirmed the surgical site, assisted in decisions during the operation, and provided real-time diagnostic feedback that would be lost under anaesthesia. Safety concerns intensified as reports accumulated, culminating in the Royal Medical and Chirurgical Society&#8217;s 1864 committee report, which catalogued 123 chloroform deaths. In the early years, the opposition to anaesthesia was widespread and grounded in the best available clinical evidence.</p><p>What none of these critics could have told you was that anaesthesia would make possible open-heart surgery, organ transplantation, neurosurgery, and the entire architecture of modern surgical specialisation. The benefit was not a more comfortable version of what surgeons had been doing. It was the appearance of a possibility space whose contents were inconceivable from inside the surgical practice <span>of 1846.</span></p></blockquote><p>Perhaps more interestingly, I received a comment with what I take to be an LLM-generated letter that could have been sent by one of those doctors opposing anaesthesia back in the 1840s.</p><blockquote><p>This is now the second public lecture in as many weeks which extols the supposed virtues of rendering our patients insensible with ether. Are we really to persuade ourselves that it is progress to exchange the sober labour of the surgeon, schooled by years of watching the suffering body, for the convenience of a chemical stupor?</p><p>I cannot agree that major operations belong among the &#8220;tasks&#8221; for which such vapours are fit. The knife was never meant to pass so lightly through flesh. Pain has been our sternest teacher: it warns us when we cut too deep, when the patient is failing, when our art has overreached our understanding. To strip it away and then congratulate ourselves on the neatness of our incisions is to mistake silence for safety.</p><p>Nor do I know how to trust the account of these new procedures. Which part of the judgement belongs to the operating surgeon, and which to the apothecary who brewed the ether? What precautions were taken? How many trials ended in collapse, in asphyxia, in deaths we will never read about in the hospital reports? When a patient survives, whose skill do we praise; when he dies, whom do we blame?</p><p>We are told that ether allows subtler, longer operations. I fear instead that it will license recklessness in those who have not earned their steadiness through the hard discipline of watching the conscious face of the wounded. If surgery is to remain a humane art, we should stand up for the craft that is learned beside the groaning bed, not the facile boldness of the man whose subject lies mute and helpless under a veil of fumes.</p><p>Until we can look a patient in the eye and say, with full honesty, that we understand what we are putting into his lungs and what miseries may follow, I for one will keep my hands&#8212;and my conscience&#8212;clean of this so&#8209;called advance.</p></blockquote><p>I quite loved this. Substitute anaesthesia with AI and you might have a keynote speech on AI use in academia :) Things will change, but they will go worse before it gets any better!</p>]]></content:encoded></item><item><title><![CDATA[On Moving From Ideal to Non-ideal theory]]></title><description><![CDATA[Recall that, a few months ago, I blogged about a paper co-authored with Dora Xu on how we can use a non-tradeability approach to establish a minimal threshold for any normative theory.]]></description><link>https://carlolc.substack.com/p/on-moving-from-ideal-to-non-ideal</link><guid isPermaLink="false">https://carlolc.substack.com/p/on-moving-from-ideal-to-non-ideal</guid><dc:creator><![CDATA[Carlo Ludovico Cordasco]]></dc:creator><pubDate>Fri, 26 Jun 2026 12:09:25 GMT</pubDate><content:encoded><![CDATA[<p>Recall that, a few months ago, I <a href="/__u/carlolc.substack.com/p/negative-dominance-and-threshold">blogged</a> about a paper co-authored with Dora Xu on how we can use a non-tradeability approach to establish a minimal threshold for any normative theory. The paper is still under review, so I can&#8217;t upload it. However, I wanted to say more about the methodology (which we call <em>minimal non-tradeability</em>) because I think the same mechanism can illuminate a longstanding problem in political philosophy, which is the division of labour between ideal and non-ideal theory and how to move from one to another.</p><p>I&#8217;ll recap how the method works first, since the application leans on it entirely, and then show how it gives a procedure for carrying a theory&#8217;s demands from the ideal case down to a world of constraints.</p><h3>How minimal non-tradeability works</h3><p>Moral and political philosophers reach constantly for minimal standards: minimal justice, minimal legitimacy, a sufficiency threshold, a human-rights floor, a decent minimum. The word &#8220;minimal&#8221; is meant to mark a level below which something has distinctively failed, and above which a theory&#8217;s basic demands have at least been met. The question the original paper asks is what &#8220;minimal&#8221; actually amounts to once you take it seriously, and the answer turns on a single condition.</p><p>Fix a normative theory and call its evaluative standpoint X. What X does is rank the available options &#8212; full social arrangements, distributions, institutional settlements, individual lives, policies, whatever the theory takes as its objects &#8212; from better to worse by its own lights. The ranking may be incomplete, in that X need not have a view about every pair. The theory cares about a collection of normative dimensions, P&#8321;, P&#8322;, &#8230;, P&#8345;, such as income, health, education, or the satisfaction of a given right, and each option has a profile across them, a level on each.</p><p>A threshold is a chosen level on one of these dimensions, say a level t on income, and an option meets the threshold when it reaches t on that dimension, on whatever scalar reading is intended, for instance the income of the worst-off person. The threshold sorts the options into those that meet it and those that fall below.</p><p>Now the condition that does all the work: <strong>a threshold on a dimension P is a genuine floor when X never strictly prefers an option that falls below the threshold to one that meets it</strong>. If there is even one pair where X ranks a below-threshold option above a meeting option, because that option does well enough on the other dimensions to compensate, then the threshold is not really a floor. It is one weighable consideration among others, and X is prepared to trade it away. A &#8220;minimal&#8221; standard that the theory itself will trade against gains elsewhere is minimal in name only, since it does not behave like a baseline within the theory&#8217;s own evaluation. So minimal non-tradeability is just this: to check whether a proposed minimum is a real floor, you ask whether the theory ever ranks something below the line above something that meets it. In the paper we run this test on sufficiency thresholds and on capability thresholds, but the machinery is general.</p><p>One feature of the test matters for what follow: whether a threshold is a floor is always relative to the set of options you test it against, since the condition only looks at the pairs available in that set. If you change the set of options, a threshold that was not a floor can become one, and a threshold that was a floor can stop being one.</p><h3>Ideal and non-ideal theory</h3><p>This is what connects the method to ideal and non-ideal theory. Ideal theory, on a familiar way of carving things up, works out what a theory demands when we set feasibility aside and consider every option it could conceivably evaluate. Non-ideal theory asks what the same theory demands of us here, among the options actually available to us. The longstanding question has been how to get from one to the other in a principled way.</p><p>In these terms the two are the same theory ranking different sets of options. Ideal theory applies X to the widest set of options; non-ideal theory applies the same X to the narrower set of feasible options. The theory, its dimensions, and its rankings do not change. Only the set of options it ranges over changes.</p><p>So the procedure for moving from one to the other is straightforward. To find what the theory demands as a floor under constraint, you run the same test over the feasible options: on each dimension, you look for the thresholds that the theory never trades away across the options that are actually available. Those are the floors the theory itself underwrites once feasibility is taken into account, and they are extracted from the theory&#8217;s own rankings rather than stipulated from outside.</p><p>A concrete case helps. Suppose a society&#8217;s commitments, considered over every arrangement anyone could imagine, would accept letting some people fall below a subsistence minimum if that were the price of some arrangement the theory rates very highly on other grounds. As long as that arrangement is in the set being ranked, &#8220;everyone is kept above subsistence&#8221; does not come out as a floor, because the theory would give it up for that arrangement. Now restrict to the arrangements we can actually bring about, and suppose the highly rated one is not among them. The test is run again over the smaller set, and this time nothing available is something the theory ranks above keeping everyone above subsistence, so the subsistence minimum comes out as a floor. Same theory, same rankings; the only thing that changed is the set of options the test ranged over.</p><p>That is the whole idea. Minimal non-tradeability gives a way of reading a theory&#8217;s floors off its own rankings, the ideal and non-ideal cases differ only in which options are on the table, and moving between them is just running the same test over the relevant set.</p><p>All of this works cleanly for demands that break down into separate thresholds on separate dimensions, income here, health there, and so on, and whether everything we want to say about justice under constraint takes that form is a further and harder question that I do not want to pretend the mechanism settles. What it offers is a single procedure that covers both the ideal and the non-ideal case, run over two different sets of options, rather than two separate enterprises with a gap between them.</p><p>As ever, this is a working idea, and I would be glad to be told why it is wrong.<br><br>Greetings from (humid) Hong Kong!</p>]]></content:encoded></item><item><title><![CDATA[On Productive and Meta-Skills]]></title><description><![CDATA[A few days ago, I argued that AI would not deskill students.]]></description><link>https://carlolc.substack.com/p/on-productive-and-meta-skills</link><guid isPermaLink="false">https://carlolc.substack.com/p/on-productive-and-meta-skills</guid><dc:creator><![CDATA[Carlo Ludovico Cordasco]]></dc:creator><pubDate>Sat, 13 Jun 2026 17:05:19 GMT</pubDate><content:encoded><![CDATA[<p>A few days ago, I argued that AI would not deskill students. The argument was simple, and I still think most of it is right. If you genuinely need a skill, I said, the world has a way of making you build it, because the system that judges your work (i.e., the exam, the client, the reader, the compiler) eventually catches you out if you have outsourced the thing you were supposed to be able to do. And if you do not need the skill, then losing it costs you nothing, in the same way that none of us mourns our inability to start a fire with flint. I added what I still think is the sharpest point in that post, that most of what we anxiously call &#8220;cheating with AI&#8221; is really a long-standing willingness to let students avoid learning things, a willingness that predates ChatGPT by decades and that the technology has merely made visible. If a student can pass your course by having a machine write their essays, the machine is not the problem; your course was already not testing what you claimed it tested.</p><p>I stand by all of that. But I also, towards the end of that post, made a concession almost in passing, and the disagreement the post generated has convinced me that the concession was quite important and that I buried it. I had granted that, during the developmental years, we might want to insulate students from AI precisely because that is when certain foundational capacities are laid down. I treated this as a minor exception to a general permissiveness. <br><br>This is a post about revising some of my views, and what I want to do is take two ideas that lodged in my mind for a couple of days, follow them carefully, and resist the temptation to let them resolve into something tidier than what I think is the truth.</p><h2>The first thing: skills and meta-skills</h2><p>Here is the distinction I was missing, or rather the distinction I had but was not making explicit. A skill can be valuable in two quite different ways. It can be a <em>productive</em> skill, valuable because it produces an output that someone wants  (e.g., the essay gets written, the code runs, the sum comes out right, the bridge stays up). Or it can be a <em>meta-skill</em>, valuable less for what it produces than for what practising it does to the person practising it: it builds judgment, or discipline, or the capacity for sustained attention, or the ability to hold an abstraction in your head long enough to do something with it, or the habit of catching your own errors before anyone else does. And the point I had not taken seriously is that the same activity <em>can be both</em> at once, and that the two kinds of value <em>can come apart</em>, so that a machine might strip out the productive value of an activity while leaving its formative value completely intact.</p><p>Writing is the example everyone reaches for, myself included, and the reason is that when I write, I produce text, and the text has value to whoever reads it; that is the productive side, and it is the side AI can increasingly take over. But writing also does something to me while I do it that has nothing to do with the reader. It forces the half-formed thought into a sentence, and the sentence is almost always worse than the thought felt, and the gap between them is information. The struggle to close that gap &#8212; to find that I did not actually believe what I thought I believed, that the argument I was pleased with has a hole in the third step, that two things I held at once quietly contradict each other &#8212; that struggle is not a side effect of writing. For a certain kind of thinking, it kind of is half the thinking.</p><p>I believe that, and yet I have learned to be suspicious of it, because notice who is saying it. I am a person who learned to think by writing, which means I experience writing and thinking as nearly the same act, which means that when I reach for an example of a formative skill, writing is the one closest to hand, and its supremacy feels obvious to me in a way that ought to make me distrust the feeling. A mathematician would tell you the same story about proof, with the same conviction, and would mean it just as sincerely. A musician would tell it about the instrument. The trouble is that the meta-skill argument, the moment you take it seriously, does not actually let any of us crown our own discipline. It raises a comparative question, and the comparative question does not politely exempt the activity we happen to love.</p><p>This is the move I made against Latin in my own head. The traditional defence of Latin is purely formative: nobody learns it to talk to Romans, they learn it because construing a Latin sentence is supposed to build precision, grammatical self-awareness, an ear for structure, a tolerance for difficulty. Maybe it does. But if the entire case for an activity is formative, and its productive value has gone to nearly zero, then it has to compete against every other activity that builds similar capacities while also producing something useful along the way, and against that field Latin has to clear a very high bar that I am not sure it clears. Fine; that is the unsentimental conclusion, and I was happy to reach it. The same blade swings back toward writing. Writing is not in Latin&#8217;s position, because it still has productive value, and its formative value is clear. But the question is <em>how much, compared to what, and when.</em></p><p>Take the comparison first. That writing forms thought tells you it belongs in an education. It does not tell you that the next hour of a student&#8217;s life is better spent writing than proving a theorem, or building an argument in symbols, or doing close empirical work, or even certain kinds of coding, all of which form thought too, and several of which also produce something the world wants. I do not know the ranking. I am fairly sure nobody does, and I am quite sure that the people most confident that writing wins are, like me, mostly writers. </p><p>And then the harder question, the one about <em>when</em>, which is where my buried concession comes back to do real work. The formative value of an activity is not a fixed quantity that the activity carries around with it. It depends enormously on the stage of the person doing it. Learning to assemble a paragraph, to follow an argument across a page, to notice when your own sentence has smuggled in a claim you cannot defend &#8212; these are transformative at fifteen and at perhaps at nineteen, when the underlying machinery is still being built, and they are, I suspect, far less transformative at twenty-five, when the machinery is mostly built and the person is now using writing rather than being formed by it. If that is right, then the value of writing is heavily front-loaded, concentrated in exactly the developmental years I had earlier waved at. It is not that writing stops playing a formative role but rather that its formative dividend is largely paid early, and the case for protecting it from AI is correspondingly strongest early and weaker later, where the question shifts from &#8220;how do I build this capacity at all&#8221; to &#8220;how do I do something with it that the machine cannot.&#8221;</p><h2>The second thing: skill is a dial, not a switch</h2><p>The second idea I owe to <a href="/__u/arnoldkling.substack.com/p/comparing-humans-and-ai-on-skills">Arnold Kling</a>, who responded to the last post with a reframing simple enough that I am slightly embarrassed not to have been using it already. My whole argument had run on a binary: is the AI good at this task or not, do you need the skill or not. Kling&#8217;s point is that this is the wrong shape, because skill is not something you have or lack. It is a scalar, an index running from nothing to mastery, and humans and machines both sit somewhere on that single scale, which means the interesting quantity is not whether the machine can do the task but how high up the scale it has climbed, and how that height compares to the level we want humans to reach.</p><p>Call the machine&#8217;s current level X. Written that way, &#8220;can AI do this?&#8221; reveals itself as a lazy compression of a much better question, because X is not a binary answer but a coordinate, and the issue is what coordinate we should be aiming the student at relative to it. And the answer is not the same everywhere. For some tasks, reaching X is plenty. With driving, if the machine becomes more competent than most people, the sensible response is not to demand that everyone surpass it but to raise the floor towards it, to make licensing more demanding over time so that human drivers are at least as safe as the automated baseline, and to accept that those who cannot reach that level perhaps should not be driving. </p><p>For other tasks, matching the machine is exactly the wrong target, because the entire human contribution lies beyond X. If an AI can produce a competent research paper, then training researchers to produce competent research papers is training them to a standard the machine already meets, which is pointless; the target has to sit above X, at the judgment and taste and sense of which-question-is-worth-asking that the machine does not have. The same logic reaches essay-writing, more uncomfortably. Once you say a person should be able to write at least a little better than the machine before they have any business writing for a living, you have set a bar that rises every year, and you have to face the possibility that for many people it will simply become too high, that there are tasks at which most of us should step aside and let the machine work because we cannot clear X and the world does not need us to.</p><p>What this does to my original argument is to expose a soft spot in &#8220;if you need a skill, the world will make you build it.&#8221; The phrase &#8220;need the skill&#8221; was hiding the moving variable, because the level you need is no longer set by the task alone but by the task relative to what the machine already does, and that level is climbing. The useful question is never again going to be &#8220;can AI do this,&#8221; which is answered yes more often every month and tells you almost nothing. It is &#8220;what level of this skill should a human still be aiming for, and why&#8221;. And the why will sometimes be productive, because we need humans above X to do work the machine cannot, and will sometimes be formative, because climbing the early part of the curve builds something in the climber even when the machine could have carried them. Notice that this is the same front-loading I arrived at with writing, coming back in Kling&#8217;s vocabulary: formation is mostly about getting up the steep early stretch of the scale, which is why the case for making students climb it themselves is strongest precisely where they are lowest on it.</p><h2>Where this leaves me: questions, and a reason to run experiments</h2><p>Hold the two ideas together and what you have are two confessions of ignorance wearing the clothes of distinctions. The first tells me a skill&#8217;s value may be productive or formative or both, that the formative part is comparative and front-loaded, and that I usually cannot tell how the comparison comes out, least of all for the activities I am personally attached to. The second tells me the right level of human mastery depends on where the machine sits and on whether the human contribution lies at the machine&#8217;s level or beyond it, and that this too is something I am mostly guessing about. Put them together and the honest summary is that I do not know which skills are worth preserving, I do not know at what level or at what age they are best cultivated, and as far as I can tell neither does anyone else, however confidently they post (which now includes the version of me from a few weeks ago).</p><p>When you do not know something, and the something is an empirical question about how a new technology shapes human development, the rational response is not to pick the answer that flatters your temperament and defend it loudly. It is to find out, and the higher education system happens to be unusually well shaped to find out, because it is not one institution but thousands, and they are under no obligation to agree. So the conclusion I did not have a few days ago is that universities should deliberately <em>not</em> converge on a single policy towards AI, and that the current rush towards sector-wide guidelines, sensible as it sounds, may be precisely the wrong instinct at precisely the wrong moment.</p><p>Let some institutions restrict AI heavily &#8212; handwritten exams, oral defences, AI kept out of coursework &#8212; because we need places where we can watch what is preserved and what is lost when students are made to climb the whole scale themselves. Let others integrate it completely, building it into everything, teaching students to work above X rather than to reproduce what sits below it, so we can see what new capacities emerge when the productive burden is lifted and, more anxiously, whether the formation survives its lifting or quietly dies. And let most do something in between, varying by discipline and by level and, above all, by developmental stage, restricting hard in the years when the foundational climbing happens and opening up later.</p><p>I want to be clear this is not a counsel of indecision, though it will be read as one. It is the opposite. The effects of AI on human learning are empirical, empirical questions are answered by varying the conditions and watching, and a sector that standardises prematurely on any single model has chosen to learn nothing, to discover in twenty years what it could have discovered in five. Institutional pluralism is not what you settle for when you cannot agree. It is what you actively want when the thing you are trying to learn can only be learned by trying several things at once and comparing them, and when one of the things you most need to learn is the comparative, stage-dependent value of activities you are currently too close to to judge.</p><p>So I will end by giving up not the argument exactly but the frame it was caught in, which is the frame most of this debate is still caught in. We keep arguing as though the choice were between preservation and surrender, the defenders of the old skills against the barbarians, the welcomers of the future against the reactionaries, everyone lined up on a side and treating the other as naive or cowardly. I no longer think that is the shape of the problem. The real task underneath the noise is to work out which human capacities still matter in a world that contains these machines, how those capacities are actually cultivated rather than how we sentimentally imagine they are, at what age the cultivating has to happen, and what part the machine should play in it, which will be a different part in different cases. None of that is a battle to be won. It is a set of questions to be answered slowly, by paying attention, and by staying willing to be wrong in public and say so.</p>]]></content:encoded></item><item><title><![CDATA[Why We Should Force Students to Run Marathons]]></title><description><![CDATA[I like long-distance running.]]></description><link>https://carlolc.substack.com/p/why-we-should-force-children-to-run</link><guid isPermaLink="false">https://carlolc.substack.com/p/why-we-should-force-children-to-run</guid><dc:creator><![CDATA[Carlo Ludovico Cordasco]]></dc:creator><pubDate>Mon, 08 Jun 2026 08:40:54 GMT</pubDate><content:encoded><![CDATA[<p>I like long-distance running. I&#8217;m pretty bad at it, but I still enjoy it, and it has taught me a lot. The fact that I can comfortably log 50&#8211;70 poorly run miles a week means I do not worry too much if I gain a bit of weight during a trip where jet lag and the workshop schedule kill my daily run but not my equally enjoyable vodka martinis. More than anything, running, and losing 30 kilograms along the way, taught me that I can go the extra mile when life requires it. I can put in the extra effort when I need to grade 300 essays in a few days, meet a submission deadline, or deal with whatever unpleasant surprise the week has in store.</p><p>I think everyone should run long distances. In fact, I think we should make it mandatory at every grade level. How else would people learn the value of suffering in pursuit of their goals? How else would they learn perseverance? How else would they learn to go the extra mile when life requires it? Ten miles a day would teach lessons that are impossible to forget. It would make the world a better place.</p><p>Granted.</p><p>Now, this is obviously silly. Long-distance running is wonderful, and I am grateful that, after fourteen years of running without rest days whenever injuries and my travel schedule permit, my knees still tolerate it. Yet it is hardly the healthiest way to stay fit. Nor is it the only way, or even necessarily the best way, to learn discipline, perseverance, or productivity. Plenty of people who have never run a marathon possess those virtues in abundance. Plenty of marathon runners do not.</p><p>In Italy, most high-school students are required to learn Latin, and ancient Greek if they are lucky enough to attend the right kind of school. One could, of course, read Virgil, Cicero, Homer, or Sophocles in Italian. One could even read them in English. But that would not do. Latin, the thought goes, teaches you how to think. Ancient Greek does too. And, naturally, there is no other, or perhaps no better, way to achieve that result.</p><p>If you do not learn Latin, you will be condemned to a life of intellectual superficiality. You may become an engineer, a physician, a scientist, a judge, an entrepreneur, or a philosopher. But somehow you will never quite learn how to think.</p><p>That&#8217;s silly, of course. Even if studying Latin helps develop valuable intellectual capacities, it hardly follows that it is the only way, or the best way, to develop them. We always need to ask:&#8221; Compared to what?&#8221;. Compared with learning a modern language, playing piano, studying mathematics, writing code, debating philosophy, reading history, or engaging with the emerging tools that increasingly shape the world we inhabit?</p><p>There is a pattern here, and once you notice it you start to see it everywhere. You take an activity that does some good. You observe, correctly, that the people who do it often turn out to have some valued capacity: discipline, rigour, the habit of close attention, the ability to think. And then, somewhere between the observation and the conclusion, &#8220;this builds the capacity&#8221; turns into &#8220;this is how the capacity is built,&#8221; and then into &#8220;without this, the capacity is never built at all.&#8221; Each step looks small. The last one is very large, and it is almost never argued for. It just gets assumed, usually by someone who acquired the capacity while doing the activity and has, very understandably, taken the one for the other.</p><p>It is a forgiving sort of reasoning, because from the inside it can never be shown wrong. The classicist learned to think while reading Thucydides, so the Greek (thanks, Eric!) gets the credit. The marathoner learned to endure well while running, so the running gets the credit. Neither of them ever got to see who they would have become without the thing they happened to do, so neither can quite believe they would have become much at all. The capacity and the activity arrived together, in the same person, at the same time, and it is very hard, after the fact, to tell which one was really responsible.</p><p>I dwell on this because it is the same reasoning now often being used against letting students use AI. Writing your own essays, the thought goes, is how you learn to think: how you learn to build an argument, weigh evidence, and find out what you actually believe by trying to say it. I suspect that is largely true. Writing did all of that for me. But look at how quickly the claim grows. We begin with &#8220;writing teaches you to think,&#8221; which I will happily grant, and end up with &#8220;writing is the only way to learn to think,&#8221; or the slightly weaker &#8220;writing is the best way,&#8221; and it is that larger, unstated claim that does the work whenever someone warns that a generation raised on these machines will never learn to think at all.</p><p>So, as ever: compared to what? Compared with reading three answers a machine has produced and working out which one is rubbish, and why. Compared with taking a model through five versions of an argument, turning down the flattering one, and holding out for the version that is actually true. Compared with having to ask a precise question in order to get a useful answer, which is already a good part of the work of thinking. All of this is difficult, and all of it calls for judgment, for discrimination, for the willingness to say &#8220;no, that is not quite right,&#8221; which may matter more to thinking well than the writing of fluent prose ever did. Whether these new activities do the job better than the old one is a real question with a real answer, and the people most alarmed about AI are the ones least willing to ask it.</p><p>None of this makes writing worthless. Writing by hand may well stay one of the best ways we have of learning to think, and for some students, at some stages, the best of all. But you only get to say so after making the comparison, not before it. The honest question is: of all the things a young person could do with their attention now, in a world that really does contain these machines, which ones best build the capacities we claim to care about? You cannot answer that by listing the merits of the old activity and declining to look at any of the others. That is just a man telling you everyone should run ten miles a day because it once worked for him.<br><br>It is worth speculating about why the comparison almost never gets run. The reasons, I think, are not trivial, which is part of why the mistake is so durable.</p><p>The first I have already touched on. The people doing the judging grew up on the old activity. They learned to think while writing essays, or while reading Cicero, and so the only model they have of someone who can think is a model of someone who did those things. When they look at a student who has not done them, they see an absence, and they fill that absence with the worst case. They are not being dishonest. They are reasoning from the only example they have, which is themselves, and they have no memory of the person they would have been on a different path, because that person never existed.</p><p>The second is that the old activity has had a very long time to build its reputation. Latin has had centuries to collect its defenders, its examiners, its stories of the great minds it supposedly formed. Essay-writing has had nearly as long. The activities that might replace some of that work are a few years old and still carry the smell of cheating. So the contest is settled before it starts. One side arrives with a long record and a great deal of institutional dignity, and the other arrives looking like a shortcut, and we mistake the difference in reputation for a difference in worth.</p><p>The third reason is that the capacities we are arguing about, judgment, discrimination, the ability to think well, are the ones we have never really known how to measure. We cannot open a student up and read off how well they think. So we have always relied on things we can see, and the essay was the most convenient of them. A student who wrote a good essay was probably thinking well, and we let the essay stand in for the thinking, because we had nothing better. The proxy worked well enough that we stopped noticing it was a proxy. And then the machine arrived and made it easy to produce the essay without the thinking, and we reacted as though the thinking itself had been taken from us, when what we had actually lost was the proxy we had quietly mistaken for it all along.</p><p>That last point is that we do not have a clean way of scoring the old activity and the new ones side by side, because we never had a clean way of scoring the thing they are both supposed to produce. It explains why the alarm is so genuine, and so genuinely misdirected. And it explains why the harder problem, the one worth our attention, is not whether students use these machines, but how we ever tell, with or without them, whether a young person has actually learned to think.</p><p>Suppose I have convinced you, then, that the comparison is worth running, and that the alarm is mostly misdirected. What follows for an actual university, the kind that has to decide before next term whether to put these machines in front of its students? Here I have to be more practical, and a good deal less sure of myself.</p><p>I should say where I stand, since it colours the rest. I am glad to work somewhere willing to experiment with all this, and I would choose such a place again. That is a preference of mine, and I do not think every institution needs to share it. Someone has to go slowly, and there is something to be said for letting others make the first mistakes. I would simply rather be among the ones making them.</p><p>There is more to it than taste, though. The skill we keep saying we care about, the ability to think well with these machines and about them, is going to be learned somewhere. If universities decline to teach it, on the honourable-sounding grounds that the decent thing is to keep the tool out of the room, it will be learned elsewhere, and elsewhere increasingly means the companies that build the tools. They have the money. They have the appetite for young talent. And they have, lately, begun hiring philosophers, of all people, which is the clearest sign I know that something serious is underway, since we are nobody&#8217;s idea of a safe hire. You can put this in blunt commercial terms and say that universities risk losing their customers. I find the commercial framing a little vulgar, but it is not wrong, and the vulgarity is part of the warning.</p><p>It matters here that we are talking about universities and not about schools. A lot of the case for keeping the tool away from students is really a case about children. When someone is still forming the basic habits of attention and argument, there may be a reason to make them do things the hard way, because the difficulty may be the lesson. But the reason fades as the capacities form, and by the time someone is sitting in a university lecture, most of it has gone. What they need to learn now is the world, and the world has these machines in it.</p><p>So a rule that made sense at ten makes much less sense at twenty. Keeping the tool from an adult, to preserve a difficulty that no longer teaches them anything, turns education into a game whose rules exist for their own sake. And there is a cost to playing it. Working well with these machines is becoming one of the central skills of the age, maybe the biggest change in how we think in a very long time, and a university that keeps it out of the classroom is leaving its students to figure it out alone, in their bedrooms, at night, with nobody to show them how. For a place that exists to teach people things, that is an odd thing to skip.</p><p>And here my confidence runs out. Saying we should bring the machines into the room and teach students to work with them is easy enough. Saying how I would grade the result is not. Every assignment I have tried to design for this, every task meant to reward working with the machine rather than just handing the work over to it, I have been able to defeat in about five minutes, usually while I am still writing it. I am not especially devious. The problem is that the thing I want to catch, whether the student thought alongside the machine or just took what it gave them, often leaves no mark on the work. The two can produce the same essay. </p><p>But marking is only half of it, and the smaller half. An assessment does not just measure what a student did. It is the main reason they did it. Most students, under deadline and pressure, do what is rewarded and not a great deal more, and there is nothing shameful in that. It is how anyone behaves when time is short and the stakes are real. The old essay worked partly because the only way to get the reward was to do the work. You could not produce the essay without the thinking, so the mark and the thinking came together. That is what AI quietly breaks. Now you can collect the reward without the work.</p><p>And the two problems turn out to be the same problem. I cannot reliably reward genuine collaboration for the same reason I cannot reliably spot it: it does not show up in the work. You cannot mark what you cannot see, and you cannot reward what you cannot mark. So the thing that defeats the grading defeats the motivation too, which is why this is harder than it looks, and why no amount of fiddling with rubrics quite fixes it.</p><p>There is one kind of assignment that works. You put the students in a room, give them the machines, and watch what they do. That gets you both things back at once: you can see the collaboration, so you can grade it, and once you can grade it, it is worth their while to do. It is also expensive and perhaps impossible to run for a course of three hundred. So I do not have a happy ending for you. We should experiment, because the alternative is to leave the most important new tool in education to the firms now hiring our colleagues. We should bring it into the room, because keeping it out gives up most of what a university is for. And we should be honest that we do not yet know how to grade what happens when we do. That is the work worth doing, and it deserves far more of our attention than the question of whether to allow the thing at all.</p>]]></content:encoded></item><item><title><![CDATA[No, AI won't deskill students]]></title><description><![CDATA[On UChicago and Claude for everyone]]></description><link>https://carlolc.substack.com/p/no-ai-wont-deskill-students</link><guid isPermaLink="false">https://carlolc.substack.com/p/no-ai-wont-deskill-students</guid><dc:creator><![CDATA[Carlo Ludovico Cordasco]]></dc:creator><pubDate>Sat, 06 Jun 2026 18:42:19 GMT</pubDate><content:encoded><![CDATA[<p>I&#8217;ve been travelling a great deal lately, which has kept me from the blog, and, if I&#8217;m honest, from finishing much of anything. I&#8217;ve started many blogposts and completed none, begun drafting several papers and closed out none of those either. I am once again in that recurring phase of life in which I try to work out where to go next while waiting, with rather less patience than I&#8217;d like to admit, to hear back from a handful of submissions whose statuses I check on the various editorial managers far more often than checking could possibly help.</p><p>This morning, though, prompted partly by an exchange with the always excellent Eric Schliesser, I found myself wanting to join the now longstanding debate on AI and deskilling, which I have found, for the most part, misleading. And misleading in a fairly specific way, which is what I want to set out here.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://carlolc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Paperclips and Other Alignment Problems! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The thing that finally made me sit down and write was a piece of news. On 2 June the University of Chicago announced that it is giving all of its students, and all of its staff and faculty, full access to Claude. Staff get it from July, and students will have it before the autumn term begins. Chicago is not the first university to do this. Dartmouth has had much the same arrangement for about a year, and Berkeley and Duke have made their own deals with other AI companies. But Chicago is a serious place, not much given to chasing fashions, and the way its president explained the decision has stuck with me. The university&#8217;s job, he said, is to teach students three things: how to think with these machines, how to think without them, and how to think about them. I think that is exactly right, and I want to come back to it at the end.</p><p>The reaction, though, was the one you would expect, and a good part of it was alarm. Give students a machine that will write their essays for them, the worry goes, and they will never learn to write, or to argue, or to think for themselves. They will hand in work they did not do and cannot understand, and a whole generation will come out the other side hollowed out. I understand the worry, and I do not think it is silly. But I think it rests on a muddle about what a skill actually is, and once that muddle is cleared up, most of the alarm has nowhere left to stand.</p><p>Start with the word everyone is leaning on: deskilling. To be deskilled is to lose a skill. So what is a skill? We value skills for all sorts of reasons, and I do not want to wave any of them away: for the independence they give us, for the plain pleasure of being good at something, for the way they tie us to a craft or a tradition. But the deskilling worry is not really about those reasons. It is a worry about losing a capacity we still need, one that still does some real work for us, and that is the kind of skill I want to take on here. Once you fix on that meaning, something strange about the worry comes into view. Handing a task to a machine is not, by itself, evidence that the underlying human capacity has lost its point. It raises the question whether that capacity is still needed, and if so where.</p><p>Think about the calculator. Nobody says, or should say, that the calculator deskilled us, even though hardly any of us could now do long division on paper with much confidence. The reason we should not say it is that the need to do long division by hand simply went away. The ability went away along with it, and we lost nothing we still had any use for. Calling that deskilling would be odd. It would mean mourning the loss of something we were rather glad to be rid of. And the same thing is true far more widely than we like to admit. When a tool takes over a job we genuinely no longer need to do ourselves, the disappearance of the old ability is not a loss at all. It is just the need leaving, and the ability quietly following it out the door.</p><h2>The dilemma</h2><p>There is a cleaner way to see all this, and it sits right at the centre of the argument. Take any task you might be tempted to hand to an AI: writing an essay, solving an equation, drafting some code, summarising a paper. Now ask one simple question about it. Is the AI good at this task, or is it not? There are two clean cases, and then one messy case where the real educational problem lies.</p><p>Suppose, first, that the AI is not good at the task. Then you still need to be able to do it yourself, or at least to tell when the machine has got it wrong. And here is the part that matters: because you need that ability, the system that assesses you has a natural way of making you build it. Picture a student who leans on a shaky AI and never bothers to learn enough to catch its mistakes. That student hands in weak, error-filled work, and is marked down for it. The wish to avoid that is what pushes the student to learn the skill properly. As long as bad work is recognised as bad work, the ordinary desire to do well keeps the skill alive, machine or no machine. So where the AI is unreliable, the skill does not vanish. It is kept safe by the plain fact that you still need it.</p><p>Now suppose, instead, that the AI is good at the task. Genuinely good, reliably good. Then, if we are honest, you probably do not need the skill anymore, in just the way that you no longer need to work out a square root with pencil and paper. The ability has become optional. You can still learn it if you enjoy it, or because it has some independent value, but nothing is lost by letting it go simply because it is no longer required. Nothing was really demanding it of you in the first place.</p><p>Put the two cases side by side and most of the worry drains away. If you still need the skill, the system that judges your work has a way of making you build it. If you no longer need it, then losing it costs you nothing. There is a third case, though, and it is the one people really have in mind when they worry about AI. A student might still need a skill and yet go through an education that lets them put off finding out they need it until it is too late to learn it. The need is there, but nothing makes them feel it in time.</p><p>Now look at what is actually causing the harm there. It is not the machine. It is a way of marking work that lets fluent, empty answers pass and never makes the student find out what they cannot yet do. And this is an old failure. Students have always copied, crammed, leaned on templates, nodded along to feedback they did not understand, and learned to give a marking scheme what it wants without ever learning the thing behind it. None of that needed AI. So when people warn that students will use AI and pass without learning, the real work in that sentence is done by two small words: and pass. If a student can do shallow work and still pass, the university already had a problem, long before any machine turned up. A lot of what we call worry about AI is really worry about how we mark and assess. AI has not made it newly possible to get through a degree without understanding. It has only made it much harder to ignore.</p><h2>Judging is its own skill</h2><p>Behind the worry about assessment lies a deeper one, which is likely the real engine of the alarm. When the machine is good, I said, the skill shifts from producing the work to judging it. This deeper worry says that this shift is a trick, because you cannot really judge work you could not have produced. To tell whether an essay is any good, the thought runs, you have to be able to write one; to catch a broken proof, you have to be able to build one. If that were right, a student who lets the machine produce would lose the power to judge along with it, and end up able to do neither.</p><p>This runs three different things together. There is the skill itself: telling a good argument from a bad one, a sound proof from a broken one. There is producing the work: coming up with the argument or the proof yourself. And there is the way we usually test the skill: the essay handed in and marked. We treat the three as one thing, and they are not.</p><p>Take producing and judging first, because that is the centre of it. They are different skills, and they come apart all the time. A football coach is often a mediocre player. A good editor can say exactly what is wrong with a novel and could never write one. A critic&#8217;s taste routinely outruns anything they could make themselves. The power to judge does not follow automatically from the power to produce, and a person can have a great deal of the one with very little of the other.</p><p>What is true is narrower. Producing is one of the ways we learn to judge. When you work out a proof of your own, you are forced to weigh each step as you take it, and that steady pressure trains the judgment. But producing is a way of training the skill, not the skill itself, and it is not the only way. You can learn to read and judge proofs without ever being made to write one, so long as something else does the forcing, so long as you are made to judge and corrected when you judge badly. Producing is one road to judgment. It is not judgment, and it is not the only road.</p><p>Even in philosophy, the essay was never the skill itself. It was a way of making thought visible, so that it could be tested, corrected, and judged. Dialogue, oral examination, seminar exchange, and written argument are different vehicles for that same underlying discipline. We forget this constantly. We watch a student stop writing essays by hand and conclude that something vital has gone, when all that may have changed is the vehicle.</p><p>So this deeper worry falls apart. When the machine takes over producing, the skill that now matters, judging what it gives back, is not pulled down with it. Judging was always a separate skill. Producing was only ever one way of teaching it, and students can be taught another way: by being made to judge the machine&#8217;s work, and corrected when they judge it badly. The task was never to keep students producing for its own sake. It is to make sure that something still makes them judge.</p><h2>The exceptions worth keeping</h2><p>Everything so far has pointed one way: offloading is, on the whole, safe, and the skills that matter look after themselves. I do not want to oversell that, though, because there are two plausible exceptions, and because they work in different ways, let me take them one at a time.</p><p>The first is about timing. Some things are simply learned better early, before a student has a machine to lean on. The worry is not that you could never pick them up later, because often you could. It is that the early years are when certain habits form most easily, and a student who hands everything to a machine from the start may never lay that groundwork down. So some early stages are worth protecting on purpose. That is not a reason to keep AI out of education. It is a reason to guard a few early stages while letting the rest become openly machine-assisted.</p><p>The second is the emergency, and its logic is different. Here the trouble is that the need for a skill can arrive faster than it could ever be learned. For a small number of jobs, the day comes when the machine fails and there is no time to catch up, so the skill has to be there already, built in the years when the machine was handling things and the person did not seem to need it. We do not deal with this by asking everyone to keep every skill, just in case, which would be both impossible and pointless. We deal with it the way we always have, by deliberately keeping certain abilities alive in the particular people who will need them. We train a few surgeons to operate when the equipment dies and pilots to fly when the automation quits, and we carry the cost because the stakes are high and the warning is short.</p><p>Different as they are, both exceptions ask for the same narrow thing: protect a few specific stages, and train a few specific people. Neither is a reason to be wary of offloading in general. They are the carve-outs we make on purpose, in the few places they plainly earn their keep, precisely so that everyone else can offload without a second thought.</p><h2>Offloading is how we build new skills</h2><p>I want to put this more strongly, because so far I have mainly argued that offloading does no harm, and, as I have argued elsewhere, I think it does positive good. Offloading is not the enemy of skill. It is the way skill keeps moving to wherever it has become useful, and that is a reason to want it, not merely to tolerate it.</p><p>When a capacity stops being needed, the effort that used to go into it does not simply disappear. It moves. The calculator did not leave the world with less mathematical ability in it. It freed people from grinding through arithmetic by hand and let them put that freed attention into harder and more interesting mathematics, and into a hundred things that were not mathematics at all. This is the ordinary pattern of every useful tool. You hand the old task to the machine precisely so that you can pick up a new one. A student who offloads the things AI now does well is not being emptied out. They are being freed to learn the things that sit on top: how to ask the right questions, how to weigh the answers, how to put strange materials together into something of their own.</p><p>So the case for handing students these machines is not just that it will not hurt them. It is that it clears the ground for them to build whatever comes next, and the more they let go of what no longer needs doing, the more room they have to do it.</p><h2>Back to Chicago</h2><p>Which brings me back to the Chicago president and his three things. Thinking with machines is the skill the world now mostly rewards, and we should teach it openly and without apology. Thinking without them is the narrow business I have just described: the few early stages worth learning the hard way, and the few jobs where someone must be able to carry on when the machine fails. And thinking about them is knowing, for any task in front of you, which of the two you are dealing with.</p><p>All of this lets us say precisely what AI does and does not do to a university, because the single word &#8220;deskilling&#8221; hides three very different claims. The first is that AI changes which skills are worth having. That is true, and it is no loss. It is just the need moving, as it always has. The second is that AI exposes assessments that were already weak measures of understanding. That is a real problem, but an old one in new clothes, and it was always going to need fixing. The third is the one that genuinely earns its worry. AI can take a weakness a university was quietly tolerating and make it acute, because it drops the cost of evasive work almost to nothing.</p><p>So the honest thing to tell Chicago, and everyone watching it, is not that the machine is safe, nor that it is dangerous. It is that the machine is a stress test. Where teaching already tracks understanding, handing students these tools pushes the work upward, toward judging, revising, and owning what the machine gives back, and that is simply education doing its job. Where teaching only ever rewarded a fluent surface, the machine will make that impossible to hide.</p><p>The sensible default is more offloading rather than less, with the burden on whoever wants to hold a skill back to show that it is one of the few genuinely worth protecting. Skill is not being destroyed here. It is moving to wherever it is still needed. Our job is to follow it there, and to make sure that our teaching and assessment can still tell whether anyone actually has it. Where a course cannot survive its students being handed a good machine, the machine was probably never the deepest problem.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://carlolc.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Paperclips and Other Alignment Problems! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The confessions of a question-driven guy]]></title><description><![CDATA[and how AI has helped me overcome a few challenges]]></description><link>https://carlolc.substack.com/p/the-honest-tails-of-a-question-driven</link><guid isPermaLink="false">https://carlolc.substack.com/p/the-honest-tails-of-a-question-driven</guid><dc:creator><![CDATA[Carlo Ludovico Cordasco]]></dc:creator><pubDate>Tue, 05 May 2026 12:07:12 GMT</pubDate><content:encoded><![CDATA[<p>AI has changed my work in ways I could not have anticipated, and I want to try to distil some of those changes in this post (I&#8217;ve started unpacking one of them <a href="/__u/carlolc.substack.com/p/carlo-but-seriously-what-are-these">here</a>). Some I expect will become legible only later, as the workflow keeps refining itself, though the changes that have surfaced so far are interesting enough to set down before they become invisible to me through familiarity.</p><p>Let me start with a basic admission. I am a master of none. My PhD is in political philosophy, though my depth of knowledge does not match that of colleagues who have spent careers within the field. I publish mostly in business ethics, despite never having been trained as a business ethicist. I find myself drawn to questions about AI and the future of work without being an AI ethicist by training. I am increasingly interested in decision theory and microeconomic theory, though I am clearly not an economist. I sometimes reach for formal modelling, and I have no mathematical training to speak of, with high school maths skills that were forgettable at best. Whatever I am doing, I am doing it without the kind of disciplinary anchoring that most of my colleagues bring to their work.</p><p>This year, I have been fortunate to place several papers, some at venues I am genuinely proud to appear in. One explanation is luck, in the form of sympathetic reviewers and well-aligned stars and the various ways an academic career can occasionally tilt in your favour without quite deserving it. The other explanation, which is more self-indulgent, is that I work orthogonally to most of my colleagues. I do not accumulate encyclopedic knowledge before looking for questions, and I do not have a method or a subject I treat as my own intellectual home. I am question-driven rather than field-driven or method-driven. The questions I find myself unable to put down happen to fall mostly within philosophy, with increasing forays into decision theory, though the questions came first and the disciplinary locations followed.</p><p>For most of my career, this way of working has been a liability as much as an asset. It produces papers that do not sit cleanly within established conversations, that have to teach the reader the relevance of the question before making the argument, and that risk being read as the work of someone who has not quite mastered any of the literatures they are drawing on, because in a sense that reading is correct. The conventional profile, which my method does not fit, is for good reasons the one the profession recognises most readily, and working outside it means accepting that the work will sometimes be read as less serious than it may be.</p><p>What AI has changed, in my experience over the past year or so, is the cost structure of working this way.</p><p>The first thing AI has changed for me is the texture of imposter syndrome. The phrase usually refers to a feeling unwarranted by the facts, where accomplished people are convinced against the evidence that they do not deserve their position. My version is different, and I suspect closer to the original meaning. My knowledge is horizontal rather than vertical, which means I am always genuinely at risk of missing chunks of literature that someone with conventional training would have absorbed years ago. I write papers that draw on fields while having read less of them than my colleagues who specialise in those areas. I make arguments that depend on a formal apparatus I learned while writing the paper rather than years before. This is a real vulnerability, and the imposter feeling that goes with it tracks something accurate.</p><p>For a long time, this kept me from writing things, or from submitting things I had written. The asymmetry between the cost of being caught missing something obvious and the cost of not writing at all felt heavy enough that the not-writing-at-all option was often the one that won. I have a folder of papers that never went out because I could not be sure I had not missed the foundational paper that would have made the argument either redundant or naive. Some of those papers were probably fine. Others probably were not. The folder was there because I could not tell which was which without the kind of literature fluency I did not have time to acquire.</p><p>What AI has changed, more than anything else in this respect, is the cost of partial reassurance. Brainstorming with an AI assistant does not give me the kind of literature coverage that years inside a subfield would. It misses things, sometimes important things, and I have been caught out by the gaps more than once. What it gives me is enough surface visibility into adjacent debates that the worst-case scenarios become much less likely, where I might write a paper that reproduces a well-known argument without realising, or make a move the literature has already shown to be wrong. The reassurance is partial and sometimes illusory, though partial reassurance turns out to be great when the alternative is no reassurance at all. The folder of unsubmitted papers has shrunk. Some of what has come out of it has been placed at venues I would not have submitted to even six months earlier, because the imposter feeling was strong enough to keep me from trying.</p><p>The second thing AI has changed for me is my relationship to formal methods. I have no mathematical training to speak of, and for most of my career this meant that the formal apparatus some of my questions seemed to require was effectively closed to me. I would gesture at imprecise probabilities or at decision-theoretic structures from outside, treating them as black boxes whose outputs I could cite without fully understanding the machinery. The papers I wrote that touched on these areas were limited by what I could understand from the surface, which was less than the questions actually demanded.</p><p>What changed, and what I want to articulate carefully, is that AI removed a particular friction that had been doing more work than I had realised. When I learn formal apparatus from a human, even a generous one, there is a threshold below which I am reluctant to ask questions. The threshold is set by embarrassment. The questions whose answers any serious student would already know, the ones that would reveal how thin the foundation actually is, get left unasked. I would try to compensate by reading more carefully or by triangulating from context, and the compensation was always partial. The foundation stayed thin, and the apparatus remained half-understood.</p><p>With AI, the threshold disappears. I ask the questions I would be embarrassed to ask a colleague, and I keep asking them until I actually understand. While writing one of my recent papers, I built my working knowledge of a formal apparatus this way, through hundreds of small exchanges in which I asked the dumb question and got the patient answer and then asked the follow-up dumb question that the answer had revealed I needed to ask. The knowledge that resulted is genuinely precarious in ways a real specialist&#8217;s would not be. I can use the apparatus for the specific arguments my paper required, and I would struggle to use it for arguments I have not already made. Even so, it was enough for the work, and the work would have been impossible without it.</p><p>The way I learn formal methods through AI is not the way someone with proper training has learned them, and the difference is huge. A trained specialist has intuitions I do not, and I am prone to mistakes a trained person would not make. For someone in my position, though, where the questions sometimes reach into formal territory while the training never went there, and where the combination of inadequate foundation and embarrassment about exposing it had previously locked me out of the apparatus entirely, the access AI provides is genuinely transformative. Better than a human teacher would be, in absolute terms, is a claim I do not want to make. Better than the no-teacher option, which was the only option <em>actually</em> available to me, certainly.</p><p>If the question-driven approach is going to be defensible rather than just a rationalisation of an unconventional profile, it needs to do work that other approaches cannot. I think it does, though the work is more specific than the broad claim suggests.</p><p>The first thing question-driven research does is start with the question rather than with an answer looking for one. The standard academic path produces researchers who develop a method, an apparatus, or a theoretical commitment over years of training, and who then look for questions the method can usefully address. The work that results is often technically excellent within its framework, though it is also often answering questions that emerged from the framework itself rather than from the world. Question-driven research inverts this. The question comes first, generated by attention to something that seems to need explanation, and the apparatus is whatever turns out to be needed. The result is less polished within any single framework, while more responsive to what actually matters.</p><p>The second thing question-driven research does is cross disciplinary boundaries naturally rather than artificially. If you start with a method, the cross-disciplinary work has to be deliberate, where you reach into another field to apply your apparatus to its questions, or you import another field&#8217;s apparatus into yours. If you start with a question, the question does not respect the boundaries the methods do. The interdisciplinarity becomes a consequence of taking the question seriously rather than a goal in itself.</p><p>The third thing is that question-driven research can identify problems disciplinary frameworks have rendered invisible. The work is no easier than conventional research, and it is located in places the disciplinary frameworks do not habitually scan. Phenomena that have been in plain sight in cases the field has been writing about for decades can remain invisible because the framings within which those cases get discussed pull attention away from structural features a different framing would surface. Question-driven work can see those features precisely because it is not committed to any single framing. The advantage is one of perspective rather than of difficulty.</p><p>I should also say that the question-driven method has a corresponding weakness. The work it produces tends not to fit cleanly into the citation networks of any single field, and it can therefore be slower to accumulate influence than work sitting firmly within an established conversation. The same feature that lets the work be reframing rather than refining is what makes it harder to be cited and built on.</p><p>There is a further difficulty I want to name, because it is specific to some of the fields I aim to publish in. Top journals in management research are typically not triple-blind, and editors at these journals manage submissions across an enormous methodological range, from formal economic theory through empirical and qualitative work and back. Desk decisions in this context have to triage, and the signals that triage relies on include features that are partly about the paper and partly about who the author is and where they sit in the network of researchers the editor recognises. This is a structural feature of how editorial work functions under those constraints rather than a critique of any particular editor or journal. What it means for someone with my profile is that the cost of being unconventional in management research is paid at the level of whether the paper gets read carefully in the first place. Philosophy journals at the top tier are mostly triple-blind, and the dynamic I am describing applies less clearly there. AI helps with both contexts in different ways, though the structural conditions that shape how unconventional work gets evaluated in management research have not changed.</p><p>What I find myself wanting to say, in conclusion, is that AI has not changed what kind of researcher I am. It has changed what kind of scholar I can sustainably be. The question-driven inductive method has always been viable for some people in some institutional contexts, though the costs of working this way have historically been high enough to keep many people who would have flourished within it from ever getting started. The reduction in those costs over the past few years is, in my experience, a structural shift in who can do this kind of work, and I suspect there are many more masters of none than the conventional profile of academic excellence has historically allowed for. Some of them will produce work the field-driven and method-driven approaches could not have produced. The shift is worth taking seriously, both for what it enables and for what it asks us to revise about who counts as a serious contributor to intellectual life.</p>]]></content:encoded></item><item><title><![CDATA[On a Bad Day for Business Ethics (and Wine, and Datacenters)]]></title><description><![CDATA[Let me rant once in a while!]]></description><link>https://carlolc.substack.com/p/on-a-bad-day-for-business-ethics</link><guid isPermaLink="false">https://carlolc.substack.com/p/on-a-bad-day-for-business-ethics</guid><dc:creator><![CDATA[Carlo Ludovico Cordasco]]></dc:creator><pubDate>Wed, 29 Apr 2026 19:06:55 GMT</pubDate><content:encoded><![CDATA[<p>It has been, as days go, not a great one. My favourite restaurant in Manchester, Climat (eighth-floor wine-led rooftop, Burgundy-leaning, the kind of place where you could pretend for two hours that you were not in fact in the rain), has closed. The data-centre stocks I optimistically loaded onto, on the theory that compute is the new oil, are taking a more pessimistic view. And the journal where I have published more papers than in any other has been removed from the FT50, which, for those of you who have the great fortune of not knowing, is the list of fifty journals the <em>Financial Times</em> deems important enough to count toward the rankings of business schools, and which most of those schools have institutionalised as the actual scoreboard for whether your work counts as research at all.</p><p>You&#8217;re meant to triage the bad news of a day by severity, but I find it helpful to triage by reversibility. Climat will not reopen. The GPUs may eventually rebound. The journal is unlikely to be reinstated. So we are dealing, on a more sober ranking, with two losses and one cyclical correction, which means the proper move is to pour something red from the Jura region and write through it.</p><h2>The institutional reality</h2><p>The Journal of Business Ethics was already ABS 3*, and Business Ethics Quarterly was demoted to ABS 3* in the 2024 list, which means that for a working philosophical or normative business ethicist embedded in a UK or European business school, there is now no reliable, ranked outlet that signals &#8220;this is research the dean should care about.&#8221; For those who don&#8217;t navigate these systems daily, this matters more than it might sound. The ABS Academic Journal Guide and the FT50 are not abstract metrics. They are the language administrators speak when they decide promotions, allocate research funding, write hiring lines, and respond to accreditation reviews. AACSB and EQUIS look at FT50 publications when they review schools. Schools look at FT50 publications when they hire. Hiring committees look at FT50 publications when they shortlist. The list is not a guide. It is the gate.</p><p>Top business schools, the Whartons and the Georgetowns and the small handful of American departments with a tradition of ethics in management, have long maintained their own bespoke lists. Their faculty handbooks count <em>Philosophical Review</em>, <em>Ethics</em>, <em>No&#251;s</em>, <em>Philosophy &amp; Public Affairs</em>, <em>Mind</em>, <em>Journal of Philosophy</em> alongside or instead of the FT50, because their administrators understand that a paper in PPA is not interchangeable with a paper in JBE about the moral psychology of board members. Those handbooks were built by deans who knew that some of the most important work in their schools wouldn&#8217;t show up in the standardised league tables, and who had the institutional confidence to define their own criteria. For everyone else, which is most business schools, the rankings are the only language that the dean speaks. If your work doesn&#8217;t appear in the rankings, your work doesn&#8217;t appear.</p><p>There is a generational subtext here that I think deserves to be made explicit, because it has implications for how the field actually evolves. The senior business ethicists who currently hold lines at most schools secured those lines when JBE was a prestigious FT50 outlet and BEQ was a 4*. The career path that produced them is being closed off below them. They will be fine. The next cohort finds itself in a peculiar position, having entered a discipline whose ladder is being dismantled rung by rung while we are halfway up.</p><h2>How JBE and BEQ contributed to this</h2><p>I should say at this point that JBE and BEQ have not been entirely innocent victims. There is a version of this story where a virtuous, philosophically rigorous outlet was crushed by philistine bibliometrics, and that version is comforting and partly true. The longer version is that both journals tried to be everything to everyone over the past decade, accepting empirical work in CSR, qualitative organisational ethnography, normative philosophy, business case studies, and a great deal of work that hovers somewhere between management theory and consulting deck. The diversification was understandable in the short term, since bigger volume means broader citation base and more constituencies, but it exposed the journals to quality control problems they would not otherwise have faced (e.g., the empirical CSR literature has its own well-known struggles with measurement, robustness, etc.), and it diluted what made them distinctive. There were already plenty of outlets for empirical management-ethics work, including <em>Academy of Management Journal</em>, <em>Strategic Management Journal</em>, <em>Organization Science</em>, <em>Journal of Management Studies</em>, and ASQ, all of which routinely publish ethics-adjacent empirical papers. There was nowhere else for normative philosophical work in business ethics to go. The journals occupied a unique slot in the ecosystem and traded part of it for a piece of a crowded one, and the trade did not pay off.</p><p>What the field needs, if there is to be a viable institutional home for analytic philosophical work in business ethics, is something like the inverse of what the journals became. A small, properly analytic philosophy journal for conceptual and normative work in management contexts, with fewer papers per year, real refereeing by working philosophers, and a willingness to engage with the problems that arise in firms, markets, and managerial decision-making while writing with the rigour that would get you taken seriously by the editors of <em>Ethics</em>. There is more than enough material to fill such a journal for a long time, including questions about AI agency in firms, about loyalty and exit, about the moral status of corporate intentions, about the conditions under which corporate apologies are sincere or merely performative, etc. The trick, the one neither JBE nor BEQ managed, is to refuse to be everything.</p><h2>What makes the timing especially galling</h2><p>What makes the timing of all this especially ironic is that philosophical expertise has, in the last few years, become genuinely commercially valuable in a way it has not been for decades. Henry Shevlin just took a senior philosopher role at Google DeepMind, focusing on machine consciousness and AGI readiness. Amanda Askell has been at Anthropic for years and is increasingly central to how the company thinks about model behaviour. Joe Carlsmith&#8217;s writing on existential risk gets read more carefully by frontier AI researchers than most management research is read by managers. Frontier AI labs are hiring philosophers in named, senior positions to think about consequences, distinctions, and what counts as a reason for what, which is to say they are hiring people to do exactly the things business schools claim to teach but increasingly do not, because the people who can teach them well are being squeezed out by ranking incentives.</p><p>The skills a good analytic philosopher brings, which include drawing distinctions that economists and lawyers don&#8217;t see, modelling decision problems formally when that helps, and refusing to accept conceptual mush as a finding, are not exotic. They overlap heavily with what good economists and good mathematicians do, which is part of why philosophers have always been useful in technical settings. What philosophers add is a willingness to take seriously the question of what a thing is before counting it, a useful corrective in a discipline where measurement frequently runs ahead of conceptualisation.</p><p>Business schools, and the bodies that produce these journal rankings, cannot afford to keep being blind to this. The market signal from AI labs, from law firms doing AI governance, from policy shops, from consultancies trying to figure out what &#8220;responsible deployment&#8221; means in operational terms, is loud and unambiguous. Philosophical training places people. Philosophical training also produces work that gets read outside the field, which is more than can be said for most of what fills the journals that have replaced JBE on the FT50.</p><h2>Rankings, signals, and who gets to design them</h2><p>Business schools are crowded places, and rankings of journals do real work as signals in the kind of market where everyone is trying to be everywhere at once. A scholar moving across institutions, a hiring committee comparing candidates trained in different subfields, a dean trying to allocate research time across a faculty of seventy people, all of them rely on lists like the FT50 and the ABS list as compressed information about what counts as good work, and the alternative to such lists is not a serene world of careful individual reading but a world in which the same compression happens through gossip and proxy. The objection to the present arrangement is not that rankings exist, since objecting to that would be silly given how much information has to be aggregated and how few hours there are in a hiring committee meeting. The objection is that rankings whose effective scope excludes whole disciplines do a disservice to the broader business and management field, because they make it institutionally irrational for schools to hire in those disciplines, and they slowly drain the kind of intellectual diversity that makes a business school more than a credentialling factory.</p><p>The good news, such as it is, is that this is a solvable problem, and the ABS has already shown it can be solved when there is institutional will to solve it. The 2024 AJG promoted journals such as <em>American Political Science Review</em> and <em>American Journal of Political Science</em> to 4*, and effectively imported the top of political science&#8217;s own ranking convention into the business-school ecosystem. Political theorists and political economists working on questions adjacent to management can now publish in their discipline&#8217;s flagship journals and have those publications counted by their business-school administrators. The procedure was straightforward, consisting in consulting the relevant scholarly community on what their top journals were and including those.</p><p>There is no good reason the same thing could not be done for philosophy. <em>Ethics</em>, <em>Philosophy &amp; Public Affairs</em>, <em>Philosophical Review</em>, <em>No&#251;s</em>, <em>Mind</em>, <em>Journal of Philosophy</em>, alongside the leading specialist outlets in moral philosophy, decision theory, and philosophy of science, all have well-understood standing in their field, and a usable list could be drawn up in an afternoon by any working analytic philosopher with a passing familiarity with the discipline&#8217;s conventions. What is stopping it, I suspect, is that this kind of inclusion has to be done by someone, and the question of who that someone is matters enormously. The ABS process relies on subject experts and methodologists whose remit is business and management, and who, to be generous, often have a limited understanding of the disciplines adjacent to their own; the categories of consultation, the closed composition of the panel, and the near-absence of working analytic philosophers in the relevant discussions all push toward a polite default in which adjacent disciplines get treated either as friendly neighbours (as in the politics case) or as suspect outsiders (as in the philosophy case), depending on factors that have very little to do with the actual quality of the work being produced.</p><p>The FT50 is, on inspection, considerably worse on this front. Its formal methodology, by the FT&#8217;s own description, is that the paper surveys the deans of business schools that pay to participate in its MBA rankings and asks them what journals they think matter, which is a method that reliably produces lists reflecting the existing prestige hierarchy of business schools rather than the actual quality of academic work, and which builds in a structural conflict of interest of the kind that would not be tolerated in any other context. The schools that are evaluated using the FT50 are the schools that get to vote on what goes into the FT50. There is no transparent academic process, no published criteria for inclusion or exclusion, and no methodology document of the kind that the ABS at least gestures at. The FT, to be clear, is in the business of producing rankings that sell newspapers and that lubricate its commercial relationships with business schools, and there is no particular reason it should be in the business of curating academic excellence.</p><p>If lists of this kind are to track what they purport to track, they have to be designed by academics from the relevant disciplines, working in consultation with those disciplines&#8217; learned societies, and willing to be transparent about their methodology. The current arrangement, in which a small panel of management scholars makes implicit decisions about the relative standing of work in fields they do not work in, and in which a financial newspaper aggregates the votes of self-interested deans, is incentive-incompatible in a fairly obvious way. It is also, again, solvable, in that the institutional capacity to solve it has already been demonstrated for politics. There is no methodological reason it cannot be demonstrated for philosophy.</p><h2>The prior question</h2><p>There is, of course, a prior question that has to be asked here, and it is one I do not want to answer too quickly either on my own behalf or on the field&#8217;s. There is a long-running and not entirely unfair perception, especially in the UK, that business schools are where academics from sociology, psychology, philosophy, and political science end up when they were not quite good enough for jobs in their home departments. The perception is partly true, in that the academic job market is thin in those disciplines and the business-school market is comparatively thicker, and it is partly unfair, in that there are excellent scholars who ended up in business schools for reasons that have nothing to do with their relative ability. What is worth asking is why the perception persists, and the honest answer has more to do with incentives than with talent.</p><p>If you want a business school to hire the best philosopher who works on management-adjacent topics, you have to give the dean a reason to do so that lines up with the metrics the dean is judged on. If publishing in <em>Ethics</em> doesn&#8217;t count for anything in the rankings the dean&#8217;s school is evaluated against, the dean will hire whoever can publish in JBE instead, and over time the field will fill with people whose comparative advantage is publishing in journals that count rather than people whose comparative advantage is doing important philosophical work. None of this is a knock on the people who got hired under the existing incentive structure, who include many of the best scholars I know. The point being made is descriptive, in that a system of misaligned incentives reliably produces, over thirty years, a workforce shaped by what was rewarded.</p><p>So before we ask how to fix the rankings, we should ask whether we want philosophers in business schools at all. I am not going to assume in this post that the answer is yes. It is genuinely up for grabs whether business schools should be in the business of housing philosophical ethicists, philosophers of science working on management methodology, philosophers of mind working on AI and organisations, or political philosophers working on the moral status of corporations. There are good arguments on either side, and I expect to come back to the question in subsequent posts. If the answer turns out to be yes, if we conclude that these conversations are valuable and that students and faculty benefit from having them in the room, then the institutional incentives have to be designed accordingly. We have to build lists that count what philosophers do, evaluated by people who can tell when philosophers do it well, and we cannot do that by delegating the ranking of ethics journals to organisational scholars or strategy theorists who do not work in moral philosophy. Doing so is how we ended up here.</p><h2>What&#8217;s left, and the reading problem</h2><p>Here, perhaps, is the part I&#8217;ve been avoiding writing because it feels uncollegial, and because the people who would read it are also the people who decide whether my next paper gets accepted somewhere. As an analytic philosopher who also does formal modelling, who reads across journal lists and tries to understand what the median paper looks like in each, I find what&#8217;s left in the rankings genuinely baffling. I will not name specific names, since that would be a bad use of a Substack post, but some of the journals on, say, the UT Dallas top-research list publish hundreds of papers a year, the bulk of which are silence. They are read by no one, cited by people who haven&#8217;t read them, summarised, when they are summarised at all, by an LLM, and used as currency for promotion cases by faculty who have never opened them either. The citation networks that hold the field together are increasingly fictional, sustained by a cooperative pretence that someone, somewhere, is doing the reading. Most of us know they&#8217;re not.</p><p>This isn&#8217;t a problem with management research as such. There are excellent papers in the top management journals, and excellent scholars publishing them, and the empirical work that gets done well in the field genuinely advances understanding of how organisations operate. The problem is with a system that has scaled production beyond the reading capacity of any human community. Too many papers, too much filler, too little actual engagement. Removing JBE from this system, on the grounds that it didn&#8217;t quite measure up, is frankly ridiculous.</p><h2>To be continued</h2><p>This is enough for one sitting, but there is more to say. The question of whether philosophers belong in business schools at all is one I want to return to with proper care, since it deserves more than the sketch I have given it here. The related questions of what individual scholars should actually do under current conditions, of how reform of the rankings might plausibly happen, and of what the international politics of these lists looks like from inside the institutions that produce them, all deserve their own treatments. None of those questions admit of quick answers, and the polemic energy that gets you through a first thousand words of grievance writing tends to dissipate around the point where you have to be constructive.</p><p>For now, this is enough. I&#8217;m going to pour some wine, in honour of Climat&#8217;s wine list, and try to figure out how to publish in SMJ :)</p><p><em>Part one of probably two.</em></p>]]></content:encoded></item></channel></rss>