<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Scaling Biotech]]></title><description><![CDATA[Exploring how processes within drug discovery are changing to take advantage of new models and technology.]]></description><link>https://scalingbiotech.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png</url><title>Scaling Biotech</title><link>https://scalingbiotech.substack.com</link></image><generator>Substack</generator><lastBuildDate>Sat, 05 Sep 2026 08:00:29 GMT</lastBuildDate><atom:link href="/__u/scalingbiotech.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Jesse Johnson]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[scalingbiotech@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[scalingbiotech@substack.com]]></itunes:email><itunes:name><![CDATA[Jesse Johnson]]></itunes:name></itunes:owner><itunes:author><![CDATA[Jesse Johnson]]></itunes:author><googleplay:owner><![CDATA[scalingbiotech@substack.com]]></googleplay:owner><googleplay:email><![CDATA[scalingbiotech@substack.com]]></googleplay:email><googleplay:author><![CDATA[Jesse Johnson]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Does AI in biotech have to come from internal teams?]]></title><description><![CDATA[To start answering the question I raised last week about what factors are slowing down tech adoption in biotech, I started looking at the history of how older technologies have been adopted by the industry.]]></description><link>https://scalingbiotech.substack.com/p/does-ai-in-biotech-have-to-come-from</link><guid isPermaLink="false">https://scalingbiotech.substack.com/p/does-ai-in-biotech-have-to-come-from</guid><dc:creator><![CDATA[Jesse Johnson]]></dc:creator><pubDate>Wed, 02 Sep 2026 14:45:46 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>To start answering </span><a href="/__u/scalingbiotech.substack.com/p/the-problems-with-tech-in-bioteh"><span>the question I raised last week</span></a><span> about what factors are slowing down tech adoption in biotech, I started looking at the history of how older technologies have been adopted by the industry. One clear pattern is that many of the tools that are standard today were developed by the same scientists who were planning to use them. This development model makes sense, but it&#8217;s also limiting. Is it the only way biotech innovation is possible?</span></p><p>I won&#8217;t be able to answer that question this week, and in fact it&#8217;s probably not exactly the right question. Instead, <span>I&#8217;m going to explore what this looked like for QSAR, one of the earliest digital technologies in drug discovery. Then I&#8217;ll discuss why this approach is limiting and how the advent of vibe coding might begin to change that calculus.</span></p><p><span>Before we start, I want to clarify what I mean by inside vs outside the industry. I&#8217;m not going to argue about who counts as a scientist and who doesn&#8217;t. Instead, I&#8217;m talking about about the role they&#8217;re in and whether their primary responsibility is to help bring drugs to market or to develop new software/tools/models. The same person could be inside or outside, depending on their job.</span></p><h4>QSAR</h4><p><span>QSAR is the idea of embedding molecules in a vector space by defining features from their molecular structures, then training ML models on the vector space. </span>The earliest QSAR algorithms used one-dimensional regression models, but the same approach is used today with much more complex models and embeddings.</p><p><span>The idea was developed by Corwin Hansch, a professor at Pomona College, in 1962. By the end of the decade, Eli Lilly had a team of chemists, mostly hired from academic computational chemistry labs, writing internal algorithms on the first transistor-based mainframes at a time when punch cards were just being phased out.</span></p><p><span>Thanks to reorgs (which I forgot to put on my list last week), Lilly pulled back on this in the early 70s, allowing Merck and a small company called Smith Klein &amp; French (later acquired by GSK) to take the lead. By the 1980s, QSAR was a common, if not standard, method across pharma.</span></p><h4>From inside or outside?</h4><p><span>The chemists who developed these tools were trying to make the existing process of small molecule optimization more efficient: Given an initial hit, they could design 100 small modifications that might make it better or worse. Previously, they would need to synthesize and test them all in vitro. With QSAR, they could digitally predict which ones were most likely to work, and only synthesize those.</span></p><p><span>So QSAR was built around an immediate need, improving an existing step in the drug discovery process, by the people who were responsible for that step, writing the code themselves. The benefit of this was that there was no question about whether or how the scientists would use these new tools. It was just a question of whether the models would work, which they did (at least eventually).</span></p><p><span>However, there are also two major limitations to this development model, compared to building new tools from the outside.</span></p><h4>The Pi-shaped person</h4><p><span>The first bottleneck is finding scientists who deeply understand both the science and machine learning/software engineering/etc. These people exist, but not so many of them, and they can&#8217;t go as deep on either side as someone who specializes in one or the other.</span></p><p><span>Vibe coding could provide at least a partial answer to this. For now at least, we still need humans who understand software engineering to build the production-grade systems, making sure the coding agents follow secure practices and don&#8217;t do anything stupid. But vibe coding allows scientists to build prototypes of their ideas that engineering teams can then incorporate into production-grade systems. So it could allow more scientists to help develop their own novel tools without needing deep engineering skills.</span></p><h4>Optimize or Disrupt?</h4><p><span>The more existential limitation is that if new digital tools can only improve and optimize existing processes then we&#8217;ll miss opportunities to make larger changes to the drug discovery process that </span><a href="/__u/scalingbiotech.substack.com/p/2b-seems-like-a-lot-to-invest-in"><span>new technologies theoretically enable</span></a><span>.</span></p><p><span>Someone whose job is to envision and build the next generation of digital tools is more likely to think in terms of those broader changes than someone who is thinking about how they can hand off a candidate to the next step faster. On the other hand, the scientist pushing drug candidates along the pipeline will have a more realistic perspective of what&#8217;s possible. So maybe we&#8217;re back to needing those hybrid scientist/engineers who can understand the full potential of the new technology.</span></p><p><span>I don&#8217;t know what the answer is, but I feel like this is at least getting closer to the right question. Stay tuned as I keep pulling on this thread, and other nearby ones.</span></p><div><hr></div><p><span>Thanks for reading! My company, Merelogic, helps biopharma teams apply the rigor they demand inside the lab to how they integrate AI and other digital tools into the drug discovery process. You can learn more at </span><a href="https://merelogic.net/">merelogic.net</a><span>.</span></p>]]></content:encoded></item><item><title><![CDATA[The problem(s) with tech in bioteh]]></title><description><![CDATA[As I&#8217;ve been shifting my consulting work more towards process change and tech adoption, it&#8217;s got me thinking about all the different reasons (or maybe excuses) that could explain why early stage drug discovery seems to have lagged behind other sectors in adopting digital tools, AI, etc.]]></description><link>https://scalingbiotech.substack.com/p/the-problems-with-tech-in-bioteh</link><guid isPermaLink="false">https://scalingbiotech.substack.com/p/the-problems-with-tech-in-bioteh</guid><dc:creator><![CDATA[Jesse Johnson]]></dc:creator><pubDate>Wed, 26 Aug 2026 14:45:45 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>As I&#8217;ve been shifting my consulting work more towards process change and tech adoption, it&#8217;s got me thinking about all the different reasons (or maybe excuses) that could explain why early stage drug discovery seems to have lagged behind other sectors in adopting digital tools, AI, etc. and how we can address these bottlenecks. To help organize my thoughts, I put together a list of potential hypotheses and this week, I want to share it. In future posts I&#8217;ll dive into individual hypotheses.</p><p><em>** But first, a quick update on the AI-for-biotech course I&#8217;m working on with Scriptome.ai: </em>The course has been approved by the Massachusetts Workforce Development Training Fund, which means your can get your tuition reimbursed if you work for a Massachusetts company with fewer than 100 employees. You can find details on <a href="https://courses.scriptome.ai/">the course page</a>.<em> **</em></p><p>It ended up being a fairly long list, so I split it into five categories to make it easier to see how they fit together. They all probably contribute at least a little to the situation, but as I dig in more, I&#8217;m hoping to get a better sense of which ones are the most significant.</p><p>Also, I reserve the right to update this list as I dig deeper. In fact, if you notice anything I missed or have any other suggestions, feel free to leave a comment, reply to this email, send me a email separate email (jesse@merelogic.net) or reach out in whatever other way you can find.</p><p>Here&#8217;s the list:</p><h4>Category 1: Our perception is off</h4><ol><li><p><strong>The grass is always greener</strong>: Tech adoption is hard in every industry. Maybe it isn&#8217;t any harder in biology, and our expectations are just too high&#8230;</p></li><li><p><strong>Would we even know?</strong>: Clinical timelines are so long and sample sizes so small, it&#8217;s possible some teams have already figured this out but we never noticed or <a href="/__u/scalingbiotech.substack.com/p/we-wont-know-when-ai-solves-drug">haven&#8217;t noticed yet</a>.</p></li></ol><h4>Category 2: It&#8217;s the science</h4><ol start="3"><li><p><strong>More exceptions than rules</strong>: Is biology just fundamentally too complex to accurately model, even with billions of parameters?</p></li><li><p><strong>Constrained by data</strong>: Maybe the cost and nature of how we&#8217;re able to collect biological data prevents models from meaningfully generalizing.</p></li><li><p><strong>Out of sample</strong>: Biopharma innovation happens at the boundaries of knowledge but ML models are mostly good at generalizing within the bounds of their training data. Are we expecting them to extrapolate beyond their limits?</p></li><li><p><strong>The reproducibility crisis</strong>: It&#8217;s well known that some (maybe a lot of) biology papers have unreliable results. Do public the datasets that are regularly used to train models have too much noise masquerading as data?</p></li></ol><h4>Category 3: It&#8217;s the culture</h4><ol start="7"><li><p><strong>Gatekeeping</strong>: Biologists over-estimate how specialized and complex biology is, so they&#8217;re reluctant to give models a chance. I mean, not *<em>all*</em> biologists, but at least a few&#8230; Is it too many?</p></li><li><p><strong>The wrong data</strong>: Until recently at least, most data was collected from experiments focused on specific, narrow questions that made the data hard to incorporate into more general training datasets. Have we just not created enough training datasets?</p></li><li><p><strong>Too many silos</strong>: Maybe the political, compartmental nature of pharma companies, and even many biotech startups, has prevented the kind of collaboration required for more fundamental changes&#8230;</p></li><li><p><strong>Alien worlds</strong>: The data scientists and AI/ML engineers who build the models just think fundamentally differently than the biologists they need to work with. Are there just not enough experts who understand both worlds?</p></li><li><p><strong>Red tape</strong>: Do the procurement processes and the long buying cycles of large pharmas kill all the innovative companies before they can make a difference?</p></li><li><p><strong>Regulatory burden</strong>: Maybe regulatory demands make real innovation too difficult, if not impossible&#8230;</p></li></ol><h4>Category 4: It&#8217;s the economics</h4><ol start="13"><li><p><strong>Boom and bust</strong>: VCs want companies to become profitable on a much shorter timeline than the 10-15 year clinical cycle allows. Are they giving up too soon, or forcing companies to try and meet unrealistic timelines?</p></li><li><p><strong>The pipeline trap</strong>: Many innovative platform companies have been pushed into developing their own pipelines, which introduces unrelated risks like picking targets that become commercially unviable when the market shifts. Has this killed too many companies with otherwise solid technology?</p></li><li><p><strong>Good enough</strong>: Existing tools and approaches are still good enough to get drugs to market, and they&#8217;re safer than betting on something new. Maybe pharma leaders are too often picking good enough over innovative-but-risky new options.</p></li><li><p><strong>Assets over innovation</strong>: Biotech investors value individual assets over innovative platforms. By reducing investments in innovation from the science side, are they forcing digital/platform companies to work with tech investors who can&#8217;t advise them from an insider&#8217;s perspective?</p></li></ol><h4>Category 5: It&#8217;s the technology</h4><ol start="17"><li><p><strong>Legacy systems</strong>: Most big pharma companies rely on a foundation of legacy infrastructure, full of technical debt. Is the technical friction this creates for new tools just too much?</p></li><li><p><strong>Black box</strong>: Many of the new models are black-box in nature while most biologists consider the causal interpretation as important or more important than the actual prediction. Maybe <a href="/__u/scalingbiotech.substack.com/p/is-model-interpretability-a-code">causality and interpretability are the issue</a>.</p></li><li><p><strong>Alignment</strong>: Many of the predictions that these models make aren&#8217;t closely aligned with the existing stages of the drug discovery process. Are we going to need to <a href="/__u/scalingbiotech.substack.com/p/2b-seems-like-a-lot-to-invest-in">fundamentally change the drug discovery process</a> to start using these tools?</p></li><li><p><strong>Interfaces</strong>: Connecting these models to the data and to each other often requires significant technical know-how and often isn&#8217;t aligned with how scientists think. Do we need <a href="/__u/scalingbiotech.substack.com/p/llms-should-be-orchestrators-not">better interfaces</a>?</p></li></ol><p>That&#8217;s my preliminary list. Did I miss anything?</p><div><hr></div><p><span>Thanks for reading! My company, Merelogic, helps biopharma teams apply the rigor they demand inside the lab to how they handle data outside the lab. You can learn more at </span><a href="https://merelogic.net/">merelogic.net</a><span>.</span></p>]]></content:encoded></item><item><title><![CDATA[GSK goes deeper on Virtual Cell Models]]></title><description><![CDATA[Just six months after announcing a $50M collaboration giving them access to NOETIK&#8217;s virtual cell models, GSK recently announced that they&#8217;re expanding their partnership with Relation Therapeutics with a $110M deal to generate perturbational data for virtual cell models.]]></description><link>https://scalingbiotech.substack.com/p/gsk-goes-deeper-on-virtual-cell-models</link><guid isPermaLink="false">https://scalingbiotech.substack.com/p/gsk-goes-deeper-on-virtual-cell-models</guid><dc:creator><![CDATA[Jesse Johnson]]></dc:creator><pubDate>Wed, 19 Aug 2026 14:46:42 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Just six months after announcing a $50M collaboration giving them access to NOETIK&#8217;s virtual cell models, GSK recently announced that they&#8217;re expanding their partnership with Relation Therapeutics with a $110M deal to generate perturbational data for virtual cell models. This deal is interesting both for how it&#8217;s structured and what it indicates about GSK&#8217;s strategy, so that&#8217;s what I want to write about this week.</p><h4>The Deal</h4><p><a href="https://www.relationrx.com/">Relation&#8217;s website</a> links to a number of articles about the partnership, though most are behind paywalls. As usual for this kind of announcement, details are limited, but we can work out the general shape of the deal. </p><p>Relation will use the $110M to generate data measuring gene expression under different cell perturbations. They&#8217;ll use the data to help train an internal model called <a href="https://www.relationrx.com/news/relation-unveils-morgan-the-next-generation-model-of-cellular-biology">MORGAN</a> that they announced at the same time as this partnership. It&#8217;s unclear if GSK will have direct access to this model or if they&#8217;ll only get indirect access through their ongoing partnered drug programs with Relation.</p><p>Regardless, GSK will also get the data itself to train their own models. So this makes the deal more of a data co-development partnership: GSK covers the cost, Relation generates the data, both companies get to use the data. Relation doesn&#8217;t offer the data or the model as an externally available product, so the data is effectively proprietary and exclusive. (At least for now - it&#8217;s unclear what happens if Relation ever wants to use the MORGAN model for a partnership with a GSK competitor&#8230;)</p><h4>GSK&#8217;s Strategy</h4><p>This is an interesting addition to GSK&#8217;s ongoing efforts related to virtual cell/perturbation models. In addition to the NOETIK deal from January, GSK has also stated publicly that they&#8217;re <a href="https://www.gsk.ai/our-work/large-perturbation-models-for-in-silico-biological-discovery/">building their own perturbation models internally</a>, suggesting that it&#8217;s trying to build internal capabilities alongside these partnerships.</p><p>As far as I can tell, this makes GSK the most active large-pharma player when it comes to virtual cell/perturbation models for early discovery (as opposed to clinical virtual patient models). I don&#8217;t know if the rest are waiting to see how effective these models prove in the long run (which may be years off) or if they just haven&#8217;t been as public about their internal development efforts. But either way, GSK seems to be making a big bet.</p><p>As I&#8217;ve stated previously, I think there&#8217;s a lot of potential for virtual cell models to fundamentally change drug discovery and address the <a href="/__u/scalingbiotech.substack.com/p/the-clinical-information-bottleneck?utm_source=publication-search">clinical information bottleneck</a>, though it&#8217;s far from guaranteed. So I see this as a high-risk, high-reward bet. I sincerely hope it pays off.</p><div><hr></div><p><span>Thanks for reading! My company, Merelogic, helps biopharma teams implement tools and practices to build a solid data foundation for whatever comes next. You can learn more at </span><a href="https://merelogic.net/">merelogic.net</a><span>.</span></p>]]></content:encoded></item><item><title><![CDATA[Biotech is now bubble-adjacent]]></title><description><![CDATA[Dimension Capital just announced its third fund, raising a total of $800M for them to invest in ai-driven biotech startups.]]></description><link>https://scalingbiotech.substack.com/p/biotech-is-now-bubble-adjacent</link><guid isPermaLink="false">https://scalingbiotech.substack.com/p/biotech-is-now-bubble-adjacent</guid><dc:creator><![CDATA[Jesse Johnson]]></dc:creator><pubDate>Wed, 12 Aug 2026 14:45:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Dimension Capital just announced its third fund, raising a total of $800M for them to invest in ai-driven biotech startups. This is more than double the size of its last fund, but it&#8217;s also less than half of the </span><a href="/__u/scalingbiotech.substack.com/p/2b-seems-like-a-lot-to-invest-in"><span>$2B that was invested in Isomorphic Labs</span></a><span> alone, from a collection of funds that typically don&#8217;t invest in biotech at all. I think it&#8217;s worth thinking of these as part of two separate phenomena: the early stages of a biotech recovery and the later stages of a much larger AI investing bubble, or at least the lead in to a bubble. This week, I want to explain why I think this, then speculate about what it means for the near-term future of biotech.</span></p><h4>Two markets</h4><p><span>The first thing we need to recognize is that Isomorphic didn&#8217;t deprive other biotech startups of the $2B they recently raised, since that money wouldn&#8217;t have gone to biotech startups otherwise. If not for Isomorphic, those investors would&#8217;ve given it to AI-enabled toasters or digital pet companions or some other questionable AI-driven hype. Meanwhile, the investors who typically invest in biotechs might&#8217;ve been willing to invest in Isomorphic, but probably at a much lower valuation.</span></p><p><span>I&#8217;m making a simplifying assumption that we can cleanly separate investors into biotech investors and pure tech investors. This isn&#8217;t true, but if we run with that for a second and add in the assumption that tech investors are open to higher valuations for AI companies than biotech investors are for the same company when viewed as a biotech, then this means that any company that can pass as an AI company can get a valuation that essentially locks out biotech investors.</span></p><p><span>So the two distinct categories of investors create two distinct categories of startups and two separate but adjacent investing markets that share a long border. </span></p><p><span>Now, what happens when one of them is a giant bubble? Let&#8217;s start by reviewing the mechanics of bubbles.</span></p><h4>The mechanics of bubbles</h4><p><span>A bubble forms when investors think that a resource is going to be worth a lot more than what it currently costs to build it, so they pump money into building the thing. It could be houses or internet companies or data centers or foundation models. As more investors pile in and the numbers get bigger, it reinforces the belief that the resource is going to be worth even more.</span></p><p><span>The problem is that once the amount invested becomes more than the resource is actually worth, someone is going to lose money. The first investors who notice and react can usually get most of their money back. After that, the longer investors wait, the more they lose. So once the sentiment starts to shift, it kicks off a race to get your money out. That&#8217;s why bubbles pop instead of slowly deflating.</span></p><p><span>That sucks for investors, and it can be even worse for companies who make long-term plans based on the assumption that they&#8217;ll be able to raise more money, then find out they can&#8217;t. But there can be a silver lining for everyone else&#8230;</span></p><h4><span>What bubbles leave behind</span></h4><p><span>The thing people often forget is that bubbles often leave behind the resource that the investors paid all that money to build. Sure, it&#8217;s worth less than the investors thought it would be. Sure, some of the companies who built it are long gone, along with some of the investors. But the thing they built is still there.</span></p><p><span>After the housing crash of 2008, there were a bunch of houses that might not have been built otherwise. Some were McMansions that no one wanted to live in, but they were there. The dot com crash of 2001 left the infrastructure that made today&#8217;s internet possible.</span></p><p><span>Whether or not AI is currently in a bubble depends on what the actual value turns out to be. But even if it isn&#8217;t a bubble yet, it certainly looks like it will make it there eventually.</span></p><p><span>The two questions for biotech are: 1) will the AI bubble popping impact the sentiment around biotech, causing biotech investing to crash as well? And 2) what will the AI bubble leave behind that biotech can benefit from?</span></p><p><span>I won&#8217;t speculate about the first one, but I will about the second one.</span></p><h4><span>Models and data</span></h4><p><span>The data centers that an AI bubble will leave behind are a mixed bag, with tremendous environmental implications, but also the possibility of much lower compute costs once supply surpasses demand. That applies across all industries.</span></p><p><span>The one that&#8217;s more interesting for biotech specifically is the bio foundation models that will be built and trained, along with the data that trained them and, perhaps more importantly, the knowledge of how to train and use them.</span></p><p><span>If Isomorphic becomes a trillion dollar company, its investors will look like geniuses. If not, the company may cease to exist. But we&#8217;ll still have the open source OpenFold model based on everything Isomorphic learned building its internal model. And we&#8217;ll still have the open source boltz model and a bunch of other models that learned from this as well. And sure, maybe even that won&#8217;t turn out to be worth $2B. But it&#8217;s better than </span>AI-enabled toasters and digital pet companions. And thanks to those tech investors, it won&#8217;t have cost the biotech community a single penny.</p><div><hr></div><p><span>Thanks for reading! My company, Merelogic, helps biopharma teams implement tools and practices to build a solid data foundation for whatever comes next. You can learn more at </span><a href="https://merelogic.net/">merelogic.net</a><span>.</span></p>]]></content:encoded></item><item><title><![CDATA[Does biopharma data need a Maslow Hierarchy?]]></title><description><![CDATA[The risks that I see most often in Data Risk Reviews seem to fall into three main categories that form a kind of Maslow-style hierarchy.]]></description><link>https://scalingbiotech.substack.com/p/does-biopharma-data-need-a-maslow</link><guid isPermaLink="false">https://scalingbiotech.substack.com/p/does-biopharma-data-need-a-maslow</guid><dc:creator><![CDATA[Jesse Johnson]]></dc:creator><pubDate>Wed, 05 Aug 2026 14:45:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The risks that I see most often in <a href="https://merelogic.net/data_risk_review">Data Risk Reviews</a> seem to fall into three main categories that form a kind of Maslow-style hierarchy. I think understanding this hierarchy could be useful for anyone who wants to do a self-review of their data risks, so that&#8217;s what this week&#8217;s post is about. In future posts, I&#8217;ll go into more detail about each of the categories.</p><p>Understanding these categories is important because early discovery research isn&#8217;t covered by GXP. Being free of GXP gives teams the flexibility to try new things and move fast, unencumbered by strict rules about to handle their data, but ignoring the principles entirely means working without a safety net. It means your team doesn&#8217;t get the security and reproducibility that GXP is meant to ensure.</p><p>So in practice, it&#8217;s up to you to decide what level of risk you&#8217;re willing to accept in return for speed and flexibility. Framing the risks within a hierarchy can help you add the right guardrails to avoid the risks that could bring your programs to a screeching halt, not just adding hurdles because it&#8217;s the &#8220;right&#8221; way to do things.</p><h4>Acute to Chronic</h4><p>Here&#8217;s how I&#8217;m currently defining the categories, though I may revisit this in the future (I&#8217;ll probably at least rename them):</p><p><strong>1. Acute, Avoidable Risks</strong>: These are risks of events with immediate negative impacts that can be avoided if you take the right measures in advance. They&#8217;re mostly types of mistakes like errors in manual data processing, picking the wrong version of a file or sharing sensitive information with partners such as CROs. They can typically be addressed by simplifying or automating manual processes, as well as adding audit trails so you can catch the mistakes earlier. If these mistakes go uncaught, your team could spend months chasing ghost results, or worse.</p><p><strong>2. Readiness Risks</strong>: These are risks of being caught unprepared for inevitable events, whether that means being unable to minimize the impact of negative events or being unable to take advantage of positive ones. These include risks that make it harder to quickly analyze new data or verify results that seem significant, as well as being able to recover from potentially cataclysmic events like losing the one person who knows where all the data is. The cost of being unprepared may be lower than the cost of acute, avoidable risks, but they come at a time when you can least afford them.</p><p><strong>3. Overhead Risks</strong>: These are issues that slowly build up over time and scale, such as small delays at each stage in a process leading to significantly longer timelines across a program, or a team that doesn&#8217;t trust its data demanding more and more validation assays until they can barely make decisions. These risks are often subtle, making it harder to estimate or identify their cost and impact. They can often be addressed with relatively small changes but those changes need to be applied consistently, which makes them much harder to address.</p><h4>The Hierarchy</h4><p>What&#8217;s nice about these categories is that addressing each one creates a foundation for dealing with the next. Addressing the acute, avoidable risks frees up time and cognitive load to start thinking about readiness risks. Then once you&#8217;re confident that you&#8217;re team will be ready for whatever comes its way, you can start thinking about addressing the more subtle overhead risks to optimize how the team works.</p><p>This is not to say that you have to address each category before you can move on to the next. Just like you can address multiple layers in Maslow&#8217;s hierarchy at the same time, you may decide to prioritize risks from these categories in any order. The point is that understanding this hierarchy can help you decide which risks you need to address first, which ones can wait, and which ones are a reasonable price for flexibility.</p><div><hr></div><p><span>Thanks for reading! My company, Merelogic, helps biopharma teams implement tools and practices to build a solid data foundation for whatever comes next. You can learn more at </span><a href="https://merelogic.net/">merelogic.net</a><span> and request a free </span><a href="https://merelogic.net/data_risk_review">Data Risk Review</a><span>.</span></p>]]></content:encoded></item><item><title><![CDATA[Open source protein models: OpenFold vs Boltz]]></title><description><![CDATA[Today, there are two open source, transformer-based protein folding models that seem to get the most attention.]]></description><link>https://scalingbiotech.substack.com/p/open-source-protein-models-openfold</link><guid isPermaLink="false">https://scalingbiotech.substack.com/p/open-source-protein-models-openfold</guid><dc:creator><![CDATA[Jesse Johnson]]></dc:creator><pubDate>Wed, 29 Jul 2026 14:45:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Today, there are two open source, transformer-based protein folding models that seem to get the most attention. They both use AlphaFold as a starting point, but with some important differences after that. This week I want to explore where these models came from and what makes them different.</p><p><em>**But first: I&#8217;ve been offering free </em><a href="https://merelogic.net/data_risk_review">Data Risk Reviews</a><em> to early discovery biopharma teams that want to build a strong data foundation for whatever comes next. If having trustworthy data matters to you, click the link to learn more.**</em></p><h4>Teams and Timelines</h4><p>The story starts when Google&#8217;s DeepMind released AlphaFold2 in November 2020. They released the model&#8217;s weights and the code for running inference as open source on GitHub, but kept the code for training the model private. This meant that while anyone could use the open source code to generate predictions, they couldn&#8217;t train their own version of the model or fine tine the weights using proprietary data.</p><p>To address this limitation, the OpenFold foundation was founded in early 2022 as a non-profit backed by both academic and industry/pharma partners. Their goal was to build a fully open source implementation of AlphaFold2, called OpenFold2. Since then, they have built a consortium that includes large pharma companies such as Bristol Myers Squibb, Novo Nordisk, Bayer, Roche, and UCB.</p><p>After DeepMind released AlphaFold3 in May 2024, the OpenFold foundation began work on OpenFold3, again with the goal of replicating the model with a fully open source implementation. (I&#8217;ll describe the technical differences between AlphaFold 2 and 3 in the next section.)</p><p>Shortly after that, in November 2024, a group of researchers at MIT released Boltz-1, a fully open source model that took the main ideas from AlphaFold3, but made changes designed to improve its performance. Then in June 2025, they released Boltz-2, a model that added capabilities that weren&#8217;t available in AlphaFold3. (More on that below too.)</p><p>In January 2026, the developers of the Boltz model founded a public benefit corporation called Boltz, headquartered in London. This allowed them to raise money from VCs while continuing to maintain the fully open source Boltz models.</p><h4>Introducing diffusion models</h4><p>To understand the technical differences between OpenFold and Boltz, we need to start with the differences between AlphaFold 2 and 3. Both models start with a <a href="/__u/scalingbiotech.substack.com/p/how-to-build-a-bio-foundation-model">transformer model where tokens represent amino acids</a> to infer spatial and physical relationships between all pairs of amino acids in a protein. Where they differ is how they translate this into a protein structure: </p><p>AlphaFold2 uses a deterministic algorithm that translates the pairwise relationships into the final positions of the amino acids. AlphaFold3 uses a diffusion model to predict the locations of individual atoms from the token embeddings. This is similar to how image generation models predict the colors of individual pixels in an image based on the token embeddings of a sentence describing the desired picture. The sequence of amino acids is the sentence. The individual atoms are the pixels.</p><p>Roughly, the diffusion model creates embedding vectors for the individual atoms in the same latent space as the amino acid tokens are embedded by the transformer, then applies an attention algorithm to update the atom embeddings based on the embeddings of both the amino acids and the other atoms. This is similar to how a transformer re-embeds tokens, as described in my <a href="/__u/scalingbiotech.substack.com/p/how-to-build-a-bio-foundation-model">post about transformers</a>. The model uses the final atom embeddings to predict their positions.</p><p>AlphaFold3 also made some changes to the transformer algorithm that put less emphasis on reference structures (MSAs) for inferring the relationships between amino acids, but that&#8217;s kind of in the weeds. The important change is the diffusion model.</p><h4>Boltz and OpenFold</h4><p>The goal of OpenFold3 is to replicate AlphaFold3 as closely as possible, so there are no major atchitectural differences between OpenFold3 and AlphaFold3. Boltz, however, is intended to improve on AlphaFold, so they made some specific changes to the architecture.</p><p>Most of these changes are too technical to go into here, but an important one is how the model interprets the outputs of the diffusion algorithm: In AlphaFold3 and OpenFold3, the predicted positions of the atoms are compared to the coordinates of the reference structure for a fixed position and rotation of the overall protein. However, that rotation and position don&#8217;t actually matter to the answer; it&#8217;s only the relative positions of the atoms that we care about. So before Boltz-1 compares the predicted positions of the atoms to the reference, it first applies a translation and rotation to the whole structure that gets the atoms as close as possible to the reference positions. That way, if the model predicts the right structure but with the wrong rotation, it still counts it as right. OpenFold3 would count it as wrong.</p><p>Because of this change and the other architectural differences, Boltz has performed slightly better than OpenFold3 on certain benchmarks, though the results aren&#8217;t completely conclusive.</p><p>The bigger difference came with Boltz-2, which introduced the ability to predict binding affinity between small molecules and a target protein. OpenFold3 can only. predict co-folding between the two, i.e. it can predict a stable position of a small molecule and a protein in close proximity. But it doesn&#8217;t predict the likelihood of the molecule getting into that position, or how strongly it will bind once there.</p><p>This means that Boltz-2 can be used directly for virtual binder screening. OpenFold3 can be part of a virtual binder screen, but only if you couple it with a model that predicts binding affinity of the predicted positions. Though as I argued recently, <a href="/__u/scalingbiotech.substack.com/p/2b-seems-like-a-lot-to-invest-in">this isn&#8217;t that valuable on it&#8217;s own</a>, and both models are <a href="/__u/scalingbiotech.substack.com/p/leash-bio-shows-that-data-beats-models">probably cheating a little</a>.</p><h4>So which one is better?</h4><p>Between OpenFold3 and Boltz-2, there isn&#8217;t a clear winner. The difference in accuracy will probably depend on the type of protein(s) you&#8217;re looking at, and in many cases may not be big enough to actually matter. The bigger differences are going to be related to what&#8217;s involved in actually getting them up and running.</p><p>OpenFold is focused on supporting large pharmas like the ones who fund their foundation. So that means providing enterprise support that can work with a large IT organization. Boltz is much more focused on smaller startups, and has worked to make their model as lightweight and cost effective as possible so anyone can get it up and running.</p><p>Which one you should use depends on a host of different factors from accuracy to application to ease of use, far too many to go into here. Luckily, since they&#8217;re both open source, you can try both and decide for yourself which one is the best fit.</p><div><hr></div><p><span>Thanks for reading! My company, Merelogic, helps biopharma teams implement tools and practices to build a solid data foundation for whatever comes next. You can learn more at </span><a href="https://merelogic.net/">merelogic.net</a><span> or by requesting a free </span><a href="https://merelogic.net/data_risk_review">Data Risk Review</a><span>.</span></p>]]></content:encoded></item><item><title><![CDATA[Every biopharma team needs a data concierge]]></title><description><![CDATA[There&#8217;s a nebulous gap in how biopharma teams approach data that I ran into when I was leading data science and engineering at my first two biotech startups, and have since noticed at many of the clients I&#8217;ve worked with.]]></description><link>https://scalingbiotech.substack.com/p/every-biopharma-team-needs-a-data</link><guid isPermaLink="false">https://scalingbiotech.substack.com/p/every-biopharma-team-needs-a-data</guid><dc:creator><![CDATA[Jesse Johnson]]></dc:creator><pubDate>Wed, 22 Jul 2026 14:45:17 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>There&#8217;s a nebulous gap in how biopharma teams approach data that I ran into when I was leading data science and engineering at my first two biotech startups, and have since noticed at many of the clients I&#8217;ve worked with. I recently landed on a way of talking about what&#8217;s missing that I think more clearly frames what it is and why it&#8217;s important, and at the center </span>is<span> a new role I think every biotech team needs, which I&#8217;m calling a </span><em><span>data concierge</span></em><span>.</span></p><p><span>In this post, I&#8217;ll explain </span>why I think we need this new role and <span>what I think it should look like.</span></p><h4>No one cares about data hygiene. Should they?</h4><p>For the rest of this post, I&#8217;m going to use the term &#8220;data hygiene&#8221; to mean all the work that goes into ensuring that the data a team needs is in the right place at the right time, in a form that they can use and trust. This includes things like data infrastructure, architecture, quality, management, curation, engineering, informatics, etc. It&#8217;s not a great term, but it sums up what all these other terms are getting at.</p><p>The big problem with the way we talk about data hygiene (whatever name we use) is that it makes it really easy for biopharma leaders to conclude that it&#8217;s low priority or unnecessary. Because of this, most companies either under-invest in it or put it off for later (which turns into never).</p><p>There are two different ways to interpret this: Either we&#8217;re presenting things accurately and it is, in fact, low priority, optional work OR we&#8217;re not presenting things right. I think it&#8217;s the second one, which means that one of the most impactful things we can do to help the industry is make a better case for data hygiene.</p><p>The problem, as I see it, is that we often frame data hygiene in terms of the long-term benefits and principled approaches to doing things the right way. If you look at job titles, we have engineers and architects building cities of data that are designed to last into the future. We have curators and stewards preparing to hand data off to the next generation of analysts.</p><p>Biopharma leaders don&#8217;t care about the next generation of analysts. They care about the next milestone. They care about <a href="https://merelogic.net/data_risk_review">eminent risks</a> and immediate value. They care about keeping their current fires under control long enough to make it across the finish line without starting too many new ones.</p><p>If data hygiene doesn&#8217;t help them do that then who can blame them for ignoring it?</p><h4>Eminent risks and immediate value</h4><p>There are real, significant, immediate risks to under-investing in data hygiene. Small errors like pasting data into the wrong cell in a spreadsheet or using the wrong version of a file can lead teams down months-long wild goose chases trying to validate spurious results. But regardless of how well we present these risks, the bigger problem is connecting the dots of how data hygiene addresses them.</p><p>The problem is that the work of data hygiene is often presented as an abstract layer of automation and control, separated from the day-to-day work of scientific teams they support. And in practice, in the places where the teams do come in contact with this infrastructure, it&#8217;s often under-used or used incorrectly because they don&#8217;t understand how it fits their needs and their workflows.</p><p>The goal of introducing the data concierge role is to reframe the work I lumped under &#8220;data hygiene&#8221; around immediate value and risk prevention by providing a tangible way to ensure teams use the available tools to do better work while minimizing risk.</p><h4>Enter the Data Concierge</h4><p>A data concierge is a person responsible for maintaining a biopharma team&#8217;s data standards and conventions, helping the team navigate them and coordinating efforts to add automation and supporting tools (though not in that order).</p><p>It isn&#8217;t a job title, necessarily, but rather a role that an FTE can take on as part of their responsibilities, or that can be outsourced to a fractional contractor.</p><p>They&#8217;re responsible for things like:</p><ul><li><p>Answering questions from team members about where to find data and where to put data. (An LLM or MCP server can probably answer most of these questions, but the concierge is responsible for keeping its configuration up to date and stepping in when the LLM gets confused or a user wants to talk to a human.)</p></li><li><p>Managing the process of defining new conventions when the team runs into new use cases. (The decisions should be made by team members, but the concierge coordinates the process and acts as referee when necessary.)</p></li><li><p>Documenting the conventions in a way that can be handed off to a new concierge at any time. (This same documentation could probably be used to power the LLM/MCP, but more importantly it means things won&#8217;t crumble if and when the person in the role of concierge moves on.)</p></li><li><p>Monitoring the data in the different systems to ensure teams are following conventions and to identify the best places to add automation or new tools.</p></li><li><p>Sketching out new data architecture and infrastructure and how it fits into the team&#8217;s workflows, then coordinating the implementation and rollout.</p></li></ul><p>Note that I listed these responsibilities from immediate, user-facing value to longer-term, more abstract value. I did that on purpose: We want to start with the things that are a clear win, making life easier for the team while addressing risks, then move towards the more abstract things that are necessary to support that immediate value.</p><h4>Hold on, &#8220;concierge&#8221; sounds expensive&#8230;</h4><p>One obvious risk with this framing is that &#8220;data concierge&#8221; sounds like a luxury service. Cash-strapped biotechs with quickly shrinking runway will think this is a frivolous nice-to-have that they can&#8217;t afford. </p><p>And I&#8217;m OK with that because they probably can&#8217;t afford a data architect either.</p><p>The thing is, even folks who can&#8217;t afford to stay at the kind of hotel that has a concierge know why the concierge is there and why they wish they had one.</p><p>Tesla started by selling expensive sports cars to show that electric cars can compete with gas-powered cars, then slowly made them less expensive until they became a mass market staple. There are obviously a few differences between Tesla and what I&#8217;m proposing, but I believe the same strategy can work: Start with something that feels like a luxury and is clearly worth it if you can afford it, then turn it into a luxury that everyone can afford.</p><p>In other words, we need to stop under-selling data hygiene (by any name).</p><h4>But it&#8217;s worth it!</h4><p>For teams that understand the value of investing a little but up front to avoid expensive mistakes that could cost many times more, investing in a data concierge should be an obvious decision for the teams that can afford it. That&#8217;s why companies pay for insurance and a bunch of other things.</p><p>And biotechs that aren&#8217;t within a few months of eminent death can afford this too. (Granted, that&#8217;s a shrinking group, but it&#8217;s not empty.)</p><p>If you&#8217;ve read this far and think this is an idea worth exploring, I&#8217;d love to hear what you think. You can reach me at <a href="mailto:jesse@merelogic.net">jesse@merelogic.net</a>.</p><div><hr></div><p>Thanks for reading! My company, Merelogic, helps biopharma teams implement tools and practices that ensure they can rely on their data, driving more confident decisions and building a solid foundation for whatever comes next. You can learn more at <a href="https://merelogic.net/">merelogic.net</a> or by requesting a free <a href="https://merelogic.net/data_risk_review">Data Risk Review</a>.</p>]]></content:encoded></item><item><title><![CDATA[$2B seems like a lot to invest in binding prediction]]></title><description><![CDATA[I&#8217;m increasingly convinced that for AI/ML/etc.]]></description><link>https://scalingbiotech.substack.com/p/2b-seems-like-a-lot-to-invest-in</link><guid isPermaLink="false">https://scalingbiotech.substack.com/p/2b-seems-like-a-lot-to-invest-in</guid><dc:creator><![CDATA[Jesse Johnson]]></dc:creator><pubDate>Wed, 15 Jul 2026 14:45:39 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I&#8217;m increasingly convinced that for AI/ML/etc. to have the impact on drug discovery/development that many people think it should, the tools can&#8217;t just be drop-in replacements for existing steps in the drug discovery process. Instead, biopharma teams will need to evolve their drug pipelines around these tools. And this week, to illustrate this, I want to specifically look at <a href="https://www.isomorphiclabs.com/articles/isomorphic-labs-announces-series-b-investment-round">Isomorphic Labs&#8217; recent $2 billion Series B round</a> and what it would take for them to have the kind of impact that would justify a $2B investment.</p><p>(A lot of this also probably applies to <a href="https://www.businesswire.com/news/home/20260713849009/en/Chai-Discovery-Announces-%24400M-Series-C-to-Advance-AI-Driven-Molecular-Design">Chai Discovery&#8217;s recent $400M raise</a> which was announced between when I wrote this post and when it went out.)</p><h4>First, the cynical take</h4><p>I want to keep this post positive, but just to get it out of the way, I feel like I should address what I suspect most biopharma insiders were thinking when they saw this announcement: </p><p>The investors in this round, all of whom where tech investors and sovereign wealth funds, probably saw the succession of multi-billion-dollar contracts that Isomorphic signed with big pharma as serious validation. But it&#8217;s not unheard of for pharma execs to drop a $30M up-front payment for the prestige of working with a big name spinout from Google/DeepMind, confident that by the time the project is quietly canceled in a few years, they will have moved on to their next role.</p><p>From that perspective, the $2B investment seems like a bad idea. But it&#8217;s also possible that in 5-10 years, Isomorphic will be worth hundreds of billions and these investors will all look like geniuses.</p><p>The rest of this post is going to explore the question of what it would take for that second scenario to happen.</p><h4>Is binding a big deal?</h4><p>I&#8217;m going to make the assumption that a large part of what Isomorphic has planned involves using AlphaFold to do binding prediction, i.e. predicting whether a drug candidate like a small molecule or a protein will bind or otherwise interact with different target proteins, including proteins without a known/experimentally verified structure. I&#8217;m sure they&#8217;re working on other models too, but Isomorphic was built around AlphaFold, so that&#8217;s what I&#8217;m going to focus on.</p><p>Binding is what determines most of what a drug candidate does in your body. So in theory, if you can determine how it binds to every protein in your system, you should be able to predict fairly accurately what the drug will or won&#8217;t do.</p><p>However, in practice, most drug discovery teams only measure binding with at most one target protein in a large-scale screen. After that, they do functional assays, genomics, etc. on the hits to understand the downstream impacts of everything else the molecule binds too. That initial binding assay is one of the least expensive parts of the process, even when you do it across millions of potential drugs.</p><p>So if you were to tell the lead scientists on a drug discovery program that you can replace that one binding experiment with a (potentially less accurate) digital model, their response would be &#8220;I&#8217;m good with the large-scale screen, thanks.&#8221; And if you tell them that you can predict whether their drug candidate binds to any protein in the genome, their response would be &#8220;What do you expect me to do with that?&#8221;</p><h4>Drop-in Replacements</h4><p>OK, that&#8217;s probably not how they would respond. Scientists tend to have plenty of ideas about how they <em>could</em> use new kinds of information to answer interesting questions. But when it comes time to actually use those answers, they mostly don&#8217;t because they don&#8217;t have a specific step in their existing process where that information answers the specific question they need answered.</p><p>And sure, Isomorphic could use the binding predictions to infer the likely outcomes of all the functional assays, genomics, etc. that answer the questions specific questions that scientists need to answer at specific steps in their process. But those models would probably be less accurate than models trained directly on the outcomes, let alone doing the actual experiments.</p><p>And even if they could predict these things faster, cheaper and more accurately than the existing lab experiments, these steps aren&#8217;t very slow or expensive compared to the later stages, particularly clinical trials. So they still wouldn&#8217;t be solving a multi-billion dollar problem.</p><h4>The clinical information bottleneck</h4><p>On the other hand, they could probably use all these binding predictions to infer a whole bunch of other things that most biologists have never even dreamed of knowing, maybe with the help of a <a href="/__u/scalingbiotech.substack.com/p/the-emerging-scientific-stack-for">virtual cell model</a>. </p><p>More importantly, this information could probably rule out drug candidates that would fail in clinical trials but that <a href="/__u/scalingbiotech.substack.com/p/the-clinical-information-bottleneck?utm_source=publication-search">existing experiments can&#8217;t rule out</a>. Lowering clinical failure rates even a little would actually justify a $2B funding round in ways that no amount of improving the existing early discovery process ever will.</p><p>However, for biologists to use that information in practice, they would need to change their overall process so that the questions that these new models answer are the specific questions that they need answered at specific times in the new process. In other words, they would need to modify the process to create the gaps that they can drop the models into.</p><h4>Evolving Drug Discovery</h4><p>So to make its investors look like geniuses, Isomorphic can&#8217;t just build better models and hand them to scientists. They&#8217;ll need to define how the drug discovery process will need to evolve to accommodate these tools, then somehow orchestrate the change management, with partners or their internal teams, to make it happen.</p><p>Otherwise, Isomorphic will become just another cautionary tale of misguided techies who over-estimated how well they understood pharma.</p><div><hr></div><p><span>Thanks for reading! I help biopharma teams evolve their processes to extend the rigor they demand inside the lab to how they handle data and leverage AI outside the lab. If this is something you&#8217;re trying to figure out, visit </span><a href="https://merelogic.net/"><span>merelogic.net</span></a><span> to learn more.</span></p>]]></content:encoded></item><item><title><![CDATA[So, Anthropic is a biotech startup now?]]></title><description><![CDATA[Anthropic and OpenAI made a number of announcements last week about tools they&#8217;re making specifically for drug discovery/biopharma.]]></description><link>https://scalingbiotech.substack.com/p/so-anthropic-is-a-biotech-startup</link><guid isPermaLink="false">https://scalingbiotech.substack.com/p/so-anthropic-is-a-biotech-startup</guid><dc:creator><![CDATA[Jesse Johnson]]></dc:creator><pubDate>Wed, 08 Jul 2026 14:45:46 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Anthropic and OpenAI made a number of announcements last week about tools they&#8217;re making specifically for drug discovery/biopharma. In this post, I want to explore the two from Anthropic that I find the most interesting:</p><ol><li><p>They released a platform called Claude Science that, from what I can tell, is mostly the existing Claude UI with some pre-installed <a href="/__u/scalingbiotech.substack.com/p/i-for-one-welcome-our-new-mcp-overlords">MCP servers</a>.</p></li><li><p>To dogfood this product, they&#8217;re starting an internal effort to try and identify drug candidates for rare and/or orphaned diseases.</p></li></ol><p>While their goal is clearly to drive adoption of Claude in pharma and biotech, I think this is much more likely to backfire and actually push these teams off of Claude. And this week, I want to explain why.</p><h4>Is Claude Science a big deal?</h4><p>My first impression when I read about Claude Science itself was mostly &#8220;meh&#8221;. Sure, they have built-in integrations with many of the hottest AI-for-biopharma vendors and they provide a single place where you can access them all centrally, but like&#8230; that&#8217;s not actually that new. </p><p>In the <a href="/__u/scalingbiotech.substack.com/p/llms-should-be-orchestrators-not">last few weeks</a>, I&#8217;ve been writing about how all these companies are already building their own <a href="/__u/scalingbiotech.substack.com/p/i-for-one-welcome-our-new-mcp-overlords">MCP servers</a> that you can hook up to Claude or any other LLM. These integrations usually just take a few clicks to configure, so it&#8217;s unclear exactly what they&#8217;re providing that you couldn&#8217;t get elsewhere.</p><p>Sure, it will be easier for big enterprises to use a pre-packaged solution from an established vendor, but even there, there are existing solutions like <a href="https://www.sigmaticsciences.com/">Sigmatic</a> (formerly Helix AI). And this slight advantage will be overwhelmed by the negatives associated with the second announcement and a broader tide that&#8217;s working against commercial LLM builders like Anthropic and OpenAI.</p><h4>The tide is going out for token-based pricing</h4><p>In the last few months, even the most die-hard AI boosters have started to notice that tokens are getting expensive. I won&#8217;t recount all the stories about obscenely high Anthropic and OpenAI bills, because I&#8217;m sure you&#8217;ve heard them.</p><p>At the same time, there&#8217;s more chatter, though maybe not quite as loud yet,  about open weight models that are almost as good, or just as good, as the commercial models. Anyone can download these and run them locally on their own Cloud instances. With unlimited tokens for the cost of the compute, some early adopters are reporting costs as low as 1/100th what they were paying for tokens.</p><p>Obviously, there&#8217;s a lot more friction involved in hosting your own LLMs, but there are open source libraries that can handle most of it, and the economics are increasingly making it worth that hassle. (If you want me to write a post about self-hosting a Claude clone, let me know&#8230;)</p><p>The problem for Claude Science is that the whole point of it is to make you use Anthropic&#8217;s models instead of hosting your own, so part of the business model is that it needs to be more expensive.</p><h4>Competing with your AI provider</h4><p>So, that was a problem for Anthropic even before last week&#8217;s announcements, perhaps an existential one. But on top of that, Anthropic has now made an announcement that its life science customers will read as &#8220;We are going to start competing with you, using what we learn from you using our product, and possibly holding back the best features for internal teams.&#8221;</p><p>Now, this is NOT what they meant to say. In fact, I think the whole point of focusing on rare and orphaned diseases, which usually aren&#8217;t commercially viable, was to seem less threatening. </p><p>They claim that the reason they&#8217;re doing this is to get first hand experience that they can use to improve the product. And I actually wouldn&#8217;t be surprised if they&#8217;re being honest, just because it would be crazy to think that an internal version of Claude Science would give them a material advantage over companies who deeply understand the other 98% of the drug discovery and development pipeline.</p><p>But it kind of doesn&#8217;t matter. Because just the risk that Anthropic might be secretly planning to do a spin-off like Isomorphic Labs for AlphaFold, or that they might change their minds to do this in the future, will make most life science customers reluctant to let Anthropic touch their data.</p><h4>But things change fast these days&#8230;</h4><p>I always promise not to make predictions about the future, and I always seem to break that promise, or at least come pretty close to it. And who knows - the landscape is changing faster than anyone can keep up, and the dynamics will probably shift again before this has any lasting impact. But I&#8217;m going to be vary curious to see how this play out.</p><div><hr></div><p><span>Thanks for reading! I help biopharma teams evolve their processes to extend the rigor they demand inside the lab to how they handle data and leverage AI outside the lab. If this is something you&#8217;re trying to figure out, visit </span><a href="https://merelogic.net/"><span>merelogic.net</span></a><span> to learn more.</span></p>]]></content:encoded></item><item><title><![CDATA[Is model interpretability a code smell?]]></title><description><![CDATA[OK, I&#8217;m finally getting back to posting about bio foundation models, informed by the other topics I&#8217;ve been writing about the last few weeks.]]></description><link>https://scalingbiotech.substack.com/p/is-model-interpretability-a-code</link><guid isPermaLink="false">https://scalingbiotech.substack.com/p/is-model-interpretability-a-code</guid><dc:creator><![CDATA[Jesse Johnson]]></dc:creator><pubDate>Wed, 01 Jul 2026 14:45:13 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>OK, I&#8217;m finally getting back to posting about bio foundation models, informed by the other topics I&#8217;ve been writing about the last few weeks. In particular, I&#8217;ve been thinking about the <a href="https://drive.google.com/file/d/11fobUT94044eMg4IoNJKMRp_fNr9Ylf8/view?usp=sharing">processes and habits</a> around data and AI in drug discovery. So when I saw <a href="https://www.linkedin.com/posts/nikhilrp_ive-watched-a-lot-of-bio-foundation-model-share-7476550123385286656-3VRq/?utm_source=share&amp;utm_medium=member_desktop&amp;rcm=ACoAAAFF22IBncBfdFysCGcM_A_GBWML2ZPxAlA">this LinkedIn post</a> by Nikhil Reddy Podduturi encouraging bio foundation model developers to create better interpretation tools, it got me wondering if this is a genuine technical need, or rather a symptom of a bigger issue related to the processes around these models. This week, I want to explain what I mean by this.</p><h4>Model Interpretability</h4><p>As AI/ML models get more complicated, it gets harder to tell how/why they came up with the predictions they did. With a linear model, you can look at which input features contributed the most to the final sum. With a multi-layer neural network, those input features get diced, folded, mixed and blended to the point where that&#8217;s no longer possible. For a <a href="/__u/scalingbiotech.substack.com/p/how-to-build-a-bio-foundation-model">transformer-based model</a> where each token is generated by an algorithm involving multiple neural networks, it&#8217;s not even in the same universe as possible.</p><p>The idea of model interpretability is to introduce another algorithm that interprets the weights, outputs, etc. from a complex model to try and explain &#8220;why&#8221; it made the predictions that it did. There are a lot of different approaches to this, but they all introduce a fundamental risk: Whenever you try to reconstruct why something happened after the fact, it&#8217;s hard to tell if you&#8217;re learning something useful or just coming up with a good sounding story.</p><h4>Code Smells</h4><p>In software engineering, there&#8217;s this idea that you&#8217;ll sometimes see a pattern that isn&#8217;t necessarily bad on its own, but is rather a sign that something else is wrong. It&#8217;s awkward or unusual enough that if that&#8217;s the only way to solve a problem, there&#8217;s probably something bigger, and perhaps harder to identify, that&#8217;s the real problem. This is called a &#8220;code smell.&#8221;</p><p>I&#8217;m starting to think model interpretability is a code smell, at least in certain situations: It isn&#8217;t inherently bad, but it suggests there&#8217;s something missing from the information that you&#8217;re getting directly from the model. And if this is the case, the risk is that an interpretability layer will just trick you into thinking that it&#8217;s providing that information, without actually telling you what you need to know.</p><h4>The Problem isn&#8217;t interpretability</h4><p>So what if we stopped framing the problem as missing interpretability and instead posed it as a problem of trying to use these models for a purpose, or a process step, that they aren&#8217;t suited for.</p><p>We see a similar phenomenon with phenotypic screening, where drug candidates are initially identified by their impact on an in vitro assay, rather than directly observing which protein target(s) they bind to. Since the classical drug discovery process starts from identifying drug candidates that bind to a specific protein, it would be natural to want the phenotypic screening results to come with an interpretability tool that identifies the exact protein(s) causing the phenotypic result&#8230; but they don&#8217;t.</p><p>In fact, figuring out the protein target(s) that drugs from phenotypic screens bind to is incredibly difficult, complex and expensive. But that hasn&#8217;t stopped plenty of startups and pharma labs from using phenotypic screening. They just have to adjust their process to use the information that they do get from the phenotypic screens. And sometimes that means not knowing the binding protein(s).</p><h4>Rethinking Processes</h4><p>To be clear, I am not saying that we should trust the predictions/outputs of foundation models without any further scrutiny or validation. I mean, it should go without saying that any predictions will need to be verified in the lab many times over before any drug makes it to the clinic. But that validation is expensive and time consuming, so we also need a way to validate predictions digitally before we validate them in the lab.</p><p>What I&#8217;m saying instead is that we shouldn&#8217;t try to treat bio foundation models, or any other kind of model, as drop-in replacements for existing pieces of the drug discovery process and fill the gaps with (the illusion of) interpretability. Instead, we should evaluate the kind of information the models are providing and the level of confidence it provides, then rethink the processes around the model to accommodate that level of risk.</p><p>I don&#8217;t know exactly what that looks like, but I&#8217;m confident we can draw from examples like phenotypic screening to figure it out.</p><div><hr></div><p>Thanks for reading! I help biopharma teams evolve their processes to extend the rigor they demand inside the lab to how they handle data and leverage AI outside the lab. If this is something you&#8217;re trying to figure out, visit <a href="https://merelogic.net/">merelogic.net</a> to learn more.</p>]]></content:encoded></item><item><title><![CDATA[The foundation for bio foundation models]]></title><description><![CDATA[I&#8217;ve been bouncing between a few different topics recently, and I thought it might be worth explaining how they fit together.]]></description><link>https://scalingbiotech.substack.com/p/the-foundation-for-bio-foundation</link><guid isPermaLink="false">https://scalingbiotech.substack.com/p/the-foundation-for-bio-foundation</guid><dc:creator><![CDATA[Jesse Johnson]]></dc:creator><pubDate>Wed, 24 Jun 2026 14:46:08 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I&#8217;ve been bouncing between a few different topics recently, and I thought it might be worth explaining how they fit together. So this week, I want to explain why the MCP servers and <a href="https://merelogic.net/data_operations_plans/">Data Operations Plans</a> that I&#8217;ve been writing about are so important to reducing clinical failure rates and supporting bio foundation models.</p><p>I&#8217;ll also explain why this has led me to start offering a service I call a <a href="https://drive.google.com/file/d/11fobUT94044eMg4IoNJKMRp_fNr9Ylf8/view?usp=sharing">Data Habits Assessment</a> that evaluates how your team actually handles data (not just how they&#8217;re supposed to) and what you can do to help them adopt better habits using the tools you already have. I&#8217;m planning to offer this for free to a handful teams while I&#8217;m validating the approach. Send an email to <a href="mailto:data-ops@merelogic.net">data-ops@merelogic.net</a> if you&#8217;re interested.</p><h4>Bio Foundation Models and the Clinical Information Bottleneck</h4><p>I started this series of posts from the hypothesis that the main cause of the high failure rates for clinical trials, which drive up the cost of drug development, is something I call the <a href="/__u/scalingbiotech.substack.com/p/the-clinical-information-bottleneck">clinical information bottleneck</a>. The reason I&#8217;m so excited about bio foundation models is they they have a tangible (though not immediate) potential to address this bottleneck by integrating clinical data into early stage discovery.</p><p>So that got me writing about how bio foundation models work and how they can fit together at different scales. I&#8217;ll get back to that topic soon, but the more I dug into it, the more I kept coming back to the fact that the models are only good as the data they&#8217;re trained on.</p><p>And the problem, I realized, is that most biopharma teams don&#8217;t trust their data.</p><h4>When you can&#8217;t trust your data&#8230;</h4><p>A major blocker, in my experience, to teams trusting their data is that most bench teams have starkly different standards for the level of rigor they demand inside the lab compared to how they handle data outside the lab. In the lab, it&#8217;s carefully designed protocols with systems dedicated to documenting and tracking the work (ELNs and LIMS). Outside the lab, it&#8217;s manually copying and pasting numbers between spreadsheets saved in arbitrary folders on your laptop.</p><p>Sure, there are more fundamental and unavoidable issues with biology data, but any gap in trust means more second guessing results, more extra validation assays, more chasing down dead ends, and generally more looking over your shoulder. </p><p>Every crack in the foundation makes you less willing to experiment and explore, and this is a big crack.</p><h4>MCP servers and AI workflows</h4><p>The reason bench scientists don&#8217;t enforce the same level of rigor outside the lab is that properly handling data has always required a level of technical complexity that is well outside the comfort zone of most bench scientists. At the same time, assay protocols and experiment designs change faster than any informatics/engineering team can keep up. So bench scientists do what they have to, like copying and pasting numbers between spreadsheets.</p><p>AI platforms connected to the right tools via MCP servers have potential to address this by allowing bench scientists to handle the technical complexity with the level of flexibility that early discovery workflows require. </p><p>I wrote about <a href="/__u/scalingbiotech.substack.com/p/llms-should-be-orchestrators-not">what this might look like</a> last week.</p><h4>Data Operations and Habits</h4><p>But even as the tools get better and become more readily available, there&#8217;s still the problem of getting everyone on the team to use them consistently, in a way that allows you to trust your data. And the introduction of new technology like MCP servers only makes this more complicated, at least in the short run,</p><p>In most teams, there isn&#8217;t an individual who&#8217;s officially responsible for training the team on data workflows, gathering feedback on what is or isn&#8217;t working, implementing follow-up mechanisms, etc.</p><p>My goal with the <a href="https://merelogic.net/data_operations_plans/">Data Operations Plans</a> and the <a href="https://drive.google.com/file/d/11fobUT94044eMg4IoNJKMRp_fNr9Ylf8/view?usp=sharing">Data Habits Assessment</a> is to help teams fill that gap.</p><p>Most teams overlook this problem. The lucky ones have someone who thinks about it and tries to fix it, even though it isn&#8217;t their job. But I think the industry needs to be more deliberate about addressing it. If you you want to be more deliberate about your own data operations and habits, send me an email at <a href="mailto:data-ops@merelogic.net">data-ops@merelogic.net</a> .</p>]]></content:encoded></item><item><title><![CDATA[LLMs should be orchestrators, not messaging busses]]></title><description><![CDATA[Ok, I promise I'll get back to writing about foundation models soon, but I fell down this MCP rabbit hole and I have a few more weeks of ideas I need to explore before I can move on.]]></description><link>https://scalingbiotech.substack.com/p/llms-should-be-orchestrators-not</link><guid isPermaLink="false">https://scalingbiotech.substack.com/p/llms-should-be-orchestrators-not</guid><dc:creator><![CDATA[Jesse Johnson]]></dc:creator><pubDate>Wed, 17 Jun 2026 14:45:33 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Ok, I promise I'll get back to writing about foundation models soon, but I fell down this MCP rabbit hole and I have a few more weeks of ideas I need to explore before I can move on.</p><p>For decades, the dream in data engineering/informatics circles has been to integrate all the different tools and software that a team uses into a single system that seamlessly and automatically transfers data across systems, replacing manual copy/paste spreadsheet analysis with reliable processes that leave an audit trail. REST APIs were supposed to make this possible, but in practice, integrations always turned out to be too complex and time consuming to make sense outside of larger teams/organizations with highly standardized processes. Most API integrations were just too brittle for the kind of exploratory work that defines early discovery.</p><p>Two weeks ago, I wrote about how MCP servers present an opportunity to fix all this by turning the LLM platform (such as Claude Code or one the growing number of alternatives) into a messaging bus that coordinates between all these different tools with significantly more flexibility at a negligible technical cost. Then last week, I posted about some of the issues that could arise from using unreliable, hallucinatory LLMs to handle data that drives incredibly expensive decisions.</p><p>After thinking about it for another week, I&#8217;ve started forming an idea about how we could address this, and turn the software that exists for biology today into an ecosystem of LLM-connected tools that safely enable the dream. My hypothesis is that the LLM platform in this future can&#8217;t be a messaging bus exactly. It has to become an orchestrator that coordinates direct, deterministic interactions between the other components.</p><p>In the rest of this post, I want to explain what I mean by this.</p><h4>The Vision</h4><p>Imagine, if you will, your experiment has just completed and the data from the last instrument has been written to whatever system it goes into. You type into your LLM a brief description of the analysis you want done, or the question you want to answer and set it off:</p><p>First, it uses an MCP connection to your ELN to find the details of how the experiment was done. Then it uses another MCP integration to find the readout data that was just written. It uses a third MCP server to kick off the primary analysis that will boil the gigabytes of data down to a single table. Then it uses a fourth MCP connection to the knowledge graph to look up contextual information about the genes and proteins involved in the experiment. A call to a fifth MCP server merges everything into a form that&#8217;s ready for final interpretation.</p><p>In practice, this would probably be multiple prompts, with a big gap in the middle while primary analysis is running, but the steps would be the same. The important part is how the LLM coordinates between the different systems via their MCP connections.</p><p>Now, you could absolutely automate this kind of workflow without an LLM, using just REST APIs (assuming the systems all have reasonable APIs, which is sometimes true.) However, to make it happen, someone with a very specific set of technical skills would need to predict more or less exactly what the workflow would look like, then spend a few days, if not weeks or months getting it to work. Even if you have someone with that particular set of skills on your team, that&#8217;s significantly more work than&#8230; typing a vague description of what you want into an LLM and waiting a few minutes.</p><h4>The Problem</h4><p>The problem that I started exploring last week is that if all these systems only talk to the LLM through their MCP servers, then all the data that passes between the systems has to be encoded then decoded by the LLM. That&#8217;s a problem for a lot of reasons: It&#8217;s expensive (how many tokens is your 20GB fastq file?) and slow, and it pretty much guarantees that you&#8217;ll have errors.</p><p>So the hub-and-spoke model in which the LLM is the messaging bus isn&#8217;t going to work. Data should only pass between and through deterministic systems that keep audit trails, etc. LLMs can generate things like code and parameters, sure, but only things that a human could reasonably review.</p><p>Instead of pure hub-and-spoke, we need direct, deterministic connections between the spokes, orchestrated by the LLM. And that probably means old fashioned APIs.</p><h4>APIs Revisited</h4><p>So sure, it sounds like we&#8217;re back to where we started - we still need to connect all our systems to each other with deterministic APIs. But what makes REST APIs brittle is that they&#8217;re designed to work with users who can&#8217;t translate what they want into technical terms. When an engineer has to do that translation in advance, you&#8217;re stuck with whatever they thought the users would need at the beginning.</p><p>In other words, the brittle nature of a REST APIs is a requirement of the context in which they&#8217;ve always been used, not a technical limitation of deterministic APIs. In a context where users can do that technical translation on the fly, APIs can be designed to be flexible. In fact, flexible API standards such as <a href="https://graphql.org/">GraphQL</a> have existed for years, but never had widespread adoption because of this translation issue.</p><p>Today, LLMs can do the technical translation for users almost instantly.</p><h4>Hub and Spokes and APIs</h4><p>So, here&#8217;s what I&#8217;m proposing: MCP servers will very soon give us lots of spokes that can be connected to whatever LLM hub you want. Almost every bio software company I&#8217;ve talked to in the last few weeks is building one, or has already built one. The rest have it on their roadmap.</p><p>Each of these individual spokes can decide on its own how it will communicate with the hub. The companies I&#8217;ve talked to are scrambling to figure that out, and that&#8217;s fine because LLMs are generally able to work with whatever they need to.</p><p>What&#8217;s missing is a way for the hub to tell the spokes how to interact with each other: Tell the workflow runner to pull the sample sheet from the ELN. Tell the visualization tool to pull the contextual information from the knowledge graph. And so on.</p><p>Doing this for a single pair of systems probably isn&#8217;t too bad. But if this dream is going to be a reality for everyone in biopharma, we need to be able to do it for every pair of systems (or close to every.)</p><p>In other words, we need the industry to agree to some kind of standard for what these appropriately flexible APIs between them should be, and how the LLMs should communicate instructions for using those connections. (GraphQL is one option, but not the only one.)</p><h4>On Standards and Adoption</h4><p>Getting industries to adopt standards like this has always been incredibly difficult and many (most?) attempts at it have failed.</p><p>They fail for many reasons, from politics and perverse incentives to the technical cost of adoption. But often, they fail because the customers who could put the most pressure on developers just don&#8217;t see the potential value. If no one&#8217;s asking for it, why would software developers do it?</p><p>This is probably the biggest reason there was never (as far as I know) an attempt to build an interoperability standard for REST APIs in early discovery biopharma: integrations were always so high-cost and low-value that no one ever asked for it.</p><p>In this case, however, there&#8217;s a real chance that the cost will be low enough and the value clear enough that customers will start asking for this kind of standardization. There will still be the politics, and potentially the perverse incentives, but user demand can do a lot to overcome those. </p><p>Either way, that&#8217;s a topic for the future. I think I&#8217;ve accomplished my goal of explaining why LLMs should become orchestrators instead of messaging busses.</p><p>If you have opinions about this and want to talk about what might be involved in making getting the industry to define and adopt a standard like this, I&#8217;d love to hear from you - leave a comment or send me an email at <a href="mailto:jesse@merelogic.net">jesse@merelogic.net</a>.</p><div><hr></div><p>Thanks for reading! I help biopharma teams evolve their processes to extend the rigor they demand inside the lab to how they handle data and leverage AI outside the lab. If this is something you&#8217;re trying to figure out, visit <a href="https://merelogic.net/">merelogic.net</a> to learn more.</p>]]></content:encoded></item><item><title><![CDATA[Data Safety in the Age of MCP/AI]]></title><description><![CDATA[Last week, I predicted that the introduction of MCP servers would increasingly turn AI platforms like Claude Code and Claude Cowork into integration busses allowing users to sync data between different external apps in ways that REST APIs are too brittle for.]]></description><link>https://scalingbiotech.substack.com/p/data-safety-in-the-age-of-mcpai</link><guid isPermaLink="false">https://scalingbiotech.substack.com/p/data-safety-in-the-age-of-mcpai</guid><dc:creator><![CDATA[Jesse Johnson]]></dc:creator><pubDate>Wed, 10 Jun 2026 14:46:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><a href="/__u/scalingbiotech.substack.com/p/i-for-one-welcome-our-new-mcp-overlords">Last week</a>, I predicted that the introduction of MCP servers would increasingly turn AI platforms like Claude Code and Claude Cowork into integration busses allowing users to sync data between different external apps in ways that REST APIs are too brittle for. But this raises some serious questions about how we can trust non-deterministic tools that are known to hallucinate and make other mistakes, with data for tasks where mistakes can be unreasonably expensive.</p><p>In this post, I want to explore what we can learn from software engineering, where the community has found safe and effective ways to use LLMs.</p><h4>Biopharma MCP Tasks</h4><p>Based on how I&#8217;ve seen MCP servers being used in other industries, it seems natural that biopharma teams will want to use them for tasks that are often covered in a <a href="https://merelogic.net/data_operations_plans/">Data Operations Plan</a>, such as:</p><ul><li><p>Reformatting raw data from instrument readouts into a form that&#8217;s ready for analysis or interpretation</p></li><li><p>Inserting that data into other databases or tools for further analysis/interpretation</p></li><li><p>Mapping and standardizing gene names and other entities to standard vocabularies/ontologies</p></li><li><p>Looking up auxiliary/contextual information from external sources and combining it with experimental data.</p></li></ul><p>Today, many biopharma teams do the first two of these manually with spreadsheets, and from what I&#8217;ve seen, most don&#8217;t do the second two at all. But they would like to have them done. And this seems to be consistent with broader industry: AI mostly isn&#8217;t replacing people, at least not in highly skilled jobs, so much as doing things that either people have always wanted done, but it was never quite worth the work it would take, or things that people had to do, but really didn&#8217;t like doing.</p><p>Of course, that doesn&#8217;t work if you can&#8217;t trust the results.</p><h4>Verifiable results</h4><p>One of the reasons that software engineers have been able to use LLMs to write code is that there are easy ways to tell whether or not it did the job correctly. Compilers, linters and unit tests are all designed to tell you (or the LLM) if the code was written correctly. So the LLM can write code that&#8217;s 98% correct, check for errors, fix whatever mistakes it made, and repeat.</p><p>Of course, this is also how human coders work, which is why compilers, linters and unit tests existed long before LLMs. In fact, human coders are a lot less than 98% accurate, even the really good ones. What makes them better than LLMs (even the next model that OpenAI or Anthropic swears is going to be too dangerous to release) is that they know where to look for errors, how much test coverage is good enough, what&#8217;s the right amount of technical debt, etc.</p><p>In other words, what makes software engineering &#8220;safe&#8221; is that the LLMs have mechanisms to verify that they did the right thing and human experts have ways to verify that these mechanisms are sufficient. (Maybe these mechanisms are the difference between vibe coding and AI-enabled software engineering...)</p><h4>Unverifiable results</h4><p>The problem with the tasks I listed above is that there aren&#8217;t separate mechanisms to verify they were done correctly. Or, at least, there aren&#8217;t established, easy to implement, mechanisms for this. </p><p>And yet, most biopharma teams are fine having humans do them manually, without any verifications. Humans who are known to make mistakes. And in fact, anyone who&#8217;s been in this business for more than a few years has at least one story of a value pasted into the wrong cell in a spreadsheet, an analysis run on the wrong file, etc. proving to be a very expensive mistake.</p><p>This makes using AI for data processing and analysis more like self driving cars than Claude Code. But unlike with self driving cars, where we have to rely on the computers eventually being safer than humans, biopharma has options for creating verification mechanisms. We just haven&#8217;t invested as much as we should in them, whether it&#8217;s LLMs or humans whose work is being verified. (Which gets back to <a href="https://merelogic.net/data_operations_plans/">Data Operations Plans</a>...)</p><h4>More code-like, less car-like</h4><p>If you give a scientist a spreadsheet of raw data and a spreadsheet that an LLM has generated from it in a new format, the scientist will not be able to tell you if the LLM did it correctly. But if the LLM generates a mapping from one format to the other and applies the mapping using a deterministic algorithm, a (properly trained) scientist should be able to look at the mapping config and tell if it was correct. The LLM could even generate small test datasets to apply the transformation to, for the user can verify.</p><p>Tools for deterministically mapping between data formats have existed for a long time, but adoption has been fairly low in biopharma because defining these mappings is often more work than just copying and pasting columns in Excel. However, using an LLM to generate the mapping is even easier, and less error prone, and leaves more of an audit trail and&#8230; The point is that LLMs could actually make these kinds of tasks MORE accurate by making it easier to the tools that we should&#8217;ve been using all along.</p><p>This was the idea behind Sphinx, for example, before it was <a href="https://www.benchling.com/blog/resync-bio-and-sphinx-bio-join-benchling">acquired by Benchling</a>, which is now presumably implementing a lot of these ideas. And similar capabilities have been added to tools like <a href="https://www.kaleidoscope.bio/">Kaleidoscope</a>, <a href="https://www.scispot.com/">SciSpot</a>, <a href="https://www.sapiosciences.com/">Sapio</a> and plenty of others.</p><p>It will be interesting to see how they make these capabilities available through the MCP servers that they&#8217;re no doubt building or have already built. (And I&#8217;m partnering with Nishant Jha from <a href="https://www.nitro.bio/">Nitro Bio</a> to help companies like this figure it out... <a href="mailto:jesse@merelogic.net">jesse@merelogic.net</a> if you want to chat about it.)</p><p>However this evolves, I think we&#8217;ll see better options in the coming years for enabling biopharma teams to safely and confidently handle data, so they can better trust the results, and waste less time trying to verify results that were actually just a copy/paste error.</p><div><hr></div><p>Thanks for reading! I help biopharma teams evolve their processes to extend the rigor they demand inside the lab to how they handle data and leverage AI outside the lab. If this is something you&#8217;re trying to figure out, visit <a href="https://merelogic.net/">merelogic.net</a> to learn more.</p>]]></content:encoded></item><item><title><![CDATA[MCP is the new SaaS]]></title><description><![CDATA[One question that keeps coming up in my conversations about building Data Operations Plans is how AI tools like Claude Cowork might fit in.]]></description><link>https://scalingbiotech.substack.com/p/i-for-one-welcome-our-new-mcp-overlords</link><guid isPermaLink="false">https://scalingbiotech.substack.com/p/i-for-one-welcome-our-new-mcp-overlords</guid><dc:creator><![CDATA[Jesse Johnson]]></dc:creator><pubDate>Wed, 03 Jun 2026 14:45:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>One question that keeps coming up in my conversations about building <a href="https://merelogic.net/data_operations_plans/">Data Operations Plans</a> is how AI tools like Claude Cowork might fit in. Many teams outside biopharma are using a protocol called MCP that lets them connect external tools like Slack and Notion to whatever AI tool they use. This is still at the early adopter phase, but it seems likely to become a common way of working, even in biopharma.</p><p>Within the next few years, every software/SaaS for bio will be expected to have an MCP server, and the quality of that MCP will be an increasingly important factor in buying decisions. So to help software-for-bio companies be ready, I&#8217;ve started partnering with Nishant Jha from <a href="https://www.nitro.bio/">Nitro Bio</a> to help developers design MCP servers that enable their users&#8217; most important workflows, while minimizing the risk of hallucinations and other reliabilty issues.</p><p>To help kick off this partnership, this week&#8217;s post will explain what MCP servers are and speculate about what MCP adoption might look like for the future of biopharma. If you want to chat with Nishant and me about an MCP server you&#8217;re building (or want to build), send me an email at <a href="mailto:jesse@merelogic.net">jesse@merelogic.net</a>.</p><p>I&#8217;m also building an MCP server for the <a href="https://merelogic.net/landscape">Biopharma AI Landscape</a> and looking for beta testers. Send me an email if you want to give it a try. </p><h4>Tools and Tool Calls</h4><p>Modern LLM-based AI systems use a mechanism called tool calls to extend the capabilities of the system. For example, LLMs are notoriously bad at doing math or following precise algorithms, but giving them the ability to call a Python interpreter means they just need to write the code, something that they are good at.</p><p>To use a tool, the LLM generates parameters, for example the code to run. The system that manages the LLM reads the parameters from the LLM&#8217;s response, uses them to run the appropriate tool, then passes the results into the next call to the LLM, along with a transcript of what the LLM has done so far. The LLM can then make more tool calls based on the results, and repeat the cycle until it decides it&#8217;s done.</p><p>LLM-based systems have had built-in tool calls for a while, to do things like searching the web. The MCP protocol allows users to give their LLM additional tools by connecting it with MCP servers. It&#8217;s a similar idea to a how REST APIs allow programs to interact with each other, but with a protocol built for how LLMs &#8220;think&#8221; about and use tools.</p><h4>The AI-powered Messaging Bus</h4><p>On it&#8217;s face, the main benefit of MCP servers is that they can make your LLM smarter by adding functionality. But the pattern I&#8217;m seeing more often (mostly outside biopharma) is that they&#8217;re allowing the LLM to become a sort of messaging bus between MCP servers - pulling information from one service and using it to update another.</p><p>For example, imagine if Claude could read an entry from your ELN, standardize the gene names based on a gene ontology, find the readout in your file system, then kick off an analysis run based on the protocol in the ELN.</p><p>That&#8217;s the kind of thing that you could theoretically do with a few REST APIs and a whole lot of code, but in practice both the APIs and the code end up being too brittle. There are so many different use cases you would need to support, it just isn&#8217;t worth writing scripts for every corner case (even with vibe coding).</p><p>Well, it turns out LLMs are quite flexible and can handle completely novel use cases with just a few carefully worded instructions.</p><p>Sure, coordinating an ELN with a bioinformatics workflow is more complicated than copying text from Notion to Slack, but LLMs are getting better at correctly handling increasingly complex tasks. Similar workflows in other domains are possible today, and it&#8217;s only a matter of time before they become possible, if not the standard, in biopharma. But the tools will need to be designed to minimize the kinds of mistakes that LLMs are prone to make.</p><h4>Local and External MCPs</h4><p>There are two ways that you can attach an MCP server to an LLM system: locally or externally.</p><p>A local MCP server runs on the same machine that the LLM system is running on. For example, if you&#8217;re using Claude Code (which runs on your own computer) and you configure it with a local MCP server, it will start the server for you and send it requests using your computer&#8217;s networking systems in such a way that all messages stay on your computer. With this approach, you don&#8217;t have to worry about security or privacy. In fact, you don&#8217;t even need to have a password on the MCP server. Local servers can also read and write files on your computer.</p><p>An external MCP server runs somewhere else on the internet. Your LLM system connects to it through the actual network rather than just pretending to make network calls like local servers. This means you don&#8217;t have to install anything on your computer and you can connect to external resources (like Notion or Slack or your ELN), but you have to deal with authentication, security, privacy, etc.</p><p>There&#8217;s also a third, sort of hybrid option, which is a local server that wraps an external REST API. The LLM system makes internal calls to your local server, which then makes actual network calls to an external REST API, then reformats the results before passing them back to the LLM.</p><p>There are already quite a few open source MCP servers for bio that use that last option. For example, <a href="https://github.com/genomoncology/biomcp">BioMCP</a> and <a href="https://github.com/openpharma-org">OpenPharma</a> run locally and provide access to different data sources, mostly by calling external REST APIs. There are fewer that are purely local, such as <a href="https://github.com/cafferychen777/ChatSpatial">ChatSpatial</a> which enables your LLM to run spatial omics analysis tools on your computer.</p><p>Purely external MCP servers are more likely to be developed by existing software/SaaS companies as a new way for users to access their software. Outside biotech, companies like Slack and Notion have already built this. Inside biotech, there are at least a few companies who have built their own MCP server, but most are still in beta.</p><p>It&#8217;s technically possible to build your own hybrid MCP server around your current ELN or any other software/SaaS. But you&#8217;ll be limited by the functionality of their REST API (which is a serious limitation for a lot of software for biopharma) and you might end up throwing it away if they eventually create their own MCP.</p><p>So in the long run, the best outcome would be for existing software/SaaS companies to build their own, official MCP servers. But given how slowly the industry has adopted REST APIs, I&#8217;m a more than a little worried.</p><h4>Hype vs Reality</h4><p>I know it can be hard to tell AI hype from reality, but this isn&#8217;t some vague promise, predicated on ChatGPT becoming sentient in the next six months. MCP servers allow the LLMs that we already have today to solve real problems that we&#8217;ve had for decades, namely connecting all the different pieces of software that should work together, but don&#8217;t. </p><p>If you&#8217;re making software today, this is something you should be paying attention to. And if you want to talk to Nishant and me about how to stay ahead of it, send me an email at <a href="mailto:jesse@merelogic.net">jesse@merelogic.net</a>.</p>]]></content:encoded></item><item><title><![CDATA[We won't know when AI solves drug discovery]]></title><description><![CDATA[Last week, I wrote a relatively optimistic post about how AI could potentially address what I see as one of the most fundamental issues in drug discovery, the clinical information bottleneck. This week&#8217;s post is going to be a bit more pessimistic because I want to talk about another, perhaps more fundamental, issue that I don&#8217;t think AI can address: the fact that you can&#8217;t really tell if a new technology has made a difference in drug development until decades after it&#8217;s introduced.]]></description><link>https://scalingbiotech.substack.com/p/we-wont-know-when-ai-solves-drug</link><guid isPermaLink="false">https://scalingbiotech.substack.com/p/we-wont-know-when-ai-solves-drug</guid><dc:creator><![CDATA[Jesse Johnson]]></dc:creator><pubDate>Wed, 27 May 2026 14:46:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Last week, I wrote a relatively optimistic post about how AI could potentially address what I see as one of the most fundamental issues in drug discovery, the <a href="/__u/scalingbiotech.substack.com/p/the-clinical-information-bottleneck">clinical information bottleneck</a>. This week&#8217;s post is going to be a bit more pessimistic because I want to talk about another, perhaps more fundamental, issue that I don&#8217;t think AI can address: the fact that you can&#8217;t really tell if a new technology has made a difference in drug development until decades after it&#8217;s introduced.</p><p>The second half of this post will explore why it will take so long to tell if we&#8217;ve fixed drug discovery, but first it&#8217;s worth stating why this is important:</p><p>We&#8217;re currently in the midst of a new wave of tech folks entering biopharma, expecting to quickly demonstrate how amazing their technology is. (I entered the field in the previous wave, which ended with the post-pandemic biotech winter around 2022.) This new wave, driven by LLM-powered AI, is potentially much bigger (at least in dollar terms) but it&#8217;s still unlikely to last long enough to see outcomes that are decades away.</p><p>The frustrating thing is that it&#8217;s possible someone already figured out how to fix drug discovery in the last wave, or the wave before that, and we haven&#8217;t noticed yet. Or more likely, they ran out of funding and scrapped the approach, so we&#8217;ll never know. </p><p>It&#8217;s unclear to me what&#8217;s stopping the same thing from happening this time.</p><h4>Error Bars</h4><p>The reason it takes so long to tell if you&#8217;ve made a difference is that the only improvement in drug development that really matters would be increasing clinical success rates. You don&#8217;t have to increase them by much, even. Just raising the current rate of 10% to a still-measly 20% would cut the cost of drug discovery in half.</p><p>However, to definitively say that you&#8217;ve done this, you would need to run enough clinical trials to know that you didn&#8217;t just get lucky. Whatever kind of error bars you use, that&#8217;s a lot of drugs that need to go through clinical trials, each taking 10-15 years from early discovery. Even if many of those trials are happening in parallel, this would take decades.</p><p>And sure, we could focus on other metrics like the cost and time of the parts of drug discovery that happen before clinical trials. But those metrics don&#8217;t work the way you might think. Because the vast majority of the time and money it takes to bring a drug to market is spent during and after the clinic, reducing the cost of the first, relatively cheap part doesn&#8217;t end up mattering that much.</p><p>In fact, if you speed up early discovery by introducing a novel and experimental (read: risky) approach, there&#8217;s a decent chance you&#8217;re just trading that early time and money for a lower success rate that will increase the cost overall. Any reasonable pharma executive would rather spend twice as much time and money on earlier discovery if it would guarantee them a better chance of success in the clinic.</p><h4>So, are we doomed?</h4><p>I don&#8217;t know what the answer to this problem is. Luckily, there are a lot of smart investors, founders, etc. who understand this better than I do, and are working to fix drug discovery with their eyes wide open. Of course, there are also plenty of people with unreasonable amounts of money, or access to VCs with that money, who don&#8217;t know what they don&#8217;t know, but are investing anyway.</p><p>Every time a wave of tech folks enter biotech and learn these lessons the hard way (I was one of them, not too long ago), I think it shifts the field towards the right balance of tech exuberance and bio reality. The question is whether we&#8217;ll ever find a balance that will give the field the patience to see if it&#8217;s working.</p><div><hr></div><p>Thanks for reading! I help biopharma teams design, implement and maintain <a href="https://merelogic.net/data_operations_plans/">Data Operations Plans</a> that extend the care and consistency they enforce inside the lab to how they handle data outside the lab. To learn more, send me an email at <a href="mailto:jesse@merelogic.net">jesse@merelogic.net</a>.</p>]]></content:encoded></item><item><title><![CDATA[The Clinical Information Bottleneck]]></title><description><![CDATA[In a few posts now, I&#8217;ve referred to this idea I call the &#8220;Clinical Information Bottleneck.&#8221; The idea is kind of abstract, but I think it&#8217;s at the heart of the rising cost of drug development, and clearing this bottleneck will be key to how AI can help bring these costs back down.]]></description><link>https://scalingbiotech.substack.com/p/the-clinical-information-bottleneck</link><guid isPermaLink="false">https://scalingbiotech.substack.com/p/the-clinical-information-bottleneck</guid><dc:creator><![CDATA[Jesse Johnson]]></dc:creator><pubDate>Wed, 20 May 2026 14:46:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In a few posts now, I&#8217;ve referred to this idea I call the &#8220;Clinical Information Bottleneck.&#8221; The idea is kind of abstract, but I think it&#8217;s at the heart of the rising cost of drug development, and clearing this bottleneck will be key to how AI can help bring these costs back down. I haven&#8217;t yet written a dedicated post that goes into detail about what I mean by this bottleneck, so that&#8217;s what I want to do this week.</p><p>There are two directions to it, which I&#8217;ll discuss in turn:</p><ol><li><p>Information from the discovery/pre-clinical phase of a drug program is not very predictive of what will happen in the clinical phase.</p></li><li><p>Information from the clinical phases of past drug programs is generally not very relevant to the discovery/pre-clinical phase of new drug programs.</p></li></ol><p>These two sides of the bottleneck are clearly related, since early discovery would probably be more predictive if it could take better advantage of past clinical trial results. And, in fact, the way to solve the first direction is probably to solve the second direction. But let&#8217;s first look at each one individually.</p><h4>The Pre-clinical &#8594; Clinical Direction</h4><p>This is the more important one, since the main driver of clinical costs is high failure rates. Only around 1 in 10 drug candidates makes it through clinical trials, which means that to get one drug to market, a pharma company would expect to have to run clinical trials on 10 candidates. That means the cost of a <strong>successful</strong> clinical trial is effectively ten times the cost of running any single trial. If you could even just make it a 1 out of 9, you&#8217;d effectively reduce the cost by 10%. Make it 1 in 5 and you&#8217;ve cut the cost in half.</p><p>In other words, lower failure rates will lower costs faster than pretty much anything you can do to make individual trials cheaper.</p><p>It&#8217;s also important to note that early discovery/preclinical work is much less expensive than clinical. So if there were more experiments that teams could run in vitro or even in vivo to improve their clinical success rate, they would absolutely be doing that already.</p><p>In other words, this is a fundamental limitation of the tools we have available for early discovery, not a matter of bad decision making.</p><h4>The Clinical &#8594; Pre-Clinical Direction</h4><p>The nature of pre-clinical (in vitro and in vivo) vs clinical data is the ultimate quality vs quantity trade-off.</p><p>In early discovery, you can quickly and inexpensively (at least relative to clinical trials) generate data on millions of drug candidates. That&#8217;s orders of magnitude more drugs than have been through clinical trials in the entire history of modern drug discovery.</p><p>If you narrow it down to clinical trial data on the disease you&#8217;re working on, you might have a few dozen related clinical trials, if you&#8217;re lucky. Then narrow it to the target you&#8217;re looking at and you might still have a few if it isn&#8217;t a novel target. But then, thanks to the patent system you probably don&#8217;t want to screen structures or sequences anywhere close to this small handful of remaining data points.</p><p>So even if you do have relevant clinical trial data (which you often don&#8217;t) and you could train a reasonable model on that small handful of data points (which you probably can&#8217;t), all the points that you would want predictions about would be so far from the training data that predictions would be completely untrustworthy.</p><p>Of course, there&#8217;s also data from patient samples collected during treatment rather than from clinical trials. This data can be useful for learning about human biology in general, but it generally can&#8217;t be applied directly to predicting clinical success of novel drug candidates.</p><h4>Foundation Models to the Rescue?</h4><p>So the situation is really really bad. But I don&#8217;t think it&#8217;s hopeless.</p><p>To understand why, I want to further break down the pre-clinical &#8594; clinical direction into two pieces. In other words, there are two reasons why in vitro assays are not predictive of clinical outcomes:</p><ol><li><p>Cells in in vitro assays react fundamentally differently from how they react in the body, so there is a fundamental limit on what can be extrapolated from them.</p></li><li><p>We just don&#8217;t know (yet) how to interpret all the information that could in principle be extrapolated into clinically relevant data.</p></li></ol><p>We can&#8217;t do anything about the first one, but it&#8217;s probably less of an issue than you might think. Even if cells react differently in vitro, there&#8217;s still information in how they react. You just have to know what to look for.</p><p>On the other hand, if you look at how we collect, analyze and interpret data today, it&#8217;s clear that part two is a big factor. Most analysis looks for relatively simple signals based on a simplified, mechanistic understanding of the in vitro biology. </p><p>However, computational biologists have been introducing AI/ML models that replace these intermediate mechanistic goals with learned phenotypes/etc. and they usually find signal that wasn&#8217;t visible before.</p><p>Today, we&#8217;re seeing increasingly powerful bio foundation models that can integrate multiple kinds of data, potentially including readouts from in vitro, in vivo and in human samples to extract even more subtle signals. And I think we&#8217;ll see more of these data sources becoming assays that are deliberately designed to maximize the clinically relevant signal that a foundation model or a purpose-built AI/ML model can extract. (Addressing both factors at once.)</p><p>This is why I&#8217;m optimistic that AI will be able to open up the clinical information bottleneck. It won&#8217;t be easy, and it won&#8217;t happen quickly. But I think it&#8217;s possible, and that there&#8217;s a much clearer path to get there than there was even a few years ago.</p><div><hr></div><p>Thanks for reading! I help biopharma teams adopt consistent, reliable data handling practices so they can trust every conclusion. Check out <a href="https://merelogic.net/">merelogic.net</a> to learn more.</p>]]></content:encoded></item><item><title><![CDATA[Paired assays and the clinical information bottleneck]]></title><description><![CDATA[Here&#8217;s another one in the category of &#8220;assay types matter less than you might think&#8221;: NOETIK recently published a blog post announcing that their new foundation model, TARIO-2 can predict spatial transcriptomics values from routinely collected tumor histology images.]]></description><link>https://scalingbiotech.substack.com/p/paired-assays-and-the-clinical-information</link><guid isPermaLink="false">https://scalingbiotech.substack.com/p/paired-assays-and-the-clinical-information</guid><dc:creator><![CDATA[Jesse Johnson]]></dc:creator><pubDate>Wed, 13 May 2026 14:45:45 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Here&#8217;s another one in the category of &#8220;assay types matter less than you might think&#8221;: NOETIK recently published <a href="https://www.noetik.blog/p/tario-2-a-whole-transcriptome-foundation">a blog post announcing</a> that their new foundation model, TARIO-2 can predict spatial transcriptomics values from routinely collected tumor histology images. This is exciting because it creates the possibility of using this highly available, inexpensive data source for better inference, or even for model training. But for me, the bigger story is how it opens the door to addressing what I call the clinical information bottleneck (which I <a href="/__u/scalingbiotech.substack.com/p/pharma-partnerships-are-the-future">wrote about a few months ago</a>).</p><p>In this post, I&#8217;ll describe what this model does, explain what it has to do with the clinical information bottleneck, then give a high-level overview of how the model works.</p><h4>Paired assays</h4><p>NOETIK&#8217;s new model understands two types of data: </p><ol><li><p>Routinely collected histology sample images that use a staining method called H&amp;E (Hematoxylin and Eosin). You take a biopsy sample, put a thin slice of it onto a slide, stain it, then run it through a digital microscope to get an image. It&#8217;s quick and relatively inexpensive, and basically any patient with a solid tumor in the last few decades has gotten at least one of them.</p></li><li><p>Spatial transcriptomics data from these same thin tumor slices. This time, instead of the digital microscope, there&#8217;s a long, complex and expensive process involving a sequencer, that ends up with the two dimensional coordinates of each RNA molecule in a large sample of the RNA from the slice. That&#8217;s obviously much more useful data, but much more expensive to collect, so there isn&#8217;t much of it.</p></li></ol><p>What NOETIK has shown is that you can use TARIO-2 to take an H&amp;E image, split it up into small squares, and predict the gene expression levels in each square. Not quite the same as the two-dimensional coordinates of each molecule, but close enough for all intents and purposes.</p><p>This type of prediction doesn&#8217;t isn&#8217;t an existing step in you&#8217;re typical drug discovery pipeline, but if you&#8217;re building a model or doing some other sort of analysis using spatial transcriptomics data, this gives you a vastly larger amount of (at least slightly lower quality) data to do it with.</p><h4>The Clinical Information Bottleneck</h4><p>I originally got excited about this result because it&#8217;s a great example of how foundation models can make any old data you have laying around more valuable than you think (particularly if it&#8217;s properly organized based on your <a href="https://merelogic.net/data_operations_plans/">data operations plan</a>). But then I started thinking about the information bottleneck between the early discovery and clinical phases that <a href="/__u/scalingbiotech.substack.com/p/pharma-partnerships-are-the-future">I wrote about back in September</a>.</p><p>The idea of the bottleneck is that you could run every available in vitro and in vivo assay on a pre-clinical drug candidate, spend billions running every conceivable experiment in the lab before filing your IND, and your drug would still have only a 10% chance of passing clinical trials and making it to market. So from a clinical standpoint, in vitro and even in vivo data isn&#8217;t quite useless, but it&#8217;s unnervingly close to it.</p><p>Meanwhile, in-human clinical data is very sparse. If you&#8217;re looking at a new target, or in a new area of chemical space, there will be little if any clinical data that&#8217;s close enough to the space you&#8217;re exploring. Little enough that it&#8217;s usually not even worth looking at.</p><p>But what if you could predict clinical results from a paired in vitro assay? In other words, if you could take a readout that is commonly collected in clinical trials, or just clinical practice, then run an in vitro assay that is appropriately paired with that clinical data, you could potentially train a model that predicts the clinical outcome directly from the in vitro assay. This is in contrast to a standard in vitro readout where you decide what a hit looks like based on some combination of your mechanistic understanding of the assay, a positive control, and vibes.</p><p>Pairing a clinical assay with an in vitro assay is a big step from what TARIO-2 does, since both H&amp;E and spatial transcriptomics come from clinical tumor samples. But NOETIK has also shown that TARIO-2 can <a href="https://www.noetik.blog/p/perturb-mars-reading-mouse-experiments">use H&amp;E images from in vivo mouse samples</a> to predict the transcriptomic profiles of human samples under the same perturbation, which is one step closer to in vitro. And previously I wrote about a company called Axiom that paired an in vitro assay with clinical data to <a href="/__u/scalingbiotech.substack.com/p/the-race-begins-for-an-fda-approved">predict liver toxicity better than animal models</a>. So maybe it&#8217;s not completely crazy&#8230;? </p><h4>Two-dimensional Transformers FTW</h4><p>TARIO-2 is a transformer model that takes as input a spatial transcriptomics readout paired with an H&amp;E image of the same sample. It&#8217;s goal is output the same thing a predicted spatial readout and H&amp;E image that match the original. Not that exciting on its own, but it gets interesting when you mask out part of the data in the inputs. The model is trained to still predict the full original data, filling in gaps based on the parts that you didn&#8217;t mask out.</p><p>What the folks at NOETIK discovered is that if you mask out all the transcriptomics data, and just give it the H&amp;E image, it can fairly accurately reproduce the original transcriptomics. Or more imporantly, if you only have the H&amp;E image, and you pretend the transcriptomics is just masked out, the model can predict what the transcriptomics data would&#8217;ve looked like.</p><p>The model is based on a transformer, so to explain how it works, I&#8217;m going to start where <a href="/__u/scalingbiotech.substack.com/p/how-to-build-a-bio-foundation-model">my post on bio foundation models</a> left off.</p><p>The first thing we need to do is define our tokens. For this, we&#8217;ll first cut the image into a grid of smaller blocks. For each block, we&#8217;ll define two tokens: One for that portion of the image, and one for the gene expression in that portion of the sample. We&#8217;ll annotate the tokens by adding vectors to indicate both what kind of token they are and where they sit in the coordinate grid.</p><p>Once the tokens are embedded, we run the transformer to update their embeddings based on the attention algorithm. Then we&#8217;ll run the final embeddings through another set of neural layers to map them back to either an image or a set of expression levels.</p><p>To mask out data, you just replace the appropriate tokens with zero vectors, then let the attention algorithm pull them to where the model thinks they would&#8217;ve been if they weren&#8217;t masked. The final neural layers will then map them to the predicted image and expression levels. If you do this with all the transcriptomics tokens set to zero, this will recover the predicted transcriptomics values based on the embeddings of the H&amp;E tokens.</p><h4>Conclusion</h4><p>There are a lot of interesting takeways from NOETIK&#8217;s work on TARIO-2, but for me the one that stands out is this: They&#8217;ve shown that the information in an uncommon, expensive spatial transcriptomics assay can also be found in a cheap, ubiquitous imaging assay, and they&#8217;ve demonstrated a way to extract this information. If others can take this general idea and run with it, who knows what they may be able to accomplish.</p><div><hr></div><p>Thanks for reading! I help biopharma teams integrate foundation models into their discovery and development pipelines. Check out <a href="https://merelogic.net/">merelogic.net</a> to learn more.</p>]]></content:encoded></item><item><title><![CDATA[Say good bye to single-assay training datasets]]></title><description><![CDATA[A few times in recent posts, I&#8217;ve claimed (without evidence) that bio foundation models can be trained on readouts from different kinds of assays to create a single model that&#8217;s more powerful than you could&#8217;ve gotten from any one assay alone.]]></description><link>https://scalingbiotech.substack.com/p/say-good-bye-to-single-assay-training</link><guid isPermaLink="false">https://scalingbiotech.substack.com/p/say-good-bye-to-single-assay-training</guid><dc:creator><![CDATA[Jesse Johnson]]></dc:creator><pubDate>Wed, 06 May 2026 14:45:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A few times in recent posts, I&#8217;ve claimed (without evidence) that bio foundation models can be trained on readouts from different kinds of assays to create a single model that&#8217;s more powerful than you could&#8217;ve gotten from any one assay alone. I like this idea because it&#8217;s a compelling reason to invest in building a <a href="https://merelogic.net/data_operations_plans/">Data Operations Plan</a> to keep track of all that data. But is it true?</p><p>Well, <a href="https://www.biorxiv.org/content/10.1101/2024.08.12.607533v3">this paper suggests it is</a>, so that&#8217;s what I&#8217;m going to write about this week. </p><p><em>*** But first, a quick reminder to check out the <a href="https://courses.scriptome.ai/">class for biologists who want to start using bio foundation models</a> that I&#8217;m co-developing with Scriptome.AI. Participants will get hands-on experience running a variety of bio foundation models. ***</em></p><p>The first author on this paper, Yuge Ji, is the founder of Reflector Bio, one of the Virtual Cell Model developers profiled in my <a href="https://merelogic.net/virtual_cell_models/academics">Buyers Guide to Virtual Cell Models</a>. Reflector is built around the model in this paper, which doesn&#8217;t quite fit the framework for transformer-based virtual cell models that <a href="/__u/scalingbiotech.substack.com/p/on-transformer-based-virtual-cell">I wrote about a few weeks ago</a>. But it makes largely the same kinds of predictions as a virtual cell model, so I think it counts. Here&#8217;s how it works:</p><h4>Inputs and Outputs</h4><p>The model predicts outcomes of different kinds of in vitro perturbational experiments. Each input to the model consists of three parts:</p><ol><li><p>An environment token indicating the cell line that the experiment was run in.</p></li><li><p>One or more treatment tokens indicating what was done to the cells. In the paper, these can either be small molecules or genes for CRISPR knock-downs, but the approach could be extended to other types.</p></li><li><p>A token indicating the type of readout/phenotype that was measured.</p></li></ol><p>The output is a set of vectors that capture the readout values from the different types of readouts/assays. The paper treats the output as a single vector that you get by concatenating all the readout vectors.</p><p>The model is based on a transformer. You can read about how they work in <a href="/__u/scalingbiotech.substack.com/p/how-to-build-a-bio-foundation-model">my post from back in March</a>. This particular transformer starts with the three from the input, and adds a fourth class token that summarizes the other three. The output/prediction comes from running the transformer algorithm to re-embedd all the tokens, then passing the final embedding of the class token through additional neural net layers to map it to the predicted readout.</p><p>This approach sits one level of abstraction above typical virtual cell models, since instead of an initial token for each gene, the environment token encapsulates all the initial expression levels. Similarly, the output is derived from a single token (the class token) rather than the final embeddings of the per-gene tokens.</p><h4>Tokens and embeddings</h4><p>What makes this approach really interesting is that all three types of input tokens are embedded in the same embedding space, even though they represent very different kinds of concepts.</p><p>The environment and treatment tokens do this with a two step process: First, each token is embedded in a type-specific intermediate space. Second, each type-specific space is mapped into the common embedding space. That way, each intermediate space can have an arbitrary number of dimensions, based on the type of token.</p><p>Cell type tokens are first embedded in a space with one vector for each gene, based on its typical expression level for the cell type. Small molecules are embedded in a space using standard structure-based embeddings found in RDKIT. Genes for CRISPR knock-downs are embedded based on a combination their co-expression matrix and their coding sequence.</p><p>Both the intermediate, type-specific embeddings and the maps into the common space are learned as part of training the overall model.</p><p>The phenotype/assay type token doesn&#8217;t use a two-step process because there isn&#8217;t an intermediary space that makes sense for it. Instead, it&#8217;s a direct embedding, learned as part of overall training.</p><h4>Using the Model</h4><p>Once you&#8217;ve trained this model on all your experimental data, you can make new predictions by swapping out one or more of the three types of input tokens from an existing experiment:</p><ul><li><p>Swap out the environment/cell line token with all the other known cell lines to predict how the treatment will affect other tissues.</p></li><li><p>Swap out the phenotype/assay type token to predict what other readouts for the same experiment would&#8217;ve looked like.</p></li><li><p>Swap out the treatment token to do virtual screening trained on all your past assays, along with public data across multiple readout types.</p></li></ul><p>Because this model, like all models, will be most accurate closer to the training data, it may not be a great idea to swap out more than one token at a time. (And for treatments, you probably don&#8217;t want to get too far away from known molecular structures&#8230;) But the paper shows promising results for swapping out each single type of token.</p><h4>Conclusion</h4><p>What I like about this approach is that it&#8217;s relatively simple but demonstrates real potential based on just publicly available data. With a few more iterations of the model architecture and a bunch of proprietary data (like, say, what <a href="/__u/scalingbiotech.substack.com/p/im-confused-about-eli-lillys-strategy">Lily Tune Labs</a> is collecting) this could be the start of something much bigger.</p><p>So maybe a few years from now, it won&#8217;t sound so crazy to say that all those billions of dollars of data that biopharma teams generate each year, then hide away on hard drives and disorganized shared drives, could actually train the foundation models that drive the next wave of drug discovery.</p><div><hr></div><p>Thanks for reading! I help biopharma teams integrate foundation models into their discovery and development pipelines. Check out <a href="https://merelogic.net/">merelogic.net</a> to learn more.</p>]]></content:encoded></item><item><title><![CDATA[Your data is probably (secretly) biased]]></title><description><![CDATA[A few weeks ago, I wrote about a paper where Leash Bio showed (among other things) that the data that is commonly used for benchmarking small molecule binding models is biased by the design quirks of individual chemists. This week, I&#8217;ve got another paper about biased benchmarking data, this time from the startup]]></description><link>https://scalingbiotech.substack.com/p/your-data-is-probably-secretly-biased</link><guid isPermaLink="false">https://scalingbiotech.substack.com/p/your-data-is-probably-secretly-biased</guid><dc:creator><![CDATA[Jesse Johnson]]></dc:creator><pubDate>Wed, 29 Apr 2026 14:45:49 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A few weeks ago, I wrote about a paper where Leash Bio showed (among other things) that the data that is commonly used for benchmarking small molecule binding models is <a href="/__u/scalingbiotech.substack.com/p/leash-bio-shows-that-data-beats-models">biased by the design quirks of individual chemists</a>. This week, I&#8217;ve got another paper about biased benchmarking data, this time from the startup <a href="https://deepflare.ai/">Deepflare</a>, looking at <a href="https://www.biorxiv.org/content/10.64898/2026.03.30.710191v2">models for predicting peptide cell markers for cancer immunotherapy</a>. In this post, I&#8217;ll explain what these experiments and models look like, and how this bias sneaks into both the benchmarks and the models.</p><p><em>*** But first, if you&#8217;re worried about bias in your own data, it might be time to define a formal Data Operations Plan. <a href="https://merelogic.net/data_operations_plans/">Check out my guide here</a> or <a href="mailto:jesse@merelogic.net">send me an email</a> to get started. ***</em></p><h4>Predicting cell surface markers</h4><p>First, the science: Every cell in your body has a bunch of peptides (small fragments of proteins) sticking out of its surface. These come from proteins inside the cell that get broken up into pieces. Some of the pieces get transported to the cell wall and pushed through, while others don&#8217;t. Your immune system uses these protein fragments to tell the difference between healthy human cells and foreign cells, as well as to detect cancer cells. (Note: I am not a biologist.)</p><p>Cancer cells are defined by mutations that create protein variants that don&#8217;t exist in healthy cells. If you can figure out which fragments of these proteins will get pushed onto the surface of the cell (and which don&#8217;t appear on healthy cells), you can try to develop a drug that attaches to those peptides and kills any cells where they appear. This is how immuno-oncology works.</p><p>The goal, then, is to train a model to take a sequence for a protein variant and predicts which fragments of it will appear on the surface of the cell.</p><h4>Measuring cell surface markers</h4><p>To build a model that predicts cell surface markers, you need training data, which means you need a way to identify the surface peptides. You&#8217;re going to have to trust me that there&#8217;s a process to get them off the surface of the cell and into a solution because I&#8217;m not qualified to explain it. (By which I mean I don&#8217;t understand it - again, not a biologist.)</p><p>Once that&#8217;s done, a mass spectrometer (mass spec) can be used to identify the proteins in the solution. This breaks up each peptide into a bunch of smaller pieces, then identifies the molecular weight of each piece. In the readout from the mass spec, each piece shows up as a peak along the scale of masses. Each protein shows up as a collection of peaks depending on its sequence.</p><p>In the mass spec readout, these peaks are all overlayed against each other, and the purpose of mass spec analysis is to deconvolute them: figure out what distribution of proteins would&#8217;ve created the overlayed collection of peaks that appear in the data. It&#8217;s possible to do this completely from scratch, but it&#8217;s much easier (and more common) to start with a &#8220;candidate list&#8221; of peptides that you expect to find. The mass spec will only look for the peptides on this candidate list.</p><p>The problem is what happens when you both use predictive models to generate the candidate list, then use the resulting data to train and benchmark more models.</p><h4>Compounding Bias</h4><p>You may see where this is going, but let&#8217;s follow it through, step by step: </p><p>First, we generate Dataset A, a list of peptides on the surfaces of different sample cells, without using a candidate peptide list. Then we train and evaluate Model A on this data.</p><p>So far so good.</p><p>Next, we generate more experimental data, Dataset B, but to make the mass spec analysis easier, we use the top peptides predicted by Model A as the candidate list. According to the Deepflare paper, we would typically use the top 2% of predictions from Model A. That means any peptide that isn&#8217;t on the candidate list won&#8217;t show up in our analysis, even if it was physically present in the sample.</p><p>We&#8217;ll then train and evaluate Model B on Dataset A and Dataset B together.</p><p>The problem, of course, is that Dataset B doesn&#8217;t include any peptides that were outside the top 2% of the predictions from Model B. In other words, if Model A was biased against certain types of sequences, then Dataset B will also be biased against them. So Model B will be even more biased against them then Model A was.</p><p>Moreover, because we&#8217;re evaluating/benchmarking Model B against a dataset with this bias, we won&#8217;t even notice it.</p><p>In fact, if you were to benchmark Model A again, based on the Dataset B (with or without Dataset A), Model A would look even better, since you would find more true positives while finding only a tiny fraction of the new false negatives.</p><p>Now repeat this over and over again: The models become increasingly biased while the benchmark values look better and better.</p><h4>The Immune Epitope Database</h4><p>The <a href="https://www.biorxiv.org/content/10.64898/2026.03.30.710191v2">Deepflare paper</a> notes that most models for this problem are trained on and benchmarked against mass spec data in the Immune Epitope Database (IEDB). They carefully reviewed the methods of 37 of the studies that produced this data and found that only around 44% of the data could be verified to use unbiased measurements, while the rest was analyzed using previously trained models.</p><p>They also ran simulations, using the clean portion of the data to verify that using candidate lists pulled from earlier models does indeed bias later models. (Not that anyone should doubt that&#8230;)</p><p>They argue that the steady increase of reported accuracy of previously defined models could potentially be a result of this compounding bias rather than actual improvements in accuracy. And, of course, they introduce their own model called DeepMHCflare, trained on only the &#8220;clean&#8221; portion of IEDB data and validated its results using an in vivo study of a cancer vaccine.</p><h4>Subtle Bias</h4><p>This particular cycle of compounding bias is specific to how mass spec data is analyzed. But if you look at it alongside the bias in small molecule binding datasets I wrote about a few weeks ago, it&#8217;s clear that bias in general can sneak into biological data in lots of different and subtle ways.</p><p>Both these examples could&#8217;ve easily gone unnoticed for years or decades, silently poisoning multi-million dollar drug discovery programs. And we might never have noticed the programs that failed and the drug candidates that were passed over as a result.</p><p>The only way to identify and fight this kind of bias is to understand both the technical and scientific aspects of data collection, and to create an environment where you can explore the data with as little friction as possible. Otherwise, it&#8217;s easy to be so focused on just getting results that you never step back to see the bigger picture.</p><div><hr></div><p>Thanks for reading! I help biopharma teams integrate foundation models into their discovery and development pipelines. Check out <a href="https://merelogic.net/">merelogic.net</a> to learn more.</p>]]></content:encoded></item><item><title><![CDATA[Making data AI-ready isn't cheap]]></title><description><![CDATA[Often, the biggest expense involved in integrating a bio foundation model into a biopharma pipeline, after the lab work itself, is getting the data into a model-ready form.]]></description><link>https://scalingbiotech.substack.com/p/make-data-ai-ready-isnt-cheap</link><guid isPermaLink="false">https://scalingbiotech.substack.com/p/make-data-ai-ready-isnt-cheap</guid><dc:creator><![CDATA[Jesse Johnson]]></dc:creator><pubDate>Wed, 22 Apr 2026 14:45:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!583h!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbaa306-035a-45a0-aa27-dc333e9d2848_498x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Often, the biggest expense involved in integrating a bio foundation model into a biopharma pipeline, after the lab work itself, is getting the data into a model-ready form. I&#8217;ve been rewriting (and adding figures to) my <a href="https://merelogic.net/data_operations_plans/">guide to significantly reducing this cost</a>, which got me thinking about all the reasons this is so hard for most biopharma labs. So this week I want to explain why this is such a hard problem in the first place, and how I&#8217;m trying to make it easier. (Next week, I&#8217;ll get back to the models themselves&#8230;)</p><p><em>*** But first, if you haven&#8217;t already, check out this upcoming <a href="https://courses.scriptome.ai/">class for biologists who want to start using bio foundation models</a> that I&#8217;m co-developing with Scriptome.AI. This will be an eight week hybrid in-person/online class in Boston/Cambridge where participants will get hands-on experience running a variety of bio foundation models.***</em></p><h4>Unspoken assumptions are often wrong</h4><p>Whenever a lab experiment completes and it&#8217;s time to gather up all the data (instrument readouts) and metadata (ELN notes, plate maps, etc.) there&#8217;s an endless list of options for how one could do this. Some of these decisions are really important, like where you put the plate map so you can find it later. Others, like whether or not to capitalize column names, are arbitrary.</p><p>Each bench scientist, if left to their own devices, will make different decisions about both the important decisions and the arbitrary ones, which is a problem when even arbitrary decisions need be made consistently.</p><p>But beyond just consistency, there are also important requirements from data scientists and computation biologists that go unstated because both teams just assume they&#8217;re so obvious that the other team must be on the same page&#8230; even though they often aren&#8217;t.</p><p>For example, I&#8217;ve known bench scientists who drafted their lab notes in word docs and only copy/pasted them into the ELN once they were perfect, because the ELN was supposed to be the golden record. And I&#8217;ve seen computational biologists shoot steam out their ears hearing this because they want that information to be in the ELN as soon as physically possible.</p><p>These are organizational and operational problems that can&#8217;t be fixed by any amount of AI or other tech.</p><h4>No one&#8217;s responsibility</h4><p>Someone just needs to make these decisions and communicate them to the rest of the team. However, in most labs, particularly early stage biotech startups, there isn&#8217;t a person who can do that.</p><p>Individual bench scientists don&#8217;t feel empowered to tell the rest of the team how to do their work, and don&#8217;t understand the downstream requirements well enough, anyway.</p><p>Data scientists and computational biologists don&#8217;t think they can tell the bench team how to do their jobs, and they often don&#8217;t have the bandwidth to understand the lab processes, document what they need, and advocate for those changes.</p><p>Anyone in leadership who&#8217;s accountable for both bench and digital teams doesn&#8217;t want to micro-manage these processes, even if they had time to get into the weeds of metadata handoffs (which they don&#8217;t).</p><p>At larger organizations, someone like a lab informatics engineer would be responsible for defining these processes. But smaller biotech startups generally can&#8217;t afford a dedicated informatics engineer. And even at larger organizations that do have them, they&#8217;re often focused on core, high-throughput assays rather than the long tail of infrequent protocols.</p><h4>Communicating the problem</h4><p>To solve this problem, I recently started creating Data Operations Plans for biopharma labs. (I had been doing projects like this a few years ago, then moved on to other kinds of work. But now that I&#8217;m seeing how important this is for AI adoption, I&#8217;m back to it.)</p><p>Since the pain of inconsistent data collection practices is often felt most acutely by the computational biologists and bench scientists who aren&#8217;t the ones signing the contracts, I wanted to give them a resource to help communicate to leadership why a project like this is necessary. So the first half of my guide to <a href="https://merelogic.net/data_operations_plans/">Data Operations Plans for Early Discovery Labs</a> is all about why building one of these plans is worth the up-front cost. </p><p>My goal is to give the folks feeling the pain the arguments to advocate for the changes that will address these problems. I just rewrote it to be (hopefully) more compelling, and added a bunch of graphics/figures to make it easier to read. I want this to be something you can share with your boss if you&#8217;re feeling the pain.</p><h4>Building a (financially) viable solution</h4><p>But I also recognize that budgets are extremely tight these days, particularly for early stage startups. Every cent needs to be carefully spent. That&#8217;s why I only build Data Operations Plans as a fixed-price service with a guaranteed outcome and a clearly defined schedule. Because my clients know exactly how much the project will cost them, regardless of the number of hours on my end, they don&#8217;t have to worry about unexpected budget overruns.</p><p>The price may still be high for some startups, but it&#8217;s a fraction of what it would cost to hire a full time lab informatics engineer, and an even smaller fraction of what they&#8217;re spending to generate data. Think of a Data Operations Plan as an insurance policy on that data.</p><p>As VCs increasingly recognize how bio foundation models are increasing the importance of data as an asset that drives biotech valuations, they&#8217;re going to look for companies that have a concrete data strategy. Having a Data Operations Plan is the first step in building that story. So I&#8217;m confident that the startups that are serious about building a moat with their data will find a way to fit it into their budget.</p><div><hr></div><p>Thanks for reading! I help biopharma teams integrate foundation models into their discovery and development pipelines. Check out <a href="https://merelogic.net/">merelogic.net</a> to learn more.</p>]]></content:encoded></item></channel></rss>