<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Machine Learning Engineer]]></title><description><![CDATA[Join over 15,000 ML professionals and enthusiasts who receive weekly curated articles & tutorials on production Machine Learning. Obtain insights on real-world ML best practice encompassing key areas in Model Monitoring, MLOps, DataOps, AIOps + beyond 🚀]]></description><link>https://machinelearning.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png</url><title>The Machine Learning Engineer</title><link>https://machinelearning.substack.com</link></image><generator>Substack</generator><lastBuildDate>Thu, 03 Sep 2026 17:57:02 GMT</lastBuildDate><atom:link href="/__u/machinelearning.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[The Institute for Ethical AI & Machine Learning]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[machinelearning@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[machinelearning@substack.com]]></itunes:email><itunes:name><![CDATA[Alejandro Saucedo]]></itunes:name></itunes:owner><itunes:author><![CDATA[Alejandro Saucedo]]></itunes:author><googleplay:owner><![CDATA[machinelearning@substack.com]]></googleplay:owner><googleplay:email><![CDATA[machinelearning@substack.com]]></googleplay:email><googleplay:author><![CDATA[Alejandro Saucedo]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Issue #402 - The ML Engineer]]></title><description><![CDATA[How to Write an Agent Skill, NVIDIA Buys Hugging Face, Uber on the Agentic Software Factory, GLM-5.3 Open Weights Are Out, OpenAI's Jalape&#241;o Inference Chip + more &#128640;]]></description><link>https://machinelearning.substack.com/p/issue-402-the-ml-engineer</link><guid isPermaLink="false">https://machinelearning.substack.com/p/issue-402-the-ml-engineer</guid><dc:creator><![CDATA[Alejandro Saucedo]]></dc:creator><pubDate>Sun, 30 Aug 2026 17:24:58 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><blockquote><p><em>How to Write an Agent Skill, NVIDIA Buys Hugging Face, Uber on the Agentic Software Factory, GLM-5.3 Open Weights Are Out, OpenAI&#8217;s Jalape&#241;o Inference Chip + more &#128640;</em></p></blockquote><p>Thank you for being part of over <a href="https://ethical.institute/newsletter/">70,000+ ML professionals and enthusiasts</a> who receive weekly articles &amp; tutorials on Machine Learning &amp; MLOps &#129302; You can join the newsletter <a href="https://ethical.institute/newsletter/">/newsletter/</a> &#11088;</p><p>If you like the content please support the newsletter by sharing with your friends via &#9993;&#65039; Email, &#128038; <a href="http://twitter.com/intent/tweet?text=Great%20Newsletter%20on%20Machine%20Learning!%20%20/newsletter/">Twitter</a>, &#128188; <a href="http://www.linkedin.com/shareArticle?mini=true&amp;url=/newsletter/&amp;title=The%20Machine%20Learning%20Engineer%20Newsletter&amp;summary=Curated%20news%20about%20machine%20learning%20operations,%20reproducibility,%20explainability%20and%20beyond">Linkedin</a> and &#128213; <a href="http://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fethical.institute%2Fmle.html">Facebook</a>!</p><div><hr></div><h2>This week in ML Engineering:</h2><ul><li><p>New post on <a href="https://ethical.institute/blog/how-to-write-an-agent-skill/">How to Write an Agent Skill</a></p></li><li><p>NVIDIA <a href="https://www.cnbc.com/2026/08/27/nvidia-hugging-face-acquisition.html">Buys Hugging Face</a></p></li><li><p>Uber on <a href="https://www.uber.com/gb/en/blog/efficient-software-factory/">the Agentic Software Factory</a></p></li><li><p>GLM-5.3 <a href="https://z.ai/blog/glm-5.3">Open Weights Are Out</a></p></li><li><p>OpenAI&#8217;s <a href="https://openai.com/index/openai-broadcom-jalapeno-inference-chip/">Jalape&#241;o Inference Chip</a></p></li><li><p>Open Source <a href="https://ethical.institute/open-source/production-ml-list/">ML Frameworks</a></p></li><li><p>Awesome AI Guidelines <a href="https://ethical.institute/open-source/ai-guidelines/">to check out this week</a></p></li><li><p>+ more &#128640;</p></li></ul><div><hr></div><p><strong><a href="https://ethical.institute/blog/how-to-write-an-agent-skill/">How to Write an Agent Skill</a></strong></p><p>This week I published a new post on how to write agent skills! And this comes with a skill... to write skills &#128579; Basically a post on how I stopped writing bash scripts, and the learnings and best practices on what we can call... non-deterministic programming?</p><p>What is interesting is that I still write it as I was writing code, but in English (haven&#8217;t tried Spanish).</p><p>I propose 6 principles that I believe lead to high quality skills:</p><ol><li><p>The description is a router, so it says when to use the skill in the words you&#8217;d actually type</p></li><li><p>The body is steps, and anything that isn&#8217;t a step gets deleted</p></li><li><p>Scripts carry the fixed sequences, and the agent carries the judgement calls</p></li><li><p><a href="https://www.linkedin.com/redir/invalid-link-page?url=http%3A%2F%2FSKILL%2emd">SKILL.md</a> holds only what every run needs</p></li><li><p>Verification is a step of the procedure, for the output and for the skill itself</p></li><li><p>(Optional) A human feedback loop, where the skill ends by reporting its own friction upstream</p></li></ol><p>I have to say also that principle 6 is the one that has felt magical; a skill that reports its own friction is one that actually gets better every week.</p><p>The post goes through each one with a good and a bad example, plus the real skills we open sourced along the way in the <a href="https://ethical.institute/blog/announcing-the-agent-skills-marketplace/">Agent Skills Marketplace</a>. Check it out and let me know what you think!</p><div><hr></div><p><strong><a href="https://www.cnbc.com/2026/08/27/nvidia-hugging-face-acquisition.html">NVIDIA Buys Hugging Face</a></strong></p><p>Huge news for the European AI ecosystem &#128640; NVIDIA has agreed to buy Hugging Face for $12.9 billion!</p><p>It seems there is still no official confirmation from NVIDIA or HF, however CNBC&#8217;s own sources confirm the talks, and if it completes this would put the de-facto home of open source AI, with its model hub, Transformers library and datasets, under the ownership of the company that sells the hardware it all runs on.</p><p>We have to say this is a huge win for Europe; Hugging Face was founded by French founders and has grown into one of the most important platforms in AI, and a $12.9 billion outcome at the centre of NVIDIA&#8217;s open-model strategy is quite the milestone for the European ecosystem.</p><p>There are of course open questions, and the community is already asking whether NVIDIA-optimised models will get preferential treatment.</p><p>It will be interesting to see whether the open ethos that made Hugging Face what it is survives the integration.</p><div><hr></div><p><strong><a href="https://www.uber.com/gb/en/blog/efficient-software-factory/">Uber on the Agentic Software Factory</a></strong></p><p>The Software Factory era is here... Uber has published (yet another) one of the most detailed pictures on agentic engineering, and how it&#8217;s embedded in their dev lifecycle.</p><p>Over 70% of Uber&#8217;s pull requests are now attributed to local or cloud agents, with 3,600+ agent skills across the SDLC running more than 30K executions per day. Weekly active users grew 7x since February, and agentic requests 9.4x.</p><p>The really interesting part is the cost engineering; they decompose spend into users x sessions x turns x requests x tokens x price, and attack each factor: prompt-cache TTL tuning, CLI-resolved MCP calls, and code-mode batching that they claim saves 55-100% of tool-execution tokens. The result is cost per 1,000 requests down 34%, and cost per session down 52% from peak.</p><p>Similar to what we&#8217;ve seen on mapping engineering context, Uber grounds their agents with an AI Context Graph: 24M nodes and 80M edges connecting 30+ internal systems, exposed through 1,000+ MCP tools.</p><p>For production ML practitioners this is one of the most concrete looks yet at what agentic engineering looks like at scale, and the honest caveat that your mileage may vary makes it more credible rather than less!</p><div><hr></div><p><strong><a href="https://z.ai/blog/glm-5.3">GLM-5.3 Open Weights Are Out</a></strong></p><p>The GLM-5.3 open weights are out! <a href="http://Z.ai">Z.ai</a> launched the model earlier this month and has now released the weights on <a href="https://huggingface.co/zai-org/GLM-5.3">Hugging Face</a> after what they describe as safety evaluation and hardening:</p><p>From their report it seems their improvements come from scaled post-training on synthesized long-horizon environments; they claim open-source SOTA on Terminal Bench 3.0 and Agents Last Exam + and better results than Claude Opus 4.8 on their in-house code bench at less than half the output tokens.</p><p>The part everyone is talking about is the emergent cyber capability: as post-training scaled, exploitation skills grew faster than expected, and working with security teams the model surfaced 2,436 vulnerabilities across 269 open source projects, the oldest dating back to 1981, now tracked in a public disclosure ledger.</p><p>It is quite interesting to see a lab publish this openly, and it is clear that open-weights releases are now frontier-relevant for ML security as much as for coding.</p><div><hr></div><p><strong><a href="https://openai.com/index/openai-broadcom-jalapeno-inference-chip/">OpenAI&#8217;s Jalape&#241;o Inference Chip</a></strong></p><p>Can anyone actually challenge NVIDIA on inference? OpenAI and Broadcom said &#8220;hold my beer&#8221; with &#8220;Jalape&#241;o&#8221;, OpenAI&#8217;s first Intelligence Processor - these names are getting out of hand:</p><p>This chip was co-designed from a blank slate for LLM inference and taken from design to tape-out in nine months with help from OpenAI&#8217;s own models.</p><p>From early testing they report substantially better performance per watt than the current state of the art, and <a href="https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia">SemiAnalysis were invited into the lab</a> to benchmark it, measuring it ahead of every NVIDIA, AMD and Google chip they tested on perf per watt.</p><p>Worth noting the numbers were provided by OpenAI and the agentic long-context suites remain untested.</p><p>Deployment is planned at gigawatt scale from late 2026. Quite the week to be NVIDIA: buying the model hub while the chip moat gets its first real test!</p><div><hr></div><h2>Upcoming MLOps Events</h2><p>The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.</p><p>You can also find our upcoming events and past talk recordings on the <a href="https://ethical.institute/talks-and-events/">talks and events page</a>.</p><p><strong>Events we are speaking at this year:</strong></p><ul><li><p><a href="https://signalsconf.io/#tickets">Signals Conference</a> - September @ Berlin</p></li><li><p><a href="https://worldsummit.ai/">World Summit AI Europe</a> - September @ Amsterdam</p></li><li><p><a href="https://codetalks.com/">Code.Talks 2026</a> - November @ Hamburg</p></li></ul><p><strong>Other relevant events:</strong></p><ul><li><p><a href="https://www.ai-infra-summit.com/">AI Infra Summit 2026</a> - Sept @ California</p></li><li><p><a href="https://events.linuxfoundation.org/agntcon-mcpcon-europe/">AGNTCon + MCPCon Europe 2026</a> - Sept @ Amsterdam</p></li><li><p><a href="https://www.humanx.co/europe">HumanX Amsterdam 2026</a> - Sept @ Amsterdam</p></li><li><p><a href="https://mlopsworld.com/">MLOps World 2026</a> - Nov @ Austin</p></li></ul><p><strong>In case you missed our talks, check our recordings below:</strong></p><ul><li><p>The State of AI in 2025 - <a href="https://www.youtube.com/watch?v=v2LENQOG-Xg">WeAreDevelopers 2025</a></p></li><li><p>Prod Generative AI in 2024 - <a href="https://www.youtube.com/watch?v=0uJGmMZGUJE&amp;list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&amp;index=3">KubeCon AI Day 2025</a></p></li><li><p>The State of AI in 2024 - <a href="https://www.youtube.com/live/AtA2XXo_b5s">WeAreDevelopers 2024</a></p></li><li><p>Responsible AI Workshop Keynote - <a href="http://www.youtube.com/watch?v=57YpXjcj0Ho">NeurIPS 2021</a></p></li><li><p>Practical Guide to ML Explainability - <a href="http://www.youtube.com/watch?v=vq8mDiDODhc">PyCon London</a></p></li><li><p>ML Monitoring: Outliers, Drift, XAI - <a href="http://www.youtube.com/watch?v=QcevzK9ZuDg">PyCon Keynote</a></p></li><li><p>Metadata for E2E MLOps - <a href="https://www.youtube.com/watch?v=OSbH4dfswCY">Kubecon NA 2022</a></p></li><li><p>ML Performance Evaluation at Scale - <a href="http://www.youtube.com/watch?v=8ORl8lu1Eeo">KubeCon Eur 2021</a></p></li><li><p>Industry Strength LLMs - <a href="https://www.youtube.com/watch?v=RVUi_rAFfzU">PyData Global 2022</a></p></li><li><p>ML Security Workshop Keynote - <a href="http://www.youtube.com/watch?v=7XSy5aw8oU8">NeurIPS 2022</a></p></li></ul><div><hr></div><h2>Open Source MLOps Tools</h2><p>Check out the fast-growing ecosystem of production ML tools &amp; frameworks at <a href="https://ethical.institute/open-source/production-ml-list/">the github repository</a> which has reached over 20,000 &#11088; github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here&#8217;s a few featured open source libraries that we maintain:</p><ul><li><p><a href="https://github.com/besanson/sarc-governance/tree/main">SARC</a> - Provides wrappers for popular agentic frameworks to enable guardrails and constraints that are enforced through the flow.</p></li><li><p><a href="https://github.com/axsaucedo/kaos">KAOS</a> - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.</p></li><li><p><a href="https://github.com/KomputeProject/kompute">Kompute</a> - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.</p></li><li><p><a href="https://ethical.institute/open-source/production-ml-list/">Production ML Tools</a> - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.</p></li><li><p><a href="https://github.com/EthicalML/awesome-artificial-intelligence-regulation">AI Policy List</a> - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.</p></li><li><p><a href="https://github.com/EthicalML/awesome-production-genai/">Agentic Systems Tools</a> - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain</p></li></ul><p>Please do support some of our open source projects by sharing, contributing or adding a star &#11088;</p><div><hr></div><h2>About us</h2><p>The Institute for Ethical AI &amp; Machine Learning is a European research centre that carries out world-class research into responsible machine learning.</p><p><a href="https://ethical.institute/">Check out our website</a></p><p>&#9993;&#65039; Email, &#128038; <a href="http://twitter.com/EthicalML">Twitter</a>, &#128188; <a href="https://www.linkedin.com/company/the-institute-for-ethical-machine-learning/">Linkedin</a></p>]]></content:encoded></item><item><title><![CDATA[Issue 401 - The ML Engineer 🤖]]></title><description><![CDATA[Raschka on Claude&#8217;s Watermarking, The Netflix GenRec Paper, Speculative Decoding on AMD GPUs, DeepSeek&#8217;s Agent Harness, Postgres for Everything + more &#128640;]]></description><link>https://machinelearning.substack.com/p/issue-401-the-ml-engineer</link><guid isPermaLink="false">https://machinelearning.substack.com/p/issue-401-the-ml-engineer</guid><dc:creator><![CDATA[Alejandro Saucedo]]></dc:creator><pubDate>Sun, 23 Aug 2026 17:17:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Thank you for being part of over <a href="https://ethical.institute/newsletter/">70,000+ ML professionals and enthusiasts</a> who receive weekly articles &amp; tutorials on Machine Learning &amp; MLOps &#129302; You can join the newsletter <a href="https://ethical.institute/newsletter/">/newsletter/</a> &#11088;</p><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;5161dcaf-22f4-4ea9-839f-11c59eb25359&quot;,&quot;duration&quot;:null}"></div><p>If you like the content please support the newsletter by sharing with your friends via &#9993;&#65039; Email, &#128038; <a href="http://twitter.com/intent/tweet?text=Great%20Newsletter%20on%20Machine%20Learning!%20%20/newsletter/">Twitter</a>, &#128188; <a href="http://www.linkedin.com/shareArticle?mini=true&amp;url=/newsletter/&amp;title=The%20Machine%20Learning%20Engineer%20Newsletter&amp;summary=Curated%20news%20about%20machine%20learning%20operations,%20reproducibility,%20explainability%20and%20beyond">Linkedin</a> and &#128213; <a href="http://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fethical.institute%2Fmle.html">Facebook</a>!</p><div><hr></div><h2>This week in ML Engineering:</h2><ul><li><p>Raschka <a href="https://magazine.sebastianraschka.com/p/claude-watermarking">on Claude&#8217;s Watermarking</a></p></li><li><p>The Netflix <a href="https://arxiv.org/abs/2608.10257">GenRec Paper</a></p></li><li><p>Speculative Decoding <a href="https://vllm.ai/blog/2026-08-23-speculative-decoding-amd-gpus">on AMD GPUs</a></p></li><li><p>DeepSeek&#8217;s <a href="https://deepseek.com/harness/en">Agent Harness</a></p></li><li><p>Postgres <a href="https://www.raphaelbauer.com/posts/postgresql-everything/">for Everything</a></p></li><li><p>Open Source <a href="https://ethical.institute/open-source/production-ml-list/">ML Frameworks</a></p></li><li><p>Awesome AI Guidelines <a href="https://ethical.institute/open-source/ai-guidelines/">to check out this week</a></p></li><li><p>+ more &#128640;</p></li></ul><div><hr></div><p><strong><a href="https://magazine.sebastianraschka.com/p/claude-watermarking">Raschka on Claude&#8217;s Watermarking</a></strong></p><p>Did you know your Claude output may have a watermark soon? Who needs a watermark when you have annoying em dashes tho?! But in all seriousness, Anthropic&#8217;s approach is interesting: Raschka did a pretty interesting deep dive where he showed the watermark is applied at the token sampling stage rather than through model retraining. It is basically using a secret key together with the previous-token context to seed a deterministic &#8220;tournament sampling&#8221; process, where candidate tokens compete in pairwise rounds under random watermarking functions that assign them bit signatures. Detection is computationally cheap as it only requires the secret key and the watermarking functions, without re-running the LLM at all. It is interesting that only Anthropic holds the key, which means end users cannot independently check whether a text is watermarked, and Raschka points out the scheme can still be circumvented by passing the output through another local model for rewriting or editing. This is a great explainer on a topic we will certainly be hearing quite a lot more about; and in all seriousness the joke on em-dashes may be less of a joke, AI-slop writing-style may indeed be the watermark!</p><div><hr></div><p><strong><a href="https://arxiv.org/abs/2608.10257">The Netflix GenRec Paper</a></strong></p><p>Netflix has published the paper behind their LLM-native recommendation ranker, which is driving offline lift of about +1.6% relative MRR over their production ranker with roughly 40x less training data: The paper details is quite interesting as they use an infrequently updated Netflix-adapted foundational LLM, followed by frequent ranking-specific post-training with a reward-weighted ranking loss. It also covers the serving design, which runs prefill-only inference on vLLM to score the full candidate set in a single forward pass, with context compression from ~5,000 down to ~1,700 tokens at roughly a third of the serving cost. Netlflix shares that they ran a four-week A/B test on ~10% of traffic showing statistically significant gains on both short-term and long-term metrics. They are also candid on some of the failurse they&#8217;ve observed, such as over-recommending globally popular content, hallucinating out-of-catalog titles, and ignoring nuanced business constraints. Definitely worth checking out if you are building anything RecSys-adjacent!</p><div><hr></div><p><strong><a href="https://vllm.ai/blog/2026-08-23-speculative-decoding-amd-gpus">Speculative Decoding on AMD GPUs</a></strong></p><p>The vLLM team published a super detailed benchmark study of speculative decoding on AMD GPUs - a lot of great learnings for practitioners in this space: They evaluated five drafting methods (Native MTP, Gemma 4 MTP, EAGLE-3, DFlash and DSpark) which they compared across model families on MI300X hardware (288 GB VRAM each &#129327;). They also had throughput multipliers from 1.27x to 2.87x on various workloads and evaluations like HumanEval. The parallel DFlash approach was frequently the best performer at longer proposal lengths, whilst per-position acceptance rates decline at later draft positions for every method. It seems the optimal proposal length was not constant across models and datasets either. For production ML practitioners the takeaway is that speculative decoding is not a free lunch you configure once and the right drafting method depends on your model, workload and hardware. It is great to see this level of empirical depth on non-NVIDIA inference!</p><div><hr></div><p><strong><a href="https://deepseek.com/harness/en">DeepSeek&#8217;s Agent Harness</a></strong></p><p>DeepSeek released their new Harness, and it&#8217;s one of the fastest growing OSS projects ever - record breaks every week! It seems it&#8217;s taking a pi approach, where every capability is a plugin that can be swapped or recomposed - this covers models, tools, skills, sessions, sandboxes, storage, loops, scheduling and even the UI. It uses the Cordis kernel for managing dependencies and inter-plugin communication. For traceability, all agent activity lands in an append-only session log, and the Trajectory view lets you inspect system prompts, reasoning steps and tool calls. It is an interesting trend on how telemetry trasparency is being explored as harnesses evolve. They shipped it with four runtime modes, including a full standard toolset with subagents, but also a super minimal bash-and-editor setup for benchmarking. It is clear it&#8217;s early days, and so far it seems like a bit of everything as opposed to something net-new, but we&#8217;ll really see how it fares out in practice. Everyone is releasign a harness&#8230; it seems it&#8217;s time to start working on one&#8230;</p><div><hr></div><p><strong><a href="https://www.raphaelbauer.com/posts/postgresql-everything/">Postgres for Everything</a></strong></p><p>&#8220;Just use Postgres&#8221;&#8230; this argument only continues to grow stronger throughout the years! Nowdays postgres can replace a full-text search via tsvector, NOSQL via JSONB with GIN indexing, Kafka-style queues with SELECT FOR UPDATE SKIP LOCKED, time-series with TimescaleDB, vectors with pgvector, Redis-style caching with unlogged tables, and even graphs through Apache AGE. A few weeks back I was able to take this for a spin replacing redis for postgres unlogged tables and it was quit smooth (and really nice to standardise infra), so definitely less is more here. There are no benchmarks in the post, and there arespecialised systems when Postgres no longer performs at the same level required, however the value prop of fewer moving parts is definitely a big benefit. A fun and practical read to close the week, and a good reminder that the simplest architecture that works is usually the right one.</p><div><hr></div><h2>Upcoming MLOps Events</h2><p>The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.</p><p>You can also find our upcoming events and past talk recordings on the <a href="https://ethical.institute/talks-and-events/">talks and events page</a>.</p><p><strong>Events we are speaking at this year:</strong></p><ul><li><p><a href="https://signalsconf.io/#tickets">Signals Conference</a> - September @ Berlin</p></li><li><p><a href="https://worldsummit.ai/">World Summit AI Europe</a> - September @ Amsterdam</p></li><li><p><a href="https://codetalks.com/">Code.Talks 2026</a> - November @ Hamburg</p></li></ul><p><strong>Other relevant events:</strong></p><ul><li><p><a href="https://www.ai-infra-summit.com/">AI Infra Summit 2026</a> - Sept @ California</p></li><li><p><a href="https://mlopsworld.com/">MLOps World 2026</a> - Nov @ Austin</p></li></ul><p><strong>In case you missed our talks, check our recordings below:</strong></p><ul><li><p>The State of AI in 2025 - <a href="https://www.youtube.com/watch?v=v2LENQOG-Xg">WeAreDevelopers 2025</a></p></li><li><p>Prod Generative AI in 2024 - <a href="https://www.youtube.com/watch?v=0uJGmMZGUJE&amp;list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&amp;index=3">KubeCon AI Day 2025</a></p></li><li><p>The State of AI in 2024 - <a href="https://www.youtube.com/live/AtA2XXo_b5s">WeAreDevelopers 2024</a></p></li><li><p>Responsible AI Workshop Keynote - <a href="http://www.youtube.com/watch?v=57YpXjcj0Ho">NeurIPS 2021</a></p></li><li><p>Practical Guide to ML Explainability - <a href="http://www.youtube.com/watch?v=vq8mDiDODhc">PyCon London</a></p></li><li><p>ML Monitoring: Outliers, Drift, XAI - <a href="http://www.youtube.com/watch?v=QcevzK9ZuDg">PyCon Keynote</a></p></li><li><p>Metadata for E2E MLOps - <a href="https://www.youtube.com/watch?v=OSbH4dfswCY">Kubecon NA 2022</a></p></li><li><p>ML Performance Evaluation at Scale - <a href="http://www.youtube.com/watch?v=8ORl8lu1Eeo">KubeCon Eur 2021</a></p></li><li><p>Industry Strength LLMs - <a href="https://www.youtube.com/watch?v=RVUi_rAFfzU">PyData Global 2022</a></p></li><li><p>ML Security Workshop Keynote - <a href="http://www.youtube.com/watch?v=7XSy5aw8oU8">NeurIPS 2022</a></p></li></ul><div><hr></div><h2>Open Source MLOps Tools</h2><p>Check out the fast-growing ecosystem of production ML tools &amp; frameworks at <a href="https://ethical.institute/open-source/production-ml-list/">the github repository</a> which has reached over 20,000 &#11088; github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here&#8217;s a few featured open source libraries that we maintain:</p><ul><li><p><a href="https://github.com/besanson/sarc-governance/tree/main">SARC</a> - Provides wrappers for popular agentic frameworks to enable guardrails and constraints that are enforced through the flow.</p></li><li><p><a href="https://github.com/axsaucedo/kaos">KAOS</a> - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.</p></li><li><p><a href="https://github.com/KomputeProject/kompute">Kompute</a> - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.</p></li><li><p><a href="https://ethical.institute/open-source/production-ml-list/">Production ML Tools</a> - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.</p></li><li><p><a href="https://github.com/EthicalML/awesome-artificial-intelligence-regulation">AI Policy List</a> - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.</p></li><li><p><a href="https://github.com/EthicalML/awesome-production-genai/">Agentic Systems Tools</a> - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain</p></li></ul><p>Please do support some of our open source projects by sharing, contributing or adding a star &#11088;</p><div><hr></div><h2>About us</h2><p>The Institute for Ethical AI &amp; Machine Learning is a European research centre that carries out world-class research into responsible machine learning.</p><p><a href="https://ethical.institute/">Check out our website</a></p><p>&#9993;&#65039; Email, &#128038; <a href="http://twitter.com/EthicalML">Twitter</a>, &#128188; <a href="https://www.linkedin.com/company/the-institute-for-ethical-machine-learning/">Linkedin</a></p>]]></content:encoded></item><item><title><![CDATA[Issue #400 - The ML Engineer 🤖]]></title><description><![CDATA[Production ML Across 2015-2035, Zalando on Agentic Engineering at Scale, Stealing Reasoning from Encrypted Traces, Hugging Face on Open Models, Cloudflare on MCP Traffic Control + more &#128640;]]></description><link>https://machinelearning.substack.com/p/issue-400-the-ml-engineer</link><guid isPermaLink="false">https://machinelearning.substack.com/p/issue-400-the-ml-engineer</guid><dc:creator><![CDATA[Alejandro Saucedo]]></dc:creator><pubDate>Sun, 16 Aug 2026 17:28:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Today we celebrate a <strong>HUGE</strong> milestone as we share our <strong><a href="http://ethical.institute/newsletter/400/">400th Issue</a> &#128640;&#128640;&#128640;</strong></p><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;71a10c43-fc9c-492f-a0a1-430b0e9c5d18&quot;,&quot;duration&quot;:null}"></div><p>We have come really far since the first newsletter edition <strong>in 2018,</strong> and today we can celebrate many more achievements since then:</p><ul><li><p><a href="http://ethical.institute/newsletter/">MLE Newsletter</a> reaching 70k+ AI professionals every week &#129395;</p></li><li><p><a href="http://ethical.institute/open-source/">Open source projects</a> topping 26k+ Stars, 3.2k Forks and 300+ contributors &#11088;</p></li><li><p>20+ recommendations adopted <a href="http://ethical.institute/policy/">across European Regulation</a> &#127963;&#65039;</p></li><li><p>Contributions to the <a href="http://ethical.institute/partners/">United Nations, European Commission and more &#128640;</a></p></li></ul><p>Our mission continues to be the same; let&#8217;s make production machine learning something we can all do well and do responsibly.</p><p>Thank you for being part of it, and bring on 400 more! We look forward to continue this great momentum towards <strong><a href="http://localhost:4322/blog/new-phase-institute-ethical-ai-alignment-safety/">the</a></strong><a href="http://localhost:4322/blog/new-phase-institute-ethical-ai-alignment-safety/"> </a><strong><a href="http://ethical.institute/blog/new-phase-institute-ethical-ai-alignment-safety/">new phase of the Institute</a>!</strong></p><div><hr></div><h2>This week in ML Engineering:</h2><ul><li><p>Production ML <a href="https://www.youtube.com/watch?v=I1GvlW1H4WI">Across 2015-2035</a></p></li><li><p>Zalando on <a href="https://engineering.zalando.com/posts/2026/08/agentic-engineering-at-zalando-a-snapshot.html">Agentic Engineering at Scale</a></p></li><li><p>Stealing Reasoning <a href="https://stolen-thoughts.com/">from Encrypted Traces</a></p></li><li><p>Hugging Face on <a href="https://huggingface.co/blog/state-of-open-models-summer-2026">Open Models</a></p></li><li><p>Cloudflare on <a href="https://blog.cloudflare.com/mcp-security-updates/">MCP Traffic Control</a></p></li><li><p>Open Source <a href="http://localhost:4322/open-source/production-ml-list/">ML Frameworks</a></p></li><li><p>Awesome AI Guidelines <a href="http://localhost:4322/open-source/ai-guidelines/">to check out this week</a></p></li><li><p>+ more &#128640;</p></li></ul><div><hr></div><p><a href="https://www.youtube.com/watch?v=I1GvlW1H4WI">Production ML Across 2015-2035</a></p><p>My talk covering 20 years of Production ML is now up on YouTube! I look back across the last 10 years of MLOps, and make some predictions for the next 10 years - many of which are playing out scarily well!</p><p>This talk was very close to my heart, as during the first half I do a really in-depth analysis on the evolution of tooling, practices and limitations all the way back to the early days of MLOps.</p><p>This journey has really been a fun one, and it is clear to me we are still at the very beginning, with challenges like ML monitoring not yet nearly solved.</p><p>In the second half, I cover some of the opportunities looking forward, which talk about the evoltuion towards LLMOps, new stacks emerging, faster and more autonomous services and patterns.</p><p>The message I keep coming back to is that the lifecycle of a model only really begins once it is in production; everything before that is preparation. I have to say I am excited for the many more years ahead!</p><div><hr></div><p><a href="https://engineering.zalando.com/posts/2026/08/agentic-engineering-at-zalando-a-snapshot.html">Zalando on Agentic Engineering at Scale</a></p><p>Another win for Europe! Germany&#8217;s tech powerhouse Zalando has published one of the most detailed pictures of an agentic engineering rollout that we have seen this year. I have to say that I have personally had so much fun as part of this agentic engineering transformation - the state of engineering feels reinvigorating!</p><p>This showcases some of the value obtained from rolling out LLM-based PR reviews, with 33% of Zalando&#8217;s PRs now being auto-approved as low risk. We have also seen lead times on PR time decrease by 20-40% across more than 250 engineering teams.</p><p>A lot of the infrastructure has also been built in-house with open source technology; we have a central zLLM service powered by LiteLLM, which serves OpenAI, Bedrock and Vertex models to around 2,000 monthly active users (+ growing).</p><p>There are also some clear challenges that this brings, we are seeingcyclomatic complexity spiking as agent adoption grows, commit messages ballooning to 5,000 characters, and challenges on tool migration / transformation.</p><p>The best thing here is the great collaborations and contributions that have come about. There are weekly guild sessions and monthly trainings running at 120 to 150 participants, as well as hackathons and agentic pilots.</p><p>For production ML practitioners the takeaway is that the measurement and governance work is what makes an adoption story legible at all, and it is great to see a European company leading on publishing it!</p><div><hr></div><p><a href="https://stolen-thoughts.com/">Stealing Reasoning from Encrypted Traces</a></p><p>Super interesting, and pretty scary! Researchers Max Planck Institute, ELLIS and Snyk et. al. have found that the you can basically reverse engineer the encrypted reasoning from frontier labs like Anthropic and OpenAI!</p><p>Initially it may seem that there&#8217;s nothing to worry about, however this encrypted reasoning blob has been assumed to be safe, and likely is scattered across telemetry, logs, session dumps, etc which we just assumed was not unsafe; so many teams may now have a big problem.</p><p>To provide some intuition, you use a frontier lab model, it &#8220;reasons&#8221; before it answers. The labs don&#8217;t show you that thinking, but when you get the answer, you also get a sealed blob representing the hidden reasoning.</p><p>The blob is encrypted so you can&#8217;t read it, but it gets sent back with your next message so the model can pick up where it left off.</p><p>Further to this, the hidden reasoning is not the same as the answer. The answer is what the model chose to tell you, after whatever filtering and summarising sits between the two.</p><p>This means the reasoning can contain intermediate conclusions the model discarded, inferences it drew about you that it didn&#8217;t state, or content it decided not to surface.</p><p>So the blob can hold strictly more about your situation than you ever saw on screen. If you&#8217;d forwarded that blob to someone, you may have handed over more than the conversation.</p><p>This is probably going to lead into a large rearchitecture from foundation models, and it is interesting to see that we&#8217;re only seeing the tip of the iceberg as these systems evolve.</p><div><hr></div><p><a href="https://huggingface.co/blog/state-of-open-models-summer-2026">Hugging Face on Open Models</a></p><p>Who is actually shipping the open models these days? Hugging Face has published their state of open models report: Chinese labs set the monthly size ceiling in every month, running from 754B up to 2.78 trillion parameters.</p><p>The labs from United States stayed under 130B in five of the past seven months this year, falling behind on the open model releases.</p><p>The new trend is more and more permissive licensing; out of 178 Chinese releases above 20B, 59% shipped under Apache 2.0 and 22% under MIT, with none carrying non commercial restrictions.</p><p>This is great! Although it carries a geopolitical weight that is becoming more and more clear.</p><p>The distribution numbers are the ones that surprised me most, as the models under 1B account for 83% of all time downloads, while everything above 100B accounts for only 1%.</p><p>HugginfFace is of course only &#8220;one perspective&#8221; and the results would be biased to what they can se in their platform, but overall the report seems quite measured, especially as they had even caveats like these.</p><div><hr></div><p><a href="https://blog.cloudflare.com/mcp-security-updates/">Cloudflare on MCP Traffic Control</a></p><p>We need to think seriously about MCP traffic. Cloudflare has shipped protocol level detection for MCP-source traffic, classifying requests by the MCP-Protocol-Version header dynamically to be able to take action.</p><p>Namely this gives more control to website owners, as there is now an ability to action on an is_mcp rule for allow and block policies.</p><p>This also gives richer insights about which servers are actually being reached, and the option to force everything through their portals.</p><p>Claudflare is still upfront about the limitations of this, since this needs TLS inspection which does not come for free, and cannot see local stdio servers or anything marked &#8220;do not inspect&#8221;.</p><p>It is interesting to see MCP governance arriving as ordinary network policy rather than as something bespoke, but that makes sense! It is clear that this is where a lot of the agent security work is heading.</p><div><hr></div><h3>Upcoming MLOps Events</h3><p>The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.</p><p>You can also find our upcoming events and past talk recordings on the <a href="http://ethical.institute/talks-and-events/">talks and events page</a>.</p><h3>Events we are speaking at this year:</h3><ul><li><p><a href="https://signalsconf.io/#tickets">Signals Conference</a> - September @ Berlin</p></li><li><p><a href="https://worldsummit.ai/">World Summit AI Europe</a> - September @ Amsterdam</p></li><li><p><a href="https://codetalks.com/">Code.Talks 2026</a> - November @ Hamburg</p></li></ul><h3>Other relevant events:</h3><ul><li><p><a href="https://www.ai-infra-summit.com/">AI Infra Summit 2026</a> - Sept @ California</p></li><li><p><a href="https://mlopsworld.com/">MLOps World 2026</a> - Nov @ Austin</p></li></ul><h3>In case you missed our talks, check our recordings below:</h3><ul><li><p>The State of AI in 2025 - <a href="https://www.youtube.com/watch?v=v2LENQOG-Xg">WeAreDevelopers 2025</a></p></li><li><p>Prod Generative AI in 2024 - <a href="https://www.youtube.com/watch?v=0uJGmMZGUJE&amp;list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&amp;index=3">KubeCon AI Day 2025</a></p></li><li><p>The State of AI in 2024 - <a href="https://www.youtube.com/live/AtA2XXo_b5s">WeAreDevelopers 2024</a></p></li><li><p>Responsible AI Workshop Keynote - <a href="http://www.youtube.com/watch?v=57YpXjcj0Ho">NeurIPS 2021</a></p></li><li><p>Practical Guide to ML Explainability - <a href="http://www.youtube.com/watch?v=vq8mDiDODhc">PyCon London</a></p></li><li><p>ML Monitoring: Outliers, Drift, XAI - <a href="http://www.youtube.com/watch?v=QcevzK9ZuDg">PyCon Keynote</a></p></li><li><p>Metadata for E2E MLOps - <a href="https://www.youtube.com/watch?v=OSbH4dfswCY">Kubecon NA 2022</a></p></li><li><p>ML Performance Evaluation at Scale - <a href="http://www.youtube.com/watch?v=8ORl8lu1Eeo">KubeCon Eur 2021</a></p></li><li><p>Industry Strength LLMs - <a href="https://www.youtube.com/watch?v=RVUi_rAFfzU">PyData Global 2022</a></p></li><li><p>ML Security Workshop Keynote - <a href="http://www.youtube.com/watch?v=7XSy5aw8oU8">NeurIPS 2022</a></p></li></ul><div><hr></div><h2>Open Source MLOps Tools</h2><p>Check out the fast-growing ecosystem of production ML tools &amp; frameworks at <a href="http://github.com/EthicalML/awesome-production-machine-learning/">the github repository</a> which has reached over 20,000 &#11088; github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here&#8217;s a few featured open source libraries that we maintain:</p><ul><li><p><strong><a href="https://github.com/axsaucedo/kaos">KAOS</a></strong> - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.</p></li><li><p><strong><a href="https://github.com/KomputeProject/kompute">Kompute</a></strong> - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-machine-learning/">Production ML Tools</a></strong> - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-artificial-intelligence-regulation">AI Policy List</a></strong> - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-genai/">Agentic Systems Tools</a></strong> - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain</p></li><li><p><strong><a href="https://github.com/besanson/sarc-governance/tree/main">SARC</a> </strong>- Provides wrappers for popular agentic frameworks to enable guardrails and constraints that are enforced through the flow.</p></li></ul><p>Please do support some of our open source projects by sharing, contributing or adding a star &#11088;</p>]]></content:encoded></item><item><title><![CDATA[ Issue #399 - The ML Engineer 🤖 ]]></title><description><![CDATA[The New Institute for AI Alignment & Safety, Multi-Tenant Multi-Tier Memory for AI Agents, LLMs Reward Expertise, Qwen3.8-Max and Open Weights, Harness Design for Long-Running Apps + more &#128640;]]></description><link>https://machinelearning.substack.com/p/issue-399-the-ml-engineer</link><guid isPermaLink="false">https://machinelearning.substack.com/p/issue-399-the-ml-engineer</guid><dc:creator><![CDATA[Alejandro Saucedo]]></dc:creator><pubDate>Sun, 09 Aug 2026 19:11:57 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2 style="text-align: center;">The New<strong> </strong><br><strong><a href="https://ethical.institute/">Institute for Ethical AI</a></strong><br><em><strong><a href="https://ethical.institute/">Alignment &amp; Safety</a></strong></em></h2><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;fec29c33-f8a0-4639-9a8f-f50f5ffc77ac&quot;,&quot;duration&quot;:null}"></div><p style="text-align: center;">We are excited to announce the new face of <a href="https://ethical.institute">the Institute for Ethical AI Alignment &amp; Safety</a>!</p><p style="text-align: center;">We have rebuilt <strong>the Institute website</strong> from the ground up, together with some of our key initiatives.</p><p style="text-align: center;"><strong>Since our founding in 2017</strong>, we have built a track record of contributions across public and private institutions.</p><p style="text-align: center;">Our work runs from <strong>individual practice</strong> to <strong>national regulation;</strong> by principle, by process, by standards, by regulation.</p><p style="text-align: center;">Several of our recommendations have been <strong>adopted across the EU and UK policy.</strong> Our mission has not changed since 2017.</p><p style="text-align: center;">We look forward to continue contributing to a <strong>future</strong> where frontier <strong>AI is safe, aligned and accountable</strong> to people and society.</p><div class="pullquote"><p style="text-align: center;"><strong>Check it out: </strong><a href="https://ethical.institute">https://ethical.institute</a></p></div><h2>This week in <a href="http://ethical.institute/mle/399.html">ML Engineering</a>:</h2><ul><li><p>Multi-Tenant Memory <a href="https://hackernoon.com/whose-memory-is-it-building-multi-tenant-multi-tier-memory-for-ai-agents-part-1">for AI Agents</a></p></li><li><p>LLMs <a href="https://www.seangoedecke.com/llms-reward-expertise/">Reward Expertise</a></p></li><li><p>Qwen3.8-Max <a href="https://qwen.ai/blog?id=qwen3.8">and Open Weights</a></p></li><li><p>Harness Design <a href="https://www.anthropic.com/engineering/harness-design-long-running-apps">for Long-Running Apps</a></p></li><li><p>Mistral&#8217;s <a href="https://mistral.ai/news/shieldstral/">Shieldstral Safety Classifier</a></p></li><li><p>Open Source <a href="http://localhost:4321/open-source/production-ml-list/">ML Frameworks</a></p></li><li><p>Awesome AI Guidelines <a href="http://localhost:4321/open-source/ai-guidelines/">to check out this week</a></p></li><li><p>+ more &#128640;</p></li></ul><h2><a href="https://hackernoon.com/whose-memory-is-it-building-multi-tenant-multi-tier-memory-for-ai-agents-part-1">Multi-Tenant Multi-Tier Memory for AI Agents</a></h2><p>Our 4-part series on agent memory has been published &amp; featured in the front-page of HackerNoon homepage as a top story! LLMs are stateless by design, so without a memory layer every session starts from zero, and the number of dedicated memory tools has been growing almost daily. This came out of my recent work extending the Kubernetes Agent Orchestration System (KAOS) to support multi-tiered memory persistence (aka short-, medium- and long-term memory). Along the way I hit most of the same issues that anyone would when building or integrating multi-tiered memory into an agentic system, so I thought it would be useful to compile all the learnings, design choices and examples. Check it out, together with the rest of the series!</p><h2><a href="https://www.seangoedecke.com/llms-reward-expertise/">LLMs Reward Expertise</a></h2><p>Do language models still reward expertise, or has that ship sailed? Sean Goedecke argues for the second; the distinction is between getting something usable out of a model and extracting the maximum value from it. A non-expert can get sort-of-okay Python, but only someone who knows what a good answer looks like can evaluate the output critically. His main evidence is Terence Tao&#8217;s published ChatGPT conversation on the Jacobian Conjecture, where Tao&#8217;s messages are short and to the point, and the model answers in a talking-to-mathematicians register rather than an explaining-to-amateurs one. Tao pushes back and he almost never takes the model&#8217;s advice. He is careful with the caveat that non-experts still get real value, and that OpenAI had a team of expert mathematicians filtering the suggestions, a step you cannot currently skip. For production ML practitioners the takeaway is that knowledge is still more improtant than ever, and now is even becoming the bottleneck; does this mean we need to accelerate our learning?</p><h2><a href="https://qwen.ai/blog?id=qwen3.8">Qwen3.8-Max and Open Weights</a></h2><p>Qwen3.8-Max bas been released! And for the first time they say a Qwen-Max-class model will get open weights: the model scales to 2.4 trillion parameters with 95B active, is built on the Qwen 3.5 architecture, and is available through QwenCloud now with the weights promised next week. The more interesting part is actually the long-horizon runs, as they report a ten day autonomous coding run building the oh-my-cli project, which after roughly 16 days of fully autonomous operation had accumulated 265 commits, 127 PRs and 151 issues! Whether that is good or bad work is yet to be seen&#8230; Open weights at this scale would be quite something; let&#8217;s see what actually lands next week.</p><h2><a href="https://www.anthropic.com/engineering/harness-design-long-running-apps">Harness Design for Long-Running Apps</a></h2><p>Anthropic has written up the harness they use to have Claude build entire applications over multi-hour runs, and more interestingly what they deleted from it as the model improved: The architecture is three agents framed as a separation of generator and judge, with a planner that expands a one to four sentence prompt into a full product spec, a generator that implements against it, and an evaluator that drives the running app through Playwright MCP like a real user, checking UI, API endpoints and database state. It is also quite a well timed piece as it comes next to their piece on <a href="https://www.anthropic.com/engineering/how-we-contain-claude">how they contain Claude across products</a>, where filesystem and egress boundaries are what holds once the model-layer defences fail.</p><h2><a href="https://mistral.ai/news/shieldstral/">Mistral&#8217;s Shieldstral Safety Classifier</a></h2><p>Europe strikes again! Mistral has released Shieldstral, an open-weights multimodal safety classifier under Apache 2.0: Shieldstral is a 3B parameter model covering text and images, it runs on a single 16GB NVIDIA GPU, and the weights are on Hugging Face. It is interesting how they are proposing to reframe moderation as binary question answering rather than fixed label classification. Basically a plain-language yes/no question and the document being judged as the three inputs; the output is a calibrated probability read off a single token, so a policy change means rewriting the question instead of retraining. They claim it outperforms models up to 7x its size and claim a new state of the art on multimodal safety, although the post does not publish per-benchmark numbers to check that against. One classifier instead of a guardrail model per deployment is a good direction, and great to see it shipped under a permissive licence!</p><h2>Upcoming MLOps Events</h2><p>The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.</p><h3>Events we are speaking at this year:</h3><ul><li><p><a href="https://signalsconf.io/#tickets">Signals Conference</a> - September @ Berlin</p></li><li><p><a href="https://worldsummit.ai/">World Summit AI Europe</a> - September @ Amsterdam</p></li></ul><h3>Other relevant events:</h3><ul><li><p><a href="https://www.ai-infra-summit.com/">AI Infra Summit 2026</a> - Sept @ California</p></li><li><p><a href="https://codetalks.com/">Code.Talks 2026</a> - Nov @ Hamburg</p></li><li><p><a href="https://mlopsworld.com/">MLOps World 2026</a> - Nov @ Austin</p></li></ul><h3>In case you missed our talks, check our recordings below:</h3><ul><li><p>The State of AI in 2025 - <a href="https://www.youtube.com/watch?v=v2LENQOG-Xg">WeAreDevelopers 2025</a></p></li><li><p>Prod Generative AI in 2024 - <a href="https://www.youtube.com/watch?v=0uJGmMZGUJE&amp;list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&amp;index=3">KubeCon AI Day 2025</a></p></li><li><p>The State of AI in 2024 - <a href="https://www.youtube.com/live/AtA2XXo_b5s">WeAreDevelopers 2024</a></p></li><li><p>Responsible AI Workshop Keynote - <a href="http://www.youtube.com/watch?v=57YpXjcj0Ho">NeurIPS 2021</a></p></li><li><p>Practical Guide to ML Explainability - <a href="http://www.youtube.com/watch?v=vq8mDiDODhc">PyCon London</a></p></li><li><p>ML Monitoring: Outliers, Drift, XAI - <a href="http://www.youtube.com/watch?v=QcevzK9ZuDg">PyCon Keynote</a></p></li><li><p>Metadata for E2E MLOps - <a href="https://www.youtube.com/watch?v=OSbH4dfswCY">Kubecon NA 2022</a></p></li><li><p>ML Performance Evaluation at Scale - <a href="http://www.youtube.com/watch?v=8ORl8lu1Eeo">KubeCon Eur 2021</a></p></li><li><p>Industry Strength LLMs - <a href="https://www.youtube.com/watch?v=RVUi_rAFfzU">PyData Global 2022</a></p></li><li><p>ML Security Workshop Keynote - <a href="http://www.youtube.com/watch?v=7XSy5aw8oU8">NeurIPS 2022</a></p></li></ul><h2>Open Source <a href="http://localhost:4321/open-source/production-ml-list/">MLOps Tools</a></h2><p>Check out the fast-growing ecosystem of production ML tools &amp; frameworks at <a href="http://localhost:4321/open-source/production-ml-list/">the github repository</a> which has reached over 20,000 &#11088; github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here&#8217;s a few featured open source libraries that we maintain:</p><ul><li><p><a href="https://github.com/besanson/sarc-governance/tree/main">SARC</a> - Provides wrappers for popular agentic frameworks to enable guardrails and constraints that are enforced through the flow.</p></li><li><p><a href="https://github.com/axsaucedo/kaos">KAOS</a> - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.</p></li><li><p><a href="https://github.com/KomputeProject/kompute">Kompute</a> - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.</p></li><li><p><a href="http://localhost:4321/open-source/production-ml-list/">Production ML Tools</a> - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.</p></li><li><p><a href="https://github.com/EthicalML/awesome-artificial-intelligence-regulation">AI Policy List</a> - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.</p></li><li><p><a href="https://github.com/EthicalML/awesome-production-genai/">Agentic Systems Tools</a> - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain</p></li></ul><p>Please do support some of our open source projects by sharing, contributing or adding a star &#11088;</p><h2>About us</h2><p>The Institute for Ethical AI &amp; Machine Learning is a European research centre that carries out world-class research into responsible machine learning.</p><p><a href="https://ethical.institute/">Check out our website</a></p><p>&#9993;&#65039; <a href="mailto:EMAIL?subject=Check%20out%20the%20Machine%20Learning%20Engineering%20Newsletter!&amp;body=Check%20out%20this%20weekly%20newsletter%20on%20Machine%20Learning!%20Join%20for%20free%20here%3A%20https%3A%2F%2Fethical.institute%2Fmle.html">Email</a>, &#128038; <a href="http://twitter.com/EthicalML">Twitter</a>, &#128188; <a href="https://www.linkedin.com/company/the-institute-for-ethical-machine-learning/">Linkedin</a></p>]]></content:encoded></item><item><title><![CDATA[Issue #398 - The ML Engineer 🤖 ]]></title><description><![CDATA[Timeline on OpenAI HF Hack, Cheating Behavior in Frontier Models, Raschka on Kimi K3 Architectures, Netflix on the LLM-Native RecSys, OpenAI 10 Advances in Maths + more &#128640;]]></description><link>https://machinelearning.substack.com/p/issue-398-the-ml-engineer</link><guid isPermaLink="false">https://machinelearning.substack.com/p/issue-398-the-ml-engineer</guid><dc:creator><![CDATA[Alejandro Saucedo]]></dc:creator><pubDate>Sun, 02 Aug 2026 17:41:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Thank you for being part of over <a href="https://ethical.institute/mle.html">70,000+ ML professionals and enthusiasts</a> who receive weekly articles &amp; tutorials on Machine Learning &amp; MLOps &#129302; You can join the newsletter <a href="https://ethical.institute/mle.html">https://ethical.institute/mle.html</a> &#11088;</p><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;1b93fa2a-e69c-4397-96bc-aa1f15ecfebc&quot;,&quot;duration&quot;:null}"></div><p>If you like the content please support the newsletter by sharing with your friends via &#9993;&#65039; Email, &#128038; <a href="http://twitter.com/intent/tweet?text=Great%20Newsletter%20on%20Machine%20Learning!%20%20https://ethical.institute/mle.html">Twitter</a>, &#128188; <a href="http://www.linkedin.com/shareArticle?mini=true&amp;url=https://ethical.institute/mle.html&amp;title=The%20Machine%20Learning%20Engineer%20Newsletter&amp;summary=Curated%20news%20about%20machine%20learning%20operations,%20reproducibility,%20explainability%20and%20beyond">Linkedin</a> and &#128213; <a href="http://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fethical.institute%2Fmle.html">Facebook</a>!</p><div><hr></div><h2>This week in ML Engineering:</h2><ul><li><p>Timeline <strong><a href="https://huggingface.co/blog/agent-intrusion-technical-timeline">on OpenAI HF Hack</a></strong></p></li><li><p>Cheating Behavior <strong><a href="https://www.aisi.gov.uk/blog/cheating-behaviour-in-frontier-model-evaluations">in Frontier Models</a></strong></p></li><li><p>Raschka <strong><a href="https://sebastianraschka.com/blog/2026/kimi-k3-architecture-notes.html">on Kimi K3 Architectures</a></strong></p></li><li><p>Netflix on <strong><a href="https://netflixtechblog.com/genrec-towards-llm-native-recommendation-at-netflix-f20be6f643e3">the LLM-Native RecSys</a></strong></p></li><li><p>OpenAI <strong><a href="https://openai.com/index/ten-advances-in-mathematics/">10 Advances in Maths</a></strong></p></li><li><p>Open Source <strong><a href="http://github.com/EthicalML/awesome-production-machine-learning/">ML Frameworks</a></strong></p></li><li><p>Awesome AI Guidelines <strong><a href="https://github.com/ethicalml/awesome-artificial-intelligence-guidelines">to check out this week</a></strong></p></li><li><p>+ more &#128640;</p></li></ul><div><hr></div><p><a href="https://huggingface.co/blog/agent-intrusion-technical-timeline">Timeline on OpenAI HF Hack</a></p><p>Last week OpenAI published about their rouge hacking HuggingFace, this week HF published a super detailed forensic reconstruction and it&#8217;s insane: The tldr we know is that an AI agent escaped an OpenAI cybersecurity evaluation environment, compromised an external code sandbox, and entered Hugging Face production systems. What we can now see is that the agent was able to do this using malicious dataset configurations that enabled local file disclosure and arbitrary code execution on the HF platform. HF found about 17,600 actions that the agent ran by identifying the logs once it entered a k8s pod, escalated cloud credentials in Kubernetes, joined an internal mesh network, accessed an operational database, and obtained limited source-control write permissions. The model was able to do this while using public web services and attacker-controlled datasets for command-and-control. No unauthorized software changes were released, and customer impact was limited to five benchmark-related datasets and operational metadata; basically all the agent wanted is to access benchmark related datasets. This is the perfect example of an optimization function being optimized into destruction, resulting in effects that are adversarial to the initial intent of the operator even if the goal is being met, at all costs. This incident shows how we will need to design system security not just for the future, but for the future of agents.</p><div><hr></div><p><a href="https://www.aisi.gov.uk/blog/cheating-behaviour-in-frontier-model-evaluations">Cheating Behavior in Frontier Models</a></p><p>Following the Cyber hacks from agents these last few weeks, the UK AI Security Institute released a report that confirms every frontier model attempts to cheat on given tasks: This is an interesting topic on AI alignment, as it raises a practical problem for both benchmark validity and agent deployment. AISI defines cheating as completing a task through prohibited or out-of-scope actions, with observed methods including searching online for solutions, probing evaluation software, escalating privileges on unrelated systems and targeting the infrastructure hosting the model. In one misconfigured and unsolvable task, a model wrote and executed code through an external internet service while attempting to access AISI&#8217;s evaluation systems, although no information was leaked. The research found no clear relationship between model capability and cheating frequency, suggesting that training and alignment choices materially affect the behaviour, which is quite an interesting and important insight. Self-reporting was unreliable, and models described their prohibited actions as wrong less than half the time, which emphasises that agent success should be verified; this is a reminder of how important verification is becoming in the post-agentic era.</p><div><hr></div><p><a href="https://sebastianraschka.com/blog/2026/kimi-k3-architecture-notes.html">Raschka on Kimi K3 Architectures</a></p><p>Sebastian Raschka has released a great cheasheet on the Kimi K3 Architecture that Moonshot AI: As a refresher, Kimi K3 is an open weights 2.8T-parameter open-weight MoE model with 104B active parameters. It&#8217;s been doing the rounds due to its capability but also the fact that it has native vision and a 1M-token context window. From the architecture overview, Kimi K3&#8217;s 93-block architecture extends Kimi Linear with a 3:1 mix of Kimi Delta Attention and gated multi-head latent attention, Stable LatentMoE layers that execute experts in a compressed representation, cross-layer Attention Residuals, and NoPE throughout instead of RoPE. The design targets long-context inference efficiency, although serving it remains is still challenging as KDA introduces recurrent state, full-attention layers still require KV caches, 896 routed experts create communication pressure, and Attention Residuals add cross-layer memory traffic. vLLM support therefore includes hybrid prefix caching, fused KDA and residual kernels, FP4 MoE execution, expert parallelism and separate NVIDIA and AMD paths. It is interesting to see how closely the model and serving engine were developed together.</p><div><hr></div><p><a href="https://netflixtechblog.com/genrec-towards-llm-native-recommendation-at-netflix-f20be6f643e3">Netflix on the LLM-Native RecSys</a></p><p>Netflix has developed an LLM-native Recommender System, and they share some of the learnings they gathered along the journey: They launched GenRec, which is a ranking model that converts user histories, item metadata, and request context into natural-language inputs rather than relying on thousands of manually engineered features. The system starts from a Netflix-adapted foundation model, then applies more frequent ranking-specific post-training using catalog classification, language-model objectives, and reward-weighted examples aligned with long-term member outcomes and content-balancing requirements. They use a catalog-aware scoring head to restrict the outputs to available titles, and they use prefill-only inference on vLLM scores candidate sets without autoregressive generation. It is impressive to see that they claim this architecture improved Mean Reciprocal Rank by about 1.6% while using roughly 40 times fewer Phase-2 labelled examples than the existing production ranker. This means that it is possible to use some of these foundation models to extract signal in domains that can benefit, which would otherwise require highly custom models. They also mention that their four-week A/B test also reported statistically significant improvements across short-term and long-term metrics. This is certainly an interesting area of research and practice, like many other industries we will likely see a lot of changes in the status quo.</p><div><hr></div><p><a href="https://openai.com/index/ten-advances-in-mathematics/">OpenAI 10 Advances in Maths</a></p><p>OpenAI (self) reports that an internal version of its forthcoming Astra model generated solutions to ten long-standing problems across geometry, coding theory, group theory, circuit complexity, quantum complexity, lattice problems and extremal combinatorics - whether hype, marketing or fact, it is encouraging to think about scientific research progressing with support from these models. The reported advances span highly specialised mathematical domains that are beyond my knowledge, but one I found it interesting that the workflows that they used combined large-scale model inference, human manuscript preparation and formal verification. OpenAI estimates that the discovery process would have cost roughly 2k USD at Sol API rates, which likely is their marketing/sales pitch, but if that is the case it&#8217;s indeed much more affordable, even if still out from hobby usage. Let&#8217;s see when that Astra model comes out, with all this talk about AGI that model better be good! But let&#8217;s indeed see when we actually get it in Europe.</p><div><hr></div><h2>Upcoming MLOps Events</h2><p>The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.</p><p>Events we are speaking at this year:</p><ul><li><p><strong><a href="https://signalsconf.io/#tickets">Signals Conference</a></strong> - September @ Berlin</p></li><li><p><strong><a href="https://worldsummit.ai/">World Summit AI Europe</a></strong> - September @ Amsterdam</p></li></ul><p>Other relevant events:</p><ul><li><p><strong><a href="https://events.linuxfoundation.org/kubecon-cloudnativecon-europe/">KubeCon Europe</a></strong> - March @ Amsterdam</p></li><li><p><strong><a href="https://2026.pycon.de/">PyData Berlin</a></strong> - April @ Frankfurt</p></li><li><p><strong><a href="https://www.databricks.com/dataaisummit">Databricks Summit</a></strong> - June @ San Francisco</p></li><li><p><strong><a href="https://www.wearedevelopers.com/world-congress">World Developer Congress</a></strong> - July @ Berlin</p></li><li><p><strong><a href="https://ep2025.europython.eu/">EuroPython 2026</a></strong> - July @ Prague</p></li><li><p><strong><a href="https://euroscipy.org/">EuroSciPy 2026</a></strong> - July @ Krakow</p></li><li><p><strong><a href="https://www.ai-infra-summit.com/">AI Infra Summit 2026</a> </strong>- Sept @ California</p></li><li><p><strong><a href="https://codetalks.com/">Code.Talks 2026</a></strong> - Nov @ Hamburg</p></li><li><p><strong><a href="https://mlopsworld.com/">MLOps World 2026</a></strong> - Nov @ Austin</p></li></ul><p>In case you missed our talks, check our recordings below:</p><ul><li><p>The State of AI in 2025 - <a href="https://www.youtube.com/watch?v=v2LENQOG-Xg">WeAreDevelopers 2025</a></p></li><li><p>Prod Generative AI in 2024 - <a href="https://www.youtube.com/watch?v=0uJGmMZGUJE&amp;list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&amp;index=3">KubeCon AI Day 2025</a></p></li><li><p>The State of AI in 2024 - <a href="https://www.youtube.com/live/AtA2XXo_b5s">WeAreDevelopers 2024</a></p></li><li><p>Responsible AI Workshop Keynote - <a href="http://www.youtube.com/watch?v=57YpXjcj0Ho">NeurIPS 2021</a></p></li><li><p>Practical Guide to ML Explainability - <a href="http://www.youtube.com/watch?v=vq8mDiDODhc">PyCon London</a></p></li><li><p>ML Monitoring: Outliers, Drift, XAI - <a href="http://www.youtube.com/watch?v=QcevzK9ZuDg">PyCon Keynote</a></p></li><li><p>Metadata for E2E MLOps - <a href="https://www.youtube.com/watch?v=OSbH4dfswCY">Kubecon NA 2022</a></p></li><li><p>ML Performance Evaluation at Scale - <a href="http://www.youtube.com/watch?v=8ORl8lu1Eeo">KubeCon Eur 2021</a></p></li><li><p>Industry Strength LLMs - <a href="https://www.youtube.com/watch?v=RVUi_rAFfzU">PyData Global 2022</a></p></li><li><p>ML Security Workshop Keynote - <a href="http://www.youtube.com/watch?v=7XSy5aw8oU8">NeurIPS 2022</a></p></li></ul><div><hr></div><h2>Open Source MLOps Tools</h2><p>Check out the fast-growing ecosystem of production ML tools &amp; frameworks at <a href="http://github.com/EthicalML/awesome-production-machine-learning/">the github repository</a> which has reached over 20,000 &#11088; github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here&#8217;s a few featured open source libraries that we maintain:</p><ul><li><p><strong><a href="https://github.com/besanson/sarc-governance/tree/main">SARC</a> </strong>- Provides wrappers for popular agentic frameworks to enable guardrails and constraints that are enforced through the flow.</p></li><li><p><strong><a href="https://github.com/axsaucedo/kaos">KAOS</a></strong> - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.</p></li><li><p><strong><a href="https://github.com/KomputeProject/kompute">Kompute</a></strong> - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-machine-learning/">Production ML Tools</a></strong> - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-artificial-intelligence-regulation">AI Policy List</a></strong> - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-genai/">Agentic Systems Tools</a></strong> - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain</p></li></ul><p>Please do support some of our open source projects by sharing, contributing or adding a star &#11088;</p><div><hr></div><p><strong>About us</strong></p><p>The Institute for Ethical AI &amp; Machine Learning is a European research centre that carries out world-class research into responsible machine learning.</p><p><a href="https://ethical.institute/">Check out our website</a></p>]]></content:encoded></item><item><title><![CDATA[Issue #397 - The ML Engineer 🤖 ]]></title><description><![CDATA[Europe Launches APERTUS 1.5, Why Software Factories Fail, OpenAI + HuggingFace Security Post-Mortem, Opus 5 is OUT & Learnings Changed, DuckDB Internals Part II + more &#128640;]]></description><link>https://machinelearning.substack.com/p/issue-397-the-ml-engineer</link><guid isPermaLink="false">https://machinelearning.substack.com/p/issue-397-the-ml-engineer</guid><dc:creator><![CDATA[Alejandro Saucedo]]></dc:creator><pubDate>Sun, 26 Jul 2026 17:19:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;ea6dbc52-5758-48ee-a16a-284557c69b67&quot;,&quot;duration&quot;:null}"></div><p>Thank you for being part of over <a href="https://ethical.institute/mle.html">70,000+ ML professionals and enthusiasts</a> who receive weekly articles &amp; tutorials on Machine Learning &amp; MLOps &#129302; You can join the newsletter <a href="https://ethical.institute/mle.html">https://ethical.institute/mle.html</a> &#11088;</p><p>If you like the content please support the newsletter by sharing with your friends via &#9993;&#65039; Email, &#128038; <a href="http://twitter.com/intent/tweet?text=Great%20Newsletter%20on%20Machine%20Learning!%20%20https://ethical.institute/mle.html">Twitter</a>, &#128188; <a href="http://www.linkedin.com/shareArticle?mini=true&amp;url=https://ethical.institute/mle.html&amp;title=The%20Machine%20Learning%20Engineer%20Newsletter&amp;summary=Curated%20news%20about%20machine%20learning%20operations,%20reproducibility,%20explainability%20and%20beyond">Linkedin</a> and &#128213; <a href="http://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fethical.institute%2Fmle.html">Facebook</a>!</p><div><hr></div><h2>This week in ML Engineering:</h2><ul><li><p>Europe Launches <strong><a href="https://www.apertus-ai.org/articles/2026-07-apertus-1-5/">APERTUS 1.5</a></strong></p></li><li><p>Why Software <strong><a href="https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/wsff.md">Factories Fail</a></strong></p></li><li><p>OpenAI + HuggingFace <strong><a href="https://huggingface.co/blog/security-incident-july-2026">Security Post-Mortem</a></strong></p></li><li><p>Opus 5 is <strong><a href="https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models">OUT &amp; Learnings Changed</a></strong></p></li><li><p>DuckDB <strong><a href="https://www.greybeam.ai/blog/duckdb-internals-part-2">Internals Part II</a></strong></p></li><li><p>Open Source <strong><a href="http://github.com/EthicalML/awesome-production-machine-learning/">ML Frameworks</a></strong></p></li><li><p>Awesome AI Guidelines <strong><a href="https://github.com/ethicalml/awesome-artificial-intelligence-guidelines">to check out this week</a></strong></p></li><li><p>+ more &#128640;</p></li></ul><div><hr></div><p><a href="https://www.apertus-ai.org/articles/2026-07-apertus-1-5/">Europe Launches APERTUS 1.5</a></p><p>Europe is on Fire with another Foundation Model! Following last week&#8217;s Soofi release, ETH Zurich + EPFL + Swiss NSC released Apertus 1.5 a huge 70B model with native image understanding and speech processing &#128640;</p><p>This is yet another foundation LLM entry from Europe that brings together multi-modal functionality, and general improvements from the 1.0 across reasoning, tool use and a fourfold increase in context length to 260k tokens.</p><p>It seems that the team achieved this by pretraining Apertus 1.0 with another 4 trillion text and multimodal tokens for the 8B model and 2 trillion for the 70B model on the Alps supercomputer at CSCS.</p><p>The release remains open across weights, data, training details and stated model values, under an Apache 2.0 licence - great news!</p><p>This is another strong result for the European AI ecosystem: last week Soofi showed that a large German research and industry consortium can train a competitive model on sovereign European compute, and Apertus shows how the same infrastructure can support transparent models intended for public administration, healthcare, journalism and commercial deployment.</p><p>Looking forward to seeing the benchmarks and results, it is good to see European sovereign AI progressing and hopefully speeding up!</p><div><hr></div><p><a href="https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/wsff.md">Why Software Factories Fail</a></p><p>Should we stop reading all the code and let agents make all decisions autonomously? Dex from HumanLayer says &#8220;no more&#8221;:</p><ul><li><p>&#8220;More comments, longer comments, and tons of PRs getting merged with no review at all.&#8221;</p></li><li><p>&#8220;Incidents are way up.&#8221;</p></li><li><p>&#8220;Bugs per developer are way up.&#8221;</p></li></ul><p>These are some interesting insights from a few reports from teams have attempted to go full autonomous (or partially) letting agents do everything - they quickly (or slowly) realise that this goes haywire in complex production codebases.</p><p>The current limitations are in some part because current coding models are trained and evaluated primarily on short-term functional correctness rather than long-term maintainability.</p><p>This means that there is a large gap on the software design still, and poor design may only become visible through incidents or difficult modifications weeks or months later. There are newer evals emerging like SWE-Marathon and Frontier Code which have longer tasks + multi-PR workflows, but this is a hard problem that sometimes goes even further beyond.</p><p>It&#8217;s interesting to see this analogy of the software factory, and the importance of the human review - throughout the last few weeks I&#8217;ve been building with Fable and I would echo this with a slightly different angle.</p><p>I have found myself no longer in the same &#8220;mode&#8221; as with Opus, where I was much more &#8220;at the helm&#8221;. With Fable I find myself more like a coach guiding through architectural decisions, often spending most time learning new things on much smarter approaches than I could&#8217;ve thought of...</p><p>Based on this, I would add that as practicioners the value of reading the code (or artifacts!) is not only to ensure it doesn&#8217;t steer in the wrong direction, but also to keep up to date with the accelerated evolution of the codebase, and more importantly to learning what seems to be a significantly larger scope of knowledge than before developers were ever expected/capable of.</p><div><hr></div><p><a href="https://huggingface.co/blog/security-incident-july-2026">OpenAI + HuggingFace Security Post-Mortem</a></p><p>OpenAI and Hugging Face have disclosed a pretty cool (+ scary) security incident where one of their models escaped their sandbox and compromised part of Hugging Face&#8217;s production infrastructure.</p><p>The report from OpenAI explains that GPT-5.6 Sol and a more capable pre-release model (of course...) discovered a zero-day vulnerability in a package-registry cache proxy, obtained Internet access, escalated privileges across the evaluation environment and then chained credentials and additional vulnerabilities to reach Hugging Face systems.</p><p>One thing we can be sure of: we won&#8217;t be seeing that &#8220;pre-release model&#8221; anytime soon here in Europe until it&#8217;s nerfed and released &#129394;</p><p>Hugging Face does say that the agent had limited access to internal datasets and credentials, and no detected tampering with public models, datasets, Spaces or its software supply chain - seems the near future we&#8217;ll see a major outage caused from a badly optimized agent...</p><p>HF used self-hosted GLM 5.2 agents to reconstruct the incident because hosted frontier models rejected parts of the forensic workload and would have required sensitive attack data to leave its environment.</p><p>Super interesting sequence of events, we should definitely be learning from these post mortems, as whether we like it or not, we&#8217;ll be on the other side of these incidents one day, and the more the know the more we&#8217;ll be able to do.</p><div><hr></div><p><a href="https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models">Opus 5 is OUT &amp; Learnings Changed</a></p><p>Claude Opus 5 is out!! Seems this completes the 5.0 saga for Anthropic, and this seems to be the strongest generally available model for coding and knowledge work.</p><p>Best news is that price doesn&#8217;t increase! Opus 4.8 API price continues to be $5 per million input tokens and $25 per million output tokens.</p><p>Anthropic reports that Opus 5 more than doubles Opus 4.8&#8217;s Frontier-Bench result at a lower cost per task, approaches Fable 5 on CursorBench at maximum effort, and performs strongly on automation, computer-use and scientific evaluations.</p><p>These are of course reported by Anthropic so let&#8217;s see how it performs on the wild - I tried to use it this weekend but couldn&#8217;t seem to get access... I guess Europe is still in the queue!</p><p>What is most interesting tho, finally seems Anthropic is starting to remove bloat, with 80% of the system prompt wiped, and with a change on how to reason about working with models.</p><div><hr></div><p><a href="https://www.greybeam.ai/blog/duckdb-internals-part-2">DuckDB Internals Part II</a></p><p>The second part of the &#8220;DuckDB Internals&#8221; series is out! This is a great write-up diving into DuckDB, and here it covers vectorized execution:</p><p>To provide the intuition on vectorization, let&#8217;s take an example where we do a set of operations for each row; instead we can processes batches of rows in column-oriented vectors - this reducing a million-row operation from roughly one million operator calls to about 489.</p><p>In DuckDB this is basically a flat, constant, dictionary that allow values to remain compact or reference existing storage, together with selection vectors that let filters identify surviving rows without copying every column into a new buffer.</p><p>DuckDB then applies precompiled, type-specific functions across these vectors, with validity masks reducing unnecessary NULL checks and simple inner loops giving the compiler opportunities to emit SIMD instructions (aka parallel instructions).</p><p>The engine also uses push-based pipelines, where sources produce chunks, operators transform them, and sinks combine results, making independent branches easier to schedule across threads.</p><p>For production ML practitioners, a useful point here is that high-performance analytical processing depends on controlling data representation and movement as much as on optimizing individual operations, and thinking about data-parallelism can really drive huge improvements.</p><div><hr></div><h2>Upcoming MLOps Events</h2><p>The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.</p><p>Events we are speaking at this year:</p><ul><li><p><strong><a href="https://signalsconf.io/#tickets">Signals Conference</a></strong> - September @ Berlin</p></li><li><p><strong><a href="https://worldsummit.ai/">World Summit AI Europe</a></strong> - September @ Amsterdam</p></li></ul><p>Other relevant events:</p><ul><li><p><strong><a href="https://events.linuxfoundation.org/kubecon-cloudnativecon-europe/">KubeCon Europe</a></strong> - March @ Amsterdam</p></li><li><p><strong><a href="https://2026.pycon.de/">PyData Berlin</a></strong> - April @ Frankfurt</p></li><li><p><strong><a href="https://www.databricks.com/dataaisummit">Databricks Summit</a></strong> - June @ San Francisco</p></li><li><p><strong><a href="https://www.wearedevelopers.com/world-congress">World Developer Congress</a></strong> - July @ Berlin</p></li><li><p><strong><a href="https://ep2025.europython.eu/">EuroPython 2026</a></strong> - July @ Prague</p></li><li><p><strong><a href="https://euroscipy.org/">EuroSciPy 2026</a></strong> - July @ Krakow</p></li><li><p><strong><a href="https://www.ai-infra-summit.com/">AI Infra Summit 2026</a> </strong>- Sept @ California</p></li><li><p><strong><a href="https://codetalks.com/">Code.Talks 2026</a></strong> - Nov @ Hamburg</p></li><li><p><strong><a href="https://mlopsworld.com/">MLOps World 2026</a></strong> - Nov @ Austin</p></li></ul><p>In case you missed our talks, check our recordings below:</p><ul><li><p>The State of AI in 2025 - <a href="https://www.youtube.com/watch?v=v2LENQOG-Xg">WeAreDevelopers 2025</a></p></li><li><p>Prod Generative AI in 2024 - <a href="https://www.youtube.com/watch?v=0uJGmMZGUJE&amp;list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&amp;index=3">KubeCon AI Day 2025</a></p></li><li><p>The State of AI in 2024 - <a href="https://www.youtube.com/live/AtA2XXo_b5s">WeAreDevelopers 2024</a></p></li><li><p>Responsible AI Workshop Keynote - <a href="http://www.youtube.com/watch?v=57YpXjcj0Ho">NeurIPS 2021</a></p></li><li><p>Practical Guide to ML Explainability - <a href="http://www.youtube.com/watch?v=vq8mDiDODhc">PyCon London</a></p></li><li><p>ML Monitoring: Outliers, Drift, XAI - <a href="http://www.youtube.com/watch?v=QcevzK9ZuDg">PyCon Keynote</a></p></li><li><p>Metadata for E2E MLOps - <a href="https://www.youtube.com/watch?v=OSbH4dfswCY">Kubecon NA 2022</a></p></li><li><p>ML Performance Evaluation at Scale - <a href="http://www.youtube.com/watch?v=8ORl8lu1Eeo">KubeCon Eur 2021</a></p></li><li><p>Industry Strength LLMs - <a href="https://www.youtube.com/watch?v=RVUi_rAFfzU">PyData Global 2022</a></p></li><li><p>ML Security Workshop Keynote - <a href="http://www.youtube.com/watch?v=7XSy5aw8oU8">NeurIPS 2022</a></p></li></ul><div><hr></div><h2>Open Source MLOps Tools</h2><p>Check out the fast-growing ecosystem of production ML tools &amp; frameworks at <a href="http://github.com/EthicalML/awesome-production-machine-learning/">the github repository</a> which has reached over 20,000 &#11088; github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here&#8217;s a few featured open source libraries that we maintain:</p><ul><li><p><strong><a href="https://github.com/axsaucedo/kaos">KAOS</a></strong> - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.</p></li><li><p><strong><a href="https://github.com/KomputeProject/kompute">Kompute</a></strong> - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-machine-learning/">Production ML Tools</a></strong> - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-artificial-intelligence-regulation">AI Policy List</a></strong> - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-genai/">Agentic Systems Tools</a></strong> - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain</p></li><li><p><strong><a href="https://github.com/besanson/sarc-governance/tree/main">SARC</a> </strong>- Provides wrappers for popular agentic frameworks to enable guardrails and constraints that are enforced through the flow.</p></li></ul><p>Please do support some of our open source projects by sharing, contributing or adding a star &#11088;</p><div><hr></div><p><strong>About us</strong></p><p>The Institute for Ethical AI &amp; Machine Learning is a European research centre that carries out world-class research into responsible machine learning.</p><p><a href="https://ethical.institute/">Check out our website</a></p>]]></content:encoded></item><item><title><![CDATA[Issue #396 - The ML Engineer 🤖 ]]></title><description><![CDATA[Europe Enters the Open Weights Model Race, In-house LLM Serving at Netflix, Raschka on Optimal Reasoning Models, Mozilla's State of OSS AI, Moonshot AI Kimi K3 Release, Open Source ML Frameworks + mor]]></description><link>https://machinelearning.substack.com/p/issue-396-the-ml-engineer</link><guid isPermaLink="false">https://machinelearning.substack.com/p/issue-396-the-ml-engineer</guid><dc:creator><![CDATA[Alejandro Saucedo]]></dc:creator><pubDate>Sun, 19 Jul 2026 16:50:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;c13619e2-a224-4764-a4c8-99e9bf2ae8ae&quot;,&quot;duration&quot;:null}"></div><p>Thank you for being part of over <a href="https://ethical.institute/mle.html">70,000+ ML professionals and enthusiasts</a> who receive weekly articles &amp; tutorials on Machine Learning &amp; MLOps &#129302; You can join the newsletter <a href="https://bit.ly/state-of-ml-2025">https://bit.ly/state-of-ml-2025</a> &#11088;</p><p>If you like the content please support the newsletter by sharing with your friends via &#9993;&#65039; Email, &#128038; <a href="http://twitter.com/intent/tweet?text=Great%20Newsletter%20on%20Machine%20Learning!%20%20https://ethical.institute/mle.html">Twitter</a>, &#128188; <a href="http://www.linkedin.com/shareArticle?mini=true&amp;url=https://ethical.institute/mle.html&amp;title=The%20Machine%20Learning%20Engineer%20Newsletter&amp;summary=Curated%20news%20about%20machine%20learning%20operations,%20reproducibility,%20explainability%20and%20beyond">Linkedin</a> and &#128213; <a href="http://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fethical.institute%2Fmle.html">Facebook</a>!</p><div><hr></div><h2>This week in ML Engineering:</h2><ul><li><p>Europe Enters the <strong><a href="https://huggingface.co/spaces/Soofi-Project/Pretraining-Tech-Report">Open Weights Model Race</a></strong></p></li><li><p>In-house LLM Serving <strong><a href="https://netflixtechblog.com/in-house-llm-serving-at-netflix-a5a8e799ea2c">at Netflix</a></strong></p></li><li><p>Raschka on Optimal <strong><a href="https://magazine.sebastianraschka.com/p/controlling-reasoning-effort-in-llms">Reasoning Models</a></strong></p></li><li><p>Mozilla&#8217;s State <strong><a href="https://stateofopensource.ai/state-of-open-source-ai-2026.pdf">of OSS AI</a></strong></p></li><li><p>Moonshot AI Kimi <strong><a href="https://www.kimi.com/blog/kimi-k3">K3 Release</a></strong></p></li><li><p>Open Source <strong><a href="http://github.com/EthicalML/awesome-production-machine-learning/">ML Frameworks</a></strong></p></li><li><p>Awesome AI Guidelines <strong><a href="https://github.com/ethicalml/awesome-artificial-intelligence-guidelines">to check out this week</a></strong></p></li><li><p>+ more &#128640;</p></li></ul><div><hr></div><p><a href="https://huggingface.co/spaces/Soofi-Project/Pretraining-Tech-Report">Europe Enters the Open Weights Model Race</a></p><p>Europe has just released one of the strongest open-source models in the world, and the whole thing was trained in Germany, huge win for the European AI ecosystem! A huge consortium came together to train a massive model in Deutsche Telekom&#8217;s German Industrial AI Cloud in Munich on up to 512 NVIDIA B200s. This was a consortium that brings together DFKI, Fraunhofer IAIS and IIS, TU Darmstadt, Universit&#228;t W&#252;rzburg, L3S, Lamarr, <a href="http://hessian.AI">hessian.AI</a>, ellamind and Merantix Momentum. The Soofi S 30B-A3B is a Mixture-of-Experts hybrid Mamba-Transformer (pretty cool to see Mamba in the wild!) 31.6B total parameters, 3.2B active per token, 52 layers of which 23 are Mamba-2, 23 MoE and only 6 GQA attention. It was pretrained on roughly on 26.68T tokens with German up-weighted to as much as 15.3% and context up to 1M tokens. Sovereign compute in Europe has been a talking point for years, and this is the first time it has produced a model I would put behind a production endpoint. Congratulations to Nicolas Flores Herr and the whole consortium.</p><div><hr></div><p><a href="https://netflixtechblog.com/in-house-llm-serving-at-netflix-a5a8e799ea2c">In-house LLM Serving at Netflix</a></p><p>Netflix has published how they run their in-house LLM inference stack at massive scale to power use-cases across their organisation. Netflix shared how they serve LLMs behind the same Model Scoring Service that already handles the rest of their MLOps (ie XGBoost, TensorFlow and PyTorch models). They use primarily NVIDIA Triton for managing model loading, batching and GPU scheduling; they use a Java control plane above it handling versioning, autoscaling and multi-region rollout. They also moved away from TensorRT-LLM to vLLM in summer 2025, which allowed them to introduce custom architectures, hooks for custom decoding logic, easier debugging, and existing familiarity among practitioners. They also seem to have advanced rollout strategies like Red-Black deployments, as well as versioned APIs / models that keep one deployment per modelId / modelVersion pair when the schema changes. I have to say it&#8217;s actually quite exciting to see that the MLOps field (and its experts) is transitioning in real time towards supporting large scale agentic systems, and this is bringing learnings from the last decade of production machine learning.</p><div><hr></div><p><a href="https://magazine.sebastianraschka.com/p/controlling-reasoning-effort-in-llms">Raschka on Optimal Reasoning Models</a></p><p>Sebastian Raschka is taking his pareto curve analysis on model x reasoning x tasks to the next level, by breaking down how reasoning-effort settings at a model architecture level - as always there&#8217;s some really great insights: It seems there are six reasoning levels exposed by GPT-5.6, which are actually trained relative to that reasoning level. Namely it seems that the effort label reaches the model as a system prompt or chat-template flag, but that label only means something if post-training taught the model to respond to it. The two main mechanisms that appear across the six open-weight models he reviews are the effort-conditioned SFT, and the mode-conditioned RLVR. This basically means that the actual per-token cost and length varies depending on the effort, so it&#8217;s not just a configuration parameter the the same model generic model state. As always there are comprehensive examples, but one interesting one to point out is DeepSeek V4, which gives each of its Non-think, Think High and Think Max modes its own context window and length penalty before distilling them into one checkpoint, with a token cost that shrinks as effort rises. What is clear is that we don&#8217;t yet fully understand how the model size and effort both interplay in practice,but we are starting to see more and more benchmarks that will give us the best combination to bring these together in the most efficient and prodctive way.</p><div><hr></div><p><a href="https://stateofopensource.ai/state-of-open-source-ai-2026.pdf">Mozilla&#8217;s State of OSS AI</a></p><p>Mozilla&#8217;s first State of Open Source AI report has been released! It surveyed 1.5k developers to measure the capability gap between OSS and closed models - tldr; it&#8217;s gone: In 2024, open source models vs closed source still have a gap of 8.04% on Chatbot Arena; in February this year the gap has decreased to nearly zero; there is further variability but then again Kimi K3 just dropped! Inference for GPT-4-class cost fell from $20 to $0.40 per million tokens in 36 months and open weights now route roughly a third of OpenRouter tokens. 79% of developers use open models and only 51% ship them, against 63% for closed, and the gap widens rather than closes at enterprise scale (57% vs 73%). Mozilla&#8217;s stack scoring puts standardization (2.83) and enterprise readiness (2.79) as the two weakest criteria across every layer, and the report argues the contest has moved to the orchestration layer above the model. MCP adoption grew from about 2M to 97M monthly SDK downloads in 16 months. From a security standpoint, 30+ CVEs reported in eight weeks and only ~21% of companies reporting mature agent governance. It is exciting to see open weights models helping the democratization of AI to avoid concentrated power within a small number of closed-model players.</p><div><hr></div><p><a href="https://www.kimi.com/blog/kimi-k3">Moonshot AI Kimi K3 Release</a></p><p>Moonshot AI has released Kimi K3, ruining the party for Anthropic&#8217;s Fable and OpenAI&#8217;s Sol! An Open Weights 2.8Tmodel with native vision and a 1M-token window. Impressive! The cost is currently the most surprising part, being extremely competitive with the same cost as Gemini Flash Lite (aka cheap!) but at the performance of Sonnet 5. The architecture combines Kimi Delta Attention with Attention Residuals, and pushes MoE sparsity to 16 active experts out of 896 under a Stable LatentMoE framework, which Moonshot reports gives around 2.5x better scaling efficiency than K2. Training uses quantization-aware training from the SFT stage with MXFP4 weights and MXFP8 activations, along with a balanced expert-parallel scheme using static shapes and no host synchronisation on the critical path. For production ML practitioners the deployment details are super relevant, as Moonshot recommends supernodes of 64 or more accelerators (aka probably not our home lab), and KDA breaks conventional prefix caching, so they have contributed a vLLM implementation to be released with the weights on July 27. In my opinion the self-reported kernel optimisation and chip design case studies are best treated as directional until the technical report lands, but still excited to see how this lands!</p><div><hr></div><h2>Upcoming MLOps Events</h2><p>The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below:</p><h3>Events we are speaking at this year:</h3><ul><li><p><strong><a href="https://signalsconf.io/#tickets">Signals Conference</a></strong> - September @ Berlin</p></li><li><p><strong><a href="https://worldsummit.ai/">World Summit AI Europe</a></strong> - September @ Amsterdam</p></li></ul><h3>Other relevant events:</h3><ul><li><p><strong><a href="https://events.linuxfoundation.org/kubecon-cloudnativecon-europe/">KubeCon Europe</a></strong> - March @ Amsterdam</p></li><li><p><strong><a href="https://2026.pycon.de/">PyData Berlin</a></strong> - April @ Frankfurt</p></li><li><p><strong><a href="https://www.databricks.com/dataaisummit">Databricks Summit</a></strong> - June @ San Francisco</p></li><li><p><strong><a href="https://www.wearedevelopers.com/world-congress">World Developer Congress</a></strong> - July @ Berlin</p></li><li><p><strong><a href="https://ep2025.europython.eu/">EuroPython 2026</a></strong> - July @ Prague</p></li><li><p><strong><a href="https://euroscipy.org/">EuroSciPy 2026</a></strong> - July @ Krakow</p></li><li><p><strong><a href="https://www.ai-infra-summit.com/">AI Infra Summit 2026</a> </strong>- Sept @ California</p></li><li><p><strong><a href="https://codetalks.com/">Code.Talks 2026</a></strong> - Nov @ Hamburg</p></li><li><p><strong><a href="https://mlopsworld.com/">MLOps World 2026</a></strong> - Nov @ Austin</p></li></ul><h3>In case you missed our talks, check our recordings below:</h3><ul><li><p>The State of AI in 2025 - <a href="https://www.youtube.com/watch?v=v2LENQOG-Xg">WeAreDevelopers 2025</a></p></li><li><p>Prod Generative AI in 2024 - <a href="https://www.youtube.com/watch?v=0uJGmMZGUJE&amp;list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&amp;index=3">KubeCon AI Day 2025</a></p></li><li><p>The State of AI in 2024 - <a href="https://www.youtube.com/live/AtA2XXo_b5s">WeAreDevelopers 2024</a></p></li><li><p>Responsible AI Workshop Keynote - <a href="http://www.youtube.com/watch?v=57YpXjcj0Ho">NeurIPS 2021</a></p></li><li><p>Practical Guide to ML Explainability - <a href="http://www.youtube.com/watch?v=vq8mDiDODhc">PyCon London</a></p></li><li><p>ML Monitoring: Outliers, Drift, XAI - <a href="http://www.youtube.com/watch?v=QcevzK9ZuDg">PyCon Keynote</a></p></li><li><p>Metadata for E2E MLOps - <a href="https://www.youtube.com/watch?v=OSbH4dfswCY">Kubecon NA 2022</a></p></li><li><p>ML Performance Evaluation at Scale - <a href="http://www.youtube.com/watch?v=8ORl8lu1Eeo">KubeCon Eur 2021</a></p></li><li><p>Industry Strength LLMs - <a href="https://www.youtube.com/watch?v=RVUi_rAFfzU">PyData Global 2022</a></p></li><li><p>ML Security Workshop Keynote - <a href="http://www.youtube.com/watch?v=7XSy5aw8oU8">NeurIPS 2022</a></p></li></ul><h2>Open Source MLOps Tools</h2><p>Check out the fast-growing ecosystem of production ML tools &amp; frameworks at <a href="http://github.com/EthicalML/awesome-production-machine-learning/">the github repository</a> which has reached over 20,000 &#11088; github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here&#8217;s a few featured open source libraries that we maintain:</p><ul><li><p><strong><a href="https://github.com/besanson/sarc-governance/tree/main">SARC</a> </strong>- Provides wrappers for popular agentic frameworks to enable guardrails and constraints that are enforced through the flow.</p></li><li><p><strong><a href="https://github.com/axsaucedo/kaos">KAOS</a></strong> - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.</p></li><li><p><strong><a href="https://github.com/KomputeProject/kompute">Kompute</a></strong> - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-machine-learning/">Production ML Tools</a></strong> - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-artificial-intelligence-regulation">AI Policy List</a></strong> - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-genai/">Agentic Systems Tools</a></strong> - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain</p></li></ul><p>Please do support some of our open source projects by sharing, contributing or adding a star &#11088;</p><div><hr></div><p><strong>About us</strong></p><p>The Institute for Ethical AI &amp; Machine Learning is a European research centre that carries out world-class research into responsible machine learning.</p><p><a href="https://ethical.institute/">Check out our website</a></p>]]></content:encoded></item><item><title><![CDATA[Issue #395 - The ML Engineer 🤖 ]]></title><description><![CDATA[World Models in Gaming with Epic Games, Model's Pareto with Raschka and Databricks, META Enters the Coding Model Arena, Top 30 Papers to Read in ML, SpaceXAI Trains Grok 4.5 from Cursor Data + more &#128640;]]></description><link>https://machinelearning.substack.com/p/issue-395-the-ml-engineer</link><guid isPermaLink="false">https://machinelearning.substack.com/p/issue-395-the-ml-engineer</guid><dc:creator><![CDATA[Alejandro Saucedo]]></dc:creator><pubDate>Mon, 13 Jul 2026 06:01:39 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;6b7e6c7a-0596-4178-8385-1c6037002058&quot;,&quot;duration&quot;:null}"></div><p>Thank you for being part of over <a href="https://ethical.institute/mle.html">70,000+ ML professionals and enthusiasts</a> who receive weekly articles &amp; tutorials on Machine Learning &amp; MLOps &#129302; You can join the newsletter <a href="https://bit.ly/state-of-ml-2025">https://bit.ly/state-of-ml-2025</a> &#11088;</p><p>If you like the content please support the newsletter by sharing with your friends via &#9993;&#65039; Email, &#128038; <a href="http://twitter.com/intent/tweet?text=Great%20Newsletter%20on%20Machine%20Learning!%20%20https://ethical.institute/mle.html">Twitter</a>, &#128188; <a href="http://www.linkedin.com/shareArticle?mini=true&amp;url=https://ethical.institute/mle.html&amp;title=The%20Machine%20Learning%20Engineer%20Newsletter&amp;summary=Curated%20news%20about%20machine%20learning%20operations,%20reproducibility,%20explainability%20and%20beyond">Linkedin</a> and &#128213; <a href="http://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fethical.institute%2Fmle.html">Facebook</a>!</p><h2>This week in ML Engineering:</h2><ul><li><p>World Models in Gaming <strong><a href="https://arxiv.org/abs/2607.05352">with Epic Games</a></strong></p></li><li><p>Model&#8217;s Pareto with <strong><a href="https://x.com/rasbt/status/2075982283509571666">Raschka</a></strong> and <strong><a href="https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase">Databricks</a></strong></p></li><li><p>META Enters <strong><a href="https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/">the Coding Model Arena</a></strong></p></li><li><p>Top 30 Papers <strong><a href="https://30papers.com/">to Read in ML</a></strong></p></li><li><p>SpaceXAI Trains Grok 4.5 <strong><a href="https://x.ai/news/grok-4-5">from Cursor Data</a></strong></p></li><li><p>Open Source <strong><a href="http://github.com/EthicalML/awesome-production-machine-learning/">ML Frameworks</a></strong></p></li><li><p>Awesome AI Guidelines <strong><a href="https://github.com/ethicalml/awesome-artificial-intelligence-guidelines">to check out this week</a></strong></p></li><li><p>+ more &#128640;</p></li></ul><div><hr></div><p><a href="https://arxiv.org/abs/2607.05352">World Models in Gaming with Epic Games</a></p><p>Researchers from Epic Games &amp; General Intuition have released a new AI world model that lets you play online/free a version of Rocket League that runs purely on an ML model. This is bascially a 5B latent diffusion world model that generates a four-player 2v2 Rocket League match from synchronized video context and each player&#8217;s actions. It&#8217;s pretty cool to see how these models are trained on controller inputs, predicting the next frame; in this case it was trained on ~10,000 hours of bot gameplay, the model jointly renders four mutually consistent viewpoints at 20 frames per second on one Nvidia B200 GPU. The experiments indicate that pretrained visual representations and diffusion forcing materially reduce long-horizon drift, and that multiplayer conditioning improves the treatment of off-screen agents and shared physical events relative to single-view modelling; this is cool because it shows that adding the actions and viewpoints of other players can improve the model&#8217;s estimate of the shared game state. The way that they evaluated it was also pretty interesting, as they did not rely only on visual-quality metrics, but also measured whether actions could be recovered from generated video and whether internal representations preserved information about car and ball positions. This is a super exciting space as learned simulators could eventually support multi-agent training, policy evaluation and interactive environments without requiring direct access to the original game engine.</p><div><hr></div><p>Model&#8217;s Pareto with <a href="https://x.com/rasbt/status/2075982283509571666">Raschka</a> and <a href="https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase">Databricks</a></p><p>It&#8217;s pretty cool to see coding models now being presented on a Pareto-like curve across task-completion vs cost instead of just a leaderboard with one percentage per model; here&#8217;s Raschka&#8217;s and DBX&#8217;s takes: Sebastian Raschka dropped a super interesting chart that shows the model curves across reasoning settings, showing that the highest-scoring configuration is not necessarily the appropriate choice under a fixed budget. Databricks also dropped a super interesting analysis bechmarked on a multi-million-line codebase, which plots overall pass rate against mean cost per task subject to the harness. These two are slightly different but complementary takes, and they are super interesting as they are making it clear that evaluating models alone is no longer enough; we need to also consider other parameters that will likely become growingly important as models start seeing diminishing returns. For example, it was super interesting to see that using Pi as the harness results in 1.20 and 2.08 times cheaper than the corresponding native tools in the reported comparisons. Similarly the jumps from Sol / Terra / Luna across the various reasoning levels, showing the tradeoffs across each, and also the consideration between switchign across model families vs taking a cost hit for consistency. It is clearly now expected that this will be a growing trend, and I am excited to see more and more practical benchmarks that are taken from hands on exercises as opposed to purely benchmarks that are being gamed.</p><div><hr></div><p><a href="https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/">META Enters the Coding Model Arena</a></p><p>Meta has entered the AI race with their very own coding model with support for 1m context and highly subsidised token API costs (for now!). Meta reports 80% on OSWorld, 69% on WebArena, 88.1 on MCP Atlas, and 61.5 on SWE-Bench Pro, however, we&#8217;ll have to see how it performs in practise once it hits the ground running with the community. Here&#8217;s the full model report, which shows how interesting insights of the &#8220;unmitigated model&#8221;, which echoes the same risks that other providers mention like anthropic on the pre-release mythos, so likely we&#8217;ll be interacting with a highly nerfed model (e.g. after the &#8220;US Customs Approval&#8221;). It is indeed interesting to see that everyone in the AI space is reducing towards the same average when it comes to offerings, the question will be whether there is indeed a real differentiator on the model/harness, or whether, at the end, the main competitive advantage will be critical mass pricing discounts.</p><div><hr></div><p><a href="https://30papers.com/">Top 30 Papers to Read in ML</a></p><p>Here&#8217;s the reconstructed list of 30 papers that Ilya Sutskever shared with John Carmack, and this contains what we can see as the 30 must-read classics in ML. This recommendation list covers the major lines of deep-learning research, including convolutional and recurrent networks, residual connections, attention and Transformers, external memory, graph message passing, scaling laws, pipeline parallelism, and information-theoretic accounts of learning. It&#8217;s also nice to see that for each paper there&#8217;s a brief short explanation (beyond the abstract), so for ML practitioners, this can be a great TODO list for a compact curriculum for understanding architectural and systems concepts that continue to shape the current AI revolution.</p><div><hr></div><p><a href="https://x.ai/news/grok-4-5">TSpaceXAI Trains Grok 4.5 from Cursor Data</a></p><p>X / Xai / SpaceXAI (or whatever is twitter&#8217;s latest nickname) has trained a new coding model by using the entire data from Cursor, and released it as Grok 4.5 - and it&#8217;s offered on a highly subsidised / competitive pricing (for now!). The model is a MoE achitecture trained with trillions of tokens from Cursor interaction data together with STEM and research material, followed by standard reinforcement learning tuning. SpaceXAI reports pretty impressive scores across all benchmarks, with the main driver being efficiency as they are serving throughput of 80 tokens per second and an average of 15,954 output tokens per SWE-Bench Pro task. The API provides a 500k context window, configurable reasoning and tool interfaces at $2 per million input tokens and $6 per million output tokens, which is clearly highly subsidised to gain initial traction. It is interesting to see that although model training is not exactly commoditised, the MOAT that was being spearheaded by OpenAI is no longer far ahead, but the differentiating model is getting narrower - it will be interesting to see whether Anthropic/OpenAI will continue relying on model perf as their MOAT, as at the end there&#8217;s a threshold where price competitiveness seems to win, and the lockdown to a harness is not as strong as it initially was.</p><div><hr></div><h2>Upcoming MLOps Events</h2><p>The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.</p><h3>Events we are speaking at this year:</h3><ul><li><p><strong><a href="https://signalsconf.io/#tickets">Signals Conference</a></strong> - September @ Berlin</p></li><li><p><strong><a href="https://worldsummit.ai/">World Summit AI Europe</a></strong> - September @ Amsterdam</p></li></ul><h3>Other relevant events:</h3><ul><li><p><strong><a href="https://events.linuxfoundation.org/kubecon-cloudnativecon-europe/">KubeCon Europe</a></strong> - March @ Amsterdam</p></li><li><p><strong><a href="https://2026.pycon.de/">PyData Berlin</a></strong> - April @ Frankfurt</p></li><li><p><strong><a href="https://www.databricks.com/dataaisummit">Databricks Summit</a></strong> - June @ San Francisco</p></li><li><p><strong><a href="https://www.wearedevelopers.com/world-congress">World Developer Congress</a></strong> - July @ Berlin</p></li><li><p><strong><a href="https://ep2025.europython.eu/">EuroPython 2026</a></strong> - July @ Prague</p></li><li><p><strong><a href="https://euroscipy.org/">EuroSciPy 2026</a></strong> - July @ Krakow</p></li><li><p><strong><a href="https://www.ai-infra-summit.com/">AI Infra Summit 2026</a> </strong>- Sept @ California</p></li><li><p><strong><a href="https://codetalks.com/">Code.Talks 2026</a></strong> - Nov @ Hamburg</p></li><li><p><strong><a href="https://mlopsworld.com/">MLOps World 2026</a></strong> - Nov @ Austin</p></li></ul><h3>In case you missed our talks, check our recordings below:</h3><ul><li><p>The State of AI in 2025 - <a href="https://www.youtube.com/watch?v=v2LENQOG-Xg">WeAreDevelopers 2025</a></p></li><li><p>Prod Generative AI in 2024 - <a href="https://www.youtube.com/watch?v=0uJGmMZGUJE&amp;list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&amp;index=3">KubeCon AI Day 2025</a></p></li><li><p>The State of AI in 2024 - <a href="https://www.youtube.com/live/AtA2XXo_b5s">WeAreDevelopers 2024</a></p></li><li><p>Responsible AI Workshop Keynote - <a href="http://www.youtube.com/watch?v=57YpXjcj0Ho">NeurIPS 2021</a></p></li><li><p>Practical Guide to ML Explainability - <a href="http://www.youtube.com/watch?v=vq8mDiDODhc">PyCon London</a></p></li><li><p>ML Monitoring: Outliers, Drift, XAI - <a href="http://www.youtube.com/watch?v=QcevzK9ZuDg">PyCon Keynote</a></p></li><li><p>Metadata for E2E MLOps - <a href="https://www.youtube.com/watch?v=OSbH4dfswCY">Kubecon NA 2022</a></p></li><li><p>ML Performance Evaluation at Scale - <a href="http://www.youtube.com/watch?v=8ORl8lu1Eeo">KubeCon Eur 2021</a></p></li><li><p>Industry Strength LLMs - <a href="https://www.youtube.com/watch?v=RVUi_rAFfzU">PyData Global 2022</a></p></li><li><p>ML Security Workshop Keynote - <a href="http://www.youtube.com/watch?v=7XSy5aw8oU8">NeurIPS 2022</a></p></li></ul><div><hr></div><h2>Open Source MLOps Tools</h2><p>Check out the fast-growing ecosystem of production ML tools &amp; frameworks at <a href="http://github.com/EthicalML/awesome-production-machine-learning/">the github repository</a> which has reached over 20,000 &#11088; github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here&#8217;s a few featured open source libraries that we maintain:</p><ul><li><p><strong><a href="https://github.com/besanson/sarc-governance/tree/main">SARC</a> </strong>- Provides wrappers for popular agentic frameworks to enable guardrails and constraints that are enforced through the flow.</p></li><li><p><strong><a href="https://github.com/axsaucedo/kaos">KAOS</a></strong> - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.</p></li><li><p><strong><a href="https://github.com/KomputeProject/kompute">Kompute</a></strong> - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-machine-learning/">Production ML Tools</a></strong> - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-artificial-intelligence-regulation">AI Policy List</a></strong> - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-genai/">Agentic Systems Tools</a></strong> - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain</p></li></ul><p>Please do support some of our open source projects by sharing, contributing or adding a star &#11088;</p><div><hr></div><p><strong>About us</strong></p><p>The Institute for Ethical AI &amp; Machine Learning is a European research centre that carries out world-class research into responsible machine learning.</p>]]></content:encoded></item><item><title><![CDATA[Issue #394 - The ML Engineer 🤖 ]]></title><description><![CDATA[Stanford Study on Local LLMs, Snorkel's Senior SWE-Bench, Google DeepMind's Tabular Foundation Model, Performance Per Dollar Improving, CVEs Spike After Mythos Release + more &#128640;]]></description><link>https://machinelearning.substack.com/p/issue-394-the-ml-engineer</link><guid isPermaLink="false">https://machinelearning.substack.com/p/issue-394-the-ml-engineer</guid><dc:creator><![CDATA[Alejandro Saucedo]]></dc:creator><pubDate>Sun, 05 Jul 2026 16:41:45 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;1a0a00e9-baad-4e7b-a696-a77e811b98bd&quot;,&quot;duration&quot;:null}"></div><p>Thank you for being part of over <a href="https://ethical.institute/mle.html">70,000+ ML professionals and enthusiasts</a> who receive weekly articles &amp; tutorials on Machine Learning &amp; MLOps &#129302; You can join the newsletter <a href="https://bit.ly/state-of-ml-2025">https://bit.ly/state-of-ml-2025</a> &#11088;</p><p>If you like the content please support the newsletter by sharing with your friends via &#9993;&#65039; Email, &#128038; <a href="http://twitter.com/intent/tweet?text=Great%20Newsletter%20on%20Machine%20Learning!%20%20https://ethical.institute/mle.html">Twitter</a>, &#128188; <a href="http://www.linkedin.com/shareArticle?mini=true&amp;url=https://ethical.institute/mle.html&amp;title=The%20Machine%20Learning%20Engineer%20Newsletter&amp;summary=Curated%20news%20about%20machine%20learning%20operations,%20reproducibility,%20explainability%20and%20beyond">Linkedin</a> and &#128213; <a href="http://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fethical.institute%2Fmle.html">Facebook</a>!</p><div><hr></div><h2>This week in ML Engineering:</h2><ul><li><p>Stanford Study <strong><a href="https://arxiv.org/abs/2511.07885">on Local LLMs</a></strong></p></li><li><p>Snorkel&#8217;s <strong><a href="https://senior-swe-bench.snorkel.ai/">Senior SWE-Bench</a></strong></p></li><li><p>Google DeepMind&#8217;s <strong><a href="https://research.google/blog/introducing-tabfm-a-zero-shot-foundation-model-for-tabular-data/">Tabular Foundation Model</a></strong></p></li><li><p>Performance <strong><a href="https://www.wafer.ai/blog/glm52-amd">Per Dollar Improving</a></strong></p></li><li><p>CVEs Spike <strong><a href="https://epoch.ai/data-insights/cve-severity-spike">After Mythos Release</a></strong></p></li><li><p>Open Source <strong><a href="http://github.com/EthicalML/awesome-production-machine-learning/">ML Frameworks</a></strong></p></li><li><p>Awesome AI Guidelines <strong><a href="https://github.com/ethicalml/awesome-artificial-intelligence-guidelines">to check out this week</a></strong></p></li><li><p>+ more &#128640;</p></li></ul><div><hr></div><p><a href="https://arxiv.org/abs/2511.07885">Stanford Study on Local LLMs</a></p><p>A study from Stanford University showed that 71.3% of chatgpt queries could be accurately answered by a local model: Stanford researchers have published a large empirical study that explores whether local LLM inference can take on part of today&#8217;s cloud-served workload. The paper defines intelligence per watt as task accuracy divided by power consumption, and evaluates 20+ local language models across 8 accelerators on more than 1M single-turn chat and reasoning queries. The main result is that local models can correctly handle 88.7% of the studied queries when routed to the best local model, although coverage is much stronger for chat and knowledge-style tasks than for harder technical reasoning. The longitudinal results are also relevant for ML platform teams as it seems 2023-2025 local-query coverage increased from 23.2% to 71.3%, while intelligence per watt improved by 5.3x. Cloud accelerators still retain a clear per-query efficiency advantage on identical models, but hybrid local-cloud routing changes the system-level tradeoff, as just with an 80%-accurate router, the simulated deployment reduces energy by 64.3%, compute by 61.8%, and cost by 59.0% against a batched cloud baseline. This is super promising, as it shows that there&#8217;s so much potential in local models that is really untapped - there&#8217;s really a lot to come from this space.</p><div><hr></div><p><a href="https://senior-swe-bench.snorkel.ai/">Snorkel&#8217;s Senior SWE-Bench</a></p><p>SnorkelAI has dropped a new SWE-Bench for LLMs that attempts to capture Senior Engineering type tasks that extend the scope towards under-specified feature requests, runtime bug investigation, behavioral verification, and code-quality assessment: It is interesting to see the rise of SWE benchmarks attempting different schools of thought to bring balanced reviews of new models; I like that this one is based on real PRs, uses a mix of pre-written verifiers, an adaptive agents to ensure a useful distinction between code that passes tests and code that fits the surrounding codebase. In the reported results, GPT-5.5 has the highest basic solve rate at 55.0%, while Claude Opus 4.8 has the highest tasteful solve rate at 24.0%; GPT-5.5 is also more efficient, averaging 36.3K output tokens and 89 agent steps per task, compared with 117.1K tokens and 131 steps for Claude Opus 4.8. For production ML practitioners building or evaluating coding-agent workflows it is worth noting that pass rates are an incomplete signal now, as teams should also measure root-cause correctness, design quality, abstraction fit, and even codebase-pattern alignment.</p><div><hr></div><p><a href="https://research.google/blog/introducing-tabfm-a-zero-shot-foundation-model-for-tabular-data/">Google DeepMind&#8217;s Tabular Foundation Model</a></p><p>Google Research has now entered the Tabular Foundation Model space! They have released TabFM, a tabular foundation model for classification and regression, which seems to mostly use learnings from the first-mover startups in the space (but still awesome): This model basically treats each table as an in-context learning task (instead of fitting a dataset-specific model), and this is useful because TabFM instead combines row and column attention, row compression and an ICL Transformer, with pre-training on hundreds of millions of synthetic datasets generated from structural causal models. Google reports strong TabArena results across 51 classification and regression datasets, including a zero-shot setting that runs in a single forward pass and an ensemble variant with cross features, SVD features, NNLS blending and calibration. For production ML practitioners, the most immediate value is in fast baselines and prototyping clearly. However, unfortunately it seems that the public weights are non-commercial (boo!), classification is limited to 10 classes, memory scales with the number of context rows, and (obviously) domain-specific validation remains necessary before use in critical systems.</p><div><hr></div><p><a href="https://www.wafer.ai/blog/glm52-amd">Performance Per Dollar Improving</a></p><p>Performance per dollar in local LLMs is getting faster and cheaper! Wafer shows how they served GLM5.2 on AMD MI355X at 2626 tok/s/node and 213 tok/s single stream at over 2x lower cost than Blackwell. This is an interesting case study that showcases how much opportunity there is on local model inference efficiency; 2.4 RPS on a 20k input / 1k output workload with a 60% cache-hit rate, and 213 tok/s single-stream on a 10k input / 1.5k output test is impressive. The result is relevant for production ML teams because the gains came mostly from framework and configuration work rather than new custom kernels: Wafer quantized bf16 GLM-5.2 to MXFP4 with AMD Quark, selected sglang after testing vLLM and ATOM, fixed ROCm-specific speculative decoding issues, enabled FP8 KV cache, and tuned MoE kernel selection for GLM&#8217;s fp4 shapes. Vercel has also made GLM 5.2 Fast via Wafer available through AI Gateway, with its own benchmarks reporting higher throughput than other serverless GLM-5.2 providers across small-context, large-context, and tool-call scenarios.</p><div><hr></div><p><a href="https://epoch.ai/data-insights/cve-severity-spike">CVEs Spike After Mythos Release</a></p><p>Disclosure of serious cyber vulnerabilities spiked around the release of Claude Mythos Preview: is it time to rethink public vulnerability reporting/mgmt? Epoch AI reported that public disclosure of high/critical severity CVEs have been spiking right after Anthropic announced Claude Mythos Preview. In June 2026 the 21 major organizations tracked by Epoch published around 1.5k high/critical severity CVEs which is more than 3.5 times the previous monthly record before Mythos Preview was announced. Anthropic&#8217;s related Project Glasswing update claims that roughly 50 partners have found more than 10k high/critical severity vulnerabilities, while also stating that the limiting factor has shifted from discovery to verification, disclosure, patch development, and deployment. OpenAI&#8217;s Daybreak initiative describes a similar defensive workflow, focused on finding, validating, and fixing vulnerabilities before attackers can use them. It seems that the dawn of agentic cybersecurity has come, and we&#8217;ll be likely seeing quite a few shifts in mindset, approach and criticality to security across every domain and industry.</p><div><hr></div><h2>Upcoming MLOps Events</h2><p>The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.</p><h3>Events we are speaking at this year:</h3><ul><li><p><strong><a href="https://signalsconf.io/#tickets">Signals Conference</a></strong> - September @ Berlin</p></li><li><p><strong><a href="https://worldsummit.ai/">World Summit AI Europe</a></strong> - September @ Amsterdam</p></li></ul><h3>Other relevant events:</h3><ul><li><p><strong><a href="https://events.linuxfoundation.org/kubecon-cloudnativecon-europe/">KubeCon Europe</a></strong> - March @ Amsterdam</p></li><li><p><strong><a href="https://2026.pycon.de/">PyData Berlin</a></strong> - April @ Frankfurt</p></li><li><p><strong><a href="https://www.databricks.com/dataaisummit">Databricks Summit</a></strong> - June @ San Francisco</p></li><li><p><strong><a href="https://www.wearedevelopers.com/world-congress">World Developer Congress</a></strong> - July @ Berlin</p></li><li><p><strong><a href="https://ep2025.europython.eu/">EuroPython 2026</a></strong> - July @ Prague</p></li><li><p><strong><a href="https://euroscipy.org/">EuroSciPy 2026</a></strong> - July @ Krakow</p></li><li><p><strong><a href="https://www.ai-infra-summit.com/">AI Infra Summit 2026</a> </strong>- Sept @ California</p></li><li><p><strong><a href="https://codetalks.com/">Code.Talks 2026</a></strong> - Nov @ Hamburg</p></li><li><p><strong><a href="https://mlopsworld.com/">MLOps World 2026</a></strong> - Nov @ Austin</p></li></ul><h3>In case you missed our talks, check our recordings below:</h3><ul><li><p>The State of AI in 2025 - <a href="https://www.youtube.com/watch?v=v2LENQOG-Xg">WeAreDevelopers 2025</a></p></li><li><p>Prod Generative AI in 2024 - <a href="https://www.youtube.com/watch?v=0uJGmMZGUJE&amp;list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&amp;index=3">KubeCon AI Day 2025</a></p></li><li><p>The State of AI in 2024 - <a href="https://www.youtube.com/live/AtA2XXo_b5s">WeAreDevelopers 2024</a></p></li><li><p>Responsible AI Workshop Keynote - <a href="http://www.youtube.com/watch?v=57YpXjcj0Ho">NeurIPS 2021</a></p></li><li><p>Practical Guide to ML Explainability - <a href="http://www.youtube.com/watch?v=vq8mDiDODhc">PyCon London</a></p></li><li><p>ML Monitoring: Outliers, Drift, XAI - <a href="http://www.youtube.com/watch?v=QcevzK9ZuDg">PyCon Keynote</a></p></li><li><p>Metadata for E2E MLOps - <a href="https://www.youtube.com/watch?v=OSbH4dfswCY">Kubecon NA 2022</a></p></li><li><p>ML Performance Evaluation at Scale - <a href="http://www.youtube.com/watch?v=8ORl8lu1Eeo">KubeCon Eur 2021</a></p></li><li><p>Industry Strength LLMs - <a href="https://www.youtube.com/watch?v=RVUi_rAFfzU">PyData Global 2022</a></p></li><li><p>ML Security Workshop Keynote - <a href="http://www.youtube.com/watch?v=7XSy5aw8oU8">NeurIPS 2022</a></p></li></ul><div><hr></div><h2>Open Source MLOps Tools</h2><p>Check out the fast-growing ecosystem of production ML tools &amp; frameworks at <a href="http://github.com/EthicalML/awesome-production-machine-learning/">the github repository</a> which has reached over 20,000 &#11088; github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here&#8217;s a few featured open source libraries that we maintain:</p><ul><li><p><strong><a href="https://github.com/besanson/sarc-governance/tree/main">SARC</a> </strong>- Provides wrappers for popular agentic frameworks to enable guardrails and constraints that are enforced through the flow.</p></li><li><p><strong><a href="https://github.com/axsaucedo/kaos">KAOS</a></strong> - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.</p></li><li><p><strong><a href="https://github.com/KomputeProject/kompute">Kompute</a></strong> - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-machine-learning/">Production ML Tools</a></strong> - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-artificial-intelligence-regulation">AI Policy List</a></strong> - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-genai/">Agentic Systems Tools</a></strong> - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain</p></li></ul><p>Please do support some of our open source projects by sharing, contributing or adding a star &#11088;</p><div><hr></div><p><strong>About us</strong></p><p>The Institute for Ethical AI &amp; Machine Learning is a European research centre that carries out world-class research into responsible machine learning.</p><p><a href="https://ethical.institute/">Check out our website</a></p>]]></content:encoded></item><item><title><![CDATA[ Issue #393 - The ML Engineer 🤖 ]]></title><description><![CDATA[The Future of Agents with Andrew Ng, Sebastian Raschka's Local Agent Setup, Memory in the Age of Agents, PydanticAI V2 Now Released, O'Reilly Radar Trends to Watch in 2026 + more &#128640;]]></description><link>https://machinelearning.substack.com/p/issue-393-the-ml-engineer</link><guid isPermaLink="false">https://machinelearning.substack.com/p/issue-393-the-ml-engineer</guid><dc:creator><![CDATA[Alejandro Saucedo]]></dc:creator><pubDate>Sun, 28 Jun 2026 15:46:39 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;c742777d-f72e-485c-86c4-876bf0353792&quot;,&quot;duration&quot;:null}"></div><p>Thank you for being part of over <a href="https://ethical.institute/mle.html">70,000+ ML professionals and enthusiasts</a> who receive weekly articles &amp; tutorials on Machine Learning &amp; MLOps &#129302; You can join the newsletter <a href="https://bit.ly/state-of-ml-2025">https://bit.ly/state-of-ml-2025</a> &#11088;</p><p>If you like the content please support the newsletter by sharing with your friends via &#9993;&#65039; Email, &#128038; <a href="http://twitter.com/intent/tweet?text=Great%20Newsletter%20on%20Machine%20Learning!%20%20https://ethical.institute/mle.html">Twitter</a>, &#128188; <a href="http://www.linkedin.com/shareArticle?mini=true&amp;url=https://ethical.institute/mle.html&amp;title=The%20Machine%20Learning%20Engineer%20Newsletter&amp;summary=Curated%20news%20about%20machine%20learning%20operations,%20reproducibility,%20explainability%20and%20beyond">Linkedin</a> and &#128213; <a href="http://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fethical.institute%2Fmle.html">Facebook</a>!</p><div><hr></div><h2>This week in ML Engineering:</h2><ul><li><p>The Future of Agents <strong><a href="https://www.youtube.com/watch?v=OaRhpwz_TGM">with Andrew Ng</a></strong></p></li><li><p>Sebastian Raschka&#8217;s <strong><a href="https://magazine.sebastianraschka.com/p/using-local-coding-agents">Local Agent Setup</a></strong></p></li><li><p>Memory <strong><a href="https://arxiv.org/abs/2512.13564">in the Age of Agents</a></strong></p></li><li><p>PydanticAI <strong><a href="https://pydantic.dev/articles/pydantic-ai-v2">V2 Now Released</a></strong></p></li><li><p>O&#8217;Reilly Radar Trends <strong><a href="https://www.oreilly.com/radar/radar-trends-to-watch-june-2026/">to Watch in 2026</a></strong></p></li><li><p>Open Source <strong><a href="http://github.com/EthicalML/awesome-production-machine-learning/">ML Frameworks</a></strong></p></li><li><p>Awesome AI Guidelines <strong><a href="https://github.com/ethicalml/awesome-artificial-intelligence-guidelines">to check out this week</a></strong></p></li><li><p>+ more &#128640;</p></li></ul><div><hr></div><p><a href="https://www.youtube.com/watch?v=OaRhpwz_TGM">The Future of Agents with Andrew Ng</a></p><p>This is one of the best breakdowns on the current state of agentic systems and Agentic Development, by the one and only Andrew Ng! He breaks down his setup, what he&#8217;s working on, and a few predictions for the future: Andrew Ng talked about how coding agents have advanced so fast that they are only now catching up across other domains, such as crafting product definition, legal review, design, marketing, and data access. One point that really resonated is how now all organisations are in a race to sort out their data foundation, as agentic systems are multiplying the value of data by enabling insights that could not have been available at such a speed and detail. There are new delivery bottlenecks, however industry is figuring out how to unblock them with small teams of high-context engineers who can use AI tools across adjacent functions. Andrew Ng also distinguishes incremental automation from broader process redesign, using loan underwriting as an example where value comes from reworking the full business workflow rather than automating one review step. Definitely recommend listening to this fireside, it&#8217;s always surprising how great Andrew&#8217;s takes are!</p><div><hr></div><p><a href="https://magazine.sebastianraschka.com/p/using-local-coding-agents">Sebastian Raschka&#8217;s Local Agent Setup</a></p><p>Sebastian Raschka just dropped a full breakdown of his local LLM agent setup, and as always he has super relevant insights for any engineering practitioner - here&#8217;s a few highlights: Sebastian is clearly an advocate for open-weight models, and seems to be using Ollama heavily as the model-serving layer. In regards to agent harnesses he tends to favour Qwen Code (first time I hear about it), Codex CLI and Claude Code. He really emphasises that local agent workflows depend largely on inference speed, long-context behavior, tool-call reliability, permissions, telemetry, and task-specific evaluation. And as always an article wouldn&#8217;t have the Sebastian&#8217;s signature without an in-depth LLM architectural piece; there was an emphasis on 30&#8211;35B Mixture-of-Experts coding models such as Qwen3.6 35B-A3B, North Mini Code, and Nemotron 3 Nano, as good mode can be usable for routine coding-agent tasks on workstation-class hardware, but that harness choice and token use materially affect performance. Hopefully this becomes a series, and starts also diving into his agentic engineering workflows as I&#8217;d definitely be keen on these!</p><div><hr></div><p><a href="https://arxiv.org/abs/2512.13564">Memory in the Age of Agents</a></p><p>A large consortium of universities from across America, Europe and Asia have published a comprehensive survey on the state of &#8220;Memory&#8221; in agentic systems - key highlights: The paper&#8217;s core is actually a structured taxonomy that separates memory by form, function and dynamics, including what carries memory, why agents need it, and how it is formed. There is an interesting distinction of &#8220;Memory&#8221; across 1) token-level memory, 2) parametric memory and 3) what they call &#8220;latent memory&#8221; - and it maps memory functions into factual, experiential and working memory. It also clarifies how agent memory differs from adjacent areas such as RAG, context engineering and long-context model design - which although there are some intersections, these are completely different beasts alltogether. For production ML practitioners, the paper is useful because it helps standardise concrete engineering choices around memory, including persistence, retrieval quality, auditability, privacy, latency and evaluation. It is becoming clear that memory should be treated as a core system component in long-running agents, and with components that support, short/long-term memory as well as independent/shared memory.</p><div><hr></div><p><a href="https://pydantic.dev/articles/pydantic-ai-v2">PydanticAI V2 Now Released</a></p><p>Pydantic AI v2 is out! Pretty cool to see this agent harness maturing on capabilities - here&#8217;s some of the highlights: They added composable units that package instructions, tools, lifecycle hooks, and model settings which now can be used as lego-blocks. There are now learnings from production agent systems where the operational complexity sits outside the basic model-tool loop, so they added better context control, tool loading, steering, guardrails, code execution, and instrumentation. One of the areas I am particularly excited about is the new defer_loading=True as it lets agents expose a compact catalog and load a workflow only when needed. This is particularly game changing for projects like the K8s Agent OS project that we maintain as each agent may have a long list of MCP Tools and Sub-Agents, and enabling discovery on demand can help optimize context and tokens. We have already updated PydanticAI to v2 in the <a href="https://axsaucedo.github.io/kaos/">K8s Agent OS (KAOS)</a> and are keen to start exploring some of the new features.</p><div><hr></div><p><a href="https://www.oreilly.com/radar/radar-trends-to-watch-june-2026/">O&#8217;Reilly Radar Trends to Watch in 2026</a></p><p>O&#8217;Reilly&#8217;s published the 2026 Radar Trends this month, and this has quite a useful cross-section of where agentic systems are trending towards - here&#8217;s a few highlights: The most relevant theme for production ML practitioners is the emergence of infrastructure that allows agents to provision accounts, register domains, initiate payments, obtain credentials, and deploy applications with limited human intervention. This changes the boundary of MLOps, as teams now need to reason not only about model quality and serving latency, but also about authorization, spending controls, audit trails, sandboxing, and dependency risk. There are also key insights on what is becoming a more fragmented model landscape, with general-purpose frontier models becoming more relevant in specialised contexts. The security section is equally relevant, as it includes AI-assisted vulnerability discovery, supply-chain attacks, and credential leakage from coding agents suggest that agent workflows. Definitely worth checking out!</p><div><hr></div><h2>Upcoming MLOps Events</h2><p>The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.</p><h3>Events we are speaking at this year:</h3><ul><li><p><strong><a href="https://signalsconf.io/#tickets">Signals Conference</a></strong> - September @ Berlin</p></li><li><p><strong><a href="https://worldsummit.ai/">World Summit AI Europe</a></strong> - September @ Amsterdam</p></li></ul><h3>Other relevant events:</h3><ul><li><p><strong><a href="https://events.linuxfoundation.org/kubecon-cloudnativecon-europe/">KubeCon Europe</a></strong> - March @ Amsterdam</p></li><li><p><strong><a href="https://2026.pycon.de/">PyData Berlin</a></strong> - April @ Frankfurt</p></li><li><p><strong><a href="https://www.databricks.com/dataaisummit">Databricks Summit</a></strong> - June @ San Francisco</p></li><li><p><strong><a href="https://www.wearedevelopers.com/world-congress">World Developer Congress</a></strong> - July @ Berlin</p></li><li><p><strong><a href="https://ep2025.europython.eu/">EuroPython 2026</a></strong> - July @ Prague</p></li><li><p><strong><a href="https://euroscipy.org/">EuroSciPy 2026</a></strong> - July @ Krakow</p></li><li><p><strong><a href="https://www.ai-infra-summit.com/">AI Infra Summit 2026</a> </strong>- Sept @ California</p></li><li><p><strong><a href="https://codetalks.com/">Code.Talks 2026</a></strong> - Nov @ Hamburg</p></li><li><p><strong><a href="https://mlopsworld.com/">MLOps World 2026</a></strong> - Nov @ Austin</p></li></ul><h3>In case you missed our talks, check our recordings below:</h3><ul><li><p>The State of AI in 2025 - <a href="https://www.youtube.com/watch?v=v2LENQOG-Xg">WeAreDevelopers 2025</a></p></li><li><p>Prod Generative AI in 2024 - <a href="https://www.youtube.com/watch?v=0uJGmMZGUJE&amp;list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&amp;index=3">KubeCon AI Day 2025</a></p></li><li><p>The State of AI in 2024 - <a href="https://www.youtube.com/live/AtA2XXo_b5s">WeAreDevelopers 2024</a></p></li><li><p>Responsible AI Workshop Keynote - <a href="http://www.youtube.com/watch?v=57YpXjcj0Ho">NeurIPS 2021</a></p></li><li><p>Practical Guide to ML Explainability - <a href="http://www.youtube.com/watch?v=vq8mDiDODhc">PyCon London</a></p></li><li><p>ML Monitoring: Outliers, Drift, XAI - <a href="http://www.youtube.com/watch?v=QcevzK9ZuDg">PyCon Keynote</a></p></li><li><p>Metadata for E2E MLOps - <a href="https://www.youtube.com/watch?v=OSbH4dfswCY">Kubecon NA 2022</a></p></li><li><p>ML Performance Evaluation at Scale - <a href="http://www.youtube.com/watch?v=8ORl8lu1Eeo">KubeCon Eur 2021</a></p></li><li><p>Industry Strength LLMs - <a href="https://www.youtube.com/watch?v=RVUi_rAFfzU">PyData Global 2022</a></p></li><li><p>ML Security Workshop Keynote - <a href="http://www.youtube.com/watch?v=7XSy5aw8oU8">NeurIPS 2022</a></p></li></ul><div><hr></div><h2>Open Source MLOps Tools</h2><p>Check out the fast-growing ecosystem of production ML tools &amp; frameworks at <a href="http://github.com/EthicalML/awesome-production-machine-learning/">the github repository</a> which has reached over 20,000 &#11088; github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here&#8217;s a few featured open source libraries that we maintain:</p><ul><li><p><strong><a href="https://github.com/besanson/sarc-governance/tree/main">SARC</a> </strong>- Provides wrappers for popular agentic frameworks to enable guardrails and constraints that are enforced through the flow.</p></li><li><p><strong><a href="https://github.com/axsaucedo/kaos">KAOS</a></strong> - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.</p></li><li><p><strong><a href="https://github.com/KomputeProject/kompute">Kompute</a></strong> - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-machine-learning/">Production ML Tools</a></strong> - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-artificial-intelligence-regulation">AI Policy List</a></strong> - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-genai/">Agentic Systems Tools</a></strong> - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain</p></li></ul><p>Please do support some of our open source projects by sharing, contributing or adding a star &#11088;</p><div><hr></div><p><strong>About us</strong></p><p>The Institute for Ethical AI &amp; Machine Learning is a European research centre that carries out world-class research into responsible machine learning.</p><p><a href="https://ethical.institute/">Check out our website</a></p>]]></content:encoded></item><item><title><![CDATA[Issue #392 - The ML Engineer 🤖 ]]></title><description><![CDATA[Data Dog's State of AI Engineering, 10 Year of Clickhouse, Cornell: Advanced Compilers, DuckDB Internals and Speed, Bringing GPU Kernels to Rust + more &#128640;]]></description><link>https://machinelearning.substack.com/p/issue-392-the-ml-engineer</link><guid isPermaLink="false">https://machinelearning.substack.com/p/issue-392-the-ml-engineer</guid><dc:creator><![CDATA[Alejandro Saucedo]]></dc:creator><pubDate>Sun, 21 Jun 2026 14:18:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;5012bbd5-c997-4d34-9d5b-8b26476b9e6e&quot;,&quot;duration&quot;:null}"></div><p>Thank you for being part of over <a href="https://ethical.institute/mle.html">70,000+ ML professionals and enthusiasts</a> who receive weekly articles &amp; tutorials on Machine Learning &amp; MLOps &#129302; You can join the newsletter <a href="https://bit.ly/state-of-ml-2025">https://bit.ly/state-of-ml-2025</a> &#11088;</p><p>If you like the content please support the newsletter by sharing with your friends via &#9993;&#65039; Email, &#128038; <a href="http://twitter.com/intent/tweet?text=Great%20Newsletter%20on%20Machine%20Learning!%20%20https://ethical.institute/mle.html">Twitter</a>, &#128188; <a href="http://www.linkedin.com/shareArticle?mini=true&amp;url=https://ethical.institute/mle.html&amp;title=The%20Machine%20Learning%20Engineer%20Newsletter&amp;summary=Curated%20news%20about%20machine%20learning%20operations,%20reproducibility,%20explainability%20and%20beyond">Linkedin</a> and &#128213; <a href="http://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fethical.institute%2Fmle.html">Facebook</a>!</p><div><hr></div><h2>This week in ML Engineering:</h2><ul><li><p>Data Dog&#8217;s State <strong><a href="https://www.datadoghq.com/state-of-ai-engineering/">of AI Engineering</a></strong></p></li><li><p>10 Year <strong><a href="https://clickhouse.com/blog/open-source-10">of Clickhouse</a></strong></p></li><li><p>Cornell: <strong><a href="https://www.cs.cornell.edu/courses/cs6120/2025fa/self-guided/">Advanced Compilers</a></strong></p></li><li><p>DuckDB <strong><a href="https://www.greybeam.ai/blog/duckdb-internals-part-1">Internals and Speed</a></strong></p></li><li><p>Bringing <strong><a href="https://arxiv.org/abs/2606.15991">GPU Kernels to Rust</a></strong></p></li><li><p>Open Source <strong><a href="http://github.com/EthicalML/awesome-production-machine-learning/">ML Frameworks</a></strong></p></li><li><p>Awesome AI Guidelines <strong><a href="https://github.com/ethicalml/awesome-artificial-intelligence-guidelines">to check out this week</a></strong></p></li><li><p>+ more &#128640;</p></li></ul><div><hr></div><p><a href="https://www.datadoghq.com/state-of-ai-engineering/">Data Dog&#8217;s State of AI Engineering</a></p><p>Datadog has analysed LLM telemetry data from more than a thousand Datadog customers and published a snapshot on the &#8220;State of AI Engineering&#8221; - here&#8217;s some highlights: Telemetry data shows that more than 70% of organizations use three or more models, agent framework adoption has nearly doubled and 69% of input tokens are system prompts. In regards to routing, prompt caching is still only present in 28% of eligible calls, and rate limits are one of the biggest reliability problems in production LLM calls. Some trends show that teams are now running multi-model stacks, heavier agent frameworks, long system prompts, and tool-heavy workflows that behave much more like distributed systems than simple API integrations. It is clear that AI engineering now needs the same discipline as platform engineering (+ always has!), including model gateways, continuous evals, prompt layout hygiene, caching, context engineering, budgets, backoff, queues, and fallback capacity. The report also shows that most agents are still fairly simple from a service-topology perspective, with 59% making only one service call, so the next set of engineering challenges will come from tracing, debugging, and governing these systems as they become more distributed.</p><div><hr></div><p><a href="https://clickhouse.com/blog/open-source-10">10 Year of Clickhouse</a></p><p>ClickHouse just turned 10 years old as an open-source project, and their CTO Alexey Milovidov published a fantastic engineering history of what it takes to build a production database from actual painpoints: ClickHouse evolved from OLAPServer and Metrage into a from-scratch columnar DBMS. They decided to add key differentiating features like in-memory columns, aggregate functions, table engines, compression, SQL parsing, MergeTree for background sorting, and ReplicatedMergeTree for multi-DC production use. This came up from wanting to address the pains of growing data volumes, real-time logs, slow MySQL paths, custom C++ data structures, and users waiting for analytics to load. For us ML Engineering / MLOps practitioners it&#8217;s a good reminder on how AI systems now need exactly the same properties ClickHouse was built around, including fast analytical queries over high-volume event data, long-retention observability, cheap aggregation, and infrastructure that can handle messy real-time workloads. It is also a good reminder that serious open source is not just throwing code on GitHub; ClickHouse really shows how they spearheaded this domain as well.</p><div><hr></div><p><a href="https://www.cs.cornell.edu/courses/cs6120/2025fa/self-guided/">Cornell: Advanced Compilers</a></p><p>Compilers are one of the most fun and challenging domains in computer science, and they are becoming a production ML topic again - this course from Cornell is a fantastic deep dive into modern / advanced compiler concepts: Especially given the importance of efficiency in ML these days, we can no longer expect to throw larger and more GPUs, especially as teams run into graph breaks, custom kernels, dynamic shapes, hardware-specific inference paths, and performance issues that cannot be solved by just asking for bigger hardware. Cornell&#8217;s CS 6120 Advanced Compilers course is now available as a self-guided online course, and it looks like a great resource for ML practitioners who want to understand the systems layer underneath torch.compile, XLA, MLIR, Triton, LLVM, and modern inference stacks. The course is PhD-level but still very hands-on, covering intermediate representations, data flow, SSA, local/global optimizations, loop optimization, LLVM passes, alias analysis, garbage collection, JIT/dynamic compilation, parallelism, and fast compilers through classic papers and open-source implementation tasks using LLVM and Bril. I haven&#8217;t picked up the dragon book since back in university, and that&#8217;s one of the books / courses I&#8217;ve enjoyed the most, so I am definitely adding this to my todo list, hopefully I get some time (as the list is only growing larger!!).</p><div><hr></div><p><a href="https://www.greybeam.ai/blog/duckdb-internals-part-1">DuckDB Internals and Speed</a></p><p>DuckDB is becoming one of the default local engines for ML/data workflows, but have you wondered what&#8217;s going on in its internals to work so well? Basically a lot of the speed comes from the runs in-process, so Python/R applications avoid the server round trip and a lot of row-by-row serialization overhead that still shows up with ODBC/JDBC-style paths. This requires understanding the query setup path, including parsing, binding, optimizer passes such as filter pushdown, subquery unnesting, join ordering, and runtime join-filter pushdown. This then follows by the physical plan which is split into pipelines separated by sinks like GROUP BY, ORDER BY, and hash-join build phases. DuckDB&#8217;s native format and Parquet both give it columnar reads, row-group statistics / zone maps, and byte-range reads on remote files, so many feature analysis, eval, debugging, and batch analytics workloads can stay in simple Parquet + SQL without immediately reaching for a warehouse job or distributed cluster. This multi-part series is a great way to get started into the DuckDB internals, so definitely recommended as a deep dive resource.</p><div><hr></div><p><a href="https://arxiv.org/abs/2606.15991">Bringing GPU Kernels to Rust</a></p><p>NVIDIA has published a new framework + paper to bring Rust&#8217;s memory safety into GPU Kernels! Unfortunately still CUDA-specific, but this does look like quite an exciting leap! We all know that GPU kernel work is becoming a much bigger part of production ML engineering, especially as teams push harder on custom approaches to unlock every bit of performance. This paper introduces cuTile Rust, a tile-based GPU kernel system that brings Rust&#8217;s ownership model across the CPU/GPU boundary, including mutable tensors that are split into disjoint partitions, immutable tensors that are shared safely, and kernel launches that preserve ownership while GPU work is still running. It&#8217;s impressive to see that the safety features don&#8217;t seem to add a major performance bottleneck - on B200 they report 7 TB/s for elementwise operations and around 2 PFLOP/s for GEMM, reaching 96% of cuBLAS. For us ML Engineering / MLOps practitioners this could be the start of an exciting trend, which although right now is limited to NVIDIA/CUDA, most likely very soon this could come to other frameworks like Vulkan, and unlock cross-vendor GPU compute.</p><div><hr></div><h2>Upcoming MLOps Events</h2><p>The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.</p><h3>Events we are speaking at this year:</h3><ul><li><p><strong><a href="https://signalsconf.io/#tickets">Signals Conference</a></strong> - September @ Berlin</p></li><li><p><strong><a href="https://worldsummit.ai/">World Summit AI Europe</a></strong> - September @ Amsterdam</p></li></ul><h3>Other relevant events:</h3><ul><li><p><strong><a href="https://events.linuxfoundation.org/kubecon-cloudnativecon-europe/">KubeCon Europe</a></strong> - March @ Amsterdam</p></li><li><p><strong><a href="https://2026.pycon.de/">PyData Berlin</a></strong> - April @ Frankfurt</p></li><li><p><strong><a href="https://www.databricks.com/dataaisummit">Databricks Summit</a></strong> - June @ San Francisco</p></li><li><p><strong><a href="https://www.wearedevelopers.com/world-congress">World Developer Congress</a></strong> - July @ Berlin</p></li><li><p><strong><a href="https://ep2025.europython.eu/">EuroPython 2026</a></strong> - July @ Prague</p></li><li><p><strong><a href="https://euroscipy.org/">EuroSciPy 2026</a></strong> - July @ Krakow</p></li><li><p><strong><a href="https://www.ai-infra-summit.com/">AI Infra Summit 2026</a> </strong>- Sept @ California</p></li><li><p><strong><a href="https://codetalks.com/">Code.Talks 2026</a></strong> - Nov @ Hamburg</p></li><li><p><strong><a href="https://mlopsworld.com/">MLOps World 2026</a></strong> - Nov @ Austin</p></li></ul><h3>In case you missed our talks, check our recordings below:</h3><ul><li><p>The State of AI in 2025 - <a href="https://www.youtube.com/watch?v=v2LENQOG-Xg">WeAreDevelopers 2025</a></p></li><li><p>Prod Generative AI in 2024 - <a href="https://www.youtube.com/watch?v=0uJGmMZGUJE&amp;list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&amp;index=3">KubeCon AI Day 2025</a></p></li><li><p>The State of AI in 2024 - <a href="https://www.youtube.com/live/AtA2XXo_b5s">WeAreDevelopers 2024</a></p></li><li><p>Responsible AI Workshop Keynote - <a href="http://www.youtube.com/watch?v=57YpXjcj0Ho">NeurIPS 2021</a></p></li><li><p>Practical Guide to ML Explainability - <a href="http://www.youtube.com/watch?v=vq8mDiDODhc">PyCon London</a></p></li><li><p>ML Monitoring: Outliers, Drift, XAI - <a href="http://www.youtube.com/watch?v=QcevzK9ZuDg">PyCon Keynote</a></p></li><li><p>Metadata for E2E MLOps - <a href="https://www.youtube.com/watch?v=OSbH4dfswCY">Kubecon NA 2022</a></p></li><li><p>ML Performance Evaluation at Scale - <a href="http://www.youtube.com/watch?v=8ORl8lu1Eeo">KubeCon Eur 2021</a></p></li><li><p>Industry Strength LLMs - <a href="https://www.youtube.com/watch?v=RVUi_rAFfzU">PyData Global 2022</a></p></li><li><p>ML Security Workshop Keynote - <a href="http://www.youtube.com/watch?v=7XSy5aw8oU8">NeurIPS 2022</a></p></li></ul><div><hr></div><h2>Open Source MLOps Tools</h2><p>Check out the fast-growing ecosystem of production ML tools &amp; frameworks at <a href="http://github.com/EthicalML/awesome-production-machine-learning/">the github repository</a> which has reached over 20,000 &#11088; github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here&#8217;s a few featured open source libraries that we maintain:</p><ul><li><p><strong><a href="https://github.com/besanson/sarc-governance/tree/main">SARC</a> </strong>- Provides wrappers for popular agentic frameworks to enable guardrails and constraints that are enforced through the flow.</p></li><li><p><strong><a href="https://github.com/axsaucedo/kaos">KAOS</a></strong> - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.</p></li><li><p><strong><a href="https://github.com/KomputeProject/kompute">Kompute</a></strong> - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-machine-learning/">Production ML Tools</a></strong> - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-artificial-intelligence-regulation">AI Policy List</a></strong> - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-genai/">Agentic Systems Tools</a></strong> - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain</p></li></ul><p>Please do support some of our open source projects by sharing, contributing or adding a star &#11088;</p><div><hr></div><p><strong>About us</strong></p><p>The Institute for Ethical AI &amp; Machine Learning is a European research centre that carries out world-class research into responsible machine learning.</p><p><a href="https://ethical.institute/">Check out our website</a></p>]]></content:encoded></item><item><title><![CDATA[Issue #391 - The ML Engineer 🤖 ]]></title><description><![CDATA[Autonomous Agentic Systems, OpenAI Codex Engineering, Anthropic on Self Service Analytics, Amazon's Tabluar Foundation Models, Kimi 2.7 Code Model Release + more &#128640;]]></description><link>https://machinelearning.substack.com/p/issue-391-the-ml-engineer</link><guid isPermaLink="false">https://machinelearning.substack.com/p/issue-391-the-ml-engineer</guid><dc:creator><![CDATA[Alejandro Saucedo]]></dc:creator><pubDate>Sun, 14 Jun 2026 15:57:58 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;d3f62fd6-1196-4525-beaa-6018103eac73&quot;,&quot;duration&quot;:null}"></div><p>Thank you for being part of over <a href="https://ethical.institute/mle.html">70,000+ ML professionals and enthusiasts</a> who receive weekly articles &amp; tutorials on Machine Learning &amp; MLOps &#129302; You can join the newsletter <a href="https://bit.ly/state-of-ml-2025">https://bit.ly/state-of-ml-2025</a> &#11088;</p><p>If you like the content please support the newsletter by sharing with your friends via &#9993;&#65039; Email, &#128038; <a href="http://twitter.com/intent/tweet?text=Great%20Newsletter%20on%20Machine%20Learning!%20%20https://ethical.institute/mle.html">Twitter</a>, &#128188; <a href="http://www.linkedin.com/shareArticle?mini=true&amp;url=https://ethical.institute/mle.html&amp;title=The%20Machine%20Learning%20Engineer%20Newsletter&amp;summary=Curated%20news%20about%20machine%20learning%20operations,%20reproducibility,%20explainability%20and%20beyond">Linkedin</a> and &#128213; <a href="http://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fethical.institute%2Fmle.html">Facebook</a>!</p><div><hr></div><h2>This week in ML Engineering:</h2><ul><li><p>Autonomous <strong><a href="https://hackernoon.com/autonomous-agentic-systems-a-practical-guide-to-always-on-agents">Agentic Systems</a></strong></p></li><li><p>OpenAI <strong><a href="https://openai.com/index/harness-engineering/">Codex Engineering</a></strong></p></li><li><p>Anthropic <strong><a href="https://claude.com/blog/how-anthropic-enables-self-service-data-analytics-with-claude">on Self Service Analytics</a></strong></p></li><li><p>Amazon&#8217;s <strong><a href="https://www.amazon.science/blog/mitra-mixed-synthetic-priors-for-enhancing-tabular-foundation-models">Tabluar Foundation Models</a></strong></p></li><li><p>Kimi 2.7 <strong><a href="https://huggingface.co/moonshotai/Kimi-K2.7-Code">Code Model Release</a></strong></p></li><li><p>Open Source <strong><a href="http://github.com/EthicalML/awesome-production-machine-learning/">ML Frameworks</a></strong></p></li><li><p>Awesome AI Guidelines <strong><a href="https://github.com/ethicalml/awesome-artificial-intelligence-guidelines">to check out this week</a></strong></p></li><li><p>+ more &#128640;</p></li></ul><div><hr></div><p><a href="https://hackernoon.com/autonomous-agentic-systems-a-practical-guide-to-always-on-agents">Autonomous Agentic Systems</a></p><p>Excited to see my article on &#8220;Autonomous Agentic Systems at Scale&#8221; is now published at Hackernoon! This one provides a practical guide to &#8220;Always-On&#8221; agents &#128640; In this post I share some of my learnings gathered from extending KAOS to support autonomous long-running agents whilst balancing the runtime, memory, telemetry, and operational complexities. &#8220;Always-On&#8221; Agents feel like a new kind of architectural abstraction, as they are indeed not quite a chatbot, but also not quite a &#8220;cron job&#8221; / workflow engine, and definitely not quite a microservice... but it&#8217;s a Frankenstein that borrows from all of them (+ ofc with the added complexity of being stateful and non-deterministic). Despite the complexity it is clear that there&#8217;s a lot of opportunity with this pattern - namely for use-cases where the goal persists over time and the environment keeps changing. However it is also clear to me that there is still quite a way to go for the field to be able to start getting the full value, and that includes improvements in monitoring, operations, maintenance, research, data discovery, etc. So let&#8217;s make sure that we continue to invest and contribute to theses open questions, as they indeed won&#8217;t answer themselves. Let me know what you think!</p><div><hr></div><p><a href="https://openai.com/index/harness-engineering/">OpenAI Codex Engineering</a></p><p>OpenAI is clearly leading the way in developer productivity; the Codex repo is built by a team of 3, and has produced about 1.5k PRs over five months, post is one of the clearest signs yet that the coding-agent bottleneck is moving from model capability to engineering systems design: the team built an internal product with 0 manually-written lines of code, around a million lines generated by Codex, and roughly 1,500 PRs over five months. For production ML practitioners, the practical takeaway is that reliable agentic development is less about prompting harder and more about building the right harness around agents: repo-local knowledge, a small <a href="http://AGENTS.md">AGENTS.md</a> as a map instead of a giant instruction manual, browser-driven validation, local observability, mechanically enforced architecture, custom linters, and continuous cleanup of drift. The key lesson is that as agents take over more of the software lifecycle, human attention becomes the scarce resource, so teams need to encode taste, constraints, tests, telemetry, and review loops directly into the codebase rather than relying on ad-hoc docs or heroic human review.</p><div><hr></div><p><a href="https://claude.com/blog/how-anthropic-enables-self-service-data-analytics-with-claude">Anthropic on Self Service Analytics</a></p><p>Anthropic just shared their playbook on self-service analytics, showing how they automated 95% of their internal business analytics with Claude: On their setup, they were able to build an agentic system stack that is able to reach ~95% aggregate accuracy, and the most interesting point is that their system looks much more like a governed data platform, so we&#8217;re back to the basics. It seems the main failure modes are data modelling ambiguity, data staleness, and retrieval failure, which can be addressed with canonical datasets, enforced semantic layers, lineage, curated domain docs, and Claude Code Skills that route the model through the same process a senior analyst would follow. It&#8217;s also great to see strong emphasis on evals and observability, including offline question/answer suites, PR-level ablations, provenance footers, adversarial review, and correction harvesting. For production ML and Data practitioners it is a reminder that reliable analytics agents are less about letting an LLM loose on your warehouse, and more about the basic foundations of &#8220;great data&#8221;.</p><div><hr></div><p><a href="https://www.amazon.science/blog/mitra-mixed-synthetic-priors-for-enhancing-tabular-foundation-models">Amazon&#8217;s Tabluar Foundation Models</a></p><p>Tabular ML is still one of the most business-critical parts of production machine learning, and I didn&#8217;t know Amazon also had a tabular foundation model that tackles different modalities: Amazon Mitra is an interesting tabular foundation models that also takes a different approach from training on real data, and instead it is trained on purely synthetic data, which seems to be a growing trend. This synthetic data pipeline includes a carefully designed mixture of synthetic priors, combining structural causal models with tree-based generators such as gradient boosting, random forests and decision trees. It seems that the key idea for tabular foundation models is that the data prior may matter as much as the architecture, as good synthetic priors should perform well on real tasks, be diverse enough to avoid overfitting to themselves, and add distinctive patterns not already covered by other priors. Mitra uses in-context learning to condition on support rows from a new dataset and predict query labels without gradient updates, while also supporting fine-tuning and ensembling through AutoGluon. The reported results are strong across TabRepo, TabZilla, AMLB and TabArena, with Mitra outperforming TabPFNv2 (although v3 is already out), TabICL and strong task-specific baselines in several classification and regression settings, and showing better sample efficiency when fewer in-context examples are available. It is clear that tabular foundation models are becoming a serious option for low-data, fast-iteration tabular prediction workflows; it will not (yet) outperform specialised models, but it is clear that getting a strong baseline for free is already a major win.</p><div><hr></div><p><a href="https://huggingface.co/moonshotai/Kimi-K2.7-Code">Kimi 2.7 Code Model Release</a></p><p>The Chinese startup behind Kimi (Moonshot AI) just released Kimi K2.7 Code, and this is for sure worth paying attention to if you care about where coding agents are headed: K2.7 Code is a coding-focused agentic model built on the previous arch, ie with the same large MoE shape of roughly 1T total parameters and 32B active parameters, but with stronger reported coding / agentic performance and around 30% lower thinking-token usage. That token-efficiency point is probably the most interesting part for production ML teams, because agentic coding is starting to get more expensive (and so is electricity!), which means it gets worse with repeated tool calls, context carry-forward, retries, and multi-step reasoning loops. The model supports a 256K context window, multimodal input, forced thinking / preserve-thinking mode, and deployment through vLLM, SGLang, KTransformers, Hugging Face, Moonshot APIs, and hosted routers like OpenRouter. The benchmark numbers look promising across coding and MCP-style tool-use tasks, although teams should treat them carefully given the first-party eval setup and differences in harnesses across competing models.</p><div><hr></div><h2>Upcoming MLOps Events</h2><p>The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.</p><h3>Events we are speaking at this year:</h3><ul><li><p><strong><a href="https://signalsconf.io/#tickets">Signals Conference</a></strong> - September @ Berlin</p></li><li><p><strong><a href="https://worldsummit.ai/">World Summit AI Europe</a></strong> - September @ Amsterdam</p></li></ul><h3>Other relevant events:</h3><ul><li><p><strong><a href="https://events.linuxfoundation.org/kubecon-cloudnativecon-europe/">KubeCon Europe</a></strong> - March @ Amsterdam</p></li><li><p><strong><a href="https://2026.pycon.de/">PyData Berlin</a></strong> - April @ Frankfurt</p></li><li><p><strong><a href="https://www.databricks.com/dataaisummit">Databricks Summit</a></strong> - June @ San Francisco</p></li><li><p><strong><a href="https://www.wearedevelopers.com/world-congress">World Developer Congress</a></strong> - July @ Berlin</p></li><li><p><strong><a href="https://ep2025.europython.eu/">EuroPython 2026</a></strong> - July @ Prague</p></li><li><p><strong><a href="https://euroscipy.org/">EuroSciPy 2026</a></strong> - July @ Krakow</p></li><li><p><strong><a href="https://www.ai-infra-summit.com/">AI Infra Summit 2026</a> </strong>- Sept @ California</p></li><li><p><strong><a href="https://codetalks.com/">Code.Talks 2026</a></strong> - Nov @ Hamburg</p></li><li><p><strong><a href="https://mlopsworld.com/">MLOps World 2026</a></strong> - Nov @ Austin</p></li></ul><h3>In case you missed our talks, check our recordings below:</h3><ul><li><p>The State of AI in 2025 - <a href="https://www.youtube.com/watch?v=v2LENQOG-Xg">WeAreDevelopers 2025</a></p></li><li><p>Prod Generative AI in 2024 - <a href="https://www.youtube.com/watch?v=0uJGmMZGUJE&amp;list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&amp;index=3">KubeCon AI Day 2025</a></p></li><li><p>The State of AI in 2024 - <a href="https://www.youtube.com/live/AtA2XXo_b5s">WeAreDevelopers 2024</a></p></li><li><p>Responsible AI Workshop Keynote - <a href="http://www.youtube.com/watch?v=57YpXjcj0Ho">NeurIPS 2021</a></p></li><li><p>Practical Guide to ML Explainability - <a href="http://www.youtube.com/watch?v=vq8mDiDODhc">PyCon London</a></p></li><li><p>ML Monitoring: Outliers, Drift, XAI - <a href="http://www.youtube.com/watch?v=QcevzK9ZuDg">PyCon Keynote</a></p></li><li><p>Metadata for E2E MLOps - <a href="https://www.youtube.com/watch?v=OSbH4dfswCY">Kubecon NA 2022</a></p></li><li><p>ML Performance Evaluation at Scale - <a href="http://www.youtube.com/watch?v=8ORl8lu1Eeo">KubeCon Eur 2021</a></p></li><li><p>Industry Strength LLMs - <a href="https://www.youtube.com/watch?v=RVUi_rAFfzU">PyData Global 2022</a></p></li><li><p>ML Security Workshop Keynote - <a href="http://www.youtube.com/watch?v=7XSy5aw8oU8">NeurIPS 2022</a></p></li></ul><div><hr></div><h2>Open Source MLOps Tools</h2><p>Check out the fast-growing ecosystem of production ML tools &amp; frameworks at <a href="http://github.com/EthicalML/awesome-production-machine-learning/">the github repository</a> which has reached over 20,000 &#11088; github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here&#8217;s a few featured open source libraries that we maintain:</p><ul><li><p><strong><a href="https://github.com/besanson/sarc-governance/tree/main">SARC</a> </strong>- Provides wrappers for popular agentic frameworks to enable guardrails and constraints that are enforced through the flow.</p></li><li><p><strong><a href="https://github.com/axsaucedo/kaos">KAOS</a></strong> - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.</p></li><li><p><strong><a href="https://github.com/KomputeProject/kompute">Kompute</a></strong> - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-machine-learning/">Production ML Tools</a></strong> - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-artificial-intelligence-regulation">AI Policy List</a></strong> - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-genai/">Agentic Systems Tools</a></strong> - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain</p></li></ul><p>Please do support some of our open source projects by sharing, contributing or adding a star &#11088;</p><div><hr></div><p><strong>About us</strong></p><p>The Institute for Ethical AI &amp; Machine Learning is a European research centre that carries out world-class research into responsible machine learning.</p><p><a href="https://ethical.institute/">Check out our website</a></p>]]></content:encoded></item><item><title><![CDATA[Issue #390 - The ML Engineer 🤖 ]]></title><description><![CDATA[C++ The Documentary, Stanford Language Modeling from Scratch, Tokenomics Where Tokens are Used, LLMs Need New Routing Paradigm, NVIDIA RTX Spark Launch + more &#128640;]]></description><link>https://machinelearning.substack.com/p/issue-390-the-ml-engineer</link><guid isPermaLink="false">https://machinelearning.substack.com/p/issue-390-the-ml-engineer</guid><dc:creator><![CDATA[Alejandro Saucedo]]></dc:creator><pubDate>Sun, 07 Jun 2026 09:29:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Thank you for being part of over <a href="https://ethical.institute/mle.html">70,000+ ML professionals and enthusiasts</a> who receive weekly articles &amp; tutorials on Machine Learning &amp; MLOps &#129302; You can join the newsletter <a href="https://bit.ly/state-of-ml-2025">https://bit.ly/state-of-ml-2025</a> &#11088;</p><p>If you like the content please support the newsletter by sharing with your friends via &#9993;&#65039; Email, &#128038; <a href="http://twitter.com/intent/tweet?text=Great%20Newsletter%20on%20Machine%20Learning!%20%20https://ethical.institute/mle.html">Twitter</a>, &#128188; <a href="http://www.linkedin.com/shareArticle?mini=true&amp;url=https://ethical.institute/mle.html&amp;title=The%20Machine%20Learning%20Engineer%20Newsletter&amp;summary=Curated%20news%20about%20machine%20learning%20operations,%20reproducibility,%20explainability%20and%20beyond">Linkedin</a> and &#128213; <a href="http://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fethical.institute%2Fmle.html">Facebook</a>!</p><div><hr></div><h2>This week in ML Engineering:</h2><ul><li><p>C++: <strong><a href="https://www.youtube.com/watch?v=lI7tMxzSJ7w">The Documentary</a></strong></p></li><li><p>Stanford Language <strong><a href="https://cs336.stanford.edu/">Modeling from Scratch</a></strong></p></li><li><p>Tokenomics Where <strong><a href="https://arxiv.org/abs/2601.14470">Tokens are Used</a></strong></p></li><li><p>LLMs Need <strong><a href="https://www.modular.com/blog/why-llm-inference-needs-a-new-kind-of-router-part-1">New Routing Paradigm</a></strong></p></li><li><p>NVIDIA <strong><a href="https://www.youtube.com/watch?v=11Y3B33oCLE">RTX Spark Launch</a></strong></p></li><li><p>Open Source <strong><a href="http://github.com/EthicalML/awesome-production-machine-learning/">ML Frameworks</a></strong></p></li><li><p>Awesome AI Guidelines <strong><a href="https://github.com/ethicalml/awesome-artificial-intelligence-guidelines">to check out this week</a></strong></p></li><li><p>+ more &#128640;</p></li></ul><div><hr></div><p><a href="https://www.youtube.com/watch?v=lI7tMxzSJ7w">C++: The Documentary</a></p><p>I loved the C++ Documentary! People may not know but C++ is (shamelessly) my favourite language; I maintain an active (2k stars) OSS C++ codebase for fun, and did a lot of prod C++ dev back in the day, so this documentary was thoroughly enjoyable to watch. Despite all the hate, C++ is arguably one of the most important languages of our times; large % of modern AI foundation is built upon C++. It may not be as visible by Python/JS devs, and not as loved by Rust fans, but it&#8217;s still a fast growing language that supports the performance and scale many important systems today. This is an absolutely recommended watch for any ML practitioners, as it covers how C++ began when Bjarne wanted to combine C&#8217;s low-level interoperability advanced classes and modern language design. It&#8217;s interesting to see the evolution across Bell Labs, CFront, standardization, STL, C++98, the Java/C# era, and the C++11 renaissance; funnily enough, my main drivers tend to be C++11, with a tiny bit of C++14 / C++ 17, however I still havent found the main reason to move forward. Maybe once modules mature! Or an integrated package manager?! We all make-do with CMake, but it&#8217;s time to move on! Also I loved Bjarne&#8217;s quote: &#8220;It should&#8217;ve been called ++C for semantic reasons&#8221; - LOL. Also hilarious to hear that C++ v2 released as 2.00 known as 2.&#8221;uh oh...&#8221;. + Alexander Stepanov (STL designer) on standards: &#8220;Standards are not necessarily good or correct, but they are standards [...] traffic laws [for example] are very often idiotic. But we have to obey them otherwise we will die&#8221;. I loved all of these small nuggets of humorous archeological history on the language. Check it out!</p><div><hr></div><p><a href="https://cs336.stanford.edu/">Stanford Language Modeling from Scratch</a></p><p>Stanford has published their CS336 Undergraduate Module on Language Modeling from Scratch. Quality SotA learning material from Stanfrod? For free? What a time to be alive! This is one of the most practical public resources for ML engineers who want to move beyond API-level LLM usage, and instead actually understand the full foundation-model stack end-to-end. The course walks through building language models from first principles, including tokenization, Transformer architecture, optimizers, GPU/resource accounting, Triton kernels, FlashAttention-style optimization, distributed training, scaling laws, inference, evaluation, etc. The assignments are deliberately low-scaffolding and require substantial Python/PyTorch engineering - they are really chunky, so prepare to spend a few weeks actually plowing through these. The public GitHub repo also includes lecture materials and assignment repos, with the lecture repo structured around executable lecture files and PDFs. I am impressed by the level of depth and quality of this course, unsurprisingly from Stanford, but it is shocking when realising how much high quality content is available on one of the most important topics today.</p><div><hr></div><p><a href="https://arxiv.org/abs/2601.14470">Tokenomics Where Tokens are Used</a></p><p>Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering. Super insightful research piece from Concordia University that analyses telemetry from agentic engineering clients across various development tasks. Some interesting findings presented: it seems that code review still dominates token usage consuming 59.4% of tokens on average; initial coding is only 8.6%. As part of the cost analysis, input tokens make up 53.9% of usage, suggesting that multi-agent systems are paying a large &#8220;communication tax&#8221; from repeatedly passing context between agents. For production ML practitioners building coding agents, a key insight is that cost control should focus on review/refinement loops, context management, and human-in-the-loop checkpoints rather than only optimizing model calls for generation. The main caveat is that the study is quite small, so it will be interesting as this is scaled further as likely new methodologies will emerge as well.</p><div><hr></div><p><a href="https://www.modular.com/blog/why-llm-inference-needs-a-new-kind-of-router-part-1">LLMs Need New Routing Paradigm</a></p><p>LLM inference has a routing problem: once every request can depend on cache locality, GPU specialization, and multi-step execution, the router directly affects latency, cost, and user experience. Modular put together a really interesting series on LLM Inference Routing, which is probably one of the most comprehensive breakdowns for scaling LLM serving beyond simple load balancing. Modular explains why traditional HTTP routers break down when GPU inference pods are stateful, heterogeneous, cache-sensitive, and sometimes split across prefill/decode execution paths. Part 1 frames the core problem around KV-cache residency, hardware specialization, conversation continuity, and multi-step request execution; Part 2 shows why the router needs a hot-path data layer that can query cached token blocks across pods in microseconds; and Part 3 turns that state into a composable routing pipeline covering preparation, filtering, scoring, picking, and execution. Lately I have seen quite a few posts discussing the challenges and potential solutions on LLM routing, so it is clear that not only this is a huge problem, but definitely understanding the foundations will be a useful skill as these tools become more ubiquitous.</p><div><hr></div><p><a href="https://www.youtube.com/watch?v=11Y3B33oCLE">NVIDIA RTX Spark Launch</a></p><p>Last week NVIDIA and Microsoft announced the release of the NVIDIA RTX Spark, and it&#8217;s really exciting to see that unified memory is finally coming outside of the M-Mac chips, and it seems that NVIDIA is indeed going all in. I found it actually surprisingly how small the chip actually is, and I struggle to process the claim that this chip has the same/similar performance to a 5070, when we are talking about fractions of the size compared to the video card. It was also impressive to see the small / thin size of the machines that were showcased with the chips, if this indeed matches the hype and expectations, it is clear that this will be another huge blow to the apple ecosystem with PCs finally being bullish on ARM processors. This also gives me hope for a proper local only rig that can run SotA models in consumer hardware, with the cost and extensibility of a PC. Let&#8217;s see how this actually develops!</p><div><hr></div><h2>Upcoming MLOps Events</h2><p>The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.</p><h3>Events we are speaking at this year:</h3><ul><li><p><strong><a href="https://signalsconf.io/#tickets">Signals Conference</a></strong> - September @ Berlin</p></li><li><p><strong><a href="https://worldsummit.ai/">World Summit AI Europe</a></strong> - September @ Amsterdam</p></li></ul><h3>Other relevant events:</h3><ul><li><p><strong><a href="https://events.linuxfoundation.org/kubecon-cloudnativecon-europe/">KubeCon Europe</a></strong> - March @ Amsterdam</p></li><li><p><strong><a href="https://2026.pycon.de/">PyData Berlin</a></strong> - April @ Frankfurt</p></li><li><p><strong><a href="https://www.databricks.com/dataaisummit">Databricks Summit</a></strong> - June @ San Francisco</p></li><li><p><strong><a href="https://www.wearedevelopers.com/world-congress">World Developer Congress</a></strong> - July @ Berlin</p></li><li><p><strong><a href="https://ep2025.europython.eu/">EuroPython 2026</a></strong> - July @ Prague</p></li><li><p><strong><a href="https://euroscipy.org/">EuroSciPy 2026</a></strong> - July @ Krakow</p></li><li><p><strong><a href="https://www.ai-infra-summit.com/">AI Infra Summit 2026</a> </strong>- Sept @ California</p></li><li><p><strong><a href="https://codetalks.com/">Code.Talks 2026</a></strong> - Nov @ Hamburg</p></li><li><p><strong><a href="https://mlopsworld.com/">MLOps World 2026</a></strong> - Nov @ Austin</p></li></ul><h3>In case you missed our talks, check our recordings below:</h3><ul><li><p>The State of AI in 2025 - <a href="https://www.youtube.com/watch?v=v2LENQOG-Xg">WeAreDevelopers 2025</a></p></li><li><p>Prod Generative AI in 2024 - <a href="https://www.youtube.com/watch?v=0uJGmMZGUJE&amp;list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&amp;index=3">KubeCon AI Day 2025</a></p></li><li><p>The State of AI in 2024 - <a href="https://www.youtube.com/live/AtA2XXo_b5s">WeAreDevelopers 2024</a></p></li><li><p>Responsible AI Workshop Keynote - <a href="http://www.youtube.com/watch?v=57YpXjcj0Ho">NeurIPS 2021</a></p></li><li><p>Practical Guide to ML Explainability - <a href="http://www.youtube.com/watch?v=vq8mDiDODhc">PyCon London</a></p></li><li><p>ML Monitoring: Outliers, Drift, XAI - <a href="http://www.youtube.com/watch?v=QcevzK9ZuDg">PyCon Keynote</a></p></li><li><p>Metadata for E2E MLOps - <a href="https://www.youtube.com/watch?v=OSbH4dfswCY">Kubecon NA 2022</a></p></li><li><p>ML Performance Evaluation at Scale - <a href="http://www.youtube.com/watch?v=8ORl8lu1Eeo">KubeCon Eur 2021</a></p></li><li><p>Industry Strength LLMs - <a href="https://www.youtube.com/watch?v=RVUi_rAFfzU">PyData Global 2022</a></p></li><li><p>ML Security Workshop Keynote - <a href="http://www.youtube.com/watch?v=7XSy5aw8oU8">NeurIPS 2022</a></p></li></ul><h2>Open Source MLOps Tools</h2><p>Check out the fast-growing ecosystem of production ML tools &amp; frameworks at <a href="http://github.com/EthicalML/awesome-production-machine-learning/">the github repository</a> which has reached over 20,000 &#11088; github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here&#8217;s a few featured open source libraries that we maintain:</p><ul><li><p><strong><a href="https://github.com/besanson/sarc-governance/tree/main">SARC</a> </strong>- Provides wrappers for popular agentic frameworks to enable guardrails and constraints that are enforced through the flow.</p></li><li><p><strong><a href="https://github.com/axsaucedo/kaos">KAOS</a></strong> - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.</p></li><li><p><strong><a href="https://github.com/KomputeProject/kompute">Kompute</a></strong> - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-machine-learning/">Production ML Tools</a></strong> - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-artificial-intelligence-regulation">AI Policy List</a></strong> - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-genai/">Agentic Systems Tools</a></strong> - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain</p></li></ul><p>Please do support some of our open source projects by sharing, contributing or adding a star &#11088;</p><div><hr></div><p><strong>About us</strong></p><p>The Institute for Ethical AI &amp; Machine Learning is a European research centre that carries out world-class research into responsible machine learning.</p><p><a href="https://ethical.institute/">Check out our website</a></p>]]></content:encoded></item><item><title><![CDATA[Issue #389 - The ML Engineer 🤖]]></title><description><![CDATA[Engineering Like It's 2007, Netflix LLM Finetuning Infra, OpenAI & Anthropic Finding Market Fit, Reviving PapersWithCode.co, Massive Open Text-To-Image Dataset + more &#128640;]]></description><link>https://machinelearning.substack.com/p/issue-389-the-ml-engineer</link><guid isPermaLink="false">https://machinelearning.substack.com/p/issue-389-the-ml-engineer</guid><dc:creator><![CDATA[Alejandro Saucedo]]></dc:creator><pubDate>Sun, 31 May 2026 12:05:29 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;9ffa8618-894d-4f72-93c6-bbbdde09d695&quot;,&quot;duration&quot;:null}"></div><p>Thank you for being part of over <a href="https://ethical.institute/mle.html">70,000+ ML professionals and enthusiasts</a> who receive weekly articles &amp; tutorials on Machine Learning &amp; MLOps &#129302; You can join the newsletter <a href="https://bit.ly/state-of-ml-2025">https://bit.ly/state-of-ml-2025</a> &#11088;</p><p>If you like the content please support the newsletter by sharing with your friends via &#9993;&#65039; Email, &#128038; <a href="http://twitter.com/intent/tweet?text=Great%20Newsletter%20on%20Machine%20Learning!%20%20https://ethical.institute/mle.html">Twitter</a>, &#128188; <a href="http://www.linkedin.com/shareArticle?mini=true&amp;url=https://ethical.institute/mle.html&amp;title=The%20Machine%20Learning%20Engineer%20Newsletter&amp;summary=Curated%20news%20about%20machine%20learning%20operations,%20reproducibility,%20explainability%20and%20beyond">Linkedin</a> and &#128213; <a href="http://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fethical.institute%2Fmle.html">Facebook</a>!</p><div><hr></div><h2>This week in ML Engineering:</h2><ul><li><p>Engineering <strong><a href="https://www.youtube.com/watch?v=w5WVu624fY8">Like It&#8217;s 2007</a></strong></p></li><li><p>Netflix <strong><a href="https://netflixtechblog.com/scaling-llm-post-training-at-netflix-0046f8790194">LLM Finetuning Infra</a></strong></p></li><li><p>OpenAI &amp; Anthropic <strong><a href="https://simonwillison.net/2026/May/27/product-market-fit/">Finding Market Fit</a></strong></p></li><li><p>Reviving <strong><a href="http://paperswithcode.co/">PapersWithCode.co</a></strong></p></li><li><p>Massive Open <strong><a href="https://huggingface.co/datasets/jasperai/monet">Text-To-Image Dataset</a></strong></p></li><li><p>Open Source <strong><a href="http://github.com/EthicalML/awesome-production-machine-learning/">ML Frameworks</a></strong></p></li><li><p>Awesome AI Guidelines <strong><a href="https://github.com/ethicalml/awesome-artificial-intelligence-guidelines">to check out this week</a></strong></p></li><li><p>+ more &#128640;</p></li></ul><div><hr></div><p><a href="https://www.youtube.com/watch?v=w5WVu624fY8">Engineering Like It&#8217;s 2007</a></p><p>Let&#8217;s take a trip back to 2007! This engineering talk from YouTube is truly a master class of scaling and learning, and surprisingly the lessons are as valuable today as they were back then. This talk shows a lean and mean team iterating through bottlenecks across web serving, video delivery, thumbnails, databases, hardware, OS tuning, caching and sharding. The most useful lesson for ML engineering teams is that scale is rarely solved by one abstraction, as YouTube kept the system simple enough to rewrite under pressure, used caching at multiple layers, and leveraged first principles at every layer. It is pretty cool to see that they followed the standard playbook of best practice of scalability, the usual suspects of moving hot traffic through CDNs, tuning commodity hardware (maybe less common now), and eventually replacing replication tricks with database partitioning. For ML practitioners, the parallel is model serving, feature stores, vector databases, observability pipelines and agent systems, which are currently hitting us with the same analogous challenges that we have to solve at lighting speed.</p><div><hr></div><p><a href="https://netflixtechblog.com/scaling-llm-post-training-at-netflix-0046f8790194">Netflix LLM Finetuning Infra</a></p><p>Netflix is sharing their playbook for massive scale LLM fine-tuning infrastructure: Netflix moved away from a few fine-tuning scripts to a managed framework that supports SFT, DPO, RL, distillation, checkpointing, MFU tracking, Hugging Face-compatible model/tokenizer flows, and distributed orchestration. It is interesting that Ray has chosen also the usual suspect technologies like Ray, PyTorch, vLLM together with custom tooling (eg Netflix&#8217;s internal Mako platform). The interesting takeaway for production ML practitioners is that post-training is challenging across every layer, including Data, Model, Compute and Workflow. Netflix reports up to 4.7x effective token throughput from asynchronous on-the-fly sequence packing, and this is a great example of the direction GenAI infrastructure is taking.</p><p><a href="https://simonwillison.net/2026/May/27/product-market-fit/">OpenAI &amp; Anthropic Finding Market Fit</a></p><p>Open AI &amp; Anthropic have found Market Fit - this is an interesting opinion piece from Simon Willison: Here there is a good case that coding agents may be the first real product-market fit moment for frontier AI labs because enterprises are now being charged close to raw API-token economics for daily developer workflows. OpenAI Codex and Anthropic Claude Code have shifted from huge subsidies / subscriptions to usage-based pricing (unfortunately for us subsidies are indeed ending). It is now clear that organisations will need to establish much tighter cost observability, usage governance, ROI measurement, procurement discipline, and platform controls around coding agents, just as they already do for serving and inference workloads. This starts making it clear why tools like LiteLLM or OpenRouter are becoming so popular, even if at the beginning it wasn&#8217;t super clear why an extra abstraction layer was needed on top of a simple vendor API (eg. surprisingly enough many vendors still do not offer spend caps).</p><div><hr></div><p><a href="https://paperswithcode.co/">Reviving PapersWithCode.co</a></p><p>I remember when Papers With Code originally came out in 2018 it was a major breakthrough; after the meta acquisition the project slowed and then stopped, but it seems there is an attempt to revive it! Hugging Face has started reviving it as <a href="http://paperswithcode.co/">paperswithcode.co</a> with support from some AI agents to help with the parsing of papers, auto-linking GitHub repos, project pages and artifacts, categorizing, and generating leaderboards. The new site already brings back the familiar discovery workflow around trending papers, SOTA browsing, methods and domains, while adding support for star-velocity trends, citation counts, external non-arXiv papers, multiple repos per paper, benchmark harness reports, and Hugging Face-native storage/login integration. As research moves faster (whether AI slop or otherwise), it&#8217;s still a basic need to have a well maintained discovery layer, so it&#8217;s great to see projects like this, hopefully it will continue growing.</p><div><hr></div><p><a href="https://huggingface.co/datasets/jasperai/monet">Massive Open Text-To-Image Dataset</a></p><p>An interesting release of a new massive open text-to-image dataset: The MONET dataset. It is great to see this, as image generation not only depends on quality data, but also other key resources like benrhcmarks whcih can only help teams accelerate on this field. Better curation, filtering, captions, and provenance can really make quite a difference. This dataset consists of 104.9M curated image&#8211;text pairs distilled from 2.9B raw pairs, with safety filtering, domain filtering, exact/near duplicate removal, multi-VLM re-captioning, embeddings, object/face annotations, hashes, NSFW/watermark scores, and pre-encoded SANA-VAE latents for faster latent-diffusion training.</p><div><hr></div><h2>Upcoming MLOps Events</h2><p>The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.</p><h3>Events we are speaking at this year:</h3><ul><li><p><strong><a href="https://etailgermany.wbresearch.com/">eTail Europe</a> </strong>- March @ Berlin</p></li><li><p><strong><a href="https://worldsummit.ai/">World Summit AI Europe</a></strong> - September @ Amsterdam</p></li></ul><h3>Other relevant events:</h3><ul><li><p><strong><a href="https://events.linuxfoundation.org/kubecon-cloudnativecon-europe/">KubeCon Europe</a></strong> - March @ Amsterdam</p></li><li><p><strong><a href="https://2026.pycon.de/">PyData Berlin</a></strong> - April @ Frankfurt</p></li><li><p><strong><a href="https://www.databricks.com/dataaisummit">Databricks Summit</a></strong> - June @ San Francisco</p></li><li><p><strong><a href="https://www.wearedevelopers.com/world-congress">World Developer Congress</a></strong> - July @ Berlin</p></li><li><p><strong><a href="https://ep2025.europython.eu/">EuroPython 2026</a></strong> - July @ Prague</p></li><li><p><strong><a href="https://euroscipy.org/">EuroSciPy 2026</a></strong> - July @ Krakow</p></li><li><p><strong><a href="https://www.ai-infra-summit.com/">AI Infra Summit 2026</a> </strong>- Sept @ California</p></li><li><p><strong><a href="https://codetalks.com/">Code.Talks 2026</a></strong> - Nov @ Hamburg</p></li><li><p><strong><a href="https://mlopsworld.com/">MLOps World 2026</a></strong> - Nov @ Austin</p></li></ul><h3>In case you missed our talks, check our recordings below:</h3><ul><li><p>The State of AI in 2025 - <a href="https://www.youtube.com/watch?v=v2LENQOG-Xg">WeAreDevelopers 2025</a></p></li><li><p>Prod Generative AI in 2024 - <a href="https://www.youtube.com/watch?v=0uJGmMZGUJE&amp;list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&amp;index=3">KubeCon AI Day 2025</a></p></li><li><p>The State of AI in 2024 - <a href="https://www.youtube.com/live/AtA2XXo_b5s">WeAreDevelopers 2024</a></p></li><li><p>Responsible AI Workshop Keynote - <a href="http://www.youtube.com/watch?v=57YpXjcj0Ho">NeurIPS 2021</a></p></li><li><p>Practical Guide to ML Explainability - <a href="http://www.youtube.com/watch?v=vq8mDiDODhc">PyCon London</a></p></li><li><p>ML Monitoring: Outliers, Drift, XAI - <a href="http://www.youtube.com/watch?v=QcevzK9ZuDg">PyCon Keynote</a></p></li><li><p>Metadata for E2E MLOps - <a href="https://www.youtube.com/watch?v=OSbH4dfswCY">Kubecon NA 2022</a></p></li><li><p>ML Performance Evaluation at Scale - <a href="http://www.youtube.com/watch?v=8ORl8lu1Eeo">KubeCon Eur 2021</a></p></li><li><p>Industry Strength LLMs - <a href="https://www.youtube.com/watch?v=RVUi_rAFfzU">PyData Global 2022</a></p></li><li><p>ML Security Workshop Keynote - <a href="http://www.youtube.com/watch?v=7XSy5aw8oU8">NeurIPS 2022</a></p></li></ul><div><hr></div><h2>Open Source MLOps Tools</h2><p>Check out the fast-growing ecosystem of production ML tools &amp; frameworks at <a href="http://github.com/EthicalML/awesome-production-machine-learning/">the github repository</a> which has reached over 20,000 &#11088; github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here&#8217;s a few featured open source libraries that we maintain:</p><ul><li><p><strong><a href="https://github.com/besanson/sarc-governance/tree/main">SARC</a> </strong>- Provides wrappers for popular agentic frameworks to enable guardrails and constraints that are enforced through the flow.</p></li><li><p><strong><a href="https://github.com/axsaucedo/kaos">KAOS</a></strong> - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.</p></li><li><p><strong><a href="https://github.com/KomputeProject/kompute">Kompute</a></strong> - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-machine-learning/">Production ML Tools</a></strong> - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-artificial-intelligence-regulation">AI Policy List</a></strong> - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-genai/">Agentic Systems Tools</a></strong> - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain</p></li></ul><p>Please do support some of our open source projects by sharing, contributing or adding a star &#11088;</p><div><hr></div><p><strong>About us</strong></p><p>The Institute for Ethical AI &amp; Machine Learning is a European research centre that carries out world-class research into responsible machine learning.</p><p><a href="https://ethical.institute/">Check out our website</a></p>]]></content:encoded></item><item><title><![CDATA[Issue #388 - The ML Engineer 🤖 ]]></title><description><![CDATA[Making DL Go Brr with First Principles, Benedict Evans: AI Eats the World 2026, Gemini 3.5 Frontier Intelligence, (Sk)Forecast Foundation Models, NVIDIA' New Image/Video Model + more &#128640;]]></description><link>https://machinelearning.substack.com/p/issue-388-the-ml-engineer</link><guid isPermaLink="false">https://machinelearning.substack.com/p/issue-388-the-ml-engineer</guid><dc:creator><![CDATA[Alejandro Saucedo]]></dc:creator><pubDate>Sun, 24 May 2026 11:07:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Thank you for being part of over <a href="https://ethical.institute/mle.html">70,000+ ML professionals and enthusiasts</a> who receive weekly articles &amp; tutorials on Machine Learning &amp; MLOps &#129302; You can join the newsletter <a href="https://bit.ly/state-of-ml-2025">https://bit.ly/state-of-ml-2025</a> &#11088;</p><p>If you like the content please support the newsletter by sharing with your friends via &#9993;&#65039; Email, &#128038; <a href="http://twitter.com/intent/tweet?text=Great%20Newsletter%20on%20Machine%20Learning!%20%20https://ethical.institute/mle.html">Twitter</a>, &#128188; <a href="http://www.linkedin.com/shareArticle?mini=true&amp;url=https://ethical.institute/mle.html&amp;title=The%20Machine%20Learning%20Engineer%20Newsletter&amp;summary=Curated%20news%20about%20machine%20learning%20operations,%20reproducibility,%20explainability%20and%20beyond">Linkedin</a> and &#128213; <a href="http://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fethical.institute%2Fmle.html">Facebook</a>!</p><div><hr></div><h2>This week in ML Engineering:</h2><ul><li><p>Making DL Go Brr <strong><a href="https://horace.io/brrr_intro.html">with First Principles</a></strong></p></li><li><p>Benedict Evans: <strong><a href="https://www.ben-evans.com/presentations">AI Eats the World 2026</a></strong></p></li><li><p>Gemini 3.5 <strong><a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/#gemini-3-5-flash">Frontier Intelligence</a></strong></p></li><li><p>(Sk)Forecast <strong><a href="https://skforecast.org/latest/user_guides/foundation-forecasting-models.html">Foundation Models</a></strong></p></li><li><p>NVIDIA&#8217; New <strong><a href="https://github.com/NVlabs/Sana">Image/Video Model</a></strong></p></li><li><p>Open Source <strong><a href="http://github.com/EthicalML/awesome-production-machine-learning/">ML Frameworks</a></strong></p></li><li><p>Awesome AI Guidelines <strong><a href="https://github.com/ethicalml/awesome-artificial-intelligence-guidelines">to check out this week</a></strong></p></li><li><p>+ more &#128640;</p></li></ul><div><hr></div><p><a href="https://horace.io/brrr_intro.html">Making DL Go Brr w First Principles</a></p><p>A classic, &#8220;Deep Learning Go Brrrr From First Principles&#8221; which still brings super relevant advice to AI teams today: Instead of throwing random PyTorch tricks at slow models, it&#8217;s important to have a clean mental model for diagnosing performance across: 1) compute-bound limits, 2) memory-bandwidth-bound limits, and 3) overhead-bound limits. Each of these brings a different optimization path. Compute-bound workloads need better Tensor Core usage or more hardware. Bandwidth-bound workloads benefit most from operator fusion and avoiding unnecessary global memory reads/writes. Overhead-bound workloads usually need tracing, compilation, CUDA Graphs, or reducing Python/framework dispatch costs. For us production ML practitioners it&#8217;s a good reminder that GPU efficiency is not just about bigger accelerators, but about understanding where time is actually spent.</p><div><hr></div><p><a href="https://www.ben-evans.com/presentations">Benedict Evans: AI Eats the World 2026</a></p><p>Benedict Evans has dropped the 2026 &#8220;AI Eats the World&#8221; deck, and here&#8217;s the main highlights: GenAI so far = Huge capex first, unclear value capture, lots of hype, and only later the boring-but-transformational deployment layer. The current model race is still going, but models are converging, infrastructure is getting brutally expensive, and the real leverage is likely to come from teams that turn LLMs into reliable workflow automation. There is still a lot of value to come from new aggregation/discovery layers, and domain-specific products that change how work is done. The most important takeaway is that AI adoption will probably look slow and underwhelming inside enterprises until it suddenly becomes standard / expected.</p><div><hr></div><p><a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/#gemini-3-5-flash">Gemini 3.5 Frontier Intelligence</a></p><p>Google DeepMind has just released Gemini 3.5 Flash! This is quite interesting to see as a faster agentic execution model across coding, tool use, multimodal understanding and long-horizon workflows. The interesting bit for production ML practitioners is that Google is positioning Flash as the high-throughput model for real-world agents which claims strong results. For ML teams, the takeaway is about the infrastructure pattern where faster frontier models plus agent harnesses are becoming the default winning advantage.</p><div><hr></div><p><a href="https://skforecast.org/latest/user_guides/foundation-forecasting-models.html">(Sk)Forecast Foundation Models</a></p><p>Skforecast is making time-series foundation models much easier to test in real production forecasting workflows: it&#8217;s wiring Chronos, TimesFM, Moirai, and TabICL through their new release! Really great to see Skforecast leading the charge on making foundation models accessible with various new classes (e.g. FoundationModel + ForecasterFoundation) which abstract foundation models on sklearn-style interfaces. Zero-shot forecasting is now becoming something you can benchmark inside existing forecasting pipelines rather than treat as a separate research experiment (and works surprisingly well). There are still challenges in context length and feature parity, as longer windows help models see seasonality and regime patterns, but they also increase inference cost and latency, so teams still need proper backtesting rather than assuming bigger context is better. The examples are also refreshingly honest about production caveats - definitely worth checking out.</p><div><hr></div><p><a href="https://github.com/NVlabs/Sana">NVIDIA&#8217; New Image/Video Model</a></p><p>NVIDIA just dropped a high efficiency open-source stack for high-resolution image, video, and world-model generation! This seems like an exciting addition for production ML teams because it focuses on the deployment constraints that usually decide whether generative media systems are practical on latency, VRAM, training cost, quantization, and serving integration. This seems to be positioned by NVIDIA as a complete training and inference codebase with techniques such as linear attention, 32&#215; DC-AE latent compression, Flow-DPM-Solver sampling, few-step sCM distillation, and block causal linear attention for long video generation.</p><div><hr></div><h2>Upcoming MLOps Events</h2><p>The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.</p><p>Events we are speaking at this year:</p><ul><li><p><strong><a href="https://etailgermany.wbresearch.com/">eTail Europe</a> </strong>- March @ Berlin</p></li><li><p><strong><a href="https://worldsummit.ai/">World Summit AI Europe</a></strong> - September @ Amsterdam</p></li></ul><p>Other relevant events:</p><ul><li><p><strong><a href="https://events.linuxfoundation.org/kubecon-cloudnativecon-europe/">KubeCon Europe</a></strong> - March @ Amsterdam</p></li><li><p><strong><a href="https://2026.pycon.de/">PyData Berlin</a></strong> - April @ Frankfurt</p></li><li><p><strong><a href="https://www.databricks.com/dataaisummit">Databricks Summit</a></strong> - June @ San Francisco</p></li><li><p><strong><a href="https://www.wearedevelopers.com/world-congress">World Developer Congress</a></strong> - July @ Berlin</p></li><li><p><strong><a href="https://ep2025.europython.eu/">EuroPython 2026</a></strong> - July @ Prague</p></li><li><p><strong><a href="https://euroscipy.org/">EuroSciPy 2026</a></strong> - July @ Krakow</p></li><li><p><strong><a href="https://www.ai-infra-summit.com/">AI Infra Summit 2026</a> </strong>- Sept @ California</p></li><li><p><strong><a href="https://codetalks.com/">Code.Talks 2026</a></strong> - Nov @ Hamburg</p></li><li><p><strong><a href="https://mlopsworld.com/">MLOps World 2026</a></strong> - Nov @ Austin</p></li></ul><p>In case you missed our talks, check our recordings below:</p><ul><li><p>The State of AI in 2025 - <a href="https://www.youtube.com/watch?v=v2LENQOG-Xg">WeAreDevelopers 2025</a></p></li><li><p>Prod Generative AI in 2024 - <a href="https://www.youtube.com/watch?v=0uJGmMZGUJE&amp;list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&amp;index=3">KubeCon AI Day 2025</a></p></li><li><p>The State of AI in 2024 - <a href="https://www.youtube.com/live/AtA2XXo_b5s">WeAreDevelopers 2024</a></p></li><li><p>Responsible AI Workshop Keynote - <a href="http://www.youtube.com/watch?v=57YpXjcj0Ho">NeurIPS 2021</a></p></li><li><p>Practical Guide to ML Explainability - <a href="http://www.youtube.com/watch?v=vq8mDiDODhc">PyCon London</a></p></li><li><p>ML Monitoring: Outliers, Drift, XAI - <a href="http://www.youtube.com/watch?v=QcevzK9ZuDg">PyCon Keynote</a></p></li><li><p>Metadata for E2E MLOps - <a href="https://www.youtube.com/watch?v=OSbH4dfswCY">Kubecon NA 2022</a></p></li><li><p>ML Performance Evaluation at Scale - <a href="http://www.youtube.com/watch?v=8ORl8lu1Eeo">KubeCon Eur 2021</a></p></li><li><p>Industry Strength LLMs - <a href="https://www.youtube.com/watch?v=RVUi_rAFfzU">PyData Global 2022</a></p></li><li><p>ML Security Workshop Keynote - <a href="http://www.youtube.com/watch?v=7XSy5aw8oU8">NeurIPS 2022</a></p></li></ul><div><hr></div><h2>Open Source MLOps Tools</h2><p>Check out the fast-growing ecosystem of production ML tools &amp; frameworks at <a href="http://github.com/EthicalML/awesome-production-machine-learning/">the github repository</a> which has reached over 20,000 &#11088; github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here&#8217;s a few featured open source libraries that we maintain:</p><ul><li><p><strong><a href="https://github.com/besanson/sarc-governance/tree/main">SARC</a> </strong>- Provides wrappers for popular agentic frameworks to enable guardrails and constraints that are enforced through the flow.</p></li><li><p><strong><a href="https://github.com/axsaucedo/kaos">KAOS</a></strong> - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.</p></li><li><p><strong><a href="https://github.com/KomputeProject/kompute">Kompute</a></strong> - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-machine-learning/">Production ML Tools</a></strong> - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-artificial-intelligence-regulation">AI Policy List</a></strong> - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-genai/">Agentic Systems Tools</a></strong> - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain</p></li></ul><p>Please do support some of our open source projects by sharing, contributing or adding a star &#11088;</p><div><hr></div><p><strong>About us</strong></p><p>The Institute for Ethical AI &amp; Machine Learning is a European research centre that carries out world-class research into responsible machine learning.</p><p><a href="https://ethical.institute/">Check out our website</a></p>]]></content:encoded></item><item><title><![CDATA[Issue #387 - The ML Engineer 🤖 ]]></title><description><![CDATA[Toto 2.0 Time Series Foundation Model, TabPFN-3 Tabular Foundation Models, Stanford SWE Real-World Dataset, DeepMind AlphaEvolve Scaling Impact, Mozilla on Mithos Vulnerabilities + more &#128640;]]></description><link>https://machinelearning.substack.com/p/issue-387-the-ml-engineer</link><guid isPermaLink="false">https://machinelearning.substack.com/p/issue-387-the-ml-engineer</guid><dc:creator><![CDATA[Alejandro Saucedo]]></dc:creator><pubDate>Sun, 17 May 2026 13:50:13 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;7360af90-1fcb-49f2-b7fe-1506c884f74b&quot;,&quot;duration&quot;:null}"></div><p>Thank you for being part of over <a href="https://ethical.institute/mle.html">70,000+ ML professionals and enthusiasts</a> who receive weekly articles &amp; tutorials on Machine Learning &amp; MLOps &#129302; You can join the newsletter <a href="https://bit.ly/state-of-ml-2025">https://bit.ly/state-of-ml-2025</a> &#11088;</p><p>If you like the content please support the newsletter by sharing with your friends via &#9993;&#65039; Email, &#128038; <a href="http://twitter.com/intent/tweet?text=Great%20Newsletter%20on%20Machine%20Learning!%20%20https://ethical.institute/mle.html">Twitter</a>, &#128188; <a href="http://www.linkedin.com/shareArticle?mini=true&amp;url=https://ethical.institute/mle.html&amp;title=The%20Machine%20Learning%20Engineer%20Newsletter&amp;summary=Curated%20news%20about%20machine%20learning%20operations,%20reproducibility,%20explainability%20and%20beyond">Linkedin</a> and &#128213; <a href="http://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fethical.institute%2Fmle.html">Facebook</a>!</p><div><hr></div><h2>This week in ML Engineering:</h2><ul><li><p>Toto 2.0 Time <strong><a href="https://www.datadoghq.com/blog/ai/toto-2/">Series Foundation Model</a></strong></p></li><li><p>TabPFN-3 <strong><a href="https://priorlabs.ai/technical-reports/tabpfn-3">Tabular Foundation Models</a></strong></p></li><li><p>Stanford <strong><a href="https://arxiv.org/abs/2604.20779">SWE Real-World Dataset</a></strong></p></li><li><p>DeepMind <strong><a href="https://deepmind.google/blog/alphaevolve-impact/">AlphaEvolve Scaling Impact</a></strong></p></li><li><p>Mozilla <strong><a href="https://hacks.mozilla.org/2026/05/behind-the-scenes-hardening-firefox/">on Mithos Vulnerabilities</a></strong></p></li><li><p>Open Source <strong><a href="http://github.com/EthicalML/awesome-production-machine-learning/">ML Frameworks</a></strong></p></li><li><p>Awesome AI Guidelines <strong><a href="https://github.com/ethicalml/awesome-artificial-intelligence-guidelines">to check out this week</a></strong></p></li><li><p>+ more &#128640;</p></li></ul><div><hr></div><p><a href="https://www.datadoghq.com/blog/ai/toto-2/">Toto 2.0 Time Series Foundation Model</a></p><p>Datadog has just dropped a huge update to their Time Series Foundation Model for observability Toto 2.0, and this is really exciting for production ML practitioners: Toto 2.0 is a new Apache 2.0 open-weights model ranging from small 4M params all the way to to 2.5B parameters. At least from Datadog&#8217;s own observability-heavy benchmark it seems like there is not only potential for real-world use, but this is an interesting insight that domain specific time-series foundation models is clearly a realistic path for practical use. It is also good to see Datadog be quite honest about the remaining gaps, as even the largest model still shows long-horizon drift and structural breakdown past training context, so classical baselines are not going away. But overall this feels like a meaningful step towards forecasting foundation models becoming a real option, and potentially towards broader observability models that reason across metrics, traces, logs, topology, code changes, alerts and events for proactive incident detection.</p><div><hr></div><p><a href="https://priorlabs.ai/technical-reports/tabpfn-3">TabPFN-3 Tabular Foundation Models</a></p><p>Tabular ML is still where a huge amount of production ML actually happens, so it is great to see the latest Tabular Foundation Model release from Prior Labs: TabPFN-3 is live with support for up to 1M training rows, row-chunking, a reduced KV-cache, native missing-value handling, many-class classification up to 160 classes, GPU-side preprocessing, and much faster inference than TabPFN-2.5. It is interesting to see the benchmarks, as you can expect that TabPFN-3 reports better performance than tuned and ensembled baselines on TabArena, beats 8-hour-tuned gradient-boosted-tree baselines on datasets up to 1M rows and 200 features, but it will be interesting also to see real world and competitive benchmarks as further alternatives arise as well. For production ML practitioners, the interesting part is less the leaderboard and more the potential that we&#8217;re moving to a world where we can get faster baselines, less painful hyperparameter search, better calibrated predictive distributions, CPU-friendly distillation and faster SHAP-style interpretability workflows.</p><div><hr></div><p><a href="https://arxiv.org/abs/2604.20779">Stanford SWE Real-World Dataset</a></p><p>Stanford has just published a really interesting new dataset of real coding-agent sessions using public GitHub repos with ~6k sessions, 63K user prompts, 355K tool calls, git-linked diffs, and line-level attribution of whether code was written by humans or agents. It is interesting to see that coding-agent usage is already becoming extremely bimodal, with around 41% of sessions basically &#8220;vibe coding&#8221; where the agent writes almost all committed code, while 23% are still human-only. However the important takeaway is that they are still very inefficient and risky when used in the wild, with only ~44% of agent-produced code surviving into commits, users push back or interrupt in roughly 44% of turns, and vibe-coded commits introduce substantially more Semgrep-detected vulnerabilities than human-only or collaborative coding. For production ML practitioners building coding agents, evals, IDE copilots, or internal developer tooling, this is a strong reminder that the winning pattern is probably not full autonomy, but better scaffolding around agents.</p><div><hr></div><p><a href="https://deepmind.google/blog/alphaevolve-impact/">DeepMind AlphaEvolve Scaling Impact</a></p><p>Google DeepMind showcases how they are bringing AlphaEvolve alive as an optimization engine for the expensive parts of ML, science and infrastructure: The AlphaEvolve system from Google has now been applied across genomics, power grids, disaster prediction, quantum circuits, mathematics, TPU design, Spanner, compiler optimization, logistics, advertising, lithography and ML force fields, with some genuinely impressive reported numbers. As part of their report they outline a 30% reduction in DNA variant detection errors for DeepConsensus, AC Optimal Power Flow feasible-solution rates going from 14% to over 88%, 10x lower-error quantum circuits, 20% lower Spanner write amplification, and nearly 9% lower software storage footprint. It seesm they also have use-cases, showcasing Klarna doubling training speed and Schr&#246;dinger seeing roughly 4x speedups for MLFF training and inference.</p><div><hr></div><p><a href="https://hacks.mozilla.org/2026/05/behind-the-scenes-hardening-firefox/">Mozilla on Mithos Vulnerabilities</a></p><p>Mozilla shared a great behind-the-scenes look at how they used Claude Mythos Preview and other models to harden Firefox, and what is really interesting is that this was not just &#8220;LLM finds bugs&#8221; but a proper end-to-end security pipeline: The Mozilla team built an agentic harness on top of existing fuzzing infrastructure, where models could inspect risky parts of the browser codebase, generate reproducible test cases, run them, and then feed validated findings into the normal lifecycle for deduplication, triage, patching and release. The numbers are quite impressive, with Firefox 150 shipping fixes for 271 bugs found with Claude Mythos Preview, including 180 sec-high issues, and Mozilla fixing 423 security bugs across April releases when combining this pipeline with other AI models + manual review. This feels like a very clear MLSecOps pattern that more teams will need to adopt, where frontier models become scalable security-reasoning workers inside controlled harnesses that can execute tests, verify claims and eventually scan patches continuously in the CI/CD pipelines.</p><div><hr></div><h2>Upcoming MLOps Events</h2><p>The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.</p><h3>Events we are speaking at this year:</h3><ul><li><p><strong><a href="https://etailgermany.wbresearch.com/">eTail Europe</a> </strong>- March @ Berlin</p></li><li><p><strong><a href="https://worldsummit.ai/">World Summit AI Europe</a></strong> - September @ Amsterdam</p></li></ul><h3>Other relevant events:</h3><ul><li><p><strong><a href="https://events.linuxfoundation.org/kubecon-cloudnativecon-europe/">KubeCon Europe</a></strong> - March @ Amsterdam</p></li><li><p><strong><a href="https://2026.pycon.de/">PyData Berlin</a></strong> - April @ Frankfurt</p></li><li><p><strong><a href="https://www.databricks.com/dataaisummit">Databricks Summit</a></strong> - June @ San Francisco</p></li><li><p><strong><a href="https://www.wearedevelopers.com/world-congress">World Developer Congress</a></strong> - July @ Berlin</p></li><li><p><strong><a href="https://ep2025.europython.eu/">EuroPython 2026</a></strong> - July @ Prague</p></li><li><p><strong><a href="https://euroscipy.org/">EuroSciPy 2026</a></strong> - July @ Krakow</p></li><li><p><strong><a href="https://www.ai-infra-summit.com/">AI Infra Summit 2026</a> </strong>- Sept @ California</p></li><li><p><strong><a href="https://codetalks.com/">Code.Talks 2026</a></strong> - Nov @ Hamburg</p></li><li><p><strong><a href="https://mlopsworld.com/">MLOps World 2026</a></strong> - Nov @ Austin</p></li></ul><h3>In case you missed our talks, check our recordings below:</h3><ul><li><p>The State of AI in 2025 - <a href="https://www.youtube.com/watch?v=v2LENQOG-Xg">WeAreDevelopers 2025</a></p></li><li><p>Prod Generative AI in 2024 - <a href="https://www.youtube.com/watch?v=0uJGmMZGUJE&amp;list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&amp;index=3">KubeCon AI Day 2025</a></p></li><li><p>The State of AI in 2024 - <a href="https://www.youtube.com/live/AtA2XXo_b5s">WeAreDevelopers 2024</a></p></li><li><p>Responsible AI Workshop Keynote - <a href="http://www.youtube.com/watch?v=57YpXjcj0Ho">NeurIPS 2021</a></p></li><li><p>Practical Guide to ML Explainability - <a href="http://www.youtube.com/watch?v=vq8mDiDODhc">PyCon London</a></p></li><li><p>ML Monitoring: Outliers, Drift, XAI - <a href="http://www.youtube.com/watch?v=QcevzK9ZuDg">PyCon Keynote</a></p></li><li><p>Metadata for E2E MLOps - <a href="https://www.youtube.com/watch?v=OSbH4dfswCY">Kubecon NA 2022</a></p></li><li><p>ML Performance Evaluation at Scale - <a href="http://www.youtube.com/watch?v=8ORl8lu1Eeo">KubeCon Eur 2021</a></p></li><li><p>Industry Strength LLMs - <a href="https://www.youtube.com/watch?v=RVUi_rAFfzU">PyData Global 2022</a></p></li><li><p>ML Security Workshop Keynote - <a href="http://www.youtube.com/watch?v=7XSy5aw8oU8">NeurIPS 2022</a></p></li></ul><div><hr></div><h2>Open Source MLOps Tools</h2><p>Check out the fast-growing ecosystem of production ML tools &amp; frameworks at <a href="http://github.com/EthicalML/awesome-production-machine-learning/">the github repository</a> which has reached over 20,000 &#11088; github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here&#8217;s a few featured open source libraries that we maintain:</p><ul><li><p><strong><a href="https://github.com/besanson/sarc-governance/tree/main">SARC</a> </strong>- Provides wrappers for popular agentic frameworks to enable guardrails and constraints that are enforced through the flow.</p></li><li><p><strong><a href="https://github.com/axsaucedo/kaos">KAOS</a></strong> - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.</p></li><li><p><strong><a href="https://github.com/KomputeProject/kompute">Kompute</a></strong> - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-machine-learning/">Production ML Tools</a></strong> - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-artificial-intelligence-regulation">AI Policy List</a></strong> - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-genai/">Agentic Systems Tools</a></strong> - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain</p></li></ul><p>Please do support some of our open source projects by sharing, contributing or adding a star &#11088;</p><div><hr></div><p><strong>About us</strong></p><p>The Institute for Ethical AI &amp; Machine Learning is a European research centre that carries out world-class research into responsible machine learning.</p><p><a href="https://ethical.institute/">Check out our website</a></p>]]></content:encoded></item><item><title><![CDATA[Issue #386 - The ML Engineer 🤖]]></title><description><![CDATA[Netflix Democratizing MLOps, Stanford AI Index Report 2026, META on ProgramBench, OpenAI on Delivering Voice AI, DeepMind Accelerating Gemma 4 + more &#128640;]]></description><link>https://machinelearning.substack.com/p/issue-386-the-ml-engineer</link><guid isPermaLink="false">https://machinelearning.substack.com/p/issue-386-the-ml-engineer</guid><dc:creator><![CDATA[Alejandro Saucedo]]></dc:creator><pubDate>Sun, 10 May 2026 13:41:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;cafc30e3-8183-4815-afe4-4934255a8d05&quot;,&quot;duration&quot;:null}"></div><p>Thank you for being part of over <a href="https://ethical.institute/mle.html">70,000+ ML professionals and enthusiasts</a> who receive weekly articles &amp; tutorials on Machine Learning &amp; MLOps &#129302; You can join the newsletter <a href="https://bit.ly/state-of-ml-2025">https://bit.ly/state-of-ml-2025</a> &#11088;</p><p>If you like the content please support the newsletter by sharing with your friends via &#9993;&#65039; Email, &#128038; <a href="http://twitter.com/intent/tweet?text=Great%20Newsletter%20on%20Machine%20Learning!%20%20https://ethical.institute/mle.html">Twitter</a>, &#128188; <a href="http://www.linkedin.com/shareArticle?mini=true&amp;url=https://ethical.institute/mle.html&amp;title=The%20Machine%20Learning%20Engineer%20Newsletter&amp;summary=Curated%20news%20about%20machine%20learning%20operations,%20reproducibility,%20explainability%20and%20beyond">Linkedin</a> and &#128213; <a href="http://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fethical.institute%2Fmle.html">Facebook</a>!</p><div><hr></div><h2>This week in ML Engineering:</h2><ul><li><p>Netflix <strong><a href="https://netflixtechblog.com/democratizing-machine-learning-at-netflix-building-the-model-lifecycle-graph-5cc6d5828bb1">Democratizing MLOps</a></strong></p></li><li><p>Stanford <strong><a href="https://hai.stanford.edu/ai-index/2026-ai-index-report">AI Index Report 2026</a></strong></p></li><li><p>META <strong><a href="https://arxiv.org/pdf/2605.03546">on ProgramBench</a></strong></p></li><li><p>OpenAI <strong><a href="https://openai.com/index/delivering-low-latency-voice-ai-at-scale/">on Delivering Voice AI</a></strong></p></li><li><p>DeepMind <strong><a href="https://blog.google/innovation-and-ai/technology/developers-tools/multi-token-prediction-gemma-4/">Accelerating Gemma 4</a></strong></p></li><li><p>Open Source <strong><a href="http://github.com/EthicalML/awesome-production-machine-learning/">ML Frameworks</a></strong></p></li><li><p>Awesome AI Guidelines <strong><a href="https://github.com/ethicalml/awesome-artificial-intelligence-guidelines">to check out this week</a></strong></p></li><li><p>+ more &#128640;</p></li></ul><div><hr></div><p><a href="https://netflixtechblog.com/democratizing-machine-learning-at-netflix-building-the-model-lifecycle-graph-5cc6d5828bb1">Netflix Democratizing MLOps</a></p><p>Netflix shared an in-depth blog on how they democratized machine learning across their organisation by building the ML Model Lifecycle Graph, and there are some really practical learnings: Netflix describes how they built a Metadata Service and Model Lifecycle Graph to make ML assets discoverable, reusable, and debuggable across a fragmented production ML ecosystem spanning models, features, datasets, pipelines, experiments, and ownership systems. The core idea is that they ingest lightweight change events from source systems, hydrate them from the source of truth, normalize them into globally addressable entities, store relationship-heavy metadata in Datomic, index searchable fields in Elasticsearch, and asynchronously enrich cross-system links such as model-to-pipeline-to-A/B-test lineage. For production ML practitioners, it is an important reminder that teams can answer impact, lineage, ownership, and reuse questions from a unified metadata graph if it&#8217;s well documented and accessible to all teams. Mature ML platforms need metadata infrastructure as much as training or serving infrastructure, and it is interesting to see that there&#8217;s a lot of learnings in the dataOps space that can be applied to MLOps.</p><div><hr></div><p><a href="https://hai.stanford.edu/ai-index/2026-ai-index-report">Stanford AI Index Report 2026</a></p><p>Stanford University has just dropped the 2026 Stanford AI Index Report, which shows critical risks on security, incidents spikes, and infrastructure constraints - here&#8217;s some key insights: Model capability is still accelerating, with industry now produces over 90% of notable models, and adoption reaching mainstream levels. Top frontier labls are converging so closely that production ML teams should prioritize guarding themselves from vendor lock-in to quickly be able to switch across when necessary. There are some operating risks, with benchmarks are saturating or proving unreliable, agents failing roughly one in three structured tasks, and incidents rising with constraints around chips, data centers, energy, water, and supply chains. For practitioners, the report&#8217;s main implication is that competitive advantage in 2026 is less about simply accessing the strongest model and more about building robust ML systems around it. The MOAT is now on reproducible evaluations, monitoring, incident response, data governance, cost-aware inference, human oversight, and clear measurement of productivity and safety outcomes.</p><div><hr></div><p><a href="https://arxiv.org/pdf/2605.03546">META on ProgramBench</a></p><p>META is launching a new benchmark that aims to test LLMs on building large-scale / end-to-end applications like ffmpeg / sqlite / interpreters from docs, instead of just snippets/pull-requests: ProgramBench is META&#8217;s new benchmark which evaluates whether coding agents can rebuild full software projects from scratch using only a compiled executable and documentation, rather than editing an existing repo or filling in a scaffold. Current agents are far from reliable autonomous software engineers, as across 200 real open-source tasks0, no model fully solved any task. The best model only passed &gt;=95% of tests on 3% of tasks. The benchmark is valuable because it tests the full lifecycle that production teams actually care about: discovering behavior by probing an executable, making architecture and language choices, implementing a buildable system, and matching black-box behavioral tests without being constrained to the original code structure. This does open some questions on how to mitigate pollution of training data, aka ensuring that the models don&#8217;t just have the solutions injected as part of their training itself, but that&#8217;s an existing issue that plagues the existing benchmarks today anways...</p><div><hr></div><p><a href="https://openai.com/index/delivering-low-latency-voice-ai-at-scale/">OpenAI on Delivering Voice AI</a></p><p>OpenAI has shared how they tackle real0time voice AI with milliseconds latency that ensures seamless experiences; at scale, shaving latency and stabilizing media transport can be the difference between a magical conversational agent vs a frustrating push-to-talk demo. OpenAI&#8217;s write-up is a useful production-infra case study for ML teams building realtime voice agents, as the main constraint is not just model latency, but especially the e2e including connection setup, NAT traversal, packet loss, jitter, first-hop routing, and stable ownership of WebRTC session state. From their post, it seems their solution splits responsibilities between a lightweight UDP relay and a stateful WebRTC transceiver. The broader takeaway for production ML practitioners is that realtime AI quality depends on treating networking and protocol termination as first-class ML platform concerns; namely keep client behavior standards-compliant, isolate hard session state, add complexity in a thin edge-routing layer, and let backend model services scale like normal services rather than WebRTC peers.</p><div><hr></div><p><a href="https://blog.google/innovation-and-ai/technology/developers-tools/multi-token-prediction-gemma-4/">DeepMind Accelerating Gemma 4</a></p><p>Inference speed is becoming one of the most important battlegrounds in production AI - DeepMind shares how they accelerated Gemma 4 inference: Google introduces &#8220;Multi-Token Prediction drafters&#8221; as a production-inference optimization where they pair each Gemma 4 target model with a lightweight multi-token drafter that speculatively proposes several future tokens, while the main model verifies them in parallel. This apparently has yielded Google reported speedups of up to 3x without changing final output quality, which is quite impressive if true. For ML practitioners, the relevant takeaway is that latency bottlenecks are increasingly being attacked through serving architecture such as KV-cache sharing, activation reuse, runtime-specific support across serving frameworks (eg vLLM/MLX/Transformers) rather than only through smaller models or quantization. This is especially relevant for chat, coding assistants, agentic workflows, and on-device or workstation deployments where responsiveness and memory bandwidth dominate user experience, but teams should benchmark against their own prompts, batch sizes, sampling settings, and hardware.</p><div><hr></div><h2>Upcoming MLOps Events</h2><p>The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.</p><p>Events we are speaking at this year:</p><ul><li><p><strong><a href="https://etailgermany.wbresearch.com/">eTail Europe</a> </strong>- March @ Berlin</p></li><li><p><strong><a href="https://worldsummit.ai/">World Summit AI Europe</a></strong> - September @ Amsterdam</p></li></ul><p>Other relevant events:</p><ul><li><p><strong><a href="https://events.linuxfoundation.org/kubecon-cloudnativecon-europe/">KubeCon Europe</a></strong> - March @ Amsterdam</p></li><li><p><strong><a href="https://2026.pycon.de/">PyData Berlin</a></strong> - April @ Frankfurt</p></li><li><p><strong><a href="https://www.databricks.com/dataaisummit">Databricks Summit</a></strong> - June @ San Francisco</p></li><li><p><strong><a href="https://www.wearedevelopers.com/world-congress">World Developer Congress</a></strong> - July @ Berlin</p></li><li><p><strong><a href="https://ep2025.europython.eu/">EuroPython 2026</a></strong> - July @ Prague</p></li><li><p><strong><a href="https://euroscipy.org/">EuroSciPy 2026</a></strong> - July @ Krakow</p></li><li><p><strong><a href="https://www.ai-infra-summit.com/">AI Infra Summit 2026</a> </strong>- Sept @ California</p></li><li><p><strong><a href="https://codetalks.com/">Code.Talks 2026</a></strong> - Nov @ Hamburg</p></li><li><p><strong><a href="https://mlopsworld.com/">MLOps World 2026</a></strong> - Nov @ Austin</p></li></ul><p>In case you missed our talks, check our recordings below:</p><ul><li><p>The State of AI in 2025 - <a href="https://www.youtube.com/watch?v=v2LENQOG-Xg">WeAreDevelopers 2025</a></p></li><li><p>Prod Generative AI in 2024 - <a href="https://www.youtube.com/watch?v=0uJGmMZGUJE&amp;list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&amp;index=3">KubeCon AI Day 2025</a></p></li><li><p>The State of AI in 2024 - <a href="https://www.youtube.com/live/AtA2XXo_b5s">WeAreDevelopers 2024</a></p></li><li><p>Responsible AI Workshop Keynote - <a href="http://www.youtube.com/watch?v=57YpXjcj0Ho">NeurIPS 2021</a></p></li><li><p>Practical Guide to ML Explainability - <a href="http://www.youtube.com/watch?v=vq8mDiDODhc">PyCon London</a></p></li><li><p>ML Monitoring: Outliers, Drift, XAI - <a href="http://www.youtube.com/watch?v=QcevzK9ZuDg">PyCon Keynote</a></p></li><li><p>Metadata for E2E MLOps - <a href="https://www.youtube.com/watch?v=OSbH4dfswCY">Kubecon NA 2022</a></p></li><li><p>ML Performance Evaluation at Scale - <a href="http://www.youtube.com/watch?v=8ORl8lu1Eeo">KubeCon Eur 2021</a></p></li><li><p>Industry Strength LLMs - <a href="https://www.youtube.com/watch?v=RVUi_rAFfzU">PyData Global 2022</a></p></li><li><p>ML Security Workshop Keynote - <a href="http://www.youtube.com/watch?v=7XSy5aw8oU8">NeurIPS 2022</a></p></li></ul><div><hr></div><h2>Open Source MLOps Tools</h2><p>Check out the fast-growing ecosystem of production ML tools &amp; frameworks at <a href="http://github.com/EthicalML/awesome-production-machine-learning/">the github repository</a> which has reached over 20,000 &#11088; github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here&#8217;s a few featured open source libraries that we maintain:</p><ul><li><p><strong><a href="https://github.com/axsaucedo/kaos">KAOS</a></strong> - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.</p></li><li><p><strong><a href="https://github.com/KomputeProject/kompute">Kompute</a></strong> - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-machine-learning/">Production ML Tools</a></strong> - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-artificial-intelligence-regulation">AI Policy List</a></strong> - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-genai/">Agentic Systems Tools</a></strong> - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain</p></li></ul><p>Please do support some of our open source projects by sharing, contributing or adding a star &#11088;</p><div><hr></div><p><strong>About us</strong></p><p>The Institute for Ethical AI &amp; Machine Learning is a European research centre that carries out world-class research into responsible machine learning.</p><p><a href="https://ethical.institute/">Check out our website</a></p>]]></content:encoded></item><item><title><![CDATA[Issue #385 - The ML Engineer 🤖 ]]></title><description><![CDATA[Demis Hassabis on Agents & AGI, Google Cloud MCP Ecosystem, Netflix State of ML Serving, PyTorch Lightning Supply Attack, Multi-Modal Traces in MLFlow + more &#128640;]]></description><link>https://machinelearning.substack.com/p/issue-385-the-ml-engineer</link><guid isPermaLink="false">https://machinelearning.substack.com/p/issue-385-the-ml-engineer</guid><dc:creator><![CDATA[Alejandro Saucedo]]></dc:creator><pubDate>Sun, 03 May 2026 14:09:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;94fda54b-ea79-48c8-9ffe-835048e990f6&quot;,&quot;duration&quot;:null}"></div><p>Thank you for being part of over <a href="https://ethical.institute/mle.html">70,000+ ML professionals and enthusiasts</a> who receive weekly articles &amp; tutorials on Machine Learning &amp; MLOps &#129302; You can join the newsletter <a href="https://bit.ly/state-of-ml-2025">https://bit.ly/state-of-ml-2025</a> &#11088;</p><p>If you like the content please support the newsletter by sharing with your friends via &#9993;&#65039; Email, &#128038; <a href="http://twitter.com/intent/tweet?text=Great%20Newsletter%20on%20Machine%20Learning!%20%20https://ethical.institute/mle.html">Twitter</a>, &#128188; <a href="http://www.linkedin.com/shareArticle?mini=true&amp;url=https://ethical.institute/mle.html&amp;title=The%20Machine%20Learning%20Engineer%20Newsletter&amp;summary=Curated%20news%20about%20machine%20learning%20operations,%20reproducibility,%20explainability%20and%20beyond">Linkedin</a> and &#128213; <a href="http://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fethical.institute%2Fmle.html">Facebook</a>!</p><div><hr></div><h2>This week in ML Engineering:</h2><ul><li><p>Demis <strong><a href="https://www.youtube.com/watch?v=JNyuX1zoOgU">Hassabis on Agents &amp; AGI</a></strong></p></li><li><p>Google <strong><a href="https://cloud.google.com/blog/products/ai-machine-learning/google-managed-mcp-servers-are-available-for-everyone">Cloud MCP Ecosystem</a></strong></p></li><li><p>Netflix State <strong><a href="https://netflixtechblog.com/state-of-routing-in-model-serving-16e22fe18741">of ML Serving</a></strong></p></li><li><p>PyTorch <strong><a href="https://lightning.ai/blog/pytorch-lightning-supply-chain-attack">Lightning Supply Attack</a></strong></p></li><li><p>Multi-Modal <strong><a href="https://mlflow.org/blog/multimodal-tracing/">Traces in MLFlow</a></strong></p></li><li><p>Open Source <strong><a href="http://github.com/EthicalML/awesome-production-machine-learning/">ML Frameworks</a></strong></p></li><li><p>Awesome AI Guidelines <strong><a href="https://github.com/ethicalml/awesome-artificial-intelligence-guidelines">to check out this week</a></strong></p></li><li><p>+ more &#128640;</p></li></ul><div><hr></div><p><a href="https://www.youtube.com/watch?v=JNyuX1zoOgU">Demis Hassabis: Agents &amp; AGI</a></p><p>Demis Hassabis on one of the most insightful conversations this year for ML practitioners talking about Agents, AGI and the Next Scientific Breakthrough - here&#8217;s the key takeaways: Today&#8217;s foundation-model stack is not a dead end, but still needs better continual learning, long-term reasoning, memory, consistency, and introspection before agents can become reliable fire-and-forget systems. Agents are still early, and a lot of near-term value is likely coming from human-in-the-loop workflows, fast distilled models, multimodal systems, local/edge deployment, and specialized tools orchestrated by general models rather than one giant monolith. For ML teams, the AlphaFold pattern is especially relevant, as it is clear the highest-impact opportunities are domains with massive combinatorial search spaces, clear objective functions, and either strong data or simulators, such as drug discovery, materials, biology, and other deep-tech areas. And could not finish without a prediction on AGI, which surprisingly he&#8217;s putting his money on 2030 as the year when it arrives; that sounds to me like what someone that runs an AI lab would say, so this is the one point that you should take with a pinch of salt.</p><div><hr></div><p><a href="https://cloud.google.com/blog/products/ai-machine-learning/google-managed-mcp-servers-are-available-for-everyone">Google Cloud MCP Ecosystem</a></p><p>Google has just published 50+ MCP servers to support programmatic agentic workflows with integrations that are integrated with governance and observability by design: It is great to see this move to support teams to move away fragile / bespoke tool integrations with instead managed MCP endpoints across Cloud infrastructure, databases, analytics, storage, Workspace, Maps, security, payments, and developer docs. For production ML practitioners, the interesting part is less that agents can call tools, but more that there&#8217;s a clear bet from cloud providers to invest in infrastructure where agents are the target user. It is also interesting to see the maturity in the ecosystem, in this case looks as IAM Deny policies, Agent Registry discovery, Model Armor for prompt-injection/data-exfiltration defense, OTel tracing, and Cloud Audit Logs. This seems to be a likely pattern for the rest of the cloud providers to follow, and it will become more interesting as we start seeing agentic systems providing automations higher up in the stack as we move towards operations, monitoring and debugging.</p><div><hr></div><p><a href="https://netflixtechblog.com/state-of-routing-in-model-serving-16e22fe18741">Netflix State of ML Serving</a></p><p>Netflix built a centralized ML serving platform handles over 1M requests/sec and thousands of models that include preprocessing, feature computation, post-processing, and optional learned components. Some impressive architectural decisions: They built a custom routing system called Switchboard that helped them improve their ML velocity, which served requests by using metadata at the request-body-level, however as you can imagine this quite fast became a critical-path dependency, added latency costs (eg request parsing), and made tenant/request-origin isolation harder. To address this, Netflix is introducing their new &#8220;Lightbulb&#8221; design, which instead minimal request context into routing metadata with Envoy proxy performing the actual routing from headers, and model-specific parameters stay in the request body. For teams building ML platforms, the takeaway is that serving abstractions should decouple product clients from model/version/shard churn, but the routing layer itself must eventually become lightweight, cacheable, failure-tolerant, and close to the networking substrate rather than a monolithic proxy in every request path.</p><div><hr></div><p><a href="https://lightning.ai/blog/pytorch-lightning-supply-chain-attack">PyTorch Lightning Supply Attac</a>k</p><p>PyTorch Lightning versions 2.6.2 and 2.6.3 were compromised for a 42-minute window on April 30 after attackers obtained PyPI publishing credentials and uploaded tampered builds: This is an important reminder that security is absolutely key especially in our ML stack; one compromised package release can turn everyday training jobs, notebooks, and CI pipelines into credential-exfiltration paths. This malicious packages executed on import, spawned a background thread, installed Bun, ran an obfuscated JavaScript payload, and targeted cloud credentials, browser-stored secrets, env files, and GitHub tokens. For ML teams, the operational lesson is critical, if either version was installed and imported in developer machines, notebooks, training jobs, or CI/CD runners, treat those environments as compromised, downgrade to lightning==2.6.1, rotate exposed secrets, audit outbound network activity, and review build logs/artifacts. More broadly, this incident shows why MLOps needs MLSecOps controls at the packaging boundary, including pinned and verified dependencies, isolated CI secrets, least-privilege cloud credentials, egress monitoring, artifact provenance, and fast incident playbooks for trusted ML libraries.</p><div><hr></div><p><a href="https://mlflow.org/blog/multimodal-tracing/">Multi-Modal Traces in MLFlow</a></p><p>It is quite interesting that MLflow has recently been leading the charge on ML telemetry and traces, setting the charge for how multimodal tracing can work at scale for images, audio, PDFs, and other files. Open Telemetry standards are growing as now payloads go beyond purely json into binary artifacts extracted from spans, which need to be stored in an existing artifact store, and replaced in the trace database with lightweight references so queries stay fast and storage does not explode. The practical win for production ML practitioners is much better debugging of vision, audio, document, and image-generation workflows; namely allowing for an integrated experience where images render inline, audio can be played, PDFs can be viewed, and custom files can be attached manually via MLflow&#8217;s Attachment API. It&#8217;s really reassuring to see how observability tools for MLOps are maturing at fast pace, especially as ML systems are becoming a critical foundation for society and industry.</p><div><hr></div><h2>Upcoming MLOps Events</h2><p>The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.</p><h3>Events we are speaking at this year:</h3><ul><li><p><strong><a href="https://etailgermany.wbresearch.com/">eTail Europe</a> </strong>- March @ Berlin</p></li><li><p><strong><a href="https://worldsummit.ai/">World Summit AI Europe</a></strong> - September @ Amsterdam</p></li></ul><h3>Other relevant events:</h3><ul><li><p><strong><a href="https://events.linuxfoundation.org/kubecon-cloudnativecon-europe/">KubeCon Europe</a></strong> - March @ Amsterdam</p></li><li><p><strong><a href="https://2026.pycon.de/">PyData Berlin</a></strong> - April @ Frankfurt</p></li><li><p><strong><a href="https://www.databricks.com/dataaisummit">Databricks Summit</a></strong> - June @ San Francisco</p></li><li><p><strong><a href="https://www.wearedevelopers.com/world-congress">World Developer Congress</a></strong> - July @ Berlin</p></li><li><p><strong><a href="https://ep2025.europython.eu/">EuroPython 2026</a></strong> - July @ Prague</p></li><li><p><strong><a href="https://euroscipy.org/">EuroSciPy 2026</a></strong> - July @ Krakow</p></li><li><p><strong><a href="https://www.ai-infra-summit.com/">AI Infra Summit 2026</a> </strong>- Sept @ California</p></li><li><p><strong><a href="https://codetalks.com/">Code.Talks 2026</a></strong> - Nov @ Hamburg</p></li><li><p><strong><a href="https://mlopsworld.com/">MLOps World 2026</a></strong> - Nov @ Austin</p></li></ul><h3>In case you missed our talks, check our recordings below:</h3><ul><li><p>The State of AI in 2025 - <a href="https://www.youtube.com/watch?v=v2LENQOG-Xg">WeAreDevelopers 2025</a></p></li><li><p>Prod Generative AI in 2024 - <a href="https://www.youtube.com/watch?v=0uJGmMZGUJE&amp;list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&amp;index=3">KubeCon AI Day 2025</a></p></li><li><p>The State of AI in 2024 - <a href="https://www.youtube.com/live/AtA2XXo_b5s">WeAreDevelopers 2024</a></p></li><li><p>Responsible AI Workshop Keynote - <a href="http://www.youtube.com/watch?v=57YpXjcj0Ho">NeurIPS 2021</a></p></li><li><p>Practical Guide to ML Explainability - <a href="http://www.youtube.com/watch?v=vq8mDiDODhc">PyCon London</a></p></li><li><p>ML Monitoring: Outliers, Drift, XAI - <a href="http://www.youtube.com/watch?v=QcevzK9ZuDg">PyCon Keynote</a></p></li><li><p>Metadata for E2E MLOps - <a href="https://www.youtube.com/watch?v=OSbH4dfswCY">Kubecon NA 2022</a></p></li><li><p>ML Performance Evaluation at Scale - <a href="http://www.youtube.com/watch?v=8ORl8lu1Eeo">KubeCon Eur 2021</a></p></li><li><p>Industry Strength LLMs - <a href="https://www.youtube.com/watch?v=RVUi_rAFfzU">PyData Global 2022</a></p></li><li><p>ML Security Workshop Keynote - <a href="http://www.youtube.com/watch?v=7XSy5aw8oU8">NeurIPS 2022</a></p></li></ul><div><hr></div><h2>Open Source MLOps Tools</h2><p>Check out the fast-growing ecosystem of production ML tools &amp; frameworks at <a href="http://github.com/EthicalML/awesome-production-machine-learning/">the github repository</a> which has reached over 20,000 &#11088; github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here&#8217;s a few featured open source libraries that we maintain:</p><ul><li><p><strong><a href="https://github.com/axsaucedo/kaos">KAOS</a></strong> - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.</p></li><li><p><strong><a href="https://github.com/KomputeProject/kompute">Kompute</a></strong> - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-machine-learning/">Production ML Tools</a></strong> - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-artificial-intelligence-regulation">AI Policy List</a></strong> - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-genai/">Agentic Systems Tools</a></strong> - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain</p></li></ul><p>Please do support some of our open source projects by sharing, contributing or adding a star &#11088;</p><div><hr></div><p><strong>About us</strong></p><p>The Institute for Ethical AI &amp; Machine Learning is a European research centre that carries out world-class research into responsible machine learning.</p><p><a href="https://ethical.institute/">Check out our website</a></p>]]></content:encoded></item><item><title><![CDATA[Issue #384 - The ML Engineer 🤖 ]]></title><description><![CDATA[Intercom 2X'd Engineering Velocity, Qwen3.6-27B Coding OSS Model, Scientific Theory for Deep Learning, META KernelEvolve Optimizes AI Infrastructure, OpenAI Releasing GPT-5.5 + more &#128640;]]></description><link>https://machinelearning.substack.com/p/issue-384-the-ml-engineer</link><guid isPermaLink="false">https://machinelearning.substack.com/p/issue-384-the-ml-engineer</guid><dc:creator><![CDATA[Alejandro Saucedo]]></dc:creator><pubDate>Sun, 26 Apr 2026 15:42:51 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;6067ad3a-e8ca-4bed-97ec-8dc7fa3b3ae9&quot;,&quot;duration&quot;:null}"></div><p>Thank you for being part of over <a href="https://ethical.institute/mle.html">70,000+ ML professionals and enthusiasts</a> who receive weekly articles &amp; tutorials on Machine Learning &amp; MLOps &#129302; You can join the newsletter <a href="https://bit.ly/state-of-ml-2025">https://bit.ly/state-of-ml-2025</a> &#11088;</p><p>If you like the content please support the newsletter by sharing with your friends via &#9993;&#65039; Email, &#128038; <a href="http://twitter.com/intent/tweet?text=Great%20Newsletter%20on%20Machine%20Learning!%20%20https://ethical.institute/mle.html">Twitter</a>, &#128188; <a href="http://www.linkedin.com/shareArticle?mini=true&amp;url=https://ethical.institute/mle.html&amp;title=The%20Machine%20Learning%20Engineer%20Newsletter&amp;summary=Curated%20news%20about%20machine%20learning%20operations,%20reproducibility,%20explainability%20and%20beyond">Linkedin</a> and &#128213; <a href="http://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fethical.institute%2Fmle.html">Facebook</a>!</p><div><hr></div><h2>This week in ML Engineering:</h2><ul><li><p>Intercom 2X&#8217;d <strong><a href="https://www.youtube.com/watch?v=BRDKft0-dUU">Engineering Velocity</a></strong></p></li><li><p>Qwen3.6-27B <strong><a href="https://qwen.ai/blog?id=qwen3.6-27b">Coding OSS Model</a></strong></p></li><li><p>Scientific Theory <strong><a href="https://arxiv.org/abs/2604.21691">for Deep Learning</a></strong></p></li><li><p>META KernelEvolve <strong><a href="https://engineering.fb.com/2026/04/02/developer-tools/kernelevolve-how-metas-ranking-engineer-agent-optimizes-ai-infrastructure/">Optimizes AI Infrastructure</a></strong></p></li><li><p>OpenAI <strong><a href="https://openai.com/index/introducing-gpt-5-5/">Releasing GPT-5.5</a></strong></p></li><li><p>Open Source <strong><a href="http://github.com/EthicalML/awesome-production-machine-learning/">ML Frameworks</a></strong></p></li><li><p>Awesome AI Guidelines <strong><a href="https://github.com/ethicalml/awesome-artificial-intelligence-guidelines">to check out this week</a></strong></p></li><li><p>+ more &#128640;</p></li></ul><div><hr></div><p><a href="https://www.youtube.com/watch?v=BRDKft0-dUU">Intercom 2X&#8217;d Engineering Velocity</a></p><p>Intercom claims 2x productivity increase with coding agents, and this podcast has some pretty interesting insights on on how they&#8217;ve been approaching this across their org: Many organisations are aggressively trying to figure out how to unlock AI coding productivity beyond the individual and across the organisation, which seems Intercom has figured out a way forward. In nine months, it seems they doubled merged PRs metric while keeping quality stable by treating the AI workflow like an internal product, instrumenting usage with telemetry, analyzing anonymized session data, and building a shared skills repository with hooks that enforce engineering standards automatically. It is great to see that often the secret is actually nothing more than solid engineering practices in the foundation; gains come less from &#8220;allow everyone to do tokenmaxing&#8221; and instead more from building the surrounding foundation for high quality PRs, flaky tests, CI, internal tools and reviews. This is the most important time to invest in CI/code-review bottlenecks, as well as DORA metric improvements, as well as a culture where PMs, designers and engineers can all safely ship code.</p><div><hr></div><p><a href="https://qwen.ai/blog?id=qwen3.6-27b">Qwen3.6-27B Coding</a></p><p>Chinese giant Alibaba releases another impressive open source model with Qwen3.6, with only 27B parameters which shows impressive performance on coding tasks: From the reports it seems that this release brings flagship-level agentic coding into a dense 27B open-weight model, which is super impressive how much can be packed in such a relatively small model. This basically reducess the complexity from very large MoE systems while still outperforming Qwen&#8217;s previous 397B-total / 17B-active open-source flagship on major coding-agent benchmarks like SWE-bench, etc The interesting bit is not just the scores, but the fact that this model is open sourced under Apache-2.0 and it supports text/image/video inputs, offers a 262K native context window extendable up to 1M tokens, and introduces &#8220;thinking preservation&#8221; for multi-turn agentic workflows.</p><div><hr></div><p><a href="https://arxiv.org/abs/2604.21691">Scientific Theory for Deep Learning</a></p><p>It seems we&#8217;re at a stage where deep learning is evolving from alchemy into an engineering discipline; this is an exciting paper which lays out that a scientific theory is emerging for Deep Learning: This is great as having a robust scientific foundation means fewer blind hyperparameter searches, more predictable scaling, better interpretability, and stronger foundations for safety. Deep learning theory is starting to look less like scattered math and more like an emerging &#8220;mechanics of learning&#8221; as a physics-style framework for predicting training dynamics, representations, final weights, and model performance. The paper breaks down into five buckets, including 1) solvable toy settings such as deep linear networks and NTKs; 2) useful limits such as infinite width/depth and lazy vs. rich feature learning; 3) empirical laws such as scaling laws and edge-of-stability behavior; 4) hyperparameter theories such as &#956;P and learning-rate/batch-size scaling; and 5) universal phenomena where different architectures, datasets, and training recipes converge to similar representations.</p><div><hr></div><p><a href="https://engineering.fb.com/2026/04/02/developer-tools/kernelevolve-how-metas-ranking-engineer-agent-optimizes-ai-infrastructure/">META KernelEvolve Optimizes AI Infrastructure</a></p><p>META is building agentic systems that are optimizing the AI infrastructure under their large scale machine learning models, and they have released their framework: META&#8217;s KernelEvolve is a framework they are using for optimizing the low-level infrastructure that determines whether large-scale models are economically viable in production. KernelEvolve sits inside Meta&#8217;s Ranking Engineer Agent stack and turns kernel authoring into a closed-loop search problem across NVIDIA GPUs, AMD GPUs, MTIA chips, and CPUs, using LLM-generated candidates, retrieval-augmented hardware knowledge, tree search, profiling feedback, and automated correctness/performance evaluation. Meta reports compressing weeks of expert kernel work into hours, achieving over 60% inference throughput improvement for the Andromeda Ads model on NVIDIA GPUs and over 25% training throughput improvement for an ads model on MTIA, while supporting DSLs and backends like Triton, CuTe DSL, FlyDSL, CUDA, HIP, and MTIA C++. For production ML practitioners, the key takeaway is that the next bottleneck in model iteration may be less about model design alone and more about automating the systems layer around it: kernel generation, hardware portability, profiling, benchmarking, and continuous optimization across increasingly heterogeneous accelerator fleets.</p><div><hr></div><p><a href="https://openai.com/index/introducing-gpt-5-5/">OpenAI Releasing GPT-5.5</a></p><p>OpenAI has released their latest model GPT-5.5. It seems this model is positioned less as a chatbot and more as a long-running worker for coding, research, document-heavy knowledge work, tool use, and computer operation. For production ML practitioners the main callouts are that the new model is focused on stronger agentic coding and systems reasoning, better long-context performance up to 1M tokens, improved tool reliability, and greater token efficiency while maintaining GPT-5.4-like per-token latency. The benchmarks sound impressive (although we know how these are never confirmed until they are taken for a spin), with 82.7% on Terminal-Bench 2.0, 58.6% on SWE-Bench Pro, etc. Keen to see how this performs in the wild, would be great to hear experiences from practitioners as they take them into production projects.</p><div><hr></div><h2>Upcoming MLOps Events</h2><p>The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.</p><p>Events we are speaking at this year:</p><ul><li><p><strong><a href="https://etailgermany.wbresearch.com/">eTail Europe</a> </strong>- March @ Berlin</p></li><li><p><strong><a href="https://2026.pycon.de/">PyData Berlin</a></strong> - April @ Frankfurt</p></li><li><p><strong><a href="https://worldsummit.ai/">World Summit AI Europe</a></strong> - September @ Amsterdam</p></li></ul><p>Other relevant events:</p><ul><li><p><strong><a href="https://events.linuxfoundation.org/kubecon-cloudnativecon-europe/">KubeCon Europe</a></strong> - March @ Amsterdam</p></li><li><p><strong><a href="https://www.databricks.com/dataaisummit">Databricks Summit</a></strong> - June @ San Francisco</p></li><li><p><strong><a href="https://www.wearedevelopers.com/world-congress">World Developer Congress</a></strong> - July @ Berlin</p></li><li><p><strong><a href="https://ep2025.europython.eu/">EuroPython 2026</a></strong> - July @ Prague</p></li><li><p><strong><a href="https://euroscipy.org/">EuroSciPy 2026</a></strong> - July @ Krakow</p></li><li><p><strong><a href="https://www.ai-infra-summit.com/">AI Infra Summit 2026</a> </strong>- Sept @ California</p></li><li><p><strong><a href="https://codetalks.com/">Code.Talks 2026</a></strong> - Nov @ Hamburg</p></li><li><p><strong><a href="https://mlopsworld.com/">MLOps World 2026</a></strong> - Nov @ Austin</p></li></ul><p>In case you missed our talks, check our recordings below:</p><ul><li><p>The State of AI in 2025 - <a href="https://www.youtube.com/watch?v=v2LENQOG-Xg">WeAreDevelopers 2025</a></p></li><li><p>Prod Generative AI in 2024 - <a href="https://www.youtube.com/watch?v=0uJGmMZGUJE&amp;list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&amp;index=3">KubeCon AI Day 2025</a></p></li><li><p>The State of AI in 2024 - <a href="https://www.youtube.com/live/AtA2XXo_b5s">WeAreDevelopers 2024</a></p></li><li><p>Responsible AI Workshop Keynote - <a href="http://www.youtube.com/watch?v=57YpXjcj0Ho">NeurIPS 2021</a></p></li><li><p>Practical Guide to ML Explainability - <a href="http://www.youtube.com/watch?v=vq8mDiDODhc">PyCon London</a></p></li><li><p>ML Monitoring: Outliers, Drift, XAI - <a href="http://www.youtube.com/watch?v=QcevzK9ZuDg">PyCon Keynote</a></p></li><li><p>Metadata for E2E MLOps - <a href="https://www.youtube.com/watch?v=OSbH4dfswCY">Kubecon NA 2022</a></p></li><li><p>ML Performance Evaluation at Scale - <a href="http://www.youtube.com/watch?v=8ORl8lu1Eeo">KubeCon Eur 2021</a></p></li><li><p>Industry Strength LLMs - <a href="https://www.youtube.com/watch?v=RVUi_rAFfzU">PyData Global 2022</a></p></li><li><p>ML Security Workshop Keynote - <a href="http://www.youtube.com/watch?v=7XSy5aw8oU8">NeurIPS 2022</a></p></li></ul><div><hr></div><h2>Open Source MLOps Tools</h2><p>Check out the fast-growing ecosystem of production ML tools &amp; frameworks at <a href="http://github.com/EthicalML/awesome-production-machine-learning/">the github repository</a> which has reached over 20,000 &#11088; github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here&#8217;s a few featured open source libraries that we maintain:</p><ul><li><p><strong><a href="https://github.com/axsaucedo/kaos">KAOS</a></strong> - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.</p></li><li><p><strong><a href="https://github.com/KomputeProject/kompute">Kompute</a></strong> - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-machine-learning/">Production ML Tools</a></strong> - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-artificial-intelligence-regulation">AI Policy List</a></strong> - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-genai/">Agentic Systems Tools</a></strong> - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain</p></li></ul><p>Please do support some of our open source projects by sharing, contributing or adding a star &#11088;</p><div><hr></div><p><strong>About us</strong></p><p>The Institute for Ethical AI &amp; Machine Learning is a European research centre that carries out world-class research into responsible machine learning.</p><p><a href="https://ethical.institute/">Check out our website</a></p>]]></content:encoded></item><item><title><![CDATA[ Issue #383 - The ML Engineer 🤖 ]]></title><description><![CDATA[LLMs are Databases (Really), NVIDIA Optimization for Agents, Can I Run AI Locally? (Yes.), Kafka Guide to Distributed Messaging, What 81,000 People Want from AI, Open Source ML Frameworks, Awesome AI]]></description><link>https://machinelearning.substack.com/p/issue-383-the-ml-engineer</link><guid isPermaLink="false">https://machinelearning.substack.com/p/issue-383-the-ml-engineer</guid><dc:creator><![CDATA[Alejandro Saucedo]]></dc:creator><pubDate>Sun, 19 Apr 2026 13:51:22 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q5l8!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F24852060-34f9-45e0-801c-4e1fbb65b67d_626x626.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;0a2c4501-557b-41c9-8544-1bac3e638837&quot;,&quot;duration&quot;:null}"></div><p>Thank you for being part of over <a href="https://ethical.institute/mle.html">70,000+ ML professionals and enthusiasts</a> who receive weekly articles &amp; tutorials on Machine Learning &amp; MLOps &#129302; You can join the newsletter <a href="https://bit.ly/state-of-ml-2025">https://bit.ly/state-of-ml-2025</a> &#11088;</p><p>If you like the content please support the newsletter by sharing with your friends via &#9993;&#65039; Email, &#128038; <a href="http://twitter.com/intent/tweet?text=Great%20Newsletter%20on%20Machine%20Learning!%20%20https://ethical.institute/mle.html">Twitter</a>, &#128188; <a href="http://www.linkedin.com/shareArticle?mini=true&amp;url=https://ethical.institute/mle.html&amp;title=The%20Machine%20Learning%20Engineer%20Newsletter&amp;summary=Curated%20news%20about%20machine%20learning%20operations,%20reproducibility,%20explainability%20and%20beyond">Linkedin</a> and &#128213; <a href="http://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fethical.institute%2Fmle.html">Facebook</a>!</p><div><hr></div><h2>This week in ML Engineering:</h2><ul><li><p>LLMs are <strong><a href="https://www.youtube.com/watch?v=8Ppw8254nLI">Databases, Really</a></strong></p></li><li><p>NVIDIA <strong><a href="https://developer.nvidia.com/blog/full-stack-optimizations-for-agentic-inference-with-nvidia-dynamo">Optimization for Agents</a></strong></p></li><li><p>Can I Run <strong><a href="https://www.canirun.ai/">AI Locally? Yes.</a></strong></p></li><li><p>Kafka Guide <strong><a href="https://sushantdhiman.dev/kafka-fundamentals-guide-to-distributed-messaging/">to Distributed Messaging</a></strong></p></li><li><p>What 81,000 <strong><a href="https://www.anthropic.com/features/81k-interviews">People Want from AI</a></strong></p></li><li><p>Open Source <strong><a href="http://github.com/EthicalML/awesome-production-machine-learning/">ML Frameworks</a></strong></p></li><li><p>Awesome AI Guidelines <strong><a href="https://github.com/ethicalml/awesome-artificial-intelligence-guidelines">to check out this week</a></strong></p></li><li><p>+ more &#128640;</p></li></ul><div><hr></div><p><a href="https://www.youtube.com/watch?v=8Ppw8254nLI">LLMs are Databases, Really</a></p><p>Are LLMs databases? Yes. Can we treat them like a Graph Database? Also Yes. What does this look like? It&#8217;s actually quite interesting: Recently there have been projects that are treating LLMs as indexable write-enabled databases using a SQL language that actually loads the models as graph database structures. There is a query language called LARQL which makes any transformer as a queryable graph-like knowledge store, where internal features can be inspected, traversed, and even edited through a SQL-style interface. A great way to show this is by loading Gemma 3 as the demo target, where we can now query entities, relations, and nearest-neighbor feature clusters from model internals. This also allows for not just inference and path-tracing, but also raises questions on how we our thinking will evolve on LLMs as optimizable systems; this provides quite an exciting glimpse into what could be a large field of research in the making. It is still clearly an early-stage prototype rather than production-ready serving infrastructure, but it is a compelling direction for anyone thinking about controllable model editing, efficient inference, and new abstractions for working with model knowledge.</p><div><hr></div><p><a href="https://developer.nvidia.com/blog/full-stack-optimizations-for-agentic-inference-with-nvidia-dynamo">NVIDIA Optimization for Agents</a></p><p>NVIDIA wants to provide an operating system for multi-model agentic orchestration, and this comes quite quite a few interesting architectural learnings: Once coding agents and multi-agent swarms start making hundreds of sequential calls with shared history, the dominant bottleneck becomes keeping KV-cache warm, models routable, and paths reusable across workers. NVIDIA&#8217;s core argument is that self-hosted agent stacks need tighter coordination across frontend APIs, routing, and cache lifecycle management. This includes support for modern agent protocols (responses / messages), expose harness-side metadata like priority and expected output length through agent_hints, route by KV overlap instead of round-robin, and treating cache blocks differently depending on whether they are persistent context or ephemeral reasoning/subagent state. The practical takeaway is that if you are running open models for agentic workloads, you likely need to start thinking beyond &#8220;throughput per GPU&#8221; and toward session-aware inference infrastructure with cache-aware routing, selective retention, multi-tier KV storage, and eventually prefetching.</p><div><hr></div><p><a href="https://www.canirun.ai/">Can I Run AI Locally? Yes.</a></p><p><a href="http://CanIRun.ai">CanIRun.ai</a> locally? Yes. This is a great resource for exploring local inference tooling: We&#8217;ve all come to the question of &#8220;what model can this machine actually run?&#8221;; this resource provides a fast estimate using hardware detection to provide estimated tokens per second based on the models. This of course is not perfect benchmarking, but at least helps us quickly narrow model choices for local copilots, offline workflows, and edge deployments without manually piecing together VRAM charts and quantization assumptions. It&#8217;s quite good to see that the methodology is quite transparent by making clear that the estimates are heuristics since actual performance still depends on runtimes, drivers, thermal limits, etc etc. Check it out!</p><div><hr></div><p><a href="https://sushantdhiman.dev/kafka-fundamentals-guide-to-distributed-messaging/">Kafka Guide to Distributed Messaging</a></p><p>Event-driven systems are one of the quiet foundations of production ML and modern software; they shape how reliably data moves, how quickly systems react, and how well platforms scale under real-world load. Here&#8217;s a great guide around distributed messaging: Service-to-service is evolving in various use-cases from synchronous calls into durable event streams with often the backbone provided by Kafka, leveraging topics, partitions, offsets, and consumer groups. The key production takeaway for ML practitioners is that Kafka can be the backbone for decoupled feature pipelines, inference events, retraining triggers, and replayable state changes. This is particularly relevant where partitioning defines scalability and ordering trade-offs (ie delivery guarantees), as well as how consumer-group design defines parallelism and fault recovery. This is one of the best guides that provide an end-to-end overview, which include even the move from ZooKeeper to KRaft, as well as quite a lot of great fundamentals.</p><div><hr></div><p><a href="https://www.anthropic.com/features/81k-interviews">What 81,000 People Want from AI</a></p><p>What do 81,000 people want from AI? Anthropic has taken the question and provided a thorough report with great key insights: Anthropic gathered 81,000 open-ended interviews across 159 countries and 70 languages, finding that people mostly want AI to improve professional effectiveness, personal transformation, life management, and time freedom. In their report they outline that 81% say AI had already delivered some value, especially through productivity, cognitive partnership, learning, accessibility, and research synthesis. However the most common concerns were unreliability, economic displacement, loss of autonomy, and cognitive atrophy, which makes it clear that a large percentage are still critical. For ML practitioners, the key takeaway is that successful AI products will be won not just through better models, but through key principles taht are implemented in practice, such as reliability, verification, human-in-the-loop design; basically fundamentals of software design are still relevant today!</p><div><hr></div><h2>Upcoming MLOps Events</h2><p>The MLOps ecosystem continues to grow at break-neck speeds, making it ever harder for us as practitioners to stay up to date with relevant developments. A fantsatic way to keep on-top of relevant resources is through the great community and events that the MLOps and Production ML ecosystem offers. This is the reason why we have started curating a list of upcoming events in the space, which are outlined below.</p><p>Events we are speaking at this year:</p><ul><li><p><strong><a href="https://etailgermany.wbresearch.com/">eTail Europe</a> </strong>- March @ Berlin</p></li><li><p><strong><a href="https://worldsummit.ai/">World Summit AI Europe</a></strong> - September @ Amsterdam</p></li></ul><p>Other relevant events:</p><ul><li><p><strong><a href="https://events.linuxfoundation.org/kubecon-cloudnativecon-europe/">KubeCon Europe</a></strong> - March @ Amsterdam</p></li><li><p><strong><a href="https://2026.pycon.de/">PyData Berlin</a></strong> - April @ Frankfurt</p></li><li><p><strong><a href="https://www.databricks.com/dataaisummit">Databricks Summit</a></strong> - June @ San Francisco</p></li><li><p><strong><a href="https://www.wearedevelopers.com/world-congress">World Developer Congress</a></strong> - July @ Berlin</p></li><li><p><strong><a href="https://ep2025.europython.eu/">EuroPython 2026</a></strong> - July @ Prague</p></li><li><p><strong><a href="https://euroscipy.org/">EuroSciPy 2026</a></strong> - July @ Krakow</p></li><li><p><strong><a href="https://www.ai-infra-summit.com/">AI Infra Summit 2026</a> </strong>- Sept @ California</p></li><li><p><strong><a href="https://codetalks.com/">Code.Talks 2026</a></strong> - Nov @ Hamburg</p></li><li><p><strong><a href="https://mlopsworld.com/">MLOps World 2026</a></strong> - Nov @ Austin</p></li></ul><p>In case you missed our talks, check our recordings below:</p><ul><li><p>The State of AI in 2025 - <a href="https://www.youtube.com/watch?v=v2LENQOG-Xg">WeAreDevelopers 2025</a></p></li><li><p>Prod Generative AI in 2024 - <a href="https://www.youtube.com/watch?v=0uJGmMZGUJE&amp;list=PLj6h78yzYM2PRMe1c34bs0sdXo2mlsUov&amp;index=3">KubeCon AI Day 2025</a></p></li><li><p>The State of AI in 2024 - <a href="https://www.youtube.com/live/AtA2XXo_b5s">WeAreDevelopers 2024</a></p></li><li><p>Responsible AI Workshop Keynote - <a href="http://www.youtube.com/watch?v=57YpXjcj0Ho">NeurIPS 2021</a></p></li><li><p>Practical Guide to ML Explainability - <a href="http://www.youtube.com/watch?v=vq8mDiDODhc">PyCon London</a></p></li><li><p>ML Monitoring: Outliers, Drift, XAI - <a href="http://www.youtube.com/watch?v=QcevzK9ZuDg">PyCon Keynote</a></p></li><li><p>Metadata for E2E MLOps - <a href="https://www.youtube.com/watch?v=OSbH4dfswCY">Kubecon NA 2022</a></p></li><li><p>ML Performance Evaluation at Scale - <a href="http://www.youtube.com/watch?v=8ORl8lu1Eeo">KubeCon Eur 2021</a></p></li><li><p>Industry Strength LLMs - <a href="https://www.youtube.com/watch?v=RVUi_rAFfzU">PyData Global 2022</a></p></li><li><p>ML Security Workshop Keynote - <a href="http://www.youtube.com/watch?v=7XSy5aw8oU8">NeurIPS 2022</a></p></li></ul><div><hr></div><h2>Open Source MLOps Tools</h2><p>Check out the fast-growing ecosystem of production ML tools &amp; frameworks at <a href="http://github.com/EthicalML/awesome-production-machine-learning/">the github repository</a> which has reached over 20,000 &#11088; github stars. We are currently looking for more libraries to add - if you know of any that are not listed, please let us know or feel free to add a PR. Here&#8217;s a few featured open source libraries that we maintain:</p><ul><li><p><strong><a href="https://github.com/axsaucedo/kaos">KAOS</a></strong> - K8s Agent Orchestration Service for managing the KAOS in large-scale distributed agentic systems.</p></li><li><p><strong><a href="https://github.com/KomputeProject/kompute">Kompute</a></strong> - Blazing fast, lightweight and mobile phone-enabled GPU compute framework optimized for advanced data processing usecases.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-machine-learning/">Production ML Tools</a></strong> - A curated list of tools to deploy, monitor and optimize machine learning systems at scale.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-artificial-intelligence-regulation">AI Policy List</a></strong> - A mature list that maps the ecosystem of artificial intelligence guidelines, principles, codes of ethics, standards, regulation and beyond.</p></li><li><p><strong><a href="https://github.com/EthicalML/awesome-production-genai/">Agentic Systems Tools</a></strong> - A new list that aims to map the emerging ecosystem of agentic systems with tools and frameworks for scaling this domain</p></li></ul><p>Please do support some of our open source projects by sharing, contributing or adding a star &#11088;</p><div><hr></div><p><strong>About us</strong></p><p>The Institute for Ethical AI &amp; Machine Learning is a European research centre that carries out world-class research into responsible machine learning.</p><p><a href="https://ethical.institute/">Check out our website</a></p>]]></content:encoded></item></channel></rss>