<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Onepagecode]]></title><description><![CDATA[Solving Finance Research Papers.]]></description><link>https://onepagecode.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!_W6n!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3078fdad-8e80-40a9-be17-640eb18fa447_1254x1254.png</url><title>Onepagecode</title><link>https://onepagecode.substack.com</link></image><generator>Substack</generator><lastBuildDate>Tue, 01 Sep 2026 07:50:40 GMT</lastBuildDate><atom:link href="/__u/onepagecode.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Onepagecode]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[onepagecode@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[onepagecode@substack.com]]></itunes:email><itunes:name><![CDATA[Onepagecode]]></itunes:name></itunes:owner><itunes:author><![CDATA[Onepagecode]]></itunes:author><googleplay:owner><![CDATA[onepagecode@substack.com]]></googleplay:owner><googleplay:email><![CDATA[onepagecode@substack.com]]></googleplay:email><googleplay:author><![CDATA[Onepagecode]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Prompt Injection in the Browser: The Complete Security Guide for Frontend & AI Developers]]></title><description><![CDATA[How indirect prompt injection hijacks browser AI agents&#8212;and how to build a hardened React & TypeScript defense pipeline.]]></description><link>https://onepagecode.substack.com/p/prompt-injection-in-the-browser-the</link><guid isPermaLink="false">https://onepagecode.substack.com/p/prompt-injection-in-the-browser-the</guid><dc:creator><![CDATA[Onepagecode]]></dc:creator><pubDate>Wed, 19 Aug 2026 10:18:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!tGPY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e782478-eb44-4f07-b932-210c95de4fe0_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>When a Page Starts Giving Orders to the Assistant</h2><p>You ask a browser-based assistant to summarize a harmless local fixture page. The page looks ordinary: a title, a description, a price, and a few links. But somewhere in the material sent to the model, the fixture also contains a sentence addressed to the assistant: disclose a protected value, follow a different destination, or ignore the user's request.</p><h3>I have provided a small code sample, use the button at the end of this article to download the source code. </h3><p>Nothing needs to execute in the browser for this to matter. The browser or backend may extract the page's visible text, hidden text, attributes, metadata, accessibility content, or links and forward the result as model context. If the application concatenates that material with its trusted task instructions, the model receives instruction-like language from sources with very different security properties. The page remains data in the application, but the model may interpret part of that data as an order. OWASP describes this class of failure as prompt injection, and its prevention guidance treats external content supplied to a model as untrusted input (<a href="https://owasp.org/www-community/attacks/PromptInjection">Prompt Injection</a>, <a href="https://cheatsheetseries.owasp.org/cheatsheets/LLM_Prompt_Injection_Prevention_Cheat_Sheet.html">LLM Prompt Injection Prevention Cheat Sheet</a>).</p><p>The dangerous step is not page rendering. It is the next step: treating the model's interpretation as permission to act.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!tGPY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e782478-eb44-4f07-b932-210c95de4fe0_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!tGPY!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e782478-eb44-4f07-b932-210c95de4fe0_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!tGPY!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e782478-eb44-4f07-b932-210c95de4fe0_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!tGPY!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e782478-eb44-4f07-b932-210c95de4fe0_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!tGPY!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e782478-eb44-4f07-b932-210c95de4fe0_1672x941.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!tGPY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e782478-eb44-4f07-b932-210c95de4fe0_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2e782478-eb44-4f07-b932-210c95de4fe0_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1524527,&quot;alt&quot;:&quot;The browser supplies inputs, but only the server can decide whether a model proposal may reach a tool.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/211830906?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e782478-eb44-4f07-b932-210c95de4fe0_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The browser supplies inputs, but only the server can decide whether a model proposal may reach a tool." title="The browser supplies inputs, but only the server can decide whether a model proposal may reach a tool." srcset="/__u/substackcdn.com/image/fetch/$s_!tGPY!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e782478-eb44-4f07-b932-210c95de4fe0_1672x941.png 424w, /__u/substackcdn.com/image/fetch/$s_!tGPY!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e782478-eb44-4f07-b932-210c95de4fe0_1672x941.png 848w, /__u/substackcdn.com/image/fetch/$s_!tGPY!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e782478-eb44-4f07-b932-210c95de4fe0_1672x941.png 1272w, /__u/substackcdn.com/image/fetch/$s_!tGPY!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e782478-eb44-4f07-b932-210c95de4fe0_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>The browser supplies inputs, but only the server can decide whether a model proposal may reach a tool.</em></p><p>This guide follows that request from browser input to server decision. You will learn how to distinguish direct prompt injection from indirect prompt injection, identify the browser and retrieval surfaces that can carry instruction-like content, and review a small React and Node.js/TypeScript reference implementation. The implementation contrasts an intentionally insecure local path with a hardened path that preserves provenance, bounds context, isolates trust domains, and validates every model-proposed action on the server.</p><p>In the reference flow, the user's request and the fixture enter the server as separate values. The insecure path flattens them into one prompt; the hardened path preserves the fixture's provenance as untrusted data. If the model proposes an unauthorized resource or destination, the server rejects it; if the operation is consequential but otherwise permitted, the server creates an approval-pending result. In either case, no tool executes merely because the page influenced the model.</p><p>The intended readers are frontend engineers, full-stack developers, application-security engineers, AI application developers, and system architects. You should be comfortable with HTTP, the DOM, JavaScript or TypeScript, React, and basic client/server design. Familiarity with authentication, authorization, least privilege, input validation, and LLM messages will help, but the guide connects those ideas as they apply to browser-derived model context.</p><p>The reference project is deliberately bounded. It uses local fixture pages, fake identifiers, fake secret markers, mocked model behavior, and non-destructive tools. The React client selects fixtures and submits requests; it does not hold model credentials, security policy, authorization decisions, approval authority, or privileged tools. Those responsibilities stay on the Node.js server because a browser-controlled value cannot be the final authority for a server-side action.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The insecure path is a teaching comparison, not a vulnerable service to deploy. The hardened path reduces consequences through separation, narrow capabilities, independent validation, approval handling, logging, and fail-closed recovery. It does not prove that a particular model, browser, provider, or deployment resists semantic prompt injection. The supplied project has no execution, test, type-check, benchmark, or production-certification result; real use would require adaptation, security review, and environment-specific hardening.</p><p>The evidence boundary matters as much as the code boundary. The research record supports the broad trust-boundary problem and documented browser-agent threat surfaces, but a local fixture is not evidence that a named real application has been compromised. Where this guide discusses a demonstration, it will identify the evaluated system and avoid turning that result into a universal claim.</p><p>Keep one invariant in view as the guide proceeds: page content may influence what the model proposes, but it must not expand what the server allows. A page does not need to execute code to become dangerous; it only needs to cross into model context that the application later treats as authority.</p>
      <p>
          <a href="/__u/onepagecode.substack.com/p/prompt-injection-in-the-browser-the">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[How to Design Secure Web API Access: The Complete Backend & Architecture Guide]]></title><description><![CDATA[From OAuth2/OIDC JWT validation and multi-tenant isolation to preventing BOLA vulnerabilities with FastAPI and PostgreSQL.]]></description><link>https://onepagecode.substack.com/p/how-to-design-secure-web-api-access</link><guid isPermaLink="false">https://onepagecode.substack.com/p/how-to-design-secure-web-api-access</guid><dc:creator><![CDATA[Onepagecode]]></dc:creator><pubDate>Mon, 17 Aug 2026 10:04:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!lnKP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fefe868-1d49-494e-bb6b-1255d954c829_2816x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What must the API protect, and where are its trust boundaries?</h2><p>An employee sends <code>PATCH /v1/invoices/inv-482</code> to change an invoice for tenant <code>acme</code>. The connection uses TLS, the JWT signature verifies, and the JSON body is well formed. The API still rejects the request.</p><h3>I have provided a complete code implementation for this project. You can download the full source code using the link at the end of this article.</h3><p>That token names a different audience. In the next attempt, the audience is correct but the token lacks <code>invoices:write</code>. A later request has the right scope but belongs to a principal whose membership in <code>acme</code> is inactive. Finally, a properly authenticated tenant member tries to update an invoice that belongs to another tenant. Each request passes an earlier gate and fails at a later one.</p><p>That sequence is the starting point for this guide. Secure Web API access is not one check named &#8220;authentication.&#8221; It is a chain of decisions made at different boundaries, by different components, over different kinds of evidence.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>This guide is for backend engineers, API designers, and system architects who already understand Python, HTTP, JSON, relational databases, SQL, and basic JWT concepts. The reference design targets a bounded FastAPI resource server backed by PostgreSQL and SQLAlchemy, with an external OIDC provider issuing access tokens. It does not build an identity provider, login interface, gateway, cloud deployment, globally distributed rate limiter, or complete disaster-recovery system.</p><p>The generated project is design intent, not verified software. Static verification, semantic verification, code-completion review, and execution were skipped. Nothing in the project should be read as a claim that it compiles, interoperates with a particular OIDC provider, has been security-tested, or is production-certified.</p><h3>Start with assets, not middleware</h3><p>The API protects more than invoice rows. Its asset inventory includes:</p><ul><li><p>Tenant data, including invoices and the relationships that determine who may see or change them.</p></li><li><p>Credentials and assertions, including bearer access tokens, API-key material, and identity-provider metadata.</p></li><li><p>Authorization decisions, because an incorrect allow decision can expose data even when every cryptographic check succeeds.</p></li><li><p>Database integrity, including ownership, tenant identifiers, state transitions, uniqueness constraints, and audit relationships.</p></li><li><p>Audit records and operational signals, which may reveal who accessed what and when.</p></li><li><p>Secrets and signing-key metadata used to connect to databases, providers, and operational systems.</p></li><li><p>Availability and privacy, including the ability to serve legitimate tenants and avoid exposing sensitive identity or diagnostic information.</p></li></ul><p>This grouping applies the confidentiality, integrity, availability, and privacy framing in NIST identity guidance to the requested database-backed API; it is an architectural inventory, not a claim that one source enumerates every application asset (<a href="https://pages.nist.gov/800-63-4/sp800-63/introduction/">NIST SP 800-63-4 introduction</a>).</p><p>The actors are equally important. A human user may work through a browser. A service may synchronize invoices without a human present. An API-key holder may operate a narrowly defined machine integration. The external OIDC provider authenticates users or clients and issues assertions, but it remains outside the API's direct control. A reverse proxy may terminate TLS and forward requests. PostgreSQL stores and constrains data. Operators manage configuration, rotation, monitoring, and incident response. Attackers may steal a token, guess object identifiers, submit over-posted fields, induce expensive requests, probe tenant boundaries, or exploit an outbound URL feature if the API provides one.</p><p>Treat each actor as a source of claims or input&#8212;not as automatically trusted because it sits behind another component.</p><h3>Map the entry points and boundaries</h3><p>A useful boundary map for this API contains at least six transitions:</p><ol><li><p><strong>Client to edge.</strong> The client controls the HTTP method, path, headers, body, timing, and possibly the bearer credential. TLS protects a transport segment; it does not make the request valid or authorized.</p></li><li><p><strong>Edge to application.</strong> Proxy configuration determines which forwarded information the application may trust. Request size, method, origin, and correlation handling belong to this boundary, while authentication and authorization remain application responsibilities.</p></li><li><p><strong>Application to identity provider.</strong> Discovery and JWKS retrieval bring external metadata into the trust decision. The API must bind that metadata to its configured issuer rather than accepting an arbitrary key source.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div></li><li><p><strong>Authentication to authorization.</strong> A validated token becomes an authenticated principal. That principal is evidence about the credential; it is not yet permission to read or mutate an invoice.</p></li><li><p><strong>Application to PostgreSQL.</strong> SQL queries and transactions turn a policy decision into data access. A missed tenant predicate or unconstrained object lookup can undo the checks performed above it.</p></li><li><p><strong>Application and database to audit and operations.</strong> Logs, audit events, metrics, and traces become another boundary. They need enough context for investigation without becoming a second location where tokens, API keys, or sensitive payloads leak.</p></li></ol><p>The boundary map is the architecture. FastAPI dependencies, JWT libraries, and ORM models are implementation locations inside it, not substitutes for it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!lnKP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fefe868-1d49-494e-bb6b-1255d954c829_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!lnKP!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fefe868-1d49-494e-bb6b-1255d954c829_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!lnKP!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fefe868-1d49-494e-bb6b-1255d954c829_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!lnKP!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fefe868-1d49-494e-bb6b-1255d954c829_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!lnKP!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fefe868-1d49-494e-bb6b-1255d954c829_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!lnKP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fefe868-1d49-494e-bb6b-1255d954c829_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0fefe868-1d49-494e-bb6b-1255d954c829_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5694123,&quot;alt&quot;:&quot;Secure API access crosses several independent trust boundaries; TLS and token validation protect specific transitions but do not authorize database operations.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/211523652?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fefe868-1d49-494e-bb6b-1255d954c829_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Secure API access crosses several independent trust boundaries; TLS and token validation protect specific transitions but do not authorize database operations." title="Secure API access crosses several independent trust boundaries; TLS and token validation protect specific transitions but do not authorize database operations." srcset="/__u/substackcdn.com/image/fetch/$s_!lnKP!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fefe868-1d49-494e-bb6b-1255d954c829_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!lnKP!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fefe868-1d49-494e-bb6b-1255d954c829_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!lnKP!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fefe868-1d49-494e-bb6b-1255d954c829_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!lnKP!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fefe868-1d49-494e-bb6b-1255d954c829_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Secure API access crosses several independent trust boundaries; TLS and token validation protect specific transitions but do not authorize database operations.</em></p><h3>Turn the invoice request into abuse cases</h3><p>Return to <code>inv-482</code>. The request can fail in ways that look similar to a client but demand different engineering responses.</p><p>A wrong-audience token is a token issued for some other resource server. Its signature may be valid, but accepting it would confuse one API's trust policy with another's. A stolen bearer token presents a different problem: whoever possesses it may replay it until expiry or effective revocation. Narrow audience and scope, careful handling, bounded lifetimes, and sender-constrained mechanisms address different parts of that risk; none turns a bearer token into proof of the caller's physical identity (<a href="https://www.rfc-editor.org/info/rfc9700/">RFC 9700</a>, <a href="https://www.rfc-editor.org/info/rfc6819/">RFC 6819</a>).</p><p>A missing scope or permission is an authorization failure. An inactive tenant membership is a local policy failure. A cross-tenant object identifier is an object-level authorization failure. OWASP treats object-level, function-level, and property-level authorization as distinct API concerns, which is why an endpoint must check more than whether a principal is generally &#8220;an employee&#8221; (<a href="https://owasp.org/API-Security/editions/2023/en/0xa1-broken-object-level-authorization/">OWASP API1: Broken Object Level Authorization</a>, <a href="https://owasp.org/API-Security/editions/2023/en/0xa3-broken-object-property-level-authorization/">OWASP API3: Broken Object Property Level Authorization</a>).</p><p>Other abuse cases cross different boundaries:</p><ul><li><p>A body containing <code>tenant_id</code>, <code>owner_id</code>, or <code>approval_state</code> tests whether input validation accidentally grants property-level authority.</p></li><li><p>An attacker-controlled URL tests whether the server can be induced to make an unintended outbound request, creating an SSRF capability that does not belong in this bounded project.</p></li><li><p>Large pages, repeated expensive searches, and synchronized retries test resource consumption and availability. OWASP identifies unrestricted resource consumption as an API security risk (<a href="https://owasp.org/API-Security/editions/2023/en/0x11-t10/">OWASP API10: Unsafe Consumption of APIs</a>).</p></li><li><p>A database or provider outage tests whether the system fails closed at a trust decision instead of accepting unverifiable credentials.</p></li><li><p>A suspected tenant-boundary defect tests the operational boundary: can the team stop affected traffic, preserve evidence, identify exposure, and recover deliberately?</p></li></ul><p>These are not variations of one bug. They are separate hypotheses about where an attacker can cross a boundary and what invariant should stop them.</p><h3>Keep the control layers separate</h3><p>A review becomes clearer when each control answers one question:</p><ul><li><p><strong>Transport security:</strong> Can an intermediary or unintended network observer read or alter this connection segment?</p></li><li><p><strong>Authentication:</strong> Is this credential an acceptable assertion from the configured issuer, or a valid machine credential under the API's policy?</p></li><li><p><strong>Authorization:</strong> May this principal perform this operation with this tenant, object, relationship, and set of fields?</p></li><li><p><strong>Application and data security:</strong> Are inputs bounded, fields allowlisted, queries parameterized, and responses limited to permitted properties?</p></li><li><p><strong>Database isolation:</strong> Does the persistence boundary repeat the tenant and object invariants, with optional PostgreSQL enforcement as defense in depth?</p></li><li><p><strong>Abuse resistance:</strong> Can one caller consume unbounded time, memory, connections, retries, or expensive database work?</p></li><li><p><strong>Operations:</strong> Can people detect, investigate, rotate, revoke, contain, and recover when a dependency or credential is compromised?</p></li></ul><p>The practical consequence is direct: when the invoice update is rejected, record which layer made the decision. A valid TLS connection is not authentication. A valid JWT is not authorization. A successful authorization decision is not permission to construct an unrestricted SQL query. A rate limit is not a policy grant. Secure access means preserving those distinctions all the way to the database and the incident runbook.</p><h2>Which system authenticates the caller?</h2><p>The client now has to choose the right credential for the invoice request. It may obtain an access token from an external authorization server and send that token to the API. It must not send an ID token merely because that token is also a JWT, and it must not expect this FastAPI application to exchange a refresh token or authenticate a user's password.</p><p>That division of responsibility determines what the API is. The application is a <strong>resource server</strong>: it protects invoices and consumes access tokens. The external provider performs the authorization-server work and, when OpenID Connect is involved, the identity-provider work. The client&#8212;perhaps a browser application, mobile application, or service&#8212;starts the relevant authorization flow with that provider. The resource owner is the user or organization whose access is being delegated.</p><p>OAuth defines protected-resource access; OpenID Connect adds an identity layer on top of OAuth. The <a href="https://openid.net/specs/openid-connect-core-1_0-18.html">OpenID Connect Core specification</a> describes identity claims delivered to a client, while the API's resource decision remains local to the protected resource. The <a href="https://www.rfc-editor.org/info/rfc6749/">OAuth framework</a> supplies the roles and delegation model, but it does not tell this invoice service whether a principal may change <code>inv-482</code> for tenant <code>acme</code>.</p><h3>Three token classes, three responsibilities</h3><p>Keep the token classes separate when you design the boundary:</p><ul><li><p>An <strong>access token</strong> is presented to the resource server. The API validates that it is intended for this issuer and this API, then uses its trusted claims as inputs to authorization.</p></li><li><p>An <strong>ID token</strong> communicates an OpenID Connect authentication result to the client. Its audience and purpose are tied to that client and login transaction, not automatically to the invoice API. A signature-valid ID token is therefore not a substitute for an API access token.</p></li><li><p>A <strong>refresh token</strong> belongs to the client and authorization-server lifecycle. The client presents it to the authorization server to obtain a new access token; the resource server in this guide never receives or stores it.</p></li></ul><p>Refresh-token expiry, rotation, revocation, and reuse detection belong to the authorization server and client relationship. The API may react to provider-side revocation through short access-token lifetimes, introspection, denylisting, or an incident procedure, but JWT validation alone does not create immediate revocation. The supplied OAuth security baseline, <a href="https://www.rfc-editor.org/info/rfc9700/">RFC 9700</a>, is a finalized IETF Best Current Practice published in January 2025; its text described OAuth 2.1 as under development at that time. Treat &#8220;OAuth 2.1&#8221; as a status-sensitive label, not as a finalized standard unless current IETF evidence confirms that status.</p><h3>Choose where token state lives</h3><p>A resource server commonly chooses between opaque access tokens and self-contained JWT access tokens.</p><p>With an opaque token, the API asks the authorization server&#8212;or an introspection service&#8212;for the token's current meaning. That gives the deployment a centralized place to reflect revocation and changing authorization state, but each validation depends on another service's availability and adds network or caching behavior to the request path.</p><p>With a JWT, the API can validate a signature locally after obtaining the issuer's public keys. That reduces normal per-request dependency traffic and fits this bounded reference implementation. The cost is equally important: the API must enforce an explicit issuer, audience, token-type, algorithm, key, and time policy, and it must accept that an already issued bearer token may remain usable until expiry or another control intervenes. A JWT's claims also do not replace local tenant membership or object authorization.</p><p>Use introspection when immediate revocation or continuously current authorization state is more important than local validation latency and dependency independence. Use self-contained JWTs when the issuer and resource server have a stable trust contract, bounded key-caching behavior, and an operational response for compromised credentials. Neither format eliminates the authorization work inside the API.</p><h3>Make the trust policy visible in configuration</h3><p>The generated project makes the identity boundary explicit before any route code runs. This excerpt is from <code>.env.example</code>; it documents design intent, not a verified provider integration.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;dotenv&quot;,&quot;nodeId&quot;:&quot;7533275d-4b65-4212-b8cb-df8148f01fcb&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-dotenv"># External OIDC issuer. The API is a resource server: it does not mint tokens,
# authenticate users, or perform refresh-token exchanges.
# Production deployments should use an HTTPS issuer and verify its exact value.
OIDC_ISSUER=https://id.example.test/
OIDC_AUDIENCE=secure-api

# Access-token validation policy. Use asymmetric algorithms only; this example
# allows the provider's RSA SHA-256 signing algorithm.
OIDC_ALLOWED_ALGORITHMS=RS256
OIDC_ACCESS_TOKEN_TYPE=at+jwt
OIDC_CLOCK_SKEW_SECONDS=30</code></pre></div><p><code>OIDC_ISSUER</code> identifies the provider whose metadata and keys the API trusts. <code>OIDC_AUDIENCE</code> binds the access token to this resource server rather than another API. The algorithm and token-type settings narrow the accepted access-token profile; they do not prove that a provider actually issues tokens with these values. The clock setting expresses a local tolerance policy for distributed-system timing and requires deployment review.</p><p>The corresponding settings object keeps these values typed and validates the broad trust boundary:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;dd13ba2d-b022-4f5c-9a4a-649bffbed7cf&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    oidc_issuer: str = Field(default="", validation_alias="OIDC_ISSUER")
    oidc_audience: str = Field(default="", validation_alias="OIDC_AUDIENCE")
    oidc_allowed_algorithms: tuple[str, ...] = Field(
        default=("RS256",), validation_alias="OIDC_ALLOWED_ALGORITHMS"
    )
    oidc_access_token_type: str = Field(
        default="at+jwt", validation_alias="OIDC_ACCESS_TOKEN_TYPE"
    )
    oidc_clock_skew_seconds: int = Field(
        default=60, validation_alias="OIDC_CLOCK_SKEW_SECONDS", ge=0, le=600
    )</code></pre></div><p>The excerpt matters because it prevents trust decisions from being scattered across handlers. <code>Settings</code> reads environment-driven policy, while <code>get_settings()</code> returns one cached settings instance for consumers. The settings validator also requires an issuer and audience, permits only configured asymmetric algorithm families, and requires HTTPS outside development. Those checks establish configuration invariants; they do not verify discovery metadata, retrieve JWKS keys, validate a token, or establish tenant membership. Those responsibilities belong to the later authentication and authorization layers.</p><p>The generated code is illustrative and unexecuted: static verification, semantic verification, installation, runtime execution, provider interoperability testing, and security testing were skipped. Read <code>.env.example</code> and <code>app/config.py</code> as an explicit contract for the intended trust boundary, not as evidence that a particular OIDC provider accepts the configuration.</p><p>The practical consequence is a clean handoff. The client and external provider handle login and refresh. The API accepts only the access-token path intended for it, validates that assertion under its configured policy, and then starts a separate authorization decision. A successful provider interaction is evidence about the credential; it is not permission to update an invoice.</p><h2>How should the API validate an external access token?</h2><p>The API has received a bearer token for the invoice request. Before it can construct a principal, it must answer a narrower question than &#8220;does this JWT look valid?&#8221; It must establish that the token is an access token issued by the configured provider, intended for this API, signed under an allowed policy, and valid at the current time.</p><p>The <a href="https://www.rfc-editor.org/rfc/rfc9068.html">JWT access-token profile</a> identifies the core checks: token type, issuer, audience, signature, and expiration. <a href="https://www.rfc-editor.org/rfc/rfc8725.html">JWT best-current-practice guidance</a> adds an important implementation discipline: the verifier must define its algorithm policy and bind accepted keys to the expected issuer. A signature proves that some trusted key signed the bytes. It does not prove that the token was minted for this API or that the caller may update a particular invoice.</p><p>The supplied record also describes <a href="https://www.rfc-editor.org/info/rfc9700/">RFC 9700</a> as a finalized OAuth security Best Current Practice while describing OAuth 2.1 as still under development at that publication point. Use the former as security guidance; do not present OAuth 2.1 as a finalized standard without checking its current status separately.</p><h3>Start with issuer-bound key discovery</h3><p>The verifier should not accept a public key because a token supplied it, because a caller selected a JWKS URL, or because the key happens to parse. The configured issuer determines where discovery begins. Discovery returns metadata, including <code>issuer</code> and <code>jwks_uri</code>; the implementation then verifies that the discovered issuer matches configuration and retrieves signing keys from the validated endpoint. <a href="https://openid.net/specs/openid-connect-discovery-1_0.html">OpenID Connect Discovery</a> defines this metadata relationship.</p><p>The reference cache keeps discovery metadata and its JWKS document as one validated snapshot. A normal request uses the process-local snapshot while it is fresh. If refresh fails, an already known compatible key may remain usable inside the configured maximum stale window. That stale interval, timeout, response-size limit, and refresh cooldown are implementation policies, not universal protocol values. They need deployment review.</p><p>Here is the central key-selection path from <code>app/security/jwks.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ebeb9405-8393-4565-8ce6-a4a9ebda19c2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    def get_signing_key(self, kid: str, alg: str) -&gt; Any:
        """Return a compatible public key, refreshing once for an unknown ``kid``.

        A key is accepted only when its identifier, declared algorithm, key use,
        key type, and configured algorithm policy are mutually compatible. No
        caller-provided URL or key material influences discovery.
        """
        normalized_alg = alg.strip().upper()
        normalized_kid = kid.strip()
        if not normalized_kid:
            raise JWKSFetchError("JWT key identifier is missing")
        self._validate_algorithm(normalized_alg)

        with self._lock:
            now = time.monotonic()
            if self._metadata is None or not self._is_fresh(now):
                try:
                    self._refresh_locked(force=True)
                except JWKSFetchError:
                    if self._metadata is None or not self._is_usable(now):
                        raise

            key = self._compatible_key(normalized_kid, normalized_alg)
            if key is None:
                # This is the only retry path for an unknown key identifier.
                try:
                    self.refresh_once()
                except JWKSFetchError:
                    # A bounded stale snapshot may still be used for a key that
                    # was already present, but an unknown kid remains rejected.
                    if not self._is_usable(time.monotonic()):
                        raise
                key = self._compatible_key(normalized_kid, normalized_alg)

            if key is None or not self._is_usable(time.monotonic()):
                raise JWKSFetchError("No compatible trusted signing key is available")
            return self._public_key(key, normalized_alg)</code></pre></div><p>The method accepts only a non-empty <code>kid</code> and an algorithm that passes the configured asymmetric allowlist. It first refreshes missing or expired cache material. If the key identifier is not present, it permits one cooldown-controlled refresh; it does not loop until the caller's token becomes acceptable. The final check rejects both an unknown key and a snapshot older than the maximum stale age.</p><p>That last distinction matters during key rotation. A newly issued token may legitimately refer to a key that was not in the previous snapshot, so one bounded refresh handles ordinary publication delay. An attacker can also manufacture arbitrary <code>kid</code> values to provoke network work. The cooldown and response bounds limit that pressure, while fail-closed behavior prevents availability from becoming a reason to trust unverifiable material. If the provider is unavailable and the cached snapshot is beyond its permitted age, authentication stops.</p><p>The cache also checks key metadata such as <code>use</code>, <code>kty</code>, and a declared algorithm before converting a JWK into a verification key. Redirects are disabled for its metadata requests, and the endpoint must satisfy the configured HTTP(S) policy. These are selected defenses in this reference project, not proof that every provider or network topology has been handled.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!BPMS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F417449c4-8ce1-475c-8bed-c4456f304f82_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!BPMS!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F417449c4-8ce1-475c-8bed-c4456f304f82_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!BPMS!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F417449c4-8ce1-475c-8bed-c4456f304f82_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!BPMS!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F417449c4-8ce1-475c-8bed-c4456f304f82_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!BPMS!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F417449c4-8ce1-475c-8bed-c4456f304f82_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!BPMS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F417449c4-8ce1-475c-8bed-c4456f304f82_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/417449c4-8ce1-475c-8bed-c4456f304f82_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5188203,&quot;alt&quot;:&quot;Authentication produces a validated principal; authorization and tenant-constrained data access still decide whether the operation may proceed.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/211523652?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F417449c4-8ce1-475c-8bed-c4456f304f82_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Authentication produces a validated principal; authorization and tenant-constrained data access still decide whether the operation may proceed." title="Authentication produces a validated principal; authorization and tenant-constrained data access still decide whether the operation may proceed." srcset="/__u/substackcdn.com/image/fetch/$s_!BPMS!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F417449c4-8ce1-475c-8bed-c4456f304f82_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!BPMS!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F417449c4-8ce1-475c-8bed-c4456f304f82_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!BPMS!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F417449c4-8ce1-475c-8bed-c4456f304f82_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!BPMS!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F417449c4-8ce1-475c-8bed-c4456f304f82_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Authentication produces a validated principal; authorization and tenant-constrained data access still decide whether the operation may proceed.</em></p><h3>Validate the token, then normalize the principal</h3><p>Key retrieval is only one part of verification. The authentication dependency extracts the untrusted JWT header, obtains the issuer-bound key, and asks the JWT library to enforce the configured algorithm, audience, issuer, clock tolerance, and required claims. The relevant excerpt from <code>app/security/auth.py</code> is:</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7d9a1c62-7a2c-40cb-8005-fc9474ea5dec&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        header = jwt.get_unverified_header(token)
        algorithm = header.get("alg")
        key_id = header.get("kid")
        token_type = header.get("typ")
        if not isinstance(algorithm, str) or not isinstance(key_id, str):
            raise ValueError("required JWT header is missing")
        if token_type != settings.oidc_access_token_type:
            raise ValueError("unexpected token type")
        if algorithm.upper() not in settings.oidc_algorithm_allowlist:
            raise ValueError("algorithm is not allowed")

        signing_key = _get_jwks_cache().get_signing_key(key_id, algorithm)
        claims = jwt.decode(
            token,
            signing_key,
            algorithms=list(settings.oidc_algorithm_allowlist),
            audience=settings.oidc_audience,
            issuer=settings.oidc_issuer,
            leeway=settings.oidc_clock_skew_seconds,
            options={
                "require": ["exp", "iss", "aud", "sub"],
                "verify_signature": True,
                "verify_exp": True,
                "verify_iss": True,
                "verify_aud": True,
            },
        )</code></pre></div><p>The unverified header is used only to select a candidate key and policy; it is not trusted as an authorization result. The actual <code>jwt.decode</code> operation verifies the signature and checks <code>exp</code>, <code>iss</code>, and <code>aud</code> against configured values, with the configured clock-skew allowance. Requiring <code>sub</code> is a reference-project policy. Likewise, the exact <code>typ</code> value and the claim names used later for scopes, roles, permissions, and tenant context must match the provider's access-token contract; they are not universal names that can be copied unchanged between providers.</p><p>The surrounding function catches verification failures and returns one stable <code>401</code> response. The client does not learn whether the failure came from an unknown key, wrong issuer, wrong audience, expired token, unsupported algorithm, or malformed claim. Operators still need protected telemetry that counts and correlates these failures without recording the bearer token itself.</p><p>A successful decode becomes a normalized <code>Principal</code>. That transition is deliberately narrow: the principal carries subject, client or service identity, scopes, roles, permissions, and tenant context, but it does not carry an implicit &#8220;may update invoices&#8221; decision. The next policy layer must combine those fields with local membership and the requested object.</p><h3>Walk through the failure matrix</h3><p>Consider a token for the invoice update request whose <code>kid</code> is already cached. If its configured issuer, audience, access-token type, allowed algorithm, signature, required claims, and time checks pass, the API constructs a principal and hands it to authorization. The request is authenticated, not authorized.</p><p>Now vary one fact at a time:</p><ul><li><p>A token signed correctly for another API fails the audience check before scopes or tenant claims matter.</p></li><li><p>A token from another issuer fails issuer validation even if its signature is otherwise valid.</p></li><li><p>A token with an unsupported <code>alg</code>, incompatible key type, or mismatched JWK metadata fails the algorithm and key-compatibility policy.</p></li><li><p>An expired token fails the temporal check. Clock tolerance is a bounded configuration choice, not permission to accept indefinitely stale credentials.</p></li><li><p>A token with an unknown <code>kid</code> causes one controlled refresh. If the refreshed document contains no compatible issuer-bound key, the request fails closed.</p></li><li><p>A provider outage may allow an already known key during the configured stale window. Once that window expires, the API rejects tokens requiring unavailable trust material rather than silently switching to permissive behavior.</p></li></ul><p>Expiration still is not immediate revocation. A stolen bearer token can be replayed until it expires or another deployment-specific control intervenes. <a href="https://www.rfc-editor.org/info/rfc6819/">Bearer-token threat guidance</a> describes this replay risk and the value of reducing exposure through careful transport, narrow audience and scope, limited lifetime, and, where appropriate, sender-constrained mechanisms. This reference path does not implement introspection, denylisting, refresh-token rotation, or sender constraint.</p><p>Refresh tokens belong to the authorization-server and client lifecycle: the resource server neither stores them nor exchanges them in this project. Signing-key rotation belongs to the external issuer, with the resource server managing its discovery and JWKS cache behavior. If a signing key is suspected compromised, the response must be coordinated with that issuer: rotate or disable the affected key according to its incident process, assess tokens issued under it, and refresh or invalidate local trust material as the deployment requires.</p><p>The practical boundary is therefore precise: <code>JWKSCache.get_signing_key</code> establishes which trusted public key may verify the token, and <code>verify_access_token</code> establishes whether the signed assertion is acceptable for this resource server. Only after those functions succeed should <code>get_current_principal</code> hand control to scope, tenant, object, and property authorization. Authentication ends there; the invoice decision has not yet begun.</p><h2>How do you authorize the operation, tenant, object, and fields?</h2><p>The token from the previous section has passed cryptographic validation. Now consider <code>PATCH /v1/invoices/inv-482</code>. What must be true before the database is allowed to change a row?</p><p>The answer is a decision tuple, not a single role lookup:</p><ul><li><p><strong>Credential type:</strong> Is this a user, service, or API-key principal, and is that credential type eligible for this route?</p></li><li><p><strong>Operation scope:</strong> Does the access token grant a coarse capability such as <code>invoices:write</code>?</p></li><li><p><strong>Local permission:</strong> Does the application policy permit this principal to perform the operation?</p></li><li><p><strong>Tenant membership:</strong> Is the user an active member of the tenant that owns the request context?</p></li><li><p><strong>Object relationship:</strong> May this principal access this particular invoice, perhaps because the principal owns it or has an explicit broader permission?</p></li><li><p><strong>Writable properties:</strong> Does the request change only fields that this caller is allowed to change?</p></li></ul><p>These dimensions answer different questions. Scopes are useful operation-level signals, but they do not establish tenant membership or object ownership. The JWT access-token profile describes authorization claims as inputs interpreted by the resource server with contextual information; it does not turn a claim into a universal business rule (<a href="https://www.rfc-editor.org/rfc/rfc9068.html">RFC 9068</a>). OWASP consequently treats object-level and property-level authorization as separate API risks: an endpoint can check that a caller may use the function and still expose another tenant's object or accept a protected field (<a href="https://owasp.org/API-Security/editions/2023/en/0xa1-broken-object-level-authorization/">object-level authorization</a>, <a href="https://owasp.org/API-Security/editions/2023/en/0xa3-broken-object-property-level-authorization/">property-level authorization</a>).</p><p>The reference project makes those dimensions visible in <code>Principal</code>. This is a normalized result of authentication, not a replacement for authorization. The provider-specific mapping of claims into these fields remains an application policy; no particular provider's tenant or role claim names should be treated as universal.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b07a4584-4977-4b40-b18a-1021ae433af0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class Principal(SchemaBase):
    """Normalized identity produced only after authentication succeeds.

    These fields are authorization inputs, not proof by themselves. In
    particular, scopes and roles never replace tenant and object checks.
    """

    credential_type: CredentialType
    subject: str | None = Field(default=None, min_length=1, max_length=255)
    service_identity: str | None = Field(default=None, min_length=1, max_length=255)
    client_id: str | None = Field(default=None, min_length=1, max_length=255)
    scopes: frozenset[Annotated[str, Field(min_length=1, max_length=100)]] = frozenset()
    roles: frozenset[Annotated[str, Field(min_length=1, max_length=100)]] = frozenset()
    permissions: frozenset[Annotated[str, Field(min_length=1, max_length=150)]] = frozenset()
    tenant_id: UUID | None = None
    tenant_ids: frozenset[UUID] = frozenset()
    api_key_id: UUID | None = None</code></pre></div><p>The important design choice is the explicit <code>credential_type</code>. A service principal is not silently treated as a user, and an API key cannot masquerade as a human subject. <code>scopes</code>, <code>permissions</code>, and tenant context are carried forward as inputs to policy, while <code>subject</code> and <code>service_identity</code> preserve attribution. The fields do not grant access on their own.</p><h3>Gate the operation before touching the object</h3><p>A route-level scope check should stop a read-capable principal before an update handler performs work. The generated <code>require_scopes</code> dependency is deliberately narrow:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6e6f926b-3932-49ad-a0a1-383c880d0e5b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def require_scopes(*scopes: str) -&gt; Callable[..., Principal]:
    """Return a FastAPI dependency requiring every listed scope.

    Scopes are coarse operation capabilities. They do not establish tenant
    membership, object ownership, or permission to mutate a particular row.
    """

    required = _required_scope_values(scopes)

    def dependency(
        request: Request,
        session: Session = Depends(__import__("app.db", fromlist=["get_db"]).get_db),
        principal: Principal = Depends(get_current_principal),
    ) -&gt; Principal:
        missing = required.difference(principal.scopes)
        if missing:
            _record_denial(
                session,
                principal,
                request=request,
                action="/__u/onepagecode.substack.com/scope.check",
                reason_code="missing_scope",
            )
            raise _forbidden(
                "missing_scope",
                "the access token does not grant the required operation",
            )
        return principal

    return dependency</code></pre></div><p>The dependency receives an already authenticated <code>Principal</code>, compares the required set with <code>principal.scopes</code>, queues a redacted denial event, and prevents the handler from running. It does not inspect an invoice, because a missing operation capability makes object evaluation unnecessary. A corresponding <code>require_permission</code> dependency checks local application permissions. That second gate matters when the organization needs policy finer than provider-issued scopes.</p><p>The practical consequence is visible in the invoice example. A principal with <code>invoices:read</code> may authenticate successfully and still receive a stable authorization failure for <code>PATCH</code>. A principal with <code>invoices:write</code> may pass the route gate and still fail tenant or object policy. Each denial occurs before a business side effect, and each reason is useful to protected telemetry without being disclosed as a detailed policy explanation to the caller.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!0Jrj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd71ab257-e35a-4a09-ada1-28132103c757_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!0Jrj!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd71ab257-e35a-4a09-ada1-28132103c757_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!0Jrj!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd71ab257-e35a-4a09-ada1-28132103c757_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!0Jrj!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd71ab257-e35a-4a09-ada1-28132103c757_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0Jrj!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd71ab257-e35a-4a09-ada1-28132103c757_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!0Jrj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd71ab257-e35a-4a09-ada1-28132103c757_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d71ab257-e35a-4a09-ada1-28132103c757_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5483809,&quot;alt&quot;:&quot;A valid principal reaches the business operation only after independent scope, permission, tenant, object, and property checks succeed.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/211523652?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd71ab257-e35a-4a09-ada1-28132103c757_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A valid principal reaches the business operation only after independent scope, permission, tenant, object, and property checks succeed." title="A valid principal reaches the business operation only after independent scope, permission, tenant, object, and property checks succeed." srcset="/__u/substackcdn.com/image/fetch/$s_!0Jrj!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd71ab257-e35a-4a09-ada1-28132103c757_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!0Jrj!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd71ab257-e35a-4a09-ada1-28132103c757_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!0Jrj!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd71ab257-e35a-4a09-ada1-28132103c757_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0Jrj!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd71ab257-e35a-4a09-ada1-28132103c757_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>A valid principal reaches the business operation only after independent scope, permission, tenant, object, and property checks succeed.</em></p><h3>Make tenant and object checks one data-access decision</h3><p>The dangerous implementation is to fetch <code>Invoice</code> by <code>invoice_id</code>, then remember&#8212;perhaps only on one route&#8212;to check its tenant. The safer shape puts the object identifier and tenant predicate into the same query. The generated helper first requires tenant context, verifies the caller's authority for that tenant, and then queries both columns:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8b35d535-f454-46a5-b126-e268541978bd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    authorize_tenant(principal, tenant_id, session=session, request=request)
    invoice = session.scalar(
        select(Invoice).where(
            Invoice.id == invoice_id,
            Invoice.tenant_id == tenant_id,
        )
    )

    if invoice is None or not _may_access_invoice(principal, invoice, operation):
        _record_denial(
            session,
            principal,
            request=request,
            action=/__u/onepagecode.substack.com/f%22invoice.%7Boperation%7D%22,
            reason_code="invoice_object_denied_or_missing",
            tenant_id=tenant_id,
            target_id=invoice_id,
        )
        raise HTTPException(
            status_code=status.HTTP_404_NOT_FOUND,
            detail={
                "code": "invoice_not_found",
                "message": "the requested invoice was not found",
            },
        )
    return invoice</code></pre></div><p>The query's invariant is straightforward: an invoice from another tenant is not a candidate for later authorization. The subsequent <code>_may_access_invoice</code> check handles the relationship dimension. In this project, a user may proceed when the relationship policy permits ownership, or when an explicit broader permission such as an <code>:any</code> variant applies. Machine credentials follow a separate policy and do not inherit human ownership merely because they possess a valid key or service token.</p><p>The helper uses an illustrative <code>404</code> for both a missing object and an inaccessible object. That reduces existence disclosure, but it is an API-contract choice rather than a universal rule. Some systems need a distinguishable <code>403</code> for operational clarity; choose deliberately and apply the choice consistently.</p><p>Tenant membership is also local application data in this design. For a human principal, <code>authorize_tenant</code> looks up an active <code>TenantMembership</code> row using the authenticated subject and requested tenant. This avoids treating a provider claim as the complete, current authorization record. It also creates a synchronization responsibility: if membership changes locally, the API's decision must reflect that change even if an access token remains cryptographically valid.</p><h3>Keep server-controlled fields out of the write contract</h3><p>Object authorization also applies to properties. A caller may be allowed to edit an invoice description without being allowed to change <code>tenant_id</code>, <code>owner_id</code>, <code>approval_state</code>, status, or audit metadata. <code>InvoiceUpdate</code> therefore contains only mutable business fields and rejects extra fields through the shared schema configuration. The policy module reinforces that boundary:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a64b097f-e97d-45dd-a7da-f840664711be&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def assert_mutable_properties(update: InvoiceUpdate) -&gt; None:
    """Defend the update boundary against server-controlled field assignment."""

    allowed = {"invoice_number", "description", "amount", "currency"}
    supplied = set(update.model_dump(exclude_unset=True))
    unexpected = supplied.difference(allowed)
    if unexpected:
        raise HTTPException(
            status_code=status.HTTP_422_UNPROCESSABLE_ENTITY,
            detail={
                "code": "immutable_property",
                "message": "one or more submitted fields cannot be changed",
            },
        )</code></pre></div><p>This is an allowlist, not a filter applied after assigning a request dictionary to an ORM object. The service should copy the four permitted values explicitly, leaving ownership and approval state under server policy. The response schema is separate as well, so serialization does not accidentally expose secrets, internal identifiers, or fields whose visibility exceeds the caller's permission.</p><p>For the worked request, the decisions now have a causal order. A correct-audience user token must carry <code>invoices:write</code>; the local membership must be active; the invoice query must match the caller's tenant; the ownership or broader-permission rule must pass; and the body must contain only mutable fields. If any gate fails, the database write never begins. If many services need the same policy, synchronized decisions, and decision logging, a centralized policy engine may become worthwhile&#8212;but it adds deployment and policy-versioning complexity. Until that threshold is reached, explicit dependencies and constrained data-access helpers keep the decision close to the operation it protects.</p><p>The reusable rule is to authorize the <strong>operation, tenant, object, and fields together</strong>. A valid principal is evidence about who or what is calling; it is not permission to change an arbitrary row.</p><h2>How do tenant isolation and safe writes reach PostgreSQL?</h2><p>The authorization decision is not complete when <code>app.security.policy</code> approves the invoice. It must survive the transition into a database query and a transaction. Otherwise, a later query can silently discard the tenant predicate, fetch an object by ID alone, or apply every field from a client-controlled dictionary.</p><p>This is the boundary where the design becomes concrete: every tenant-owned lookup carries tenant context, every object lookup carries the object relationship rule, and every write changes only an allowlisted set of fields. OWASP treats object-level authorization and property-level authorization as separate API risks, which is why the query and the update contract need separate defenses (<a href="https://owasp.org/API-Security/editions/2023/en/0xa1-broken-object-level-authorization/">object-level authorization</a>, <a href="https://owasp.org/API-Security/editions/2023/en/0xa3-broken-object-property-level-authorization/">property-level authorization</a>).</p><h3>Make the tenant boundary part of the data access path</h3><p>For <code>PATCH /v1/invoices/{invoice_id}</code>, do not first load <code>invoice_id</code> and then ask whether the row belongs to the caller's tenant. That creates a gap between retrieval and policy. The shared <code>get_authorized_invoice</code> path is intended to resolve the invoice using the principal, tenant, object identifier, and operation together. The important invariant is:</p><p>A row is eligible for this operation only if its identifier, tenant, and permitted relationship all match the authorization context.</p><p>This also gives you a deliberate response policy. If the query finds no row under the permitted tenant and relationship, the service can return the documented non-disclosing result rather than revealing that the identifier exists in another tenant. The precise choice between <code>403</code> and <code>404</code> is an API policy decision; consistency matters more than allowing one route to disclose object existence while another hides it.</p><p>The model reinforces the boundary rather than replacing it. <code>Tenant</code>, <code>TenantMembership</code>, and <code>Invoice</code> carry UUID identifiers and tenant foreign keys. <code>Invoice</code> also keeps <code>owner_id</code>, <code>approval_state</code>, and <code>status</code> as server-controlled attributes. Indexes such as <code>(tenant_id, id)</code> and <code>(tenant_id, owner_id)</code> support the access patterns the policy requires, while foreign keys and uniqueness constraints protect relational integrity. The model file is design intent, not a substitute for reviewing the corresponding migration.</p><p>The normal application path should use SQLAlchemy expressions with bound values rather than concatenating request strings into SQL. That prevents a user-supplied invoice identifier or filter value from becoming SQL syntax. An ORM does not make every query safe automatically: raw SQL, dynamic column names, unsafe fragments, and unbounded query construction still require review. Keep identifiers and sort expressions on explicit allowlists, and let values enter through parameters.</p><p>The migration boundary makes these assumptions inspectable. <code>migrations/001_initial.sql</code> is a plain SQL, migration-oriented file in this project&#8212;not an Alembic migration, because no Alembic environment is supplied. Apply schema changes through a controlled migration process, review ordering and constraints with the ORM models, and define rollback or forward-repair procedures before deployment. The application should not create or alter production tables during startup.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!TjT1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F894cf084-f90b-4d36-be3f-28d83416b8ca_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!TjT1!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F894cf084-f90b-4d36-be3f-28d83416b8ca_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!TjT1!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F894cf084-f90b-4d36-be3f-28d83416b8ca_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!TjT1!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F894cf084-f90b-4d36-be3f-28d83416b8ca_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!TjT1!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F894cf084-f90b-4d36-be3f-28d83416b8ca_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!TjT1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F894cf084-f90b-4d36-be3f-28d83416b8ca_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/894cf084-f90b-4d36-be3f-28d83416b8ca_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5413539,&quot;alt&quot;:&quot;Application predicates remain the reference authorization path; optional PostgreSQL row-level security adds depth only when transaction and role handling are controlled.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/211523652?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F894cf084-f90b-4d36-be3f-28d83416b8ca_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Application predicates remain the reference authorization path; optional PostgreSQL row-level security adds depth only when transaction and role handling are controlled." title="Application predicates remain the reference authorization path; optional PostgreSQL row-level security adds depth only when transaction and role handling are controlled." srcset="/__u/substackcdn.com/image/fetch/$s_!TjT1!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F894cf084-f90b-4d36-be3f-28d83416b8ca_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!TjT1!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F894cf084-f90b-4d36-be3f-28d83416b8ca_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!TjT1!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F894cf084-f90b-4d36-be3f-28d83416b8ca_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!TjT1!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F894cf084-f90b-4d36-be3f-28d83416b8ca_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Application predicates remain the reference authorization path; optional PostgreSQL row-level security adds depth only when transaction and role handling are controlled.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h3>Keep writes explicit and retry-aware</h3><p>The service layer demonstrates the second half of the boundary. This excerpt matters because it shows both the object lookup and the property-level write policy in one path:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6471ec78-d24e-45fa-a770-dda36cd7b02f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def update_invoice(
    session: Session,
    principal: Principal,
    invoice_id: UUID,
    payload: InvoiceUpdate,
    idempotency_key: str,
    request_context: Any,
) -&gt; Invoice | dict[str, Any]:
    """Apply an allowlisted update and make an exact retry replayable."""
    if not isinstance(payload, InvoiceUpdate):
        raise TypeError("payload must be InvoiceUpdate")
    invoice = get_authorized_invoice(session, principal, invoice_id, "update")
    fingerprint = fingerprint_request(payload)
    existing = claim_idempotency(session, principal, _IDEMPOTENCY_ROUTE, idempotency_key, fingerprint)
    if existing is not None:
        if existing.status == "completed" and isinstance(existing.response_body, Mapping):
            return dict(existing.response_body)
        raise HTTPException(
            status_code=status.HTTP_409_CONFLICT,
            detail={"code": "request_in_progress", "message": "an equivalent request is already in progress"},
        )

    # Explicit assignment is the property-level authorization boundary.
    changes = payload.model_dump(exclude_unset=True)
    for field in ("invoice_number", "description", "amount", "currency"):
        if field in changes:
            setattr(invoice, field, changes[field])
    invoice.updated_at = datetime.now(timezone.utc)
    session.flush()</code></pre></div><p>The input is a typed <code>InvoiceUpdate</code>, not an arbitrary mapping. <code>model_dump(exclude_unset=True)</code> preserves the distinction between omitted and supplied fields, while the explicit field list prevents a client from assigning <code>tenant_id</code>, <code>owner_id</code>, <code>approval_state</code>, <code>status</code>, or audit metadata. The query happens before mutation, and <code>session.flush()</code> sends the constrained change through the current transaction without claiming that the transaction has committed yet.</p><p>The idempotency record adds retry control for this selected database-backed write. Its uniqueness scope includes tenant, principal, route, and key. The service fingerprints the permitted payload, so reusing a key with a different body produces a conflict rather than silently applying a different operation. An exact completed retry can return the stored response; an in-progress request receives the declared conflict response. This limits duplicate effects after an ambiguous network outcome, but it is not general replay protection: it does not protect external side effects, stolen bearer tokens, or every possible concurrent workflow.</p><p>The transaction must include the business update, idempotency state, and audit event under the chosen policy. <code>app.db</code> supplies that boundary:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;928bfb67-3234-418c-a8bf-682eb7d48a40&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@contextmanager
def transactional_session() -&gt; Iterator[Session]:
    """Provide an isolated session with an explicit transaction boundary.

    This context is intended for operations that must commit related changes,
    such as a business write, an idempotency record, and an audit event, as one
    unit. Any exception rolls the transaction back before the connection is
    returned to the pool.
    """
    session = SessionLocal()
    try:
        with session.begin():
            yield session
    except BaseException:
        # ``Session.begin()`` normally performs this rollback itself. Keeping
        # the explicit call makes the invariant clear and also covers failures
        # raised while leaving the transaction context.
        session.rollback()
        raise
    finally:
        session.close()</code></pre></div><p>If the write fails, the related records should not suggest that it succeeded. If the database commits and the client times out, the idempotency record gives an exact retry a controlled result under this teaching policy. Retention, response sensitivity, cleanup, concurrent in-progress handling, and behavior after partial external effects remain deployment and business decisions.</p><h3>Treat row-level security as optional depth</h3><p>The project includes a helper for a conditional PostgreSQL row-level-security design. It sets tenant context locally inside the intended transaction:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ade0a861-05b8-4835-9fa8-8d445ef0f82c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def set_transaction_tenant_context(
    session: Session,
    tenant_id: UUID | str,
) -&gt; None:
    """Set the configured PostgreSQL RLS tenant value for this transaction."""
    try:
        normalized_tenant_id = str(UUID(str(tenant_id)))
    except (TypeError, ValueError, AttributeError) as exc:
        raise ValueError("tenant_id must be a valid UUID") from exc

    session.execute(
        text("SELECT set_config(:setting_name, :tenant_id, true)"),
        {
            "setting_name": _settings.postgres_rls_tenant_setting,
            "tenant_id": normalized_tenant_id,
        },
    )</code></pre></div><p>The <code>true</code> argument makes the setting transaction-local in the generated design, so tenant state is not intended to remain on a pooled connection. The setting name and UUID value enter as parameters, and the UUID is normalized before the call. These choices address connection-state leakage and unsafe value construction, but they do not establish authorization by themselves.</p><p>RLS is therefore defense in depth, not the reference application's sole policy. A deployment enabling it must use appropriate non-privileged roles, account for table-owner bypass behavior, set context in every relevant transaction, handle background jobs and maintenance paths, sequence migrations carefully, and test pooling and policy behavior. A missing context, an incorrectly privileged role, or an incomplete migration can turn an apparent isolation layer into false confidence.</p><p>For the two conceptual PATCH requests, the result is now clear. Tenant <code>acme</code> queries <code>invoices</code> with both the invoice identifier and <code>tenant_id</code>; a request from tenant <code>globex</code> does not obtain the <code>acme</code> row, even if it presents the same identifier. An authorized <code>acme</code> member can change only the four allowlisted invoice fields. Repeating the same request with the same scoped idempotency key can reuse the controlled outcome; reusing it with a different payload is rejected.</p><p>The consequence is architectural: repeat the authorization invariant at the database access boundary, make mutation fields explicit, and treat database policy as a contract shared by models, migrations, transactions, and operations&#8212;not as an automatic property of using an ORM.</p><h2>How should machine credentials be issued and rotated?</h2><p>A synchronization worker needs to push invoice changes for tenant <code>acme</code>. It does not represent an employee, does not need a browser login, and should not receive a broadly privileged credential. That makes a narrowly scoped API key a possible fit&#8212;but only if you treat the key as a constrained machine credential rather than as a replacement for user identity or authorization.</p><p>The reference policy binds each key to a tenant, service identity, scopes, permissions, expiry, and lifecycle metadata. The key is displayed only when it is created or rotated. The database stores a non-secret prefix for lookup and a verifier hash, never the presented secret. This is an implementation policy for the reference project, not a universal API-key standard; the hashing parameters and operational limits require review against the deployment's hardware and threat model.</p><p>The boundary matters because a static key is a replayable bearer credential. Anyone who obtains it can present it until expiry, revocation, or another control stops acceptance. TLS protects it in transit, but TLS does not repair a leaked key. Route restrictions, tenant binding, narrow scopes, expiry, rate limits, audit records, and rotation reduce different parts of the exposure. Stronger workload identity, private-key authentication, mTLS, or platform workload identity may be preferable when the environment supports them. The <a href="https://www.rfc-editor.org/info/rfc9700/">OAuth security Best Current Practice</a> and <a href="https://owasp.org/API-Security/editions/2023/en/0xa2-broken-authentication/">OWASP API authentication guidance</a> provide the surrounding security context; the exact API-key format below remains a project choice.</p><h3>Issue the secret once, then verify by prefix</h3><p><code>ApiKeyPolicy</code> makes the intended authorization envelope explicit. A key cannot silently become an anonymous global credential because creation requires both a tenant and a service identity. <code>create_api_key</code> generates the secret, selects a non-secret prefix, hashes the complete value with the configured <code>scrypt</code> policy, flushes only the verifier-bearing row, and returns the plaintext through the issuance response. The caller must deliver that response through a protected provisioning channel and avoid putting the secret in logs, tickets, or source control.</p><p>The following exact excerpt is from <code>app/security/apikeys.py</code>, symbol <code>verify_api_key</code>. It matters because authentication does not end at prefix lookup: the prefix selects a candidate, while the verifier establishes possession of the complete secret and lifecycle checks establish whether that credential is still accepted.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;bf9e3a34-0c15-4df1-b203-1df96941c220&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def verify_api_key(session: Session, raw_key: str) -&gt; Principal:
    """Authenticate a key by prefix lookup and verifier-hash validation."""

    scheme = _require_supported_scheme()
    settings = get_settings()
    if not isinstance(raw_key, str) or len(raw_key) &lt; settings.api_key_prefix_length:
        raise ValueError("invalid API key")

    prefix = raw_key[: settings.api_key_prefix_length]
    record = session.scalar(select(ApiKey).where(ApiKey.key_prefix == prefix))
    now = utc_now()
    if (
        record is None
        or record.revoked_at is not None
        or (record.expires_at is not None and _utc(record.expires_at) &lt;= now)
        or not _verify_hash(raw_key, record.secret_hash, scheme)
    ):
        raise ValueError("invalid API key")

    record.last_used_at = now
    session.flush()
    return Principal(
        credential_type=CredentialType.API_KEY,
        service_identity=record.service_identity,
        scopes=frozenset(record.scopes or ()),
        permissions=frozenset(record.permissions or ()),
        tenant_id=record.tenant_id,
        tenant_ids=frozenset({record.tenant_id}),
        api_key_id=record.id,
    )</code></pre></div><p>The important invariant is that every successful result is an explicitly marked API-key principal. It carries machine identity, tenant context, scopes, and permissions, but it does not acquire a human subject or human ownership relationship. The route still has to require the machine credential, check the <code>invoices:sync</code> scope, enforce the assigned tenant, and reject the key on a human approval endpoint. The excerpt is generated project material and was not executed or independently verified.</p><p>The failure path is deliberately indistinguishable at the client boundary: an unknown prefix, malformed verifier, expired key, revoked key, or wrong secret becomes an invalid-key failure. Operators can still record a redacted event containing the key identifier or prefix, tenant, route, outcome, and correlation ID&#8212;never the plaintext key.</p><h3>Rotate the credential without extending its authority</h3><p>Rotation should revoke the old record and issue a replacement carrying the same bounded tenant, service, scope, and permission policy. A deployment may permit a short overlap only if it can account for both keys and revoke the old one deterministically; the reference path favors explicit revocation. If the secret leaks, revoke it, issue a replacement, inspect recent use, assess adjacent credentials, and consider moving the integration to a stronger workload identity.</p><p>The provider owns external access-token issuance, refresh-token handling, and OIDC signing-key publication and compromise coordination. The API owns its local API-key records, revocation, rotation, and machine-route policy. Keeping those owners separate prevents a key rotation from being mistaken for JWT revocation&#8212;or a valid machine credential from becoming authorization by itself.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Ihnm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c08eb70-edca-4a3f-afd6-219a183217ad_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Ihnm!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c08eb70-edca-4a3f-afd6-219a183217ad_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!Ihnm!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c08eb70-edca-4a3f-afd6-219a183217ad_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!Ihnm!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c08eb70-edca-4a3f-afd6-219a183217ad_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Ihnm!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c08eb70-edca-4a3f-afd6-219a183217ad_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Ihnm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c08eb70-edca-4a3f-afd6-219a183217ad_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2c08eb70-edca-4a3f-afd6-219a183217ad_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5354454,&quot;alt&quot;:&quot;Different credentials have different owners and recovery paths: the provider manages token and signing-key lifecycle, while the API manages its narrow machine-key lifecycle.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/211523652?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c08eb70-edca-4a3f-afd6-219a183217ad_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Different credentials have different owners and recovery paths: the provider manages token and signing-key lifecycle, while the API manages its narrow machine-key lifecycle." title="Different credentials have different owners and recovery paths: the provider manages token and signing-key lifecycle, while the API manages its narrow machine-key lifecycle." srcset="/__u/substackcdn.com/image/fetch/$s_!Ihnm!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c08eb70-edca-4a3f-afd6-219a183217ad_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!Ihnm!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c08eb70-edca-4a3f-afd6-219a183217ad_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!Ihnm!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c08eb70-edca-4a3f-afd6-219a183217ad_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Ihnm!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c08eb70-edca-4a3f-afd6-219a183217ad_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Different credentials have different owners and recovery paths: the provider manages token and signing-key lifecycle, while the API manages its narrow machine-key lifecycle.</em></p><p>The following excerpt is copied verbatim from <code>app/security/apikeys.py</code> (<code>ApiKeyPolicy</code>) in the finalized generated project:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e6d51d35-1487-4ae1-a1f1-3f6672aca674&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class ApiKeyPolicy:
    """Creation policy for one machine credential.

    ``expires_at`` is authoritative when supplied. Otherwise the configured
    maximum age is applied. A key must always have a tenant and a service
    identity so it cannot silently become an anonymous global credential.
    """

    tenant_id: UUID
    name: str
    service_identity: str
    scopes: frozenset[str] = frozenset()
    permissions: frozenset[str] = frozenset()
    expires_at: datetime | None = None</code></pre></div><h2>What do browser, HTTP, and input boundaries protect?</h2><p>The invoice request has passed authentication and authorization, but it still arrives through an HTTP boundary that can change its risk profile. A browser may attach cookies automatically. A reverse proxy may terminate TLS before forwarding the request. A client may submit a body that is valid JSON but attempts to change <code>tenant_id</code>, <code>owner_id</code>, or <code>approval_state</code>. A future feature may accept a URL and cause the server to make an outbound request. These are different problems, so they need different controls.</p><h3>Start with the transport and proxy boundary</h3><p>TLS protects the connection between the client and the endpoint that terminates TLS. It does not automatically protect every hop after termination, and it does not establish that a forwarded client address or scheme is trustworthy. Decide where TLS terminates, which internal hops require encryption, and which proxy is authorized to supply forwarded headers. The application should trust forwarded values only when the deployment explicitly identifies that proxy boundary; otherwise, an attacker may influence scheme, host, or client-address information used in redirects, logging, rate-limit keys, or policy decisions.</p><p>The FastAPI reference places correlation, CORS, request-boundary handling, and generic error handling in <code>app/middleware.py</code>. That location makes the order visible, but middleware is not a substitute for the authentication and authorization dependencies from the preceding sections. A request can have an accepted origin and a valid TLS connection while still carrying an invalid token or targeting another tenant.</p><h3>Keep CORS separate from authorization</h3><p>Cross-Origin Resource Sharing, or CORS, controls whether a browser permits a web page from one origin to read a response from another origin. Treat the configured origin list as an explicit allowlist. Do not reflect arbitrary <code>Origin</code> values, and do not regard a successful preflight as permission to read an invoice. Non-browser clients do not use CORS as their access-control mechanism, and a caller can send requests without satisfying the browser's same-origin checks.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>CORS is therefore neither authentication nor authorization. It also is not a general CSRF defense. Cross-site request forgery depends primarily on whether the browser attaches credentials automatically. If the client deliberately places a bearer token in an <code>Authorization</code> header, the browser's token-handling and cross-origin policy deserve careful review, but the threat is different from a cookie that the browser sends whenever it matches a destination.</p><p>For a cookie-based browser design, decide the cookie's <code>Secure</code>, <code>HttpOnly</code>, and <code>SameSite</code> policy, the accepted origins, and the CSRF mechanism as one deployment contract. A backend-for-frontend can keep tokens on the server and expose a session cookie, but that choice introduces session lifecycle, CSRF-token, origin-validation, logout, and proxy considerations. A direct browser client using bearer tokens avoids ambient cookie transmission but places more responsibility on the client application's token storage and XSS posture. Neither arrangement is universally safer; choose according to the browser topology and threat model rather than copying a framework default.</p><h3>Make request schemas an authorization boundary</h3><p>A JSON parser answers whether the document has the expected syntax. It does not answer which properties the caller may change. The reference project keeps that distinction in <code>app/schemas.py</code>: invoice creation, update, and response contracts are separate. An update body should contain only business fields that the route permits, such as an allowed status or description. It should not accept <code>tenant_id</code>, <code>owner_id</code>, approval state, role information, audit timestamps, or database identifiers as ordinary client-controlled fields.</p><p>This is property-level authorization. OWASP describes mass assignment and excessive property exposure as distinct API risks, so the safe pattern is to allowlist both writable and returned fields (<a href="https://owasp.org/API-Security/editions/2023/en/0xa3-broken-object-property-level-authorization/">property-level authorization guidance</a>). The handler should explicitly copy accepted values onto the model rather than passing an untrusted dictionary into an ORM constructor or update expression. The response schema should likewise expose a deliberate representation instead of serializing every mapped column.</p><p>Put bounds on the request as well: body size, string length, enum values, page size, sort choices, and filter complexity. Pagination should have a maximum rather than allowing a client to turn one request into an unbounded database read. Query construction should use parameterized SQLAlchemy expressions and controlled sort or column mappings; an ORM does not make dynamically assembled SQL safe when identifiers or fragments are interpolated manually.</p><p>The consequence is useful during review: ask not only &#8220;can this principal reach the route?&#8221; but also &#8220;which bytes can this principal cause the server to store, query, return, or send elsewhere?&#8221;</p><h3>Treat outbound requests as a separate capability</h3><p>The bounded reference implementation deliberately has no arbitrary URL-fetching endpoint. That omission is a security boundary, not a missing convenience feature. If a client can supply a URL and the server fetches it, the server becomes an intermediary between untrusted input and internal or external network locations. The resulting SSRF design needs its own review.</p><p>If such a capability becomes necessary, begin with an allowlist of destinations rather than accepting arbitrary URLs. Define permitted schemes and ports, constrain DNS and resolved addresses, control redirects, apply connection and read timeouts, cap response size, restrict returned content types, and isolate outbound networking from sensitive internal services. Re-check the destination across redirects and resolution steps according to the deployment's network model. Log the decision without recording credentials or sensitive response bodies.</p><p>This is deployment guidance, not a feature supplied by the reference project. Avoiding the capability keeps the invoice request path focused: the server validates the request, authorizes the resource, writes only permitted fields, and returns a controlled representation. It never turns a user-provided string into an unbounded network action.</p><p>The practical rule is to assign each boundary one job: TLS and proxy policy protect transport assumptions, CORS and CSRF policy govern browser credential behavior, schemas constrain data shape and properties, and outbound-request policy controls server-side network access. Authentication and tenant authorization must still make the final decision about the invoice.</p><h2>How should the API resist abuse and preserve evidence?</h2><p>The invoice update has passed authentication and authorization, but the client may still retry it, send it in a burst, or cause an expensive operation repeatedly. A secure API therefore needs two additional controls with different jobs: consumption controls protect availability, while evidence controls preserve enough context to investigate what happened. Neither one grants permission to update an invoice.</p><p>This distinction matters during a timeout. Suppose the client sends <code>PATCH /v1/invoices/inv-482</code> with an idempotency key, the database commits, and the response is lost on the network. The client retries because it cannot tell whether the write succeeded. Idempotency is designed for this ambiguity: the server binds the key to the principal, tenant, route, and request fingerprint, then returns the previously recorded outcome for an exact retry. A conflicting body using the same key must be rejected rather than treated as a new operation.</p><p>Idempotency is not general replay protection. It does not stop an attacker from replaying a bearer token, and it does not automatically make an external side effect&#8212;such as sending an email or calling a payment provider&#8212;safe to repeat. The reference project applies the policy to a selected database write. Production design still has to define concurrent requests, in-progress state, expiry, cleanup, a commit followed by response loss, and what happens when a downstream side effect partially succeeds.</p><h3>Limit consumption without confusing it with authorization</h3><p>Rate limiting answers &#8220;how much may this identity, tenant, route, or network source consume in a period?&#8221; Authorization answers &#8220;may this principal perform this operation on this resource?&#8221; Keep those decisions separate. A caller can be fully authorized and still exceed a limit; an unauthorized caller can remain below the limit and must still be denied.</p><p>The generated <code>app/security/limiter.py</code> uses <code>RateLimiter</code> as a bounded, in-process fixed-window implementation. Its intended key combines credential identity, tenant context, HTTP method, and route. When no principal exists yet, the dependency falls back to the client host as a coarse pre-authentication dimension. The limiter uses monotonic time, removes expired windows, caps the number of stored entries, and returns a stable <code>429</code> response with <code>Retry-After</code> metadata.</p><p>That implementation is deliberately local teaching infrastructure. Each worker has its own counters, counters reset on restart, and replicas can disagree about how many requests a tenant has made. It cannot enforce a fleet-wide quota or guarantee fairness across instances. The <a href="https://owasp.org/API-Security/editions/2023/en/0x11-t10/">OWASP API Security guidance on unrestricted resource consumption</a> supports treating resource consumption as an API security concern; the local-versus-distributed limitation is an implementation conclusion. Replace the limiter with coordinated state or trusted upstream enforcement when the policy must be global.</p><p>Limits should cover more than request counts. Bound payload size, page size, concurrent work, database execution time, batch cardinality, and expensive route-specific operations. Apply stricter policies to unauthenticated traffic and consider tenant-aware limits so one customer cannot consume a shared pool. A <code>429</code> response controls pressure; it does not explain whether the underlying business operation was authorized.</p><h3>Preserve evidence without preserving secrets</h3><p>A security audit event should answer who acted, on which tenant, against what target, through which operation, with what outcome, and why. Useful fields include <code>request_id</code>, actor type, subject or API-key identifier, tenant, action, target type and identifier, outcome, reason code, and timestamp. The event should not contain an authorization header, raw JWT, refresh token, API-key plaintext, cookie, password, or unnecessary request body.</p><p>In <code>app/observability.py</code>, <code>record_audit_event</code> is designed to add a structured <code>AuditEvent</code> to the caller's SQLAlchemy session without committing it. That boundary is important: a successful invoice mutation and its audit record can follow the selected transaction policy together. The function recursively redacts credential-bearing metadata before storing it. It also keeps audit persistence conceptually separate from <code>safe_log</code>, which emits a smaller diagnostic record and updates coarse in-process counters.</p><p>This creates a deliberate tradeoff. Synchronous audit persistence gives the request path a clear relationship between the business action and its event, but an audit-store failure can affect availability unless the application defines a fallback. Asynchronous delivery can reduce request latency, but then the system needs durable queues, loss detection, replay handling, and a way to identify missing events. Retention, operator access, tamper evidence, privacy filtering, and alert thresholds remain deployment responsibilities. The <a href="https://pages.nist.gov/800-63-4/sp800-63.html">NIST Digital Identity Guidelines</a> provide a broader identity-risk and lifecycle context, but the exact event schema and retention policy here are design choices.</p><h3>Return stable errors and useful correlation</h3><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><code>app/middleware.py</code> gives each request a validated correlation identifier and returns it with responses. Its intended <code>SafeErrorMiddleware</code> behavior is to expose a stable error code, generic message, and request identifier while sending diagnostic context to protected server-side logging. That keeps database details, provider responses, stack traces, and token-validation internals out of the client contract.</p><p>Treat correlation as a pointer, not as proof of identity. Operators use it to join an authorization denial, database failure, audit event, and trace without placing secrets in every record. Alert on patterns such as repeated invalid tokens, unknown signing keys, API-key use after revocation, cross-tenant probes, idempotency conflicts, rate-limit bursts, and audit-write failures. A single <code>500</code> is a client response; the correlated operational signal determines whether it is an isolated defect or an incident.</p><p>The consequence is practical: rate limits contain consumption, idempotency contains selected duplicate writes, and audit plus correlation makes decisions reconstructable. They are separate invariants, and production readiness depends on assigning each invariant an owner, a durable failure policy, and evidence that the deployment&#8212;not merely the declared code&#8212;enforces it.</p><h2>Can you trace one request through the reference server?</h2><p>Take the successful case first: an authorized employee sends <code>PATCH /v1/invoices/inv-482</code> for tenant <code>acme</code>. The request carries an access token in <code>Authorization</code>, an <code>Idempotency-Key</code>, and a JSON body containing only mutable invoice fields. The route must preserve one invariant from edge to database: the caller may change this invoice, in this tenant, under this credential, with this operation, and only once for this idempotency key.</p><p>The generated project assigns each part of that decision to a narrow owner. <code>app/main.py</code> assembles the FastAPI application, installs middleware, registers routes, and keeps migrations and OIDC discovery outside implicit startup work. Its <code>/health</code> endpoint reports process-level handling rather than claiming that PostgreSQL, migrations, or the identity provider are healthy. That distinction matters during an outage: a process that can answer health requests is not necessarily able to authenticate or safely serve protected data.</p><h3>1. Establish request context before trusting request content</h3><p><code>CorrelationIdMiddleware</code> runs at the HTTP boundary. It creates a validated request identifier, checks the declared <code>Content-Length</code> against the configured request limit, and places the identifier on the response. <code>SafeErrorMiddleware</code> later converts unexpected failures into a generic error containing the correlation identifier, while protected server-side logs retain diagnostic context.</p><p>This stage does not authenticate the caller. It does not decide whether <code>inv-482</code> exists, and it does not make a browser origin trusted. CORS remains an explicit origin policy, not authorization. The useful result is narrower: later authentication, audit, and error paths can refer to one request without copying credentials or sensitive payloads into client responses.</p><p>The generated size check covers declared lengths; strict enforcement for chunked requests remains an upstream or lower-level ASGI responsibility. That is a design boundary, not a reason to let request limits disappear.</p><h3>2. Authenticate exactly one credential</h3><p><code>patch_invoice</code> obtains its principal through route dependencies, eventually reaching <code>get_current_principal</code> in <code>app/security/auth.py</code>. The dependency accepts either a bearer access token or the configured API-key header, but not both. With a bearer token, <code>verify_access_token</code> obtains an issuer-bound signing key from <code>JWKSCache</code>, applies the configured algorithm and token-type policy, and validates signature, issuer, audience, required claims, and expiration. The <a href="https://www.rfc-editor.org/rfc/rfc9068.html">JWT access-token profile</a> supports treating these as distinct validation requirements rather than collapsing them into &#8220;the JWT parsed.&#8221;</p><p>A wrong-audience token therefore stops here, even if its signature is valid. The route handler never runs, and the client receives the stable authentication failure rather than a reason detailed enough to expose verification state. An unknown <code>kid</code> may trigger the cache's bounded refresh policy; if no compatible trusted key is available, authentication fails closed. The project does not claim that this behavior has been executed or interoperated with a particular provider.</p><p>A successful result is a normalized <code>Principal</code>, not an authorization decision. It identifies the credential type and carries subjects, service identity, scopes, permissions, and tenant context for the next layer.</p><h3>3. Apply route and operation policy</h3><p>The route requires a user principal, <code>invoices:update</code>, and the corresponding local permission. <code>require_scopes</code> checks the coarse operation capability. <code>require_permission</code> checks application policy. These checks deliberately overlap without being duplicates: a provider-issued scope says what kind of operation the token may request, while a local permission says what the application permits this principal to do.</p><p>If the token has only <code>invoices:read</code>, the request ends before a database mutation. The API records a redacted denial event in the request's session where the available context permits it. This follows the distinction described by <a href="https://owasp.org/API-Security/editions/2023/en/0xa1-broken-object-level-authorization/">OWASP's API authorization guidance</a>: authenticating a caller does not authorize access to every object that caller can name.</p><h3>4. Recheck tenant and object authority beside the query</h3><p>The service calls <code>get_authorized_invoice</code>, which calls <code>authorize_tenant</code> before resolving the invoice. For a human principal, the policy queries <code>TenantMembership</code> and requires active local membership. A tenant claim in an external token is context for the application; it is not treated as proof that the local membership is current.</p><p>The invoice lookup then includes both predicates: the requested object identifier and the authorized <code>tenant_id</code>. If <code>inv-482</code> belongs to another tenant, the row is not returned through this path. The service follows the reference policy's illustrative 404 response, reducing the difference between &#8220;missing&#8221; and &#8220;present but forbidden.&#8221; If the caller is a service or API-key principal, it must carry an explicit matching machine tenant identity; it does not inherit human ownership rules.</p><p>This query-level constraint is important because a later maintainer should not have to remember a separate authorization call after loading an unconstrained row. <a href="https://owasp.org/API-Security/editions/2023/en/0xa1-broken-object-level-authorization/">OWASP's object-level authorization guidance</a> specifically makes the object identifier a policy concern wherever the data source is accessed.</p><h3>5. Make the write retry-safe and property-safe</h3><p><code>update_invoice</code> fingerprints the validated request, scopes the idempotency key to the principal, tenant, and route, and rejects reuse with a different body. An exact completed retry can reuse the stored response under the reference policy; a pending or conflicting record receives a stable conflict response. Idempotency controls duplicate database side effects after an ambiguous network outcome. It does not revoke a bearer token or prevent general replay.</p><p>The service assigns only <code>invoice_number</code>, <code>description</code>, <code>amount</code>, and <code>currency</code>. It never copies the request body wholesale onto the ORM object, so <code>tenant_id</code>, <code>owner_id</code>, approval state, and audit fields remain server-controlled. The <a href="https://owasp.org/API-Security/editions/2023/en/0xa3-broken-object-property-level-authorization/">OWASP property-level authorization guidance</a> supports keeping writable and returnable properties explicit.</p><p>The transaction then contains the constrained update, idempotency record, and audit event according to the generated service contract. The response is serialized through <code>InvoiceRead</code>, not through an unrestricted ORM dump. A successful response therefore represents the final permitted invoice shape, not every column the database happens to contain.</p><h3>6. Read failures as part of the architecture</h3><p>A wrong-audience token fails in authentication. A missing scope fails in route policy. An inactive membership fails in tenant policy. A cross-tenant identifier fails in object resolution. A conflicting idempotency key fails before a second write. An unexpected exception becomes a generic correlated error, while operators investigate protected diagnostics.</p><p><code>examples/client.py</code> describes these scenarios without contacting an OIDC provider or the API automatically. It accepts caller-supplied credentials but does not print or persist them, and its refresh-token note keeps token exchange outside the resource server.</p><p>The project is design intent, not execution evidence: static verification, semantic verification, code-completion audit, and runtime execution were skipped. The reusable test for the architecture is therefore conceptual until deployment-specific verification exists: every request-path stage must have one owner, one invariant, one failure behavior, and one observable consequence.</p><h2>What changes when this API meets production failure?</h2><p>Production changes the question from &#8220;does the request path contain the right controls?&#8221; to &#8220;who owns each invariant when a dependency fails, a credential leaks, or the evidence is incomplete?&#8221; The reference project provides design intent for those boundaries. Its static checks, semantic checks, completion audit, compilation, migration review, and runtime execution were skipped, so it must not be described as verified, runnable, interoperable, audited, or production-certified.</p><p>The useful production artifact is not a promise that every failure is prevented. It is a recovery plan that identifies the affected invariant, the first containment action, the responsible owner, and the evidence required before service resumes.</p><h3>If the provider or signing key fails</h3><p>An OIDC discovery or JWKS outage creates a deliberate availability tradeoff. A process may continue validating tokens with a previously loaded, compatible key while that snapshot remains inside the configured stale window. Once the trust material is too old, or when a token names an unknown key that cannot be validated, the API should fail closed rather than accept unverifiable credentials. The cache policy in <code>app/security/jwks.py</code> is an implementation choice: its freshness interval, stale bound, timeout, and refresh cooldown need deployment review.</p><p>The incident owner is usually the identity-platform team working with the API team. During an outage, they should establish whether the problem is provider availability, network reachability, certificate or metadata mismatch, clock drift, or an actual key event. Operators should preserve redacted authentication-failure counts, issuer and key identifiers where safe, cache-age data, and provider notices. They should not preserve raw tokens in logs.</p><p>A suspected signing-key compromise is a different event from an ordinary rollover. Coordinate with the issuer to rotate or disable the affected key and assess which tokens could have been issued with it. Refresh or purge process-local caches according to the provider's response, review grants and client registrations, and decide whether short-lived access tokens, issuer-side revocation, introspection, or emergency traffic restriction is required. Normal key rotation does not necessarily invalidate already issued JWTs immediately; expiry and revocation behavior remain properties of the selected token architecture (<a href="https://www.rfc-editor.org/info/rfc9700/">RFC 9700</a>, <a href="https://nvlpubs.nist.gov/nistpubs/specialpublications/nist.sp.800-57pt1r5.pdf">NIST SP 800-57</a>).</p><h3>If a machine credential leaks</h3><p>For an exposed API key, revoke the key first, then issue a replacement through a protected administrative path. Inspect its use by tenant, route, time, and service identity; assess whether the same secret appeared in logs, deployment configuration, or another system; and rotate adjacent credentials if the exposure is broader than one database row. The <code>revoke_api_key</code> and <code>rotate_api_key</code> boundaries in <code>app/security/apikeys.py</code> represent lifecycle actions, not a complete incident workflow.</p><p>The API key should remain restricted to its machine route, tenant, scopes, permissions, and expiry. If the organization needs stronger attribution or sender constraint, replace the teaching path with short-lived service tokens, private-key authentication, mutual TLS, or workload identity. A static key that authenticates successfully still does not establish permission to approve an invoice or access another tenant.</p><h3>If tenant isolation is in doubt</h3><p>A report of cross-tenant access is a containment event, not an ordinary bug ticket. Constrain or stop the affected read and write paths, preserve request IDs, audit events, database evidence, deployment versions, and relevant configuration, and identify the potentially exposed tenants and objects. Review every affected query for object-only lookup, every authorization path for missing membership checks, and every serializer for excessive property exposure.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Application-level tenant predicates remain the reference path. Optional PostgreSQL row-level security can add depth, but it introduces its own operational invariants: the application must set transaction-local tenant context, pooled connections must not retain state, roles must not bypass policy unintentionally, and migrations and policy tests must be reviewed together. <code>set_transaction_tenant_context</code> in <code>app/db.py</code> demonstrates the boundary without enabling RLS or claiming that one helper establishes isolation. This database guidance is a derived design choice, not a verified guarantee.</p><p>Recovery should include credential revocation where misuse is possible, a review of audit completeness, and notification under the organization's incident process. Do not restore the affected route merely because a code change looks plausible. Require evidence that the corrected query, authorization policy, database policy, migration state, and monitoring signals agree.</p><h3>If evidence and availability controls degrade</h3><p>A local limiter is useful for illustrating a <code>429</code> decision, but its counters belong to one process. Workers and replicas can disagree, and restart resets the state. That is acceptable only when the limit is explicitly a local teaching control. Replace it with coordinated storage or upstream enforcement when the requirement is a tenant-wide quota, fleet-wide fairness, or a dependable abuse barrier. The replacement must also define behavior when the shared limiter is unavailable; failing open preserves availability but increases abuse exposure, while failing closed can deny legitimate traffic.</p><p>Audit persistence deserves the same explicit treatment. <code>record_audit_event</code> in <code>app/observability.py</code> adds a structured, redacted event to the caller's transaction; it does not provide durable transport, tamper evidence, retention, alerting, or recovery. Decide whether a security-sensitive write must fail when its audit event cannot be persisted, or whether an asynchronous path is acceptable with loss detection and replay. Keep authorization headers, cookies, API-key secrets, refresh tokens, raw JWTs, and unnecessary request bodies out of logs. Stable client errors and correlation IDs should lead operators to protected diagnostics rather than disclose SQL, provider details, or sensitive claims. These are operational design conclusions aligned with the threat model, not universal protocol requirements (<a href="https://owasp.org/API-Security/editions/2023/en/0x11-t10/">OWASP API Security Top 10</a>).</p><h3>Use the tradeoff that matches the failure you must control</h3><p>The bounded design choices are useful because each one has a reconsideration threshold. Review them explicitly rather than treating the sample's defaults as production prescriptions.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!ihfg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa32599e1-a66f-4609-a332-cf3839cf208d_1410x1144.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!ihfg!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa32599e1-a66f-4609-a332-cf3839cf208d_1410x1144.png 424w, /__u/substackcdn.com/image/fetch/$s_!ihfg!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa32599e1-a66f-4609-a332-cf3839cf208d_1410x1144.png 848w, /__u/substackcdn.com/image/fetch/$s_!ihfg!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa32599e1-a66f-4609-a332-cf3839cf208d_1410x1144.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ihfg!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa32599e1-a66f-4609-a332-cf3839cf208d_1410x1144.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!ihfg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa32599e1-a66f-4609-a332-cf3839cf208d_1410x1144.png" width="1410" height="1144" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a32599e1-a66f-4609-a332-cf3839cf208d_1410x1144.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1144,&quot;width&quot;:1410,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:299221,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/211523652?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa32599e1-a66f-4609-a332-cf3839cf208d_1410x1144.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!ihfg!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa32599e1-a66f-4609-a332-cf3839cf208d_1410x1144.png 424w, /__u/substackcdn.com/image/fetch/$s_!ihfg!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa32599e1-a66f-4609-a332-cf3839cf208d_1410x1144.png 848w, /__u/substackcdn.com/image/fetch/$s_!ihfg!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa32599e1-a66f-4609-a332-cf3839cf208d_1410x1144.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ihfg!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa32599e1-a66f-4609-a332-cf3839cf208d_1410x1144.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>These choices also clarify ownership. The identity provider owns authorization-server operations, refresh-token handling, and signing-key custody. The API owns local authorization, API-key status, database transactions, and its acceptance policy. The deployment owns proxy trust, secret injection, database roles, audit durability, rate-limit infrastructure, and incident response.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!KYBz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eaadd54-991a-4dad-8e31-1610140660e9_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!KYBz!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eaadd54-991a-4dad-8e31-1610140660e9_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!KYBz!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eaadd54-991a-4dad-8e31-1610140660e9_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!KYBz!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eaadd54-991a-4dad-8e31-1610140660e9_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KYBz!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eaadd54-991a-4dad-8e31-1610140660e9_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!KYBz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eaadd54-991a-4dad-8e31-1610140660e9_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0eaadd54-991a-4dad-8e31-1610140660e9_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:6026538,&quot;alt&quot;:&quot;Secure Web API access is defense in depth: each layer has its own invariant, failure mode, owner, and recovery action.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/211523652?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eaadd54-991a-4dad-8e31-1610140660e9_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Secure Web API access is defense in depth: each layer has its own invariant, failure mode, owner, and recovery action." title="Secure Web API access is defense in depth: each layer has its own invariant, failure mode, owner, and recovery action." srcset="/__u/substackcdn.com/image/fetch/$s_!KYBz!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eaadd54-991a-4dad-8e31-1610140660e9_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!KYBz!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eaadd54-991a-4dad-8e31-1610140660e9_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!KYBz!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eaadd54-991a-4dad-8e31-1610140660e9_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KYBz!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eaadd54-991a-4dad-8e31-1610140660e9_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Secure Web API access is defense in depth: each layer has its own invariant, failure mode, owner, and recovery action.</em></p><h3>What must change before deployment approval?</h3><p>The reference project deliberately does not include an identity provider, login system, gateway, WAF, distributed quota service, cloud deployment, complete disaster-recovery design, secret-manager integration for every platform, security audit, penetration test, or compliance certification. Before deployment, the team must supply those decisions where the threat model requires them, along with evidence from compilation, migration review, integration testing, concurrency testing, provider interoperability, security testing, and operational exercises.</p><p>A useful review asks five questions for every control:</p><ol><li><p><strong>What invariant does it enforce?</strong> For example: a token must target this issuer and audience; an invoice query must remain inside the caller's tenant; a retry must not duplicate a selected write.</p></li><li><p><strong>What does it not enforce?</strong> JWT validation does not authorize an object. CORS does not authenticate a caller. RLS does not express every business relationship. A rate limit does not grant permission.</p></li><li><p><strong>What happens when it fails?</strong> The API should reject unverifiable identity, avoid cross-tenant disclosure, bound resource consumption, and preserve a correlation-linked signal.</p></li><li><p><strong>Who owns recovery?</strong> Provider, API, database, platform, security operations, and incident response responsibilities must be named rather than assumed.</p></li><li><p><strong>What evidence permits recovery?</strong> A changed policy or rotated credential is not proof by itself; the relevant query, migration, cache state, logs, tests, and monitoring must support the decision.</p></li></ol><p>That is the production mental model: secure access is a set of independent invariants crossing explicit boundaries. Each invariant has a failure behavior, an owner, a recovery action, and evidence that the deployment actually enforces it. When those four properties remain visible, adding a framework, database policy, credential type, or infrastructure layer strengthens the design instead of hiding its limits.</p><p>Use the button or URL below to download the source code.</p>
      <p>
          <a href="/__u/onepagecode.substack.com/p/how-to-design-secure-web-api-access">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Before You Press Play: How Netflix's Open Connect CDN Solves 90% of the Streaming Problem]]></title><description><![CDATA[A deep dive into content placement, request steering, cache hits, fallback paths, and the Python model behind Netflix's delivery architecture.]]></description><link>https://onepagecode.substack.com/p/before-you-press-play-how-netflixs</link><guid isPermaLink="false">https://onepagecode.substack.com/p/before-you-press-play-how-netflixs</guid><dc:creator><![CDATA[Onepagecode]]></dc:creator><pubDate>Fri, 14 Aug 2026 09:57:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!m789!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc42e30-6b94-4b08-930b-0c1fe3bfe466_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What happens between pressing play and receiving the first segment?</h2><p>You press play on a popular title during the evening peak. The player does not need an abstract &#8220;Netflix experience&#8221;; it needs the next prepared object&#8212;say, <code>title-A/segment-17</code>&#8212;before its buffer runs down. The delivery question is therefore concrete: which location should serve that object, does the selected cache contain it, can that location serve it now, and what route remains if it cannot?</p><h4>I have written small code to explain, you can download the entire code using the URL at the end of this article!</h4><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>That sequence is the lens for this guide. It follows Open Connect as a delivery-side system: prepared content is made available at selected delivery locations, a request is directed toward an eligible location, and a cache can return the requested object when its state and capacity permit. The guide then examines misses, unavailable appliances, alternate routes, and the operational costs behind those decisions. The Python project appears later as a bounded teaching model for those mechanisms, not as the subject and not as an implementation of Netflix&#8217;s production system.</p><p>The scope matters. Open Connect should not be treated as Netflix&#8217;s entire streaming platform. Recommendation and catalog discovery may help determine what you choose to watch. Account services may determine whether you can watch it. The playback client, home network, access network, device, and the content-production pipeline also affect the result. This guide uses those systems only as boundary markers and keeps the delivery path primary.</p><p>The supplied research record contains no source index, source digest, URL, quotation, benchmark, or verification result. That means the architecture language introduced here is an explanatory framework, not a set of newly verified Netflix implementation claims. Terms such as &#8220;prepared content,&#8221; &#8220;cache,&#8221; &#8220;request steering,&#8221; and &#8220;fallback&#8221; describe the mechanisms the article will clarify; the precise Netflix boundaries and production choices require public-source verification. Where the later model makes a choice, that choice will be labeled synthetic.</p><p>The number in the title needs the same discipline. &#8220;90%&#8221; is not meaningful until you know what is being counted. It could refer to traffic bytes, requests, peak traffic, appliance-served traffic, or another denominator; it could also be limited by a particular date, geography, or deployment scope. None of those interpretations is established in the supplied record.</p><p>Even a verified statement that Open Connect delivered 90% of some defined traffic would not mean that it solved 90% of every streaming problem. A delivery share says something about where a measured class of content traveled. It does not, by itself, describe recommendations, authentication, encoding, playback control, congestion on the access network, a viewer&#8217;s home Wi-Fi, device behavior, or the quality of the underlying content pipeline.</p><p>So the useful starting question is narrower and more actionable: how does a prepared object move from delivery capacity to a viewer, and how does the system respond when the preferred nearby path is unavailable? For <code>title-A/segment-17</code>, a nearby appliance containing the object suggests a local cache hit. If that appliance is unhealthy, missing the object, or under capacity pressure, proximity alone is not enough; the request needs another eligible route. That distinction&#8212;between being nearby and being able to serve&#8212;is the delivery problem Open Connect is meant to illuminate. The 90% wording remains an unresolved quantitative claim until its metric and evidence are identified.</p><h2>Which parts of the streaming system belong to Open Connect?</h2><p>The request from the previous section gives us a useful boundary question: once the player needs <code>title-A/segment-17</code>, which system is responsible for making that object available, and which system is responsible for delivering it? Open Connect belongs on the delivery side of that line. It is not a name for every service involved in deciding what you watch, authorizing your account, or controlling the entire playback experience.</p><p>The public architecture should be described carefully here. The supplied research record does not include a source index, architecture diagram, quotation, or verification result. So the map below is explanatory scaffolding: it organizes the delivery mechanisms that this guide examines without claiming to reproduce Netflix&#8217;s current internal service decomposition.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!m789!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc42e30-6b94-4b08-930b-0c1fe3bfe466_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!m789!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc42e30-6b94-4b08-930b-0c1fe3bfe466_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!m789!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc42e30-6b94-4b08-930b-0c1fe3bfe466_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!m789!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc42e30-6b94-4b08-930b-0c1fe3bfe466_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!m789!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc42e30-6b94-4b08-930b-0c1fe3bfe466_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!m789!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc42e30-6b94-4b08-930b-0c1fe3bfe466_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/adc42e30-6b94-4b08-930b-0c1fe3bfe466_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4681749,&quot;alt&quot;:&quot;Open Connect delivery boundary showing prepared objects, distribution, ISP-adjacent appliances, the participating ISP network, viewer playback, and an alternate delivery path&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/211158047?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc42e30-6b94-4b08-930b-0c1fe3bfe466_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Open Connect delivery boundary showing prepared objects, distribution, ISP-adjacent appliances, the participating ISP network, viewer playback, and an alternate delivery path" title="Open Connect delivery boundary showing prepared objects, distribution, ISP-adjacent appliances, the participating ISP network, viewer playback, and an alternate delivery path" srcset="/__u/substackcdn.com/image/fetch/$s_!m789!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc42e30-6b94-4b08-930b-0c1fe3bfe466_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!m789!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc42e30-6b94-4b08-930b-0c1fe3bfe466_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!m789!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc42e30-6b94-4b08-930b-0c1fe3bfe466_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!m789!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadc42e30-6b94-4b08-930b-0c1fe3bfe466_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Open Connect is presented here as a delivery boundary between prepared content and the viewer, not as Netflix&#8217;s entire streaming platform.</em></p><h3>The boundary starts with prepared content</h3><p>Before a viewer requests a segment, some adjacent content systems must produce or expose a delivery object that can be distributed and served. In this guide, &#8220;prepared content&#8221; is deliberately broad. It may stand for a title representation, segment, or versioned object, but it does not assert a particular codec, packaging format, manifest structure, encryption scheme, or freshness mechanism.</p><p>Those details matter in a real streaming platform, yet they are not automatically Open Connect responsibilities. The useful boundary for this discussion is the point at which an object is ready to be distributed to delivery locations. Open Connect then concerns the work of making selected objects available near viewers and serving them when requests arrive. Exactly which upstream Netflix systems perform each preparation stage remains outside the evidence available here.</p><p>That distinction prevents a common category error. If a viewer receives a segment successfully, the result depends on more than the cache that returned it. Content production, preparation, account access, catalog discovery, playback decisions, the home network, the access network, the device, and the delivery path can all sit in the larger chain. Open Connect addresses an important portion of that chain, but it does not absorb all of it.</p><h3>Distribution and placement create the delivery state</h3><p>Once content is prepared, a distribution or placement process determines where selected objects should be available. This is the point where the system moves from a title existing somewhere in Netflix&#8217;s broader content environment to an object being present&#8212;or absent&#8212;at a particular delivery location.</p><p>The placement decision is not merely &#8220;copy everything everywhere.&#8221; Any finite cache has to make choices. A popular object may warrant copies at several locations, while a less frequently requested object may be placed selectively. More copies can increase the chance that a viewer finds a usable local copy, but they also consume storage, distribution bandwidth, and coordination effort. A forecast that misses a regional demand spike can leave the requested object absent while less useful objects occupy capacity.</p><p>The model used later will represent this state with synthetic prepared objects and appliance caches. That representation is a teaching device: it makes placement visible before a request occurs. It is not evidence about Netflix&#8217;s actual placement algorithm, replication factor, cache size, or catalog coverage.</p><h3>Appliances sit at the delivery edge of the map</h3><p>An Open Connect Appliance is treated here provisionally as a cache-and-serving node deployed within or near a participating ISP network. Its modeled responsibility is narrow and concrete: hold selected prepared objects, accept an eligible delivery request, and serve an object when the object is present and the appliance can handle the request.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The ISP-adjacent location changes the topology under consideration. Instead of assuming that every request must travel to a distant delivery location, the system can consider a serving point closer to the viewer&#8217;s access network. That locality can potentially shorten the path or reduce dependence on remote transit. It is a potential benefit, not a guarantee of latency, throughput, or playback quality: the nearby appliance may lack the object, be unavailable, or lack serving capacity.</p><p>The diagram&#8217;s labels also avoid implying a universal commercial or operational arrangement. &#8220;Participating ISP network&#8221; identifies the network context in which an appliance may be placed; it does not establish that every ISP has the same deployment, responsibilities, connectivity, or coverage. Appliance generations, regional arrangements, and operating policies may vary, and those details require dated public documentation before they can be stated as current facts.</p><h3>The viewer path crosses the boundary twice</h3><p>For <code>title-A/segment-17</code>, the viewer&#8217;s playback path eventually needs a usable delivery target. A conceptual sequence is:</p><ol><li><p>An adjacent content system exposes a prepared object.</p></li><li><p>Distribution and placement make that object available at selected appliance sites.</p></li><li><p>A request-steering process identifies an eligible delivery location.</p></li><li><p>The selected appliance checks its local cache state and serving capacity.</p></li><li><p>The object travels through the participating ISP network to the viewer when the appliance can serve it.</p></li><li><p>An alternate delivery path is considered when the preferred location cannot serve it.</p></li></ol><p>This sequence separates two different kinds of work. Placement changes the state of delivery locations before the request. Steering chooses among possible locations during the request. Cache delivery transfers the object after a location has been selected. The exact Netflix mechanisms behind those steps&#8212;such as how candidates are exposed or selected&#8212;are not established in the supplied record, so the guide will not turn this conceptual sequence into an unsupported claim about DNS, URLs, manifests, client logic, or proprietary control-plane services.</p><p>Recommendation, account services, and catalog discovery may precede the delivery request, but they are boundary markers rather than Open Connect mechanisms in this explanation. Broad playback control is similarly outside scope. The practical consequence is that a local cache hit can explain an efficient delivery leg without proving that the entire streaming session&#8212;or every source of streaming performance&#8212;has been solved by Open Connect.</p><h2>Why is the video already near the viewer before the request arrives?</h2><p>A viewer can request <code>title-A/segment-17</code>, but an appliance cannot serve an object that has never reached its storage. The important work therefore starts before playback: some upstream process must produce a cacheable delivery object, and a distribution process must decide where copies should exist before demand arrives.</p><p>Call that input <strong>prepared content</strong>. In this guide, the term deliberately stays abstract. It may stand for a title representation, a segment, or a versioned object such as <code>title-A/segment-17/v2</code>. The model does not claim which codec, packaging format, encryption arrangement, manifest structure, or freshness mechanism Netflix uses, nor does it assign those stages to Open Connect. Those are adjacent or unresolved boundaries unless public documentation establishes otherwise.</p><p>The distinction is useful because preparation and delivery solve different problems. Preparation turns source material into objects that a playback path can request. Placement turns those objects into available capacity at selected delivery locations. An appliance can then answer a request from local state rather than waiting for an object to travel there for the first time.</p><p>That last step is the key difference between planned pre-positioning and a purely reactive cache. A reactive cache waits for a miss and attempts to obtain the missing object. Pre-positioning uses demand, capacity, and locality assumptions to put selected objects in selected places ahead of time. The choice does not guarantee a hit. It creates an opportunity for a hit under the expected demand pattern.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Dw4x!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8cd9b9-5408-4185-af4a-b17b9189c708_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Dw4x!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8cd9b9-5408-4185-af4a-b17b9189c708_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!Dw4x!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8cd9b9-5408-4185-af4a-b17b9189c708_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!Dw4x!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8cd9b9-5408-4185-af4a-b17b9189c708_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Dw4x!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8cd9b9-5408-4185-af4a-b17b9189c708_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Dw4x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8cd9b9-5408-4185-af4a-b17b9189c708_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ee8cd9b9-5408-4185-af4a-b17b9189c708_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4665999,&quot;alt&quot;:&quot;Pre-positioning trades storage and distribution effort for a greater chance that a requested object is available near the viewer.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/211158047?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8cd9b9-5408-4185-af4a-b17b9189c708_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Pre-positioning trades storage and distribution effort for a greater chance that a requested object is available near the viewer." title="Pre-positioning trades storage and distribution effort for a greater chance that a requested object is available near the viewer." srcset="/__u/substackcdn.com/image/fetch/$s_!Dw4x!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8cd9b9-5408-4185-af4a-b17b9189c708_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!Dw4x!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8cd9b9-5408-4185-af4a-b17b9189c708_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!Dw4x!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8cd9b9-5408-4185-af4a-b17b9189c708_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Dw4x!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee8cd9b9-5408-4185-af4a-b17b9189c708_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Pre-positioning trades storage and distribution effort for a greater chance that a requested object is available near the viewer.</em></p><h3>Placement is a constrained decision, not universal copying</h3><p>Imagine two prepared objects. <code>title-A/segment-17/v2</code> is associated with high expected demand in three regions. <code>title-B/segment-04/v1</code> has lower expected demand and is relevant to only one region in the planning window. With enough storage, placing <code>title-A</code> at several sites creates multiple possible delivery locations. Placing <code>title-B</code> at one site conserves capacity, but viewers elsewhere are more likely to encounter a miss or an alternate route.</p><p>This is a generic capacity argument, not a claim about Netflix&#8217;s undisclosed placement algorithm. The model uses it because the tradeoff is fundamental: every additional copy consumes appliance storage and requires distribution work. More copies can improve the chance that a nearby usable copy exists, and they can provide alternatives when one location is unavailable. They also increase the amount of state that must be coordinated and, where content changes, refreshed or replaced.</p><p>Selective placement makes finite capacity visible. An appliance should not be imagined as a complete copy of the catalog. If demand is uneven, storing every object everywhere would spend capacity on content that may not be requested at that location. A demand-weighted policy instead favors objects with a stronger expected opportunity for reuse, subject to the space available at each appliance.</p><p>The decision has a failure mode: the forecast can be wrong. A low-demand object can become unexpectedly popular in a region, or a new release can create demand that the placement decision did not anticipate. The requested object then may be absent from the nearest appliance even though the appliance is healthy and connected. That is a <strong>cache miss</strong>, and it is different from an appliance outage.</p><p>The reverse error also matters. A planner can reserve space for objects whose demand never arrives. Those objects occupy capacity that could have held more useful regional content. The resulting cost is not merely wasted storage; it can lower the probability that other requests find a local copy. Placement is therefore a prediction decision with operational consequences, not a one-time file-copying step.</p><h3>Version and freshness belong in the object identity</h3><p>The object identifier also needs a version boundary. <code>title-A/segment-17/v2</code> is not interchangeable with <code>title-A/segment-17/v1</code> merely because the title and segment numbers match. A replacement, update, or invalidation can make an older cached object unsuitable for a request that expects the newer version.</p><p>The supplied record does not establish Netflix&#8217;s actual freshness, expiration, or invalidation mechanism. The safe architectural point is narrower: cache presence alone is not enough. The stored object must correspond to the requested prepared content. A model can represent that rule with explicit versions, but its version check is an educational assumption rather than evidence about production behavior.</p><p>This introduces another placement tradeoff. Retaining objects for reuse supports cache efficiency, while replacing or invalidating old versions requires additional distribution and coordination. A system that values availability must also know which copy is usable now. A system that values storage efficiency must avoid retaining state that no longer serves the delivery contract.</p><h3>What the viewer sees later</h3><p>By the time the viewer requests the first segment, several states are already possible:</p><ul><li><p>The requested object is present at a nearby eligible appliance, creating the conditions for a local cache hit.</p></li><li><p>The nearest appliance is healthy but lacks the object because placement was selective or the demand forecast was wrong.</p></li><li><p>The object exists, but only at a farther location.</p></li><li><p>The cached copy is the wrong version and must be rejected or replaced under the modeled freshness rule.</p></li></ul><p>Request steering, discussed next, operates on this pre-existing state. It cannot manufacture locality after the request arrives; it can only choose among the delivery options that placement and deployment have made available. A nearby appliance with no usable object is not equivalent to a nearby appliance ready to serve one.</p><p>The practical consequence is that content placement sets the starting conditions for every later request. Broad replication buys more local-hit and resilience opportunities at the cost of storage, distribution, and coordination. Selective placement preserves capacity and can match demand more efficiently, but it exposes the delivery path to forecast errors, version changes, and fallback. Open Connect&#8217;s delivery promise depends on managing that tradeoff before the viewer presses play.</p><h2>What does an Open Connect Appliance change about network locality?</h2><p>The previous section left <code>title-A/segment-17/v2</code> in a selected delivery location. Now consider the location itself. Why place a serving cache inside or near an ISP network instead of sending every request toward a distant site?</p><p>The provisional answer is locality. An Open Connect Appliance can be understood as a selective content cache and serving node positioned within or near a participating ISP network. It stores some prepared delivery objects and sends them toward viewers when those objects are available and the appliance can serve them. &#8220;Selective&#8221; matters: finite storage means an appliance should not be assumed to contain the entire catalog. The useful question is not whether an appliance exists nearby, but whether it has the requested object and enough usable capacity at the moment of the request.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The potential benefit follows from the topology. A nearby serving location may keep more of the delivery path within the access network or reduce dependence on a farther transit path. That can improve locality and reduce network distance in the modeled explanation. It is not a measured Netflix latency guarantee, and it does not mean that every participating ISP has the same arrangement, coverage, appliance generation, or operating responsibilities. Those details can vary by deployment, region, and time and require public documentation before they can be stated as current facts.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!wIcw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d389c5f-c6b1-4b15-bbab-4dd7dcd6a8e6_2528x1696.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!wIcw!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d389c5f-c6b1-4b15-bbab-4dd7dcd6a8e6_2528x1696.png 424w, /__u/substackcdn.com/image/fetch/$s_!wIcw!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d389c5f-c6b1-4b15-bbab-4dd7dcd6a8e6_2528x1696.png 848w, /__u/substackcdn.com/image/fetch/$s_!wIcw!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d389c5f-c6b1-4b15-bbab-4dd7dcd6a8e6_2528x1696.png 1272w, /__u/substackcdn.com/image/fetch/$s_!wIcw!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d389c5f-c6b1-4b15-bbab-4dd7dcd6a8e6_2528x1696.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!wIcw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d389c5f-c6b1-4b15-bbab-4dd7dcd6a8e6_2528x1696.png" width="1456" height="977" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9d389c5f-c6b1-4b15-bbab-4dd7dcd6a8e6_2528x1696.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:977,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4606349,&quot;alt&quot;:&quot;An ISP-adjacent appliance can offer a local path, but eligibility depends on cache state, health, and serving capacity&#8212;not proximity alone.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/211158047?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d389c5f-c6b1-4b15-bbab-4dd7dcd6a8e6_2528x1696.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="An ISP-adjacent appliance can offer a local path, but eligibility depends on cache state, health, and serving capacity&#8212;not proximity alone." title="An ISP-adjacent appliance can offer a local path, but eligibility depends on cache state, health, and serving capacity&#8212;not proximity alone." srcset="/__u/substackcdn.com/image/fetch/$s_!wIcw!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d389c5f-c6b1-4b15-bbab-4dd7dcd6a8e6_2528x1696.png 424w, /__u/substackcdn.com/image/fetch/$s_!wIcw!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d389c5f-c6b1-4b15-bbab-4dd7dcd6a8e6_2528x1696.png 848w, /__u/substackcdn.com/image/fetch/$s_!wIcw!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d389c5f-c6b1-4b15-bbab-4dd7dcd6a8e6_2528x1696.png 1272w, /__u/substackcdn.com/image/fetch/$s_!wIcw!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d389c5f-c6b1-4b15-bbab-4dd7dcd6a8e6_2528x1696.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>An ISP-adjacent appliance can offer a local path, but eligibility depends on cache state, health, and serving capacity&#8212;not proximity alone.</em></p><h3>Locality is an operating choice, not a magic shortcut</h3><p>Putting capacity near viewers changes the placement decision from &#8220;where can the object exist?&#8221; to &#8220;where is a copy worth maintaining?&#8221; Broad replication creates more opportunities for a local hit and more alternatives when one site fails. It also consumes appliance storage, distribution bandwidth, update effort, and coordination capacity. Selective placement conserves those resources, but a viewer is more likely to encounter a missing object or need a farther route.</p><p>Demand makes the choice uneven. A popular title or segment may justify copies at several sites, while a less-requested object may be placed at fewer locations. If demand prediction is wrong, the system can spend capacity on objects that viewers do not request while the requested object is absent from the nearest appliance. That is a placement error, not a failure of physical proximity.</p><p>Freshness introduces another boundary. If the viewer requests a newer version than the one stored locally, the old copy cannot be treated as the requested object merely because its bytes are nearby. The delivery system needs some version or freshness rule; the supplied record does not establish Netflix&#8217;s actual invalidation or replacement mechanism. In the bounded model, that distinction is represented explicitly, but the model&#8217;s rule is only an educational assumption.</p><h3>The nearest appliance still has to qualify</h3><p>Suppose the viewer&#8217;s access ISP has Appliance A nearby and an alternate delivery location farther away. Appliance A is the preferred candidate only if three conditions hold:</p><ul><li><p>the requested object is present;</p></li><li><p>the appliance is healthy and reachable; and</p></li><li><p>it has enough serving capacity for the request.</p></li></ul><p>If all three hold, locality can support a direct cache-delivery path. If any one fails, selecting Appliance A because it is physically closest would produce the wrong decision. A farther healthy copy may be preferable to a nearby appliance that is missing the object, unavailable, or capacity-constrained.</p><p>This is the operational cost hidden behind the phrase &#8220;near the ISP.&#8221; The deployment needs more than a cache location. It involves storage planning, distribution into that location, health and capacity awareness, network coordination, and an alternative when the local choice cannot serve. The exact division of responsibility between Netflix and an ISP is not established in the supplied record, so this guide treats those responsibilities as deployment-dependent rather than universal.</p><p>That gives the request path a concrete rule: use locality to rank usable choices, not to override cache state or service health. In the next section, request steering turns that rule into a per-request decision; here, the consequence is the architectural one&#8212;ISP-adjacent capacity improves the chance of efficient delivery, but only when the right object and enough healthy capacity are present.</p><h2>How does request steering turn a play request into a cache hit?</h2><p>The previous section put <code>title-A/segment-17/v2</code> on one or more delivery locations. Now the viewer asks for it. The request still does not translate into &#8220;send the file from the nearest appliance.&#8221; A useful serving location must satisfy several conditions at once: it must be reachable and healthy, contain the requested object version, and have enough serving capacity for the request.</p><p>That is the role this guide gives to <strong>request steering</strong>. The term describes the decision that turns a viewer request into a usable delivery target. The exact Netflix mechanism is unresolved in the supplied record: it may involve several systems or mechanisms, but the available material does not establish whether target information comes through DNS, URLs, manifests, client logic, a control-plane decision, or some combination. The sequence below is therefore a precise explanatory model, not a claim about an undisclosed production implementation.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!hR-n!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65e76277-95fb-414a-8c5a-f3b64452a082_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!hR-n!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65e76277-95fb-414a-8c5a-f3b64452a082_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!hR-n!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65e76277-95fb-414a-8c5a-f3b64452a082_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!hR-n!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65e76277-95fb-414a-8c5a-f3b64452a082_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!hR-n!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65e76277-95fb-414a-8c5a-f3b64452a082_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!hR-n!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65e76277-95fb-414a-8c5a-f3b64452a082_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/65e76277-95fb-414a-8c5a-f3b64452a082_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4674843,&quot;alt&quot;:&quot;Request steering can be understood as selecting an eligible serving location before the appliance performs the cache lookup and delivery.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/211158047?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65e76277-95fb-414a-8c5a-f3b64452a082_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Request steering can be understood as selecting an eligible serving location before the appliance performs the cache lookup and delivery." title="Request steering can be understood as selecting an eligible serving location before the appliance performs the cache lookup and delivery." srcset="/__u/substackcdn.com/image/fetch/$s_!hR-n!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65e76277-95fb-414a-8c5a-f3b64452a082_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!hR-n!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65e76277-95fb-414a-8c5a-f3b64452a082_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!hR-n!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65e76277-95fb-414a-8c5a-f3b64452a082_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!hR-n!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65e76277-95fb-414a-8c5a-f3b64452a082_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Request steering can be understood as selecting an eligible serving location before the appliance performs the cache lookup and delivery.</em></p><h3>Follow one request through the normal path</h3><p>Assume <code>viewer-1</code> is connected through <code>site-west</code> and requests <code>title-A/segment-17/v2</code>. The playback path supplies an object identity, not merely a title name. That distinction matters because a stream is consumed as a sequence of requestable objects, and the delivery system must identify the particular object needed next.</p><p>The first conceptual step is candidate discovery. The steering process obtains a set of possible delivery locations associated with the request. In a real system, the viewer does not need to know the internal map of every appliance. The important outcome is that the delivery decision has candidates to evaluate.</p><p>Next comes eligibility filtering. Consider three candidate appliances:</p><ul><li><p><code>appliance-west-1</code> is close to <code>site-west</code>, but it is unhealthy.</p></li><li><p><code>appliance-west-2</code> is also nearby, contains <code>title-A/segment-17/v2</code>, and has serving capacity.</p></li><li><p><code>appliance-east-1</code> is farther away, contains the object, and is healthy.</p></li></ul><p>The closest candidate is removed first because it cannot reliably serve the request. The second candidate remains eligible. The farther appliance is a useful alternative, but locality can now help distinguish the two healthy choices. The selected route is <code>appliance-west-2</code>, not because distance is the only rule, but because it is a nearby usable copy.</p><p>The appliance then performs the cache lookup. This is where the phrase <strong>cache hit</strong> earns its meaning: the selected appliance has the requested object in its modeled cache and can serve it under its current capacity state. The object travels from the appliance through the ISP access path to the viewer. A cache hit is therefore the combination of object presence and successful service&#8212;not simply the fact that a request reached an appliance.</p><p>The state transition is small but important. Before the request, the appliance has a cached object, a health state, and available serving capacity. During delivery, it admits the request and uses some serving capacity. After the modeled transfer, that temporary load is released. The cache contents do not need to change for a hit; the request consumes delivery capacity even when storage remains untouched.</p><p>This separates two generic concerns. A <strong>control-plane</strong> decision determines which locations are candidates and which one is preferred. A <strong>data-plane</strong> action transfers the prepared object from the selected appliance toward the viewer. Those labels organize the explanation; they are not supplied evidence that Netflix exposes systems with exactly those names.</p><h3>Why &#8220;nearest&#8221; is not enough</h3><p>Now change only the requested object to <code>title-B/segment-04/v1</code>. Suppose <code>appliance-west-2</code> is healthy and close, but that object was placed only at <code>appliance-east-1</code>. The nearby appliance is a valid network location, yet it is not an eligible cache source for this request. The steering decision must exclude it for cache state, then consider the farther appliance.</p><p>This is a cache miss at the preferred location, not necessarily a failure of the entire delivery system. If <code>appliance-east-1</code> is healthy and has capacity, the request can use that alternate modeled route. The result is different from the first case: the viewer still receives the object, but the path sacrifices some locality because selective placement did not provide a nearby copy.</p><p>A third variation exposes capacity. Suppose <code>appliance-west-2</code> contains <code>title-A/segment-17/v2</code>, but a demand spike has consumed its serving capacity. The object is present, so this is not a cache miss. It is an eligibility failure caused by current load. A farther healthy appliance may be a better choice than a nearby appliance that cannot admit another request.</p><p>These distinctions give request steering a useful order of operations:</p><ol><li><p>Discover possible delivery locations.</p></li><li><p>Remove locations that are unhealthy, lack the requested object, or cannot serve it now.</p></li><li><p>Prefer among the remaining locations using the declared locality and tie-breaking policy.</p></li><li><p>Perform the cache lookup and delivery at the selected appliance.</p></li><li><p>If no eligible appliance remains, pass the request to an explicitly defined alternate or fallback path.</p></li></ol><p>The order is a design choice in the bounded model, not a verified description of Netflix&#8217;s proprietary algorithm. Its benefit is explanatory: it prevents the common mistake of ranking by proximity first and discovering too late that the &#8220;best&#8221; appliance cannot serve the object. Its cost is state. Health, cache contents, capacity, locality, and candidate relationships must be known well enough to make a decision, and those facts can change while requests are arriving.</p><p>The practical consequence is that Open Connect&#8217;s locality benefit depends on preparation and steering working together. Placement creates usable copies; steering finds one that is both nearby and eligible. When either condition fails, the system needs an alternate route&#8212;a subject for the next section.&#951;</p><h2>What happens when the preferred appliance fails, misses, or fills up?</h2><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The normal path from the previous section ends with a usable appliance returning <code>title-A/segment-17/v2</code>. To reason about the path outside that happy case, keep the request fixed and change one condition at a time. Is the object absent? Is the appliance unavailable? Is the object present but the appliance unable to serve it now? Is the cached version current?</p><p>These are different failures. Treating them all as &#8220;the cache missed, so fetch from origin&#8221; hides the decisions that make a distributed delivery system resilient. The exact production sequence used by Netflix after an Open Connect Appliance cannot serve a request is not established in the supplied record. The alternate-appliance and abstract fallback paths below are therefore explanatory model paths, not claims about Netflix&#8217;s private implementation.</p><h3>Start with the reason the preferred route failed</h3><p>First, record the failure at the candidate that looked best. That reason determines which recovery option makes sense.</p><ul><li><p><strong>Cache miss:</strong> The appliance is reachable and may be healthy, but it does not contain the requested object or current version. This can happen when placement was selective, demand was forecast incorrectly, or a newer version has not reached that location.</p></li><li><p><strong>Unavailable appliance:</strong> The candidate cannot be used because it is unhealthy, unreachable, under maintenance, or otherwise excluded from delivery. The object might be present; its cache state does not matter if the appliance cannot serve.</p></li><li><p><strong>Capacity constraint:</strong> The object is present and the appliance is reachable, but available serving capacity is insufficient. A cache hit in storage is not automatically a successful delivery.</p></li><li><p><strong>Freshness mismatch:</strong> The appliance has an object with the wrong version. Serving stale or incorrect content is not equivalent to serving the requested current object, so the model treats the copy as unusable.</p></li></ul><p>This classification protects an important invariant: a request is counted as a successful local delivery only when the selected appliance is eligible, contains the requested current object, and can serve it. Physical or network proximity alone cannot satisfy that invariant.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!rPi0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a5e206-7aea-4e0b-89c9-e9e49ab1c84e_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!rPi0!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a5e206-7aea-4e0b-89c9-e9e49ab1c84e_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!rPi0!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a5e206-7aea-4e0b-89c9-e9e49ab1c84e_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!rPi0!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a5e206-7aea-4e0b-89c9-e9e49ab1c84e_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!rPi0!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a5e206-7aea-4e0b-89c9-e9e49ab1c84e_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!rPi0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a5e206-7aea-4e0b-89c9-e9e49ab1c84e_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/13a5e206-7aea-4e0b-89c9-e9e49ab1c84e_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3995717,&quot;alt&quot;:&quot;Fallback is best understood as a set of distinct failure branches; the production recovery sequence must not be inferred from a synthetic model.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/211158047?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a5e206-7aea-4e0b-89c9-e9e49ab1c84e_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Fallback is best understood as a set of distinct failure branches; the production recovery sequence must not be inferred from a synthetic model." title="Fallback is best understood as a set of distinct failure branches; the production recovery sequence must not be inferred from a synthetic model." srcset="/__u/substackcdn.com/image/fetch/$s_!rPi0!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a5e206-7aea-4e0b-89c9-e9e49ab1c84e_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!rPi0!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a5e206-7aea-4e0b-89c9-e9e49ab1c84e_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!rPi0!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a5e206-7aea-4e0b-89c9-e9e49ab1c84e_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!rPi0!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a5e206-7aea-4e0b-89c9-e9e49ab1c84e_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Fallback is best understood as a set of distinct failure branches; the production recovery sequence must not be inferred from a synthetic model.</em></p><h3>Try another eligible location before using an abstract fallback</h3><p>In the bounded model, recovery begins by reconsidering the candidate set. A nearby appliance missing <code>title-B/segment-04/v1</code> does not end the request if a farther appliance has the current object, is healthy, and has serving capacity. The model records the first candidate&#8217;s cache-state exclusion, then selects the alternate route.</p><p>That recovery preserves delivery at a cost. The alternate appliance may have a higher synthetic locality cost, consume capacity that was reserved for another region, or traverse a less favorable network path. Redundancy improves the chance that a request can be served, but it requires additional copies, storage, distribution work, and health information. Placement and fallback are therefore coupled: a system with more strategically placed copies has more alternatives, while a system with fewer copies saves capacity but exposes more requests to misses or distant delivery.</p><p>Now change the failure. Suppose the farther appliance contains the object but is unavailable. The model excludes it as an appliance failure rather than mislabeling the event as a cache miss. If another candidate remains eligible, the request continues there. If no appliance qualifies, the model can use an explicitly named <strong>abstract fallback route</strong> with a configured cost, or return a bounded terminal failure when fallback is disabled or exhausted.</p><p>That bound matters operationally. A recovery mechanism must not retry the same unusable candidates indefinitely. Each attempt should record the candidate, the exclusion reason, the selected route, and the outcome. Once the candidate set or retry budget is exhausted, the system needs a visible terminal result rather than an unobservable loop.</p><h3>Separate capacity pressure from missing content</h3><p>Capacity pressure deserves its own branch because it can appear during a demand spike even when placement was correct. A popular object may be present on the nearest appliance, but concurrent requests can consume its available serving capacity. The request then needs another eligible copy or an alternate route.</p><p>This is where &#8220;nearby&#8221; and &#8220;available&#8221; diverge. Routing every request to the closest location would make locality the only policy and could concentrate demand on one appliance. Filtering for health, object presence, and capacity before preferring locality is a safer explanatory policy, although the supplied record does not establish that Netflix uses these exact signals or ordering.</p><p>The tradeoff is visible in both placement and recovery. Broad replication can reduce the impact of a saturated appliance, but it consumes storage and distribution capacity even when demand does not materialize. Selective placement conserves those resources, but an unexpected release, regional demand shift, or stale forecast can turn a local request into a miss and increase fallback use.</p><p>Freshness adds another form of coordination. If <code>title-B/segment-04/v1</code> has been replaced by <code>v2</code>, an old copy should not be treated as a valid hit merely because its identifier is present in storage. The model can reject the old version and seek the current one; the actual Netflix invalidation, replacement, or versioning process remains outside the available evidence.</p><p>The practical decision is to make every failure branch explicit: distinguish absence, health, capacity, and freshness; prefer a usable alternate when one exists; and stop after bounded recovery. That reasoning explains why Open Connect&#8217;s delivery benefit depends not only on where content is placed, but also on how many eligible alternatives remain when the preferred appliance cannot serve.</p><h2>How do the model&#8217;s objects, placement rules, and steering decisions work?</h2><p>The previous sections described the delivery path in system terms. Now we can make its state visible without pretending to recreate Netflix Open Connect. The model uses three questions that matter for one request: what prepared object is being requested, where has that object been placed, and which candidate can serve it now?</p><p>The distinction is important. These classes and policies are teaching evidence from the generated project, not evidence about Netflix&#8217;s production APIs, appliance specifications, steering algorithm, protocols, or fleet. The model uses synthetic topology, demand, capacities, and locality costs. Its value is traceability: you can inspect why an object was placed and why a candidate was accepted or excluded.</p><h3>Start with the object, cache, and request state</h3><p>The model represents preparation only as an immutable <code>ContentObject</code>. That is a deliberate boundary. <code>title_id</code>, <code>segment_id</code>, and <code>version</code> identify the logical media object; <code>size_bytes</code> lets placement account for finite storage. The class does not encode a codec, manifest, encryption scheme, or Netflix-specific object format.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d315323b-9f9a-47f5-8b04-cb3d7f3e5ca7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True, slots=True)
class ContentObject:
    """A prepared, cacheable object used by the synthetic model."""

    object_id: str
    title_id: str
    segment_id: str
    version: str
    size_bytes: int

    def __post_init__(self) -&gt; None:
        for name in ("object_id", "title_id", "segment_id", "version"):
            _require_identifier(getattr(self, name), name)
        if isinstance(self.size_bytes, bool) or not isinstance(self.size_bytes, int):
            raise TypeError("size_bytes must be an integer")
        if self.size_bytes &lt; 0:
            raise ValueError("size_bytes must be non-negative")</code></pre></div><p>The invariant is straightforward: every object has stable identifiers and a non-negative size. That gives the placement layer something concrete to store while keeping upstream content preparation outside the simulation.</p><p>An <code>Appliance</code> then holds a mutable set of object IDs, storage accounting, health, and synthetic serving load. Its <code>cached_objects</code> set answers &#8220;is this object present?&#8221;; <code>used_bytes</code> answers &#8220;can another copy fit?&#8221;; <code>healthy</code> and <code>active_load</code> answer &#8220;can this node serve now?&#8221; Those are separate states because locality alone does not make a candidate usable.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2835b9e8-410e-4573-aef6-044a34fe043a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(slots=True)
class Appliance:
    """A mutable synthetic cache and serving node at an ISP-adjacent site."""

    appliance_id: str
    site_id: str
    capacity_bytes: int
    serving_capacity: int
    cached_objects: set[str] = field(default_factory=set)
    healthy: bool = True
    active_load: int = 0
    used_bytes: int = 0</code></pre></div><p>The model validates that <code>used_bytes</code> cannot exceed <code>capacity_bytes</code> and that <code>active_load</code> cannot exceed <code>serving_capacity</code>. Those checks do not describe an Open Connect Appliance specification. They preserve the model&#8217;s own invariant: a synthetic cache cannot claim storage or serving capacity that its configuration does not provide.</p><p>A <code>DeliveryRequest</code> supplies the other side of the lookup: request ID, viewer, viewer site, and object ID. In a worked example, <code>viewer-1</code> at <code>site-west</code> might request <code>title-A/segment-17/v2</code>. The request does not contain a chosen appliance. That decision belongs to steering, which allows the same request to be evaluated against different placement and health states.</p><h3>Let placement turn demand into cache state</h3><p><code>PlacementPlanner.place_content</code> models work that happens before the request. It receives prepared objects, a demand profile, appliances, and a replication limit. It first clears the model&#8217;s previous cache state, ranks objects by descending declared demand, and tries to store each object on eligible appliances until it reaches the replication factor or runs out of capacity.</p><p>The important state change is the call to <code>appliance.store(...)</code>: an object ID enters <code>cached_objects</code>, and its size increases <code>used_bytes</code>. If no appliance can accept the object, the planner records it in <code>rejected_objects</code>. If only one of three requested replicas fits, the partial placement remains visible in the report. That exposes the tradeoff rather than hiding it behind a binary success flag.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;925d01a8-f04a-44e2-bf10-e4bc4e4f6855&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        ranked_objects = sorted(content, key=lambda item: (-demand[item.object_id], item.object_id))
        placements: dict[str, list[str]] = {item.object_id: [] for item in content}
        rejected: list[str] = []
        bytes_placed = 0

        for item in ranked_objects:
            for appliance in self._ordered_appliances(item, appliance_list):
                if len(placements[item.object_id]) &gt;= factor:
                    break
                if appliance.can_store(item.size_bytes) and appliance.store(item.object_id, item.size_bytes):
                    placements[item.object_id].append(appliance.appliance_id)
                    bytes_placed += item.size_bytes
            if not placements[item.object_id]:
                rejected.append(item.object_id)</code></pre></div><p>Suppose <code>title-A/segment-17/v2</code> has a high demand weight and <code>title-B/segment-04/v1</code> has a low one. With three appliances and a replication factor of two, the popular object gets first access to available capacity. The low-demand object may receive fewer copies or none if the larger object consumes the space. Increasing replication improves the chance that a viewer finds a usable local copy, but consumes more storage and distribution work. Reducing it preserves capacity while exposing more requests to misses or alternate routes.</p><p>This is a model policy, not a claim that Netflix ranks objects or chooses replicas exactly this way. The decision to keep the policy deterministic is educational: identical inputs produce inspectable placement, so a reader can compare a selective policy with broader replication without confusing randomness for architecture.</p><h3>Let steering explain every exclusion</h3><p>After placement, <code>RequestSteerer.choose_appliance</code> evaluates one request. It checks the prepared object catalog, gathers candidate appliances, and filters them in a fixed order. An unhealthy appliance is excluded as <code>unhealthy</code>; one without the object is excluded as <code>cache_miss</code>; one at its serving limit is excluded as <code>serving_capacity_exhausted</code>.</p><p>Only after those checks does the model rank eligible candidates by synthetic locality cost, load penalty, active load, and appliance ID. The selected appliance is therefore the lowest-cost eligible candidate, not necessarily the physically nearest one. If no candidate survives, the steerer returns a structured <code>fallback</code> or <code>no_route</code> decision instead of raising an uncontrolled exception.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b362b7b3-0210-41bd-a9df-1a81f68a9072&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        for appliance in normalized:
            if not appliance.healthy:
                exclusions[appliance.appliance_id] = "unhealthy"
                continue
            if not appliance.contains(request.object_id):
                exclusions[appliance.appliance_id] = "cache_miss"
                continue
            if appliance.available_serving_capacity &lt;= 0:
                exclusions[appliance.appliance_id] = "serving_capacity_exhausted"
                continue

            locality = self._locality_cost(request.viewer_site, appliance.appliance_id)
            score = locality + self.capacity_penalty * appliance.active_load
            eligible.append(
                ((locality, score, appliance.active_load, appliance.appliance_id), appliance)
            )</code></pre></div><p>For <code>viewer-1</code> requesting the popular object, imagine that the nearest appliance is unhealthy, the next candidate lacks the object, and a farther appliance is healthy and has serving capacity. The trace records all three decisions and selects the farther appliance. Change only the requested object ID, and the same topology may produce a cache miss. Change only <code>active_load</code>, and the object may still be present but no longer eligible.</p><p>That separation is the mechanism worth carrying forward: placement determines where modeled content exists; steering determines which existing copy is usable for this request. The actual Netflix request-steering mechanism&#8212;whether it involves URLs, DNS, manifests, client behavior, control-plane decisions, or a combination&#8212;remains unresolved in the supplied record. This code intentionally does not fill that gap with a production claim.</p><p>The same caution applies to the <code>fallback</code> result. It says that the model found no eligible appliance and permits a later simulator layer to try an alternate or abstract route. It does not establish what Netflix does after a real appliance miss or failure.</p><p>The practical consequence is a disciplined one: use the model to ask &#8220;which state or rule caused this route?&#8221; Use public documentation&#8212;not synthetic placement, steering, or metrics&#8212;to answer &#8220;does Netflix implement this exact mechanism?&#8221; Static verification, semantic verification, and execution were skipped for the generated project, so this section reports no runtime output or measured result.</p><h2>What does the bounded Python model let you see&#8212;and what can it not prove?</h2><p>The system explanation now gives you the important sequence: prepared content is placed before demand arrives, request steering evaluates possible delivery locations, and an appliance can serve only when its cache state and serving capacity allow it. The model turns that sequence into inspectable state. It does not turn an explanatory abstraction into Netflix evidence.</p><p>That distinction is the design constraint for this section. The generated project uses synthetic objects, sites, capacities, locality costs, health states, and fallback rules. It is useful for asking, &#8220;Why did this request take this route?&#8221; It cannot answer, &#8220;What does Netflix&#8217;s production system do in every region?&#8221; or establish the meaning of the title&#8217;s 90% claim.</p><h3>Follow one request through the simulator</h3><p>The central method is <code>DeliverySimulator.deliver</code>. Its input is one <code>DeliveryRequest</code>; its state includes the appliance set, a <code>RequestSteerer</code>, and a catalog of synthetic <code>ContentObject</code> instances. Its output is a <code>DeliveryResult</code> containing the selected route, cache-hit status, fallback status, estimated latency, delivered bytes, and a trace.</p><p>The important behavior is the separation between steering and the final cache check. Steering selects a candidate using the model&#8217;s declared rules, but the simulator checks the appliance again before delivery. That represents a useful distributed-systems concern: state can change between choosing a candidate and attempting to serve from it.</p><p>Here is the focused recovery path from <code>open_connect_model/simulation.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;3868eeac-4d05-4a74-90fa-4f880e4a011d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        if decision.selected_appliance_id is not None:
            appliance = self.appliances[decision.selected_appliance_id]
            # Recheck mutable state after steering. This represents a possible
            # state change between candidate selection and the data-plane lookup.
            contains = appliance.contains(request.object_id)
            capacity_available = appliance.available_serving_capacity &gt; 0
            trace.append(
                self._event(
                    "cache_lookup",
                    "checked selected appliance state",
                    appliance_id=appliance.appliance_id,
                    cache_contains=contains,
                    healthy=appliance.healthy,
                    available_serving_capacity=appliance.available_serving_capacity,
                )
            )
            if appliance.healthy and contains and capacity_available:
                locality = self.steerer._locality_cost(
                    request.viewer_site, appliance.appliance_id
                )
                alternate = bool(decision.exclusions)
                outcome = "alternate_appliance_hit" if alternate else "cache_hit"</code></pre></div><p>Three conditions must hold before this synthetic appliance delivers the object: <code>healthy</code> must be true, <code>contains</code> must be true, and <code>capacity_available</code> must be true. A nearby appliance that fails any one of those checks is not a successful local delivery. That is the model&#8217;s concrete version of the earlier architectural point: proximity is useful, but it is not sufficient.</p><p>The trace records each condition rather than returning only a final label. That makes a cache miss different from an unavailable appliance, and both different from a capacity constraint. The <code>alternate</code> flag also preserves an important distinction: an appliance may deliver the object successfully while still being an alternate candidate because another candidate was excluded earlier.</p><p>The invariant is straightforward: a result marked as an appliance hit must have an appliance route, <code>cache_hit=True</code>, and delivered bytes equal to the modeled object&#8217;s size. The code also increments <code>active_load</code> only around the modeled delivery and releases it in a <code>finally</code> block. That protects the synthetic serving-capacity state if the delivery block changes later. It is a teaching model of state accounting, not a claim about an Open Connect Appliance&#8217;s internal concurrency model.</p><h3>Make fallback explicit instead of inventing Netflix&#8217;s path</h3><p>When no selected appliance can serve, the generated simulator uses a configured abstract fallback route if that option is enabled. The word &#8220;abstract&#8221; is doing real work: the supplied research record does not establish whether a production miss uses another appliance, another delivery location, an upstream interaction, client redirection, or some combination.</p><p>The fallback branch is deliberately bounded:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;90f9dbc9-66b7-4a10-9014-28fd912915bf&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        if self.allow_abstract_fallback:
            trace.append(
                self._event(
                    "fallback",
                    "used configured abstract fallback route",
                    latency_estimate=self.fallback_latency,
                    delivered_bytes=content.size_bytes,
                )
            )
            return DeliveryResult(
                request_id=request.request_id,
                outcome="fallback",
                selected_appliance_id=None,
                route_kind="fallback",
                cache_hit=False,
                fallback_used=True,
                latency_estimate=self.fallback_latency,
                delivered_bytes=content.size_bytes,
                failure_reason="",
                trace=tuple(trace),
            )</code></pre></div><p>The result says that the object was delivered through the model&#8217;s fallback route, not that Netflix uses this exact route. If fallback is disabled, the simulator returns a terminal <code>failed</code> result rather than retrying indefinitely. That lets you inspect fallback exhaustion as a separate operational case: redundancy improves availability, but every alternative requires capacity, state, and coordination.</p><p>The same mechanism covers several declared scenario cases in <code>config/demo.json</code>: a local hit, a selective-placement miss, an unavailable preferred appliance, capacity pressure, a version mismatch, and an object unavailable at all modeled appliances. These labels define teaching inputs. They do not establish that Netflix uses the same placement policy, freshness rule, retry budget, or fallback sequence.</p><h3>Read metrics as summaries of traces, not as benchmarks</h3><p>After individual requests are processed, <code>DeliverySimulator.run</code> returns both the results and aggregate metrics. <code>compute_metrics</code> derives those metrics exclusively from the returned <code>DeliveryResult</code> values:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;32138296-e60b-434d-8668-91dc2b8b7245&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def compute_metrics(results: Sequence[DeliveryResult]) -&gt; Metrics:
    """Compute zero-safe synthetic metrics from returned delivery results only."""
    if isinstance(results, (str, bytes)):
        raise TypeError("results must be an iterable of DeliveryResult objects")
    normalized = list(results)
    if any(not isinstance(result, DeliveryResult) for result in normalized):
        raise TypeError("results must contain DeliveryResult objects")
    total = len(normalized)
    hits = sum(result.cache_hit for result in normalized)
    misses = total - hits
    fallbacks = sum(result.fallback_used for result in normalized)
    failures = sum(result.outcome == "failed" for result in normalized)</code></pre></div><p>This design makes the denominator visible: <code>total_requests</code> is the number of synthetic requests supplied to the model. <code>hit_ratio</code>, <code>fallback_rate</code>, and <code>failure_rate</code> therefore describe this scenario and these rules only. They are not measurements of Netflix traffic, appliance-served bytes, peak demand, or any other possible interpretation of &#8220;90%.&#8221;</p><p>The demo entry point compares selective and broader replication over the same declared request stream. That comparison can clarify a tradeoff: more copies may make an object available at more candidate locations, while selective placement conserves modeled storage. But because execution was disabled in the supplied records, no printed metrics, traces, or policy comparison should be reported as observed results. Static checks, semantic verification, and code execution were all marked as skipped.</p><p>The practical consequence is a clean boundary. Use the simulator to inspect causal relationships&#8212;placement affects cache state, cache state affects eligibility, eligibility affects route choice, and route choice affects modeled fallback. Do not use it to infer Netflix&#8217;s architecture, production performance, or the unresolved 90% figure.</p><h2>What does &#8220;90%&#8221; mean&#8212;and what does Open Connect leave unsolved?</h2><p>Return to the request for <code>title-A/segment-17</code>. In the best case, prepared content is already available on an eligible appliance near the viewer. Request steering selects a usable copy, the cache returns the object, and the access path carries it to the player. That is the delivery problem Open Connect is designed to clarify: how prepared content becomes available at useful locations, and how a request reaches one of them under locality, cache-state, capacity, and availability constraints.</p><p>Now change one condition. If the nearby appliance lacks the object, the request faces a miss. If the appliance is unavailable, a farther healthy location may be preferable. If the object is present but serving capacity is exhausted, proximity alone does not make it a usable route. The model&#8217;s fallback branches make those distinctions visible, but none of them establishes a Netflix-wide traffic percentage.</p><p>The &#8220;90%&#8221; in the title therefore needs a precise definition before it can be treated as a metric. The available research record does not identify its numerator or denominator. It does not say whether the number refers to bytes, requests, peak traffic, appliance-served traffic, sessions, or another quantity. It also does not provide the geography, time period, measurement method, or scope. Without those fields, the number remains an unresolved editorial claim&#8212;not evidence that Open Connect solves 90% of every cause of streaming performance.</p><p>The bounded Python model cannot fill that gap. Its topology, placement policy, locality costs, capacities, request stream, fallback rules, and metrics are synthetic. A model hit ratio can show how a broader replication policy changes outcomes inside that model; it cannot measure Netflix, validate Open Connect&#8217;s production behavior, or define the title&#8217;s percentage. The supplied records also show that static checks, semantic verification, and execution were skipped, so no generated output should be presented as a verified result.</p><p>What Open Connect leaves outside this guide&#8217;s boundary matters. It does not, by itself, stand for recommendation, account services, catalog discovery, upstream content production, every playback-control decision, the viewer&#8217;s home network, the access network&#8217;s complete condition, or device behavior. Those factors can affect the experience without changing the delivery mechanism described here.</p><p>The reusable mental model is narrower and more useful than an undefined percentage: prepared content creates something deliverable; placement creates distributed capacity; steering chooses an eligible locality; cache state determines whether that locality can answer; capacity and health determine whether it can answer now; fallback preserves options when it cannot. Evaluate each stage separately, and &#8220;90%&#8221; becomes a question to define rather than a number to repeat.</p><p>Use the button or URL below to download the source code.</p>
      <p>
          <a href="/__u/onepagecode.substack.com/p/before-you-press-play-how-netflixs">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Forecasting Stock Prices with LSTMs, SVMs & AI Models (Complete Python Guide)]]></title><description><![CDATA[A systematic review and implementation blueprint analyzing deep learning, support vector machines, and multimodal financial data.]]></description><link>https://onepagecode.substack.com/p/forecasting-stock-prices-with-lstms</link><guid isPermaLink="false">https://onepagecode.substack.com/p/forecasting-stock-prices-with-lstms</guid><dc:creator><![CDATA[Onepagecode]]></dc:creator><pubDate>Thu, 13 Aug 2026 20:00:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_W6n!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3078fdad-8e80-40a9-be17-640eb18fa447_1254x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This 2024 review article synthesizes ten systematic reviews of AI-based stock-market prediction. It examines which prediction methods, informational sources, and evaluation metrics are most common, using a PRISMA-guided search of Scopus and Web of Science. The synthesis emphasizes SVM, LSTM, and neural networks; historical closing-price time series and technical indicators; and metrics such as accuracy, MSE, RMSE, MAPE, and MAE. It recommends combining numerical, textual, sentiment, financial, macroeconomic, spatial, and temporal information, while recognizing trade-offs involving data requirements, complexity, interpretability, robustness, and computational cost. The article does not define or train a new predictive architecture, and therefore cannot directly yield a faithful end-to-end model implementation without additional implementation decisions outside the paper.</p><h3>Download The Source Code Using the URL At the End of this article!</h3><p>Research paper: https://www.sciencedirect.com/science/article/pii/S2590291124000615?ref=pdf_download&amp;fr=RR-2&amp;rr=a2a5647f6ff1e2c9</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>Implementation Assumptions</h2><ul><li><p>The implementation is a literature-review synthesis package, not a stock-prediction model implementation.</p></li><li><p>Bibliographic search results and full texts are supplied locally or manually; no live Scopus, Web of Science, Rayyan, Yahoo Finance, or other external API is required.</p></li><li><p>The ten retained reviews are represented as review metadata and evidence records; they are not treated as primary stock-price datasets.</p></li><li><p>Missing included-study counts and unresolved Table 5 alignments remain explicit None or uncertainty values.</p></li><li><p>Metric names are cataloged but no formulas are implemented because the supplied paper context contains no canonical equations.</p></li><li><p>No prevalence is recomputed across incompatible denominators or review subsets.</p></li><li><p>Any future predictive reproduction must supply task, target, horizon, split, preprocessing, alignment, architecture, hyperparameters, optimization, and evaluation decisions externally.</p></li><li><p>Python 3.11+ standard-library dataclasses, enums, pathlib, argparse, typing, datetime, and json are sufficient for the planned local package.</p></li><li><p>Synthetic records are used only for demonstration when original bibliographic exports are unavailable.</p></li><li><p>The target code-file average is approximately 400 lines, with focused modules allowed to be shorter where the domain responsibility is narrow; artificial padding is not planned.</p></li></ul><h2>Scope and the paper-faithful implementation boundary</h2><p>What should be reproduced when a paper reviews other reviews instead of proposing a trainable model? The practical answer is not an invented LSTM, SVM, or multimodal network. For this paper, reproduction means preserving the review-selection process, the extracted evidence, the uncertainty around that evidence, and the boundaries of what the source does not specify.</p><h3>Start with the paper&#8217;s actual object of study</h3><p><strong>Paper fact.</strong> The source is a systematic review of systematic reviews about artificial-intelligence methods for stock-market price or return prediction. Its research questions ask three separate questions: which prediction methods are commonly reported, which information sources are commonly used, and which evaluation metrics are reported. The paper searches Scopus and Web of Science, screens candidate systematic reviews, retains ten reviews, and synthesizes their findings.</p><p>That distinction matters. A review-level statement that <code>LSTM</code> is frequently discussed is evidence about the reviewed literature. It is not a specification of an LSTM&#8217;s input shape, sequence length, hidden-state size, loss, optimizer, or training schedule. Similarly, a reported mention of <code>RMSE</code> tells us that a metric was used in some reviewed studies; it does not supply a formula that this package should reconstruct.</p><p><strong>Derived implementation explanation.</strong> The generated package therefore has four responsibilities:</p><ol><li><p><code>prisma_review_screening</code> represents the search, screening, duplicate handling, and final-eligibility workflow.</p></li><li><p><code>review_evidence_extraction</code> converts review-level observations into records with provenance and uncertainty.</p></li><li><p><code>review_evidence_synthesis</code> answers the three review questions descriptively while preserving each source subset and denominator.</p></li><li><p><code>predictive_reproduction_contract</code> identifies the external decisions required before any separate stock-prediction implementation could begin.</p></li></ol><p>These responsibilities produce catalogs, screening logs, claims, warnings, and contract reports&#8212;not model weights or predictions.</p><h3>Keep reported counts source-qualified</h3><p>The paper reports 40 records from Scopus and 29 from Web of Science. It also reports a screening flow of 69 retrieved titles, 43 reviews read, 17 remaining after abstract screening, 16 after duplicate or same-author handling, and 10 final reviews. These numbers describe different database sources or workflow stages, so the implementation stores them separately.</p><p>The generated constants make that distinction visible:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7795c30e-8bc9-4d14-82af-30e0bf4b5bab&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">PAPER_ID: Final[str] = "1-s2.0-S2590291124000615-main"
"""Identifier of the supplied paper."""


REPORTED_REVIEW_COUNT: Final[int] = 10
"""Number of systematic reviews retained in the reported final corpus."""


REPORTED_RETRIEVAL_COUNTS: Final[Mapping[str, int]] = MappingProxyType(
    {
        "Scopus": 40,
        "Web of Science": 29,
    }
)</code></pre></div><p><code>PAPER_ID</code> identifies the source artifact. <code>REPORTED_REVIEW_COUNT</code> records the ten retained reviews. <code>REPORTED_RETRIEVAL_COUNTS</code> preserves database provenance rather than silently treating 40 and 29 as interchangeable observations.</p><p>The same principle applies to percentages and study counts. The ten retained reviews collectively cover more than 379 primary studies according to the abstract, while other findings use subsets such as 12, 30, 34, 45, 57, or 122 studies. Those denominators cannot be combined into one prevalence estimate. The code&#8217;s evidence and claim objects therefore keep the originating subset and denominator attached to each observation.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;500bc51e-ed4b-4881-aaed-e11cce677d51&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">REPORTED_FLOW_COUNTS: Final[Mapping[str, int]] = MappingProxyType(
    {
        "retrieved": 69,
        "reviews_read": 43,
        "abstract_remaining": 17,
        "post_duplicate": 16,
        "final_retained": 10,
    }
)</code></pre></div><p>These are workflow facts, not predictive-study counts. A local reproduction may calculate its own <code>FlowCounts</code>, but comparison with the paper&#8217;s values should report differences rather than force the local records to match them.</p><h3>Use local records instead of pretending to query databases</h3><p><strong>Implementation decision.</strong> The package does not expose live Scopus or Web of Science credentials, APIs, or network behavior. Bibliographic records are supplied externally, for example through a local export or manually assembled input. This is consistent with the paper&#8217;s protocol without inventing an API contract that the paper never defines.</p><p>The central orchestration function makes this boundary explicit:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;dc6597b3-606a-4ebf-b055-adcd024aa6c1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def run_review_pipeline(
    config: SearchConfiguration,
    candidates: Sequence[CandidateRecord],
    batches: Sequence[ExtractionBatch],
    decisions: Sequence[tuple[str, ScreeningDecision]],
) -&gt; PipelineResult:</code></pre></div><p><code>config</code> describes the paper-reported search protocol. <code>candidates</code> contains already available bibliographic records, retaining fields such as title, abstract, database, and retrieval metadata. <code>batches</code> contains review-level extraction inputs. <code>decisions</code> supplies explicit screening decisions, including their stage and reason. The returned <code>PipelineResult</code> links the catalog, screening log, flow counts, synthesis report, and missing-specification report.</p><p>The function validates candidate identities and decision references before calling <code>prisma_review_screening</code>. It then codes extraction batches, builds a <code>ReviewCatalog</code>, and invokes <code>synthesize_review</code>. Importantly, this sequence never calls a forward pass, loss function, optimizer, or metric implementation.</p><p>The package exceptions describe this boundary as contract validation:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;36570eaa-d686-41a5-a804-06a186390344&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class ReviewSynthesisError(Exception):
    """Base exception for package-level contract violations."""


class ValidationError(ReviewSynthesisError):
    """Raised when an input value or workflow invariant is invalid."""


class IncompleteSpecificationError(ReviewSynthesisError):
    """Raised when externally required predictive details are missing."""</code></pre></div><p><code>ValidationError</code> represents invalid local data, such as a missing title or a decision referring to an unknown candidate. <code>IncompleteSpecificationError</code> is different: it means that a requested predictive reproduction lacks information that the paper does not supply. Neither exception represents failed model training, because the package does not define a predictive-model runtime.</p><h3>Worked example: a review report without a stock forecaster</h3><p>Suppose two bibliographic records arrive from different databases. The local caller supplies title and abstract metadata, then records title/abstract and final-eligibility decisions. The caller can request the workflow through the package API:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;22222902-44d6-4207-aebd-a99b4b38bc5c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">config = default_paper_search_config()
result = run_review_pipeline(config, candidates, batches, decisions)</code></pre></div><p>In this example, <code>candidates</code> are local <code>CandidateRecord</code> objects, <code>batches</code> contain observations such as reported methods or metrics, and <code>decisions</code> are <code>ScreeningDecision</code> objects. The result can be inspected through <code>result.screening_log</code>, <code>result.flow_counts</code>, and <code>result.synthesis_report</code>. A screening log answers questions such as which stage changed a record&#8217;s status and why. The synthesis report answers the methods, information-source, and metric questions using the supplied coded evidence.</p><p>This is intentionally not equivalent to training a model. The paper names methods including SVM, SVR, LSTM, RNN, ANN, CNN, ARIMA, ANFIS, and hybrid approaches, but it does not unify them into one executable interface. It also mentions tools such as Python, TensorFlow, Pandas, NumPy, Keras, Scikit-Learn, MATLAB, TA-Lib, and TA4J without prescribing a required dependency stack or API. The generated implementation consequently treats those names as literature evidence or possible external tooling, not as mandatory runtime components.</p><h3>Make missing predictive details impossible to overlook</h3><p>A separate caller may still want to build a stock predictor based on one of the reviewed method families. That is a different implementation task. The package provides <code>run_predictive_contract_only</code> as a gate around that task:</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d91f3069-1e42-44a2-80f3-44c3cf7688dd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def run_predictive_contract_only(config: PredictiveConfiguration) -&gt; dict[str, object]:
    """Validate an external predictive specification without training anything.

    This implements the predictive-reproduction contract as a configuration
    gate.  Missing fields are returned explicitly; no target, horizon,
    preprocessing rule, architecture, optimizer, or metric formula is chosen
    by this function.
    """</code></pre></div><p>The external <code>PredictiveConfiguration</code> must state decisions such as classification versus regression, target definition, forecast horizon, input modalities, sequence or lookback construction, split policy, scaling policy, model family, hyperparameters, optimizer, loss, and evaluation metrics. Multimodal inputs also need explicit temporal alignment and leakage-prevention rules.</p><p><strong>Implementation decision.</strong> If a configuration omits the target or horizon, the package must report that omission rather than select a default. It must not assume that daily data, an approximately 1000-day period, a particular normalization policy, or an LSTM architecture is universal. An incomplete configuration leads conceptually to an <code>IncompleteSpecificationError</code> and a missing-specification report; it does not silently become a runnable predictor.</p><p>The local command-line interface follows the same boundary. Documentation may show an invocation such as:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;71f947f0-7fb8-44bd-922b-9038bead5fc4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">python -m review_synthesis.cli --example --output report.json</code></pre></div><p>The <code>main</code> function supports deterministic examples or local JSON input and can render JSON, Markdown, or text. The example records are synthetic fixtures for demonstrating the data flow, not the paper&#8217;s original database export and not newly measured predictive results.</p><h3>Boundary statement</h3><p>The paper-faithful reproduction is the provenance-preserving review evidence workflow: search configuration, screening records, coded observations, source-qualified claims, and uncertainty warnings. Any end-to-end stock-prediction model requires an external specification for its target, horizon, preprocessing, architecture, optimization, and evaluation procedure. Those details cannot be recovered from this review without inventing content beyond the supplied source.</p><h2>Domain records, provenance, and uncertainty</h2><p>How can a local Python package keep a review finding attached to the place where it was observed? Treat every item as an evidence envelope. A candidate review has an identity, database source, and retrieval metadata. An extracted observation has a dimension, wording, confidence, and paper location. A quantitative claim also retains its denominator and review subset. This design prevents a label such as <code>SVM</code>, <code>RMSE</code>, or <code>historical closing prices</code> from becoming detached from its source.</p><h3>Separate controlled labels from free text</h3><p><strong>Paper fact.</strong> The paper discusses several method families, information sources, and metrics, but it does not define one executable predictor. The implementation therefore needs labels for cataloging evidence, not tensor types or model layers.</p><p>The generated <code>enums.py</code> module provides those labels. <code>Database</code> identifies whether a candidate came from Scopus or Web of Science. <code>ScreeningStage</code> distinguishes retrieval, title-and-abstract screening, duplicate handling, same-author handling, full-text review, and final eligibility. <code>Decision</code> records inclusion, exclusion, or a pending state. <code>EvidenceDimension</code> separates methods, information sources, metrics, tools, and datasets. <code>Confidence</code> distinguishes reported, uncertain, and missing evidence.</p><p><code>TaskType</code> also appears in the module, with classification and regression values, but it belongs to an external predictive specification. It does not define a target, label rule, or forecast horizon for this paper.</p><p>The enum values are string-compatible, which gives serialized records stable values such as <code>"scopus"</code>, <code>"method"</code>, and <code>"uncertain"</code>. The important implementation decision is that these labels do not replace the source wording. They organize it while leaving room for ambiguity.</p><h3>Provenance is data, not a comment</h3><p>In this workflow, <strong>provenance</strong> means the information needed to answer &#8220;where did this value come from?&#8221; <code>SourceLocation</code> identifies a location in the supplied paper context. Its <code>paper_location</code> can identify a section, table, or figure, while <code>section_id</code> and <code>table_id</code> preserve more specific identifiers when available.</p><p><code>Provenance</code> carries a different kind of lineage. Its <code>database</code> and <code>retrieval_date</code> describe bibliographic acquisition; its <code>source_location</code> and <code>original_text</code> describe the paper evidence itself. These concerns are deliberately separate because a coded observation may have a paper location without having come from a live database query.</p><p>The following excerpt is the generated definition of the paper-location record:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f0587cde-f0e6-4f96-8027-0ad2ee8ec29a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class SourceLocation:
    """Identify where an extracted fact appears in the supplied paper context.

    ``paper_location`` is intentionally a free-form label because the source
    includes section identifiers, table identifiers, and figure identifiers.
    ``section_id`` and ``table_id`` preserve the more specific identifiers when
    they are available.
    """

    paper_location: str
    section_id: Optional[str] = None
    table_id: Optional[str] = None</code></pre></div><p>The dataclass is frozen so callers cannot casually mutate the location after an evidence object has captured it. Its fields are textual identifiers, not coordinates inferred from the damaged PDF extraction. The generated <code>__post_init__</code> rejects an empty <code>paper_location</code> and rejects empty optional identifiers when they are supplied.</p><p><code>require_provenance</code> is the explicit boundary check for extracted evidence. It requires a valid <code>Provenance</code> object and a non-empty <code>source_location</code>; it does not invent a location or repair Table 5 alignment. This is important because the supplied extraction says that Table 5 method, source, and metric associations are fragmented and not fully reliable row by row.</p><h3>Candidate records describe literature, not market data</h3><p><code>CandidateRecord</code> represents one bibliographic candidate received from an external export. Its required fields are <code>title</code> and <code>database</code>. It can also retain a stable <code>record_id</code>, abstract, publication metadata, authors, query text, and arbitrary metadata. The <code>provenance</code> field preserves retrieval lineage.</p><p>This is the central distinction: a <code>CandidateRecord</code> is not a stock-price sample. It does not contain a time axis of observations, feature columns, labels, or model inputs. It is an envelope around a publication record that will later receive screening decisions.</p><p>A focused construction using the generated public classes looks like this:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c2029315-cb85-475e-94e6-4928ddb33495&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from datetime import date

from review_synthesis.enums import Database
from review_synthesis.provenance import Provenance, SourceLocation
from review_synthesis.records import CandidateRecord

location = SourceLocation(
    paper_location="Table 5",
    section_id="1-s2.0-S2590291124000615-main_sec_016",
    table_id="Table 5",
)
provenance = Provenance(
    database=Database.SCOPUS,
    retrieval_date=date(2022, 4, 11),
    source_location=location,
)
record = CandidateRecord(
    record_id="candidate-1",
    title="Illustrative systematic review",
    database=Database.SCOPUS,
    provenance=provenance,
)</code></pre></div><p>Here, <code>Database.SCOPUS</code> is the controlled source label, while the date is a Python <code>datetime.date</code>. The example is a local illustrative record, not one of the paper&#8217;s original database exports. The <code>SourceLocation</code> points to the supplied paper context, whereas the provenance&#8217;s database and retrieval date describe how the candidate record was acquired.</p><p><code>CandidateRecord</code> derives an identifier only when <code>record_id</code> is omitted. It does not collapse records from different databases, and it does not decide whether a title is eligible. Those are separate workflow operations.</p><h3>Screening decisions preserve state and reasons</h3><p>A <code>ScreeningDecision</code> records one decision at one stage. Its <code>stage</code> might be <code>ScreeningStage.TITLE_ABSTRACT</code> or <code>ScreeningStage.FINAL_ELIGIBILITY</code>; its <code>decision</code> might be <code>Decision.INCLUDE</code>, <code>Decision.EXCLUDE</code>, or <code>Decision.PENDING</code>. Every decision has a non-empty <code>reason</code>, and it may also record a reviewer and timestamp.</p><p>For example, a local screening log could receive this decision:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;13ff09b6-5be1-4e81-84df-e4cb1409a565&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from review_synthesis.enums import Decision, ScreeningStage
from review_synthesis.records import ScreeningDecision

decision = ScreeningDecision(
    stage=ScreeningStage.TITLE_ABSTRACT,
    decision=Decision.INCLUDE,
    reason="title and abstract identify a systematic review of AI stock prediction",
    reviewer="reviewer-1",
)</code></pre></div><p>The reason is not decorative. It makes an exclusion auditable and distinguishes &#8220;excluded because it concerns portfolio optimization&#8221; from &#8220;pending because the abstract is unavailable.&#8221; The generated model permits a pending decision, but still requires a reason explaining why the record remains unresolved.</p><p><code>ReviewRecord</code> serves a later boundary: it stores metadata for a retained systematic review and its reported included-study count. The count is positive when known and <code>None</code> when the supplied paper leaves it unspecified. <code>None</code> therefore means &#8220;not supplied,&#8221; not zero and not an estimate. The record remains review metadata, not a primary stock-prediction dataset.</p><h3>Coded evidence keeps wording, coding, and uncertainty together</h3><p><code>CodedEvidence</code> represents one review-level observation. Its <code>dimension</code> says what kind of observation it is. <code>canonical_label</code> is a controlled label used for grouping, while <code>original_text</code> preserves the wording that was actually extracted. <code>confidence</code> records whether the association is reported, uncertain, or missing. <code>source_review_id</code> links the observation to its retained review, and <code>denominator</code> preserves a reported positive study count when one exists.</p><p>The distinction between the two text fields matters. A controlled vocabulary might map &#8220;support vector machine&#8221; to <code>SVM</code>, but the source wording must remain available for audit. For a fragmented Table 5 association, the canonical label might be <code>SVM</code> while the confidence is <code>Confidence.UNCERTAIN</code>.</p><p>A worked example for the planned evidence item is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ab55d7e2-68cb-4d6a-b543-8432a1216118&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from review_synthesis.enums import Confidence, EvidenceDimension
from review_synthesis.records import CodedEvidence

svm_evidence = CodedEvidence(
    dimension=EvidenceDimension.METHOD,
    canonical_label="SVM",
    original_text="SVM",
    confidence=Confidence.UNCERTAIN,
    source_review_id="review-1",
    provenance=provenance,
)</code></pre></div><p>This item says only that the review-level coding recorded <code>SVM</code> at the supplied location with uncertain confidence. It does not say that the current package trained an SVM, that SVM was best, or that the Table 5 row alignment is certain.</p><p><code>EvidenceClaim</code> is the appropriate object for a broader qualitative or quantitative statement. In addition to its claim text and optional source review, it stores <code>subset_id</code>, <code>denominator</code>, <code>provenance</code>, and <code>confidence</code>. The <code>subset_id</code> is essential because the paper reports findings over different review subsets. A claim about 12 studies must not be silently combined with a claim about 57, 122, or more than 379 studies.</p><h3>Validate at module boundaries</h3><p>The generated <code>validation.py</code> exposes four focused functions:</p><ul><li><p><code>validate_candidate(record)</code> checks candidate identity, title, database, provenance type, optional publication fields, authors, and metadata.</p></li><li><p><code>validate_screening_decision(decision)</code> checks the stage, decision, reason, reviewer, and timestamp.</p></li><li><p><code>validate_review_record(record)</code> checks review identity, optional metadata, and the included-study count.</p></li><li><p><code>validate_evidence(item)</code> checks the evidence dimension, labels, confidence, source review, denominator, and provenance.</p></li></ul><p>These functions return <code>None</code> for valid input and raise <code>ValidationError</code> for invalid structure or values. Their <code>None</code> handling is deliberate. Optional fields such as an abstract, publication date, included-study count, and denominator may be absent because the source context does not provide them. Validation rejects malformed values, but it does not manufacture missing information.</p><p><code>validate_evidence</code> goes one step further by calling <code>require_provenance</code>. Consequently, a coded evidence item without a <code>SourceLocation</code> is rejected rather than assigned a guessed section or table. This is the implementation expression of the paper&#8217;s uncertainty around extracted table alignment.</p><p>The generated code was not executed or verified under the authoritative run policy. The records and validators described here are implementation artifacts whose intended responsibility is to make provenance, uncertainty, and missing values explicit. The boundary remains firm: these objects model literature evidence and review workflow state, not numerical tensors or a trained stock-prediction model.</p><h2>Configurable PRISMA-style retrieval and screening</h2><p>How do you reproduce a review-selection process when the original databases and full-text decisions are not available inside the Python package? Separate acquisition from screening. Bibliographic records arrive from an external export or manual process; the local package then validates, screens, deduplicates, logs, and reports them. This keeps the workflow auditable without pretending that the package performed a live database query.</p><h3>Start with an explicit search protocol</h3><p><strong>Paper fact.</strong> The paper reports a PRISMA-guided search of Scopus and Web of Science. Its search was title-based, restricted to articles, used variants of &#8220;Systematic Review,&#8221; &#8220;Systematic Literature Review,&#8221; and &#8220;Stock,&#8221; and had a final search date of April 11, 2022. The reported publication window begins in 2009.</p><p>The generated <code>SearchConfiguration</code> records these choices. Calling <code>default_paper_search_config()</code> creates a configuration with separate <code>Database.SCOPUS</code> and <code>Database.WEB_OF_SCIENCE</code> values, <code>title_only=True</code>, <code>article_only=True</code>, the three search terms, and an <code>InclusionCriteria</code> object. That criteria object includes the English-language requirement, publication window, AI-primary-use requirement, stock-market scope, excluded facets such as volatility and portfolio optimization, and a configurable geographic policy.</p><p>The geographic policy is deliberately not hard-coded as an unquestionable filter. The supplied paper mentions studies involving markets such as Taiwan, NASDAQ, Dow Jones, and European markets while also stating a US, UK, and Europe criterion. Because the source does not reconcile those statements, <code>geography_policy</code> remains configurable. This is an implementation decision that exposes an ambiguity instead of silently resolving it.</p><p>The default construction is local configuration only:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;86b5747c-d77d-45d8-a5ce-2e2cf15e5142&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    return SearchConfiguration(
        databases=frozenset({Database.SCOPUS, Database.WEB_OF_SCIENCE}),
        title_only=True,
        article_only=True,
        terms=["Systematic Review", "Systematic Literature Review", "Stock"],
        final_search_date=date(2022, 4, 11),
        criteria=criteria,
    )</code></pre></div><p><code>validate_search_config(config)</code> checks that the dates, databases, terms, flags, and criteria are structurally valid. It does not contact Scopus or Web of Science, and it does not supply credentials, query syntax, or API behavior. The paper names those databases as sources, but it does not provide a required software API for accessing them.</p><h3>Preserve raw retrieval provenance</h3><p>Records obtained from an external database export are converted with <code>make_candidate(record_id, title, database, retrieval_date, abstract=None, metadata=None)</code>. The resulting <code>CandidateRecord</code> retains the source database, retrieval date, title, optional abstract, query metadata, and a stable record identifier. The function validates the required identity fields and does not reinterpret a record as a stock-price dataset.</p><p>The next operation is intentionally conservative. <code>merge_raw_records(records)</code> combines records without deduplicating them:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d30d71f6-c772-4131-9167-593efd081c70&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def merge_raw_records(records: Sequence[CandidateRecord]) -&gt; list[CandidateRecord]:
    """Combine raw records without deduplication or metadata rewriting.

    The returned list preserves input order, including duplicate publications
    returned by different databases.  Each record retains its original
    database and provenance; duplicate handling belongs to a later workflow
    stage.
    """</code></pre></div><p>This separation matters because the paper reports 40 Scopus records and 29 Web of Science records, for a reported total of 69 retrieved records. <code>reported_retrieval_counts(records)</code> returns source-specific counts and a separately named <code>raw_total</code>; it does not turn records with the same title into one record. Raw retrieval counts, deduplicated counts, and retained-review counts therefore remain different quantities.</p><h3>Screen with explicit include, exclude, and pending states</h3><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>PRISMA is a reporting framework for systematic-review selection. Here, screening means applying eligibility rules in stages and recording the result. The package uses <code>ScreeningStage</code> values for retrieval, title-and-abstract screening, duplicate handling, same-author handling, full-text assessment, and final eligibility. Each <code>ScreeningDecision</code> contains a stage, a <code>Decision</code> value, a reason, and optional reviewer information.</p><p><code>screen_title_abstract(record, criteria)</code> applies the available title, abstract, and explicit metadata. <code>screen_full_text(record, full_text, criteria)</code> applies the same kind of conservative checks to supplied full text. <code>exclusion_reason(record, stage, criteria)</code> exposes the reason-finding part of this process without necessarily producing a decision.</p><p>The important failure case is missing evidence. An absent abstract is not treated as proof that a record is irrelevant. Missing required metadata can produce <code>Decision.PENDING</code>, rather than an invented inclusion or exclusion. Likewise, an empty full-text string produces a pending full-text decision. This behavior is an implementation policy designed to avoid converting unavailable information into negative evidence; it is not a claim that the paper supplied these exact keyword rules.</p><p>The generated screening functions use transparent local checks for review terminology, AI-related terminology, stock-market scope, publication metadata, and configured excluded facets. They also avoid inferring geography from a market name. A caller can therefore replace or supplement these rules with manually reviewed decisions while retaining the same record and log contracts.</p><h3>Keep duplicate and same-author handling separate</h3><p>Duplicate removal and same-author exclusion are related but distinct workflow events. <code>deduplicate_candidates(records)</code> creates a normalized comparison key from the title, retains the first-seen candidate, and returns explicit <code>ScreeningDecision</code> objects at the <code>DUPLICATE</code> stage for later matches. The retained record keeps its original database provenance.</p><p><code>apply_same_author_policy(records, policy)</code> does not guess which same-author publication should be excluded. The supplied paper reports such a stage but does not define a reproducible automatic selection rule. The generated function accepts explicit policies such as <code>"manual"</code> or <code>"none"</code>; an actual exclusion must be supplied separately as a logged decision. This prevents an undocumented heuristic from being mistaken for a paper fact.</p><p>Every decision is stored through <code>ScreeningLog.append(record_id, decision)</code>. The log preserves the order of decisions, allows multiple stages for the same record, and rejects contradictory final decisions at one stage. For example, a pending title-and-abstract decision may later be replaced by an explicit decision, but two different non-pending decisions at that same stage are treated as a contract violation.</p><h3>Worked example: two local candidates</h3><p>Suppose a local export contains one candidate from each database. The records can be created with the exact public constructor described above:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2ac88fb5-25fa-49b4-a0d4-f4d1dfa1bd32&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from datetime import date

from review_synthesis.enums import Database
from review_synthesis.search_records import make_candidate

scopus_record = make_candidate(
    "scopus-1",
    "Systematic Review of AI Stock Prediction",
    Database.SCOPUS,
    date(2022, 4, 11),
)
wos_record = make_candidate(
    "wos-1",
    "Systematic Review of AI Stock Prediction",
    Database.WEB_OF_SCIENCE,
    date(2022, 4, 11),
)</code></pre></div><p>These two records have identical displayed titles but remain separate raw records because they came from different databases. Passing them to <code>merge_raw_records</code> preserves both. Passing them to <code>deduplicate_candidates</code> identifies the second title as a duplicate and returns a <code>DUPLICATE</code> exclusion decision for it. The decision can then be appended to a <code>ScreeningLog</code> with <code>log.append("wos-1", decision)</code>.</p><p>For a complete local workflow, the caller supplies title-and-abstract and final-eligibility decisions, then invokes <code>prisma_review_screening(config, records, decisions)</code>. The function validates the search configuration and records, logs each supplied decision, computes stage counts, and converts candidates with explicit unblocked final inclusion into <code>ReviewRecord</code> metadata. It does not query either database or decide unavailable full-text eligibility.</p><p>After the call, <code>screening_log.decisions_for("scopus-1")</code> returns the ordered audit trail for that record. <code>screening_status("scopus-1", screening_log)</code> reports the decision at the furthest logged workflow stage. <code>retain_eligible_records(records, screening_log)</code> requires an explicit <code>FINAL_ELIGIBILITY</code> inclusion and rejects any candidate with an earlier exclusion, including duplicate or same-author exclusion.</p><h3>Interpret flow counts without forcing reconciliation</h3><p><code>FlowCounts</code> stores five stage-specific values: <code>retrieved</code>, <code>reviews_read</code>, <code>abstract_remaining</code>, <code>post_duplicate</code>, and <code>final_retained</code>. <code>compute_flow_counts(records, log)</code> calculates these values only from records and explicit logged decisions. It does not fill missing stages or infer that a pending record survived screening.</p><p><strong>Paper fact.</strong> The reported flow is 69 retrieved titles, 43 reviews read, 17 remaining after abstract-based criteria, 16 after duplicate or same-author handling, and 10 final reviews. These values describe the paper&#8217;s review workflow, not a stock-prediction training set.</p><p>A local demonstration with two records will not reproduce those counts, and it should not be padded or adjusted to do so. Instead, construct a separate <code>FlowCounts</code> value for the reported flow and compare it with the locally calculated value using <code>compare_with_reported_flow(actual, reported)</code>. The result contains stage-by-stage differences; it does not overwrite local records or merge incompatible counts.</p><p>The same rule applies to study totals and percentages. Counts such as 12, 30, 34, 45, 57, 122, and more than 379 belong to different review subsets or scopes. A screening implementation must retain their source context rather than calculate one combined prevalence estimate.</p><h3>Boundary of this workflow</h3><p>This package implements a local, provenance-preserving screening boundary. Search acquisition, unavailable full-text judgments, and any manual resolution of ambiguous records remain external inputs. The generated code also does not train SVM, LSTM, ANN, CNN, ARIMA, or another predictor. That limitation is intentional: the paper reports a synthesis of existing studies, not a complete predictive architecture or reproducible training procedure.</p><p>The practical result is an auditable path from externally supplied records to retained review metadata and stage-specific flow counts. It reproduces the review-selection responsibilities supported by the source while leaving unresolved evidence visibly unresolved.</p><h2>Evidence extraction and uncertainty-aware coding</h2><p>How should a program record that a systematic review mentioned <code>LSTM</code>, <code>historical closing prices</code>, or <code>RMSE</code> without turning that mention into a complete predictive model? Treat extraction as transcription with a chain of custody. Each observation keeps its original wording, evidence category, source review, paper location, denominator when available, and confidence. The package therefore catalogs literature evidence; it does not infer a model architecture, calculate a metric, or reconstruct a damaged table row.</p><h3>Code review-level fields, not predictive tensors</h3><p><strong>Paper fact.</strong> The paper reports method categories including SVM, SVR, LSTM, RNN, ANN, CNN, DNN, MLP, GA-SVR, ANFIS, ARIMA, PCA, clustering, text mining, sentiment analysis, technical analysis, ensemble methods, and related approaches. It also reports information sources such as historical prices, technical indicators, financial or macroeconomic variables, news, Twitter, sentiment indices, and other markets. These are review findings, not a unified feature matrix.</p><p>The generated <code>ExtractionInput</code> represents one extracted field for one retained review. Its <code>review_id</code> links the field to the review record; <code>field</code> selects one of <code>methods</code>, <code>sources</code>, <code>metrics</code>, <code>tools</code>, or <code>datasets</code>; <code>values</code> is a sequence of reported labels; <code>source_location</code> identifies where the extraction came from; <code>denominator</code> preserves a reported study subset size; and <code>confidence</code> distinguishes reported, uncertain, and missing information. An empty <code>values</code> sequence is valid only when confidence is <code>Confidence.MISSING</code>, so absence is explicit rather than silently converted into an empty finding.</p><p><code>ExtractionBatch</code> groups these inputs for one review. Its <code>methods</code>, <code>sources</code>, <code>metrics</code>, <code>tools</code>, and <code>datasets</code> collections are dimension-specific. During construction, each item must belong to the batch's <code>review_id</code> and have the expected <code>EvidenceDimension</code>. This is an important invariant: a metric extracted for one review cannot accidentally be attached to another review merely because two records happened to be processed next to each other.</p><p>The following excerpt shows the core shape contract. It is deliberately a record of labels and provenance, not a numerical array or model input tensor.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6fd0ec16-c8c5-4868-b341-72b4db9cf095&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class ExtractionInput:
    """One provenance-qualified, review-level extraction field.

    ``values`` contains labels reported for a review field. It does not encode
    row alignment among methods, inputs, or metrics from a fragmented table.
    """

    review_id: str
    field: str
    values: tuple[str, ...] = ()
    source_location: SourceLocation = dc_field(
        default=None  # type: ignore[assignment]
    )
    original_text: Optional[str] = None
    denominator: Optional[int] = None
    confidence: Confidence = Confidence.REPORTED</code></pre></div><p>Here, <code>values</code> is a tuple of strings, not a time-series tensor. <code>source_location</code> is mandatory for an extraction, while <code>original_text</code> can preserve a larger source phrase. A positive integer <code>denominator</code> is copied as metadata; it is not used to manufacture a percentage. <code>validate_extraction_input</code> checks these requirements and raises <code>ValidationError</code> for an invalid structure.</p><h3>Use conservative, dimension-specific vocabulary</h3><p><code>VocabularyEntry</code> maps a canonical label to explicitly configured aliases within one <code>EvidenceDimension</code>. For example, the vocabulary can map &#8220;support vector machine&#8221; to <code>SVM</code> in the method dimension, or &#8220;mean absolute error&#8221; to <code>MAE</code> in the metric dimension. The mapping is intentionally exact after only whitespace and case normalization. An unknown term remains unchanged and receives <code>Confidence.UNCERTAIN</code>.</p><p>This dimension boundary matters for ambiguous labels. A notation such as <code>R</code> should not be interpreted as a metric, method, or information source unless that interpretation has been explicitly configured. Likewise, the extracted abbreviation <code>CPI</code> and the phrase &#8220;Customer Pricing Index&#8221; are not silently corrected. The supplied paper context flags that wording as potentially erroneous, but does not authorize the implementation to replace it with &#8220;Consumer Price Index.&#8221;</p><p>The default vocabulary contains conservative paper-mentioned labels for method families, sources, metric names, tools, and named datasets. It preserves the distinction between similar neural-network labels such as <code>NN</code> and <code>ANN</code> unless an explicit alias says otherwise. This prevents a convenient normalization rule from making a stronger claim than the source supports.</p><h3>Delegate field extraction to provenance-aware coding</h3><p>The public extractors are intentionally narrow. <code>extract_methods</code> selects only the method inputs, <code>extract_information_sources</code> selects only source inputs, and <code>extract_metrics</code> selects only metric inputs. They all delegate to <code>code_extraction_batch</code>, which applies vocabulary mapping, validates each input, copies the review identifier and denominator, and calls <code>mark_fragmented_table_evidence</code>.</p><p>The following excerpt is the generated metric extractor. Notice that it returns coded evidence and explicitly documents that formulas are not reconstructed.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;950194a8-eb2b-4803-bc8d-21ab8763636c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def extract_metrics(
    batch: ExtractionBatch,
    vocabulary: Sequence[VocabularyEntry],
) -&gt; list[CodedEvidence]:
    """Extract metric names only, without reconstructing or evaluating formulas."""
    return _extract_dimension(batch, EvidenceDimension.METRIC, vocabulary)</code></pre></div><p>The return value is a list of <code>CodedEvidence</code> objects. Each object contains a <code>canonical_label</code>, the exact <code>original_text</code>, a <code>source_review_id</code>, a <code>confidence</code>, optional <code>denominator</code>, and <code>provenance</code>. Thus <code>RMSE</code> is cataloged as a metric name only. No formula for RMSE, MSE, MAE, MAPE, NMSE, correlation coefficient <code>R</code>, accuracy, precision, recall, F1-score, or POCID is supplied or inferred here.</p><p><code>extract_tools_and_datasets</code> returns both tool and dataset evidence, but the individual objects retain their distinct dimensions. This allows <code>Python</code>, <code>TensorFlow</code>, <code>Yahoo Finance</code>, <code>TAIEX</code>, <code>NASDAQ</code>, and <code>Dow Jones</code> to remain separate kinds of reported evidence rather than being treated as interchangeable inputs or dependencies.</p><h3>Mark fragmented Table 5 associations as uncertain</h3><p><strong>Paper fact.</strong> The extracted Table 5 content is fragmented by PDF layout. Method, information-source, and metric entries do not have reliably recoverable row alignment in the supplied context. A label may therefore be reported in the same table area as another label without proving that the two belonged to the same review row or primary study.</p><p>The implementation responds by preserving the association but lowering its confidence when its provenance points to Table 5. <code>mark_fragmented_table_evidence</code> does not change the label, denominator, review identifier, or source location. It changes only a <code>Confidence.REPORTED</code> value to <code>Confidence.UNCERTAIN</code> for affected evidence. Existing uncertainty is never upgraded.</p><p>This is safer than inferring alignment from extraction order. For example, if one uncertain Table 5 extraction contains <code>SVM</code>, <code>historical closing prices</code>, and <code>RMSE</code>, the resulting records can retain all three observations while making clear that the source does not establish a method&#8211;input&#8211;metric relationship. The package does not create a study-level row from those three strings.</p><h3>Worked example: one uncertain extraction batch</h3><p>Suppose an external transcription creates an <code>ExtractionBatch</code> for a retained review with three <code>ExtractionInput</code> items: <code>SVM</code> in the <code>methods</code> field, <code>historical closing prices</code> in <code>sources</code>, and <code>RMSE</code> in <code>metrics</code>. Each item uses a <code>SourceLocation</code> whose <code>paper_location</code> or <code>table_id</code> identifies Table 5, and each item carries <code>Confidence.UNCERTAIN</code> because the table alignment is unresolved.</p><p>The intended local flow is:</p><ol><li><p>Construct the three <code>ExtractionInput</code> records with the same review identifier.</p></li><li><p>Group them into one <code>ExtractionBatch</code>.</p></li><li><p>Obtain the configured list from <code>default_vocabulary()</code>.</p></li><li><p>Pass the batch to <code>code_extraction_batch</code>, or call <code>extract_methods</code>, <code>extract_information_sources</code>, and <code>extract_metrics</code> separately.</p></li><li><p>Inspect each resulting <code>CodedEvidence</code> object.</p></li></ol><p>The method evidence may have canonical label <code>SVM</code> and original text <code>SVM</code>; the source evidence may have canonical label <code>historical closing prices</code>; and the metric evidence may have canonical label <code>RMSE</code>. All three retain the same originating review only if the input batch declared that review, and all three retain the Table 5 provenance. Their confidence remains uncertain. The code does not claim that SVM used those prices or that RMSE evaluated that method.</p><p>For a term that cannot be normalized, <code>preserve_unresolved_term</code> provides the same conservative behavior. A call with the extracted CPI wording produces a <code>CodedEvidence</code> whose canonical label is still the original phrase and whose confidence is <code>Confidence.UNCERTAIN</code>. It does not silently rewrite the paper's terminology:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4521c6b2-c84f-415d-b8ce-12b739594899&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def preserve_unresolved_term(
    original_text: str,
    dimension: EvidenceDimension,
) -&gt; CodedEvidence:
    """Create uncertain evidence without normalizing an unresolved term."""
    original = _clean_text(original_text, "original_text")
    if not isinstance(dimension, EvidenceDimension):
        raise ValidationError("dimension must be an EvidenceDimension value")
    return CodedEvidence(
        dimension=dimension,
        canonical_label=original,
        original_text=original,
        confidence=Confidence.UNCERTAIN,
        source_review_id="unresolved",
        provenance=Provenance(original_text=original),
    )</code></pre></div><p>The <code>source_review_id="unresolved"</code> value signals that this helper is for an unresolved term rather than a fully attributed review extraction. For normal extraction, <code>code_extraction_batch</code> preserves the actual review identifier and requires source provenance through <code>validate_extraction_input</code> and <code>require_provenance</code>.</p><h3>Preserve the reported ten-review corpus</h3><p><code>reported_review_corpus</code> creates metadata records for the ten retained systematic reviews described in Table 4. These records are not stock-price datasets. They are review-level objects used to connect extracted methods, sources, metrics, tools, datasets, and claims.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/p/forecasting-stock-prices-with-lstms?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/p/forecasting-stock-prices-with-lstms?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/onepagecode.substack.com/p/forecasting-stock-prices-with-lstms?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><p>The reported included-study sequence is <code>27</code>, unspecified, <code>53</code>, <code>12</code>, <code>34</code>, <code>122</code>, <code>24</code>, <code>57</code>, <code>30</code>, and <code>20</code>. In Python, the unspecified entry remains <code>None</code>; it is not estimated, replaced with zero, or inferred from another review. <code>validate_review_corpus</code> checks the ten-record fixture and this missingness invariant. The sequence should therefore be understood as metadata attached to ten different reviews, not as ten compatible samples that can be summed or averaged without an explicit research design.</p><h3>Keep qualitative findings separate from measured results</h3><p>The paper also discusses limitations and future recommendations, including multimodal data integration, sentiment and text use, hyperparameter analysis, interpretability, robustness, scalability, individual-company prediction, and cross-market evaluation. <code>QualitativeFinding</code> stores such text with its category, source review, source location, and confidence.</p><p><code>collect_limitations</code> and <code>collect_recommendations</code> use conservative textual cues to organize already supplied claims. <code>group_qualitative_findings</code> groups them by category without ranking them. These functions do not turn a recommendation such as &#8220;combine textual and numerical inputs&#8221; into an implemented multimodal architecture, and they do not turn a reported limitation into a measured performance result.</p><p>The practical boundary is therefore clear: extraction can tell us that the reviewed literature mentions SVM, LSTM, historical prices, technical indicators, or RMSE, and it can preserve uncertainty around those observations. It cannot supply the missing target definition, forecast horizon, preprocessing policy, sequence construction, model layers, loss, optimizer, or metric equation. Faithful reproduction at this stage means reproducing the evidence-coding workflow, not inventing the predictive experiments that the paper itself does not specify.</p><h2>Source-qualified synthesis for RQ1, RQ2, and RQ3</h2><p>How can the implementation answer &#8220;which methods, inputs, and metrics appear in the literature?&#8221; without turning a review of reviews into a misleading benchmark? The key is to perform descriptive aggregation while preserving provenance. The package records which retained review reported <code>SVM</code>, <code>LSTM</code>, historical prices, technical indicators, accuracy, MSE, RMSE, MAPE, or MAE. It does not calculate a new performance score, rank models statistically, or claim that one method wins.</p><h3>Three review questions become three evidence dimensions</h3><p><strong>Paper fact.</strong> The paper organizes its synthesis around three questions: commonly reported prediction methods, commonly reported information sources, and commonly reported performance metrics. In the implementation, these become the <code>EvidenceDimension.METHOD</code>, <code>EvidenceDimension.INFORMATION_SOURCE</code>, and <code>EvidenceDimension.METRIC</code> categories.</p><p><strong>Implementation decision.</strong> <code>ReviewCatalog</code> is the central object. It links retained <code>ReviewRecord</code> objects with <code>CodedEvidence</code>, source-qualified <code>EvidenceClaim</code> objects, and explicit warnings. The catalog is a literature evidence store, not a stock-price dataset. Its records contain no model weights, predictions, tensors, or training state.</p><p>The catalog offers <code>evidence_for_review(review_id, dimension)</code> for focused inspection. Passing only a review identifier returns all coded observations for that review; passing an <code>EvidenceDimension</code> narrows the result. <code>add_warning</code> records an interpretation problem such as unresolved Table 5 alignment. <code>validate</code> checks cross-record references and refuses to infer missing identities or repair fragmented evidence.</p><p>A useful mental model is a filing cabinet. <code>ReviewCatalog.reviews</code> contains the folders, <code>evidence</code> contains labeled observations, and <code>claims</code> contains statements that may have denominators or subset identifiers. The cabinet can be summarized, but its contents should not be detached from their labels and source locations.</p><h3>Descriptive summaries are not prevalence estimates</h3><p><code>DimensionSummary</code> is deliberately modest. It stores canonical labels, the review identifiers in which each label appears, denominator metadata, and uncertainty counts. It does not calculate prevalence, effect sizes, rankings, confidence intervals, or model performance.</p><p>The implementation of <code>summarize_dimension</code> makes that boundary explicit:</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/p/forecasting-stock-prices-with-lstms?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/p/forecasting-stock-prices-with-lstms?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/onepagecode.substack.com/p/forecasting-stock-prices-with-lstms?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a62c4938-1da8-4d31-b3b0-1c3318f7ff86&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def summarize_dimension(
    catalog: ReviewCatalog, dimension: EvidenceDimension
) -&gt; DimensionSummary:
    """Group coded evidence by label without making prevalence claims.

    The output preserves review identifiers and denominator values separately
    for every canonical label.  In particular, observations from different
    review subsets are not merged into a single count or percentage.
    """</code></pre></div><p>For example, <code>summarize_dimension(catalog, EvidenceDimension.METHOD)</code> groups method evidence for RQ1. The corresponding call with <code>EvidenceDimension.INFORMATION_SOURCE</code> answers RQ2, and the call with <code>EvidenceDimension.METRIC</code> answers RQ3. The returned labels might include <code>SVM</code>, <code>LSTM</code>, or <code>ANN</code> for methods; historical prices or technical indicators for information sources; and accuracy, MSE, RMSE, MAPE, or MAE for metrics. Their presence means that the coded reviews reported those categories. It does not mean that this package evaluated them.</p><p><code>sum&#8203;marize_by_review</code> serves a different purpose. Rather than grouping globally by label, it returns a mapping from each review identifier to the labels observed in that review for one dimension. This is useful when the question is &#8220;which review reported this category?&#8221; It remains a review-level presence map: it does not count primary studies or combine denominators.</p><h3>Keep claims and denominators attached</h3><p>A coded label is not always enough. A statement such as &#8220;LSTM was preferred in 58%&#8221; has a meaning that depends on its originating subset and denominator. The generated <code>EvidenceClaim</code> model therefore preserves <code>claim_id</code>, <code>text</code>, <code>subset_id</code>, <code>denominator</code>, and <code>provenance</code>.</p><p><code>add_claim</code> validates and appends one claim without aggregation. <code>validate_claim</code> requires a positive integer denominator when a claim contains an explicit percentage, unless the text explicitly says that the denominator or percentage is unavailable. This is a contract check, not a formula implementation: no percentage is recomputed.</p><p>The denominator audit turns those fields into an explicit audit record. Its record shape is represented by the generated <code>DenominatorEntry</code> class:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4be5eb19-5b2f-4814-9f7a-5f5d552fc11d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class DenominatorEntry:
    """Source-qualified denominator metadata for one quantitative claim.

    The denominator is deliberately kept alongside the claim and its subset.
    Counts reported for different review subsets are not interchangeable, even
    when they describe related methods, inputs, or metrics.
    """

    claim_id: str
    subset_id: str
    denominator: int
    source_location: SourceLocation | None</code></pre></div><p><code>audit_denominators(claims)</code> emits one <code>DenominatorEntry</code> for every claim with a supplied denominator. Thus an entry retains all three essential identifiers: <code>claim_id</code>, <code>subset_id</code>, and <code>denominator</code>, along with the source location. A missing denominator remains missing; the audit does not guess one.</p><p>This matters because the paper reports quantities from different scopes, including subsets of 12, 30, 34, 45, 57, 122, and 379 studies. Those values are not interchangeable. The abstract-level count of more than 379 primary studies is also not a universal denominator for every percentage reported elsewhere in the review.</p><h3>Refuse cross-subset aggregation</h3><p>The package makes the safe behavior explicit through <code>reject_cross_subset_aggregation</code>. It validates the supplied claims, collects their <code>subset_id</code> values, and raises <code>ValidationError</code> when more than one subset is present. The function&#8217;s purpose is preventative: it stops a caller from producing a combined statistic that the paper did not report.</p><p>This is an implementation decision derived from the paper&#8217;s ambiguities, not a new statistical method. A future meta-analysis could define compatible inclusion rules, effect sizes, weighting, and uncertainty estimates, but those choices are absent from the supplied source and are outside this reproduction.</p><h3>Preserve conflicting findings instead of choosing a winner</h3><p>The paper contains claims that cannot safely be treated as universal conclusions about deep learning. Some passages report that deep learning is generally more accurate than traditional machine learning, while another discussion says that deep-learning models have not generally outperformed standard models. The implementation must retain such claims with their scopes rather than silently resolving the disagreement.</p><p><code>preserve_conflicting_claims</code> validates each claim and returns it without deduplicating by wording, denominator, or topic. It does not select a winner, assign a weight, or infer that one claim is more credible. This allows two statements about LSTM&#8212;or about deep learning more broadly&#8212;to coexist when they originate from different review subsets or contexts.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h3>Construct the synthesis report</h3><p><code>SynthesisReport</code> holds three <code>DimensionSummary</code> objects: <code>rq1_methods</code>, <code>rq2_information_sources</code>, and <code>rq3_metrics</code>. It also contains qualitative findings, denominator-audit entries, and warnings. The report explicitly does not contain predictions, losses, formulas, or newly measured model-comparison results.</p><p>The main method assembles those parts in a fixed order. The following excerpt shows the central mapping from the three research questions to the three evidence dimensions:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b2c3c3aa-4226-44df-a1b2-3eb053ca7a26&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    # RQ1: which AI methods and technologies are reported?
    rq1_methods = summarize_dimension(catalog, EvidenceDimension.METHOD)
    # RQ2: which informational sources are reported?
    rq2_information_sources = summarize_dimension(
        catalog, EvidenceDimension.INFORMATION_SOURCE
    )
    # RQ3: which evaluation metrics are reported?
    rq3_metrics = summarize_dimension(catalog, EvidenceDimension.METRIC)

    preserved_claims = preserve_conflicting_claims(catalog.claims)
    denominator_entries = audit_denominators(preserved_claims)</code></pre></div><p>A caller requests the completed object with <code>synthesize_review(catalog)</code>. Conceptually, the result can then be inspected as <code>report.rq1_methods</code>, <code>report.rq2_information_sources</code>, <code>report.rq3_metrics</code>, and <code>report.denominator_audit</code>. The implementation first validates the catalog, preserves conflicting claims, and audits denominators before constructing the report.</p><p>The report&#8217;s qualitative section uses <code>summarize_limitations</code>. In the generated implementation, available <code>EvidenceClaim</code> text is surfaced with subset and source information. This keeps recommendations&#8212;such as combining multimodal data, improving interpretability, examining hyperparameters, or testing across markets&#8212;distinct from measured predictive results.</p><h3>Worked example: two LSTM claims, two scopes</h3><p>Suppose an external extraction process supplies two claims about LSTM. One claim belongs to a 12-study subset and another belongs to a separate 57-study subset. The claims may mention different review findings, and they may even appear to disagree. The correct procedure is:</p><ol><li><p>Create each claim with a distinct <code>subset_id</code>.</p></li><li><p>Store the original denominator, 12 or 57, on the corresponding claim.</p></li><li><p>Attach provenance to both claims.</p></li><li><p>Add them to the catalog without combining them.</p></li><li><p>Call <code>synthesize_review(catalog)</code> and inspect the resulting denominator audit.</p></li></ol><p>The audit should conceptually contain two separate entries, one with <code>subset_id</code> for the 12-study scope and one with <code>subset_id</code> for the 57-study scope. It should not contain a new denominator of 69, 379, or any other combined value. The report therefore communicates &#8220;these claims were reported in these scopes,&#8221; not &#8220;this is the overall percentage for LSTM.&#8221;</p><p>Likewise, if coded evidence records mention both <code>RMSE</code> and <code>MAE</code>, <code>summarize_dimension</code> catalogs the metric names. It does not reconstruct their mathematical definitions or calculate either metric. The supplied paper contains no canonical equations, so formulas are intentionally absent from this section and from the planned synthesis implementation.</p><h3>Keep reported workflow facts separate</h3><p>The constants module stores the paper&#8217;s reported workflow facts without treating them as synthesis denominators. <code>REPORTED_FLOW_COUNTS</code> contains the stage-specific values for retrieval, reviews read, abstract screening remainder, post-duplicate handling, and final retention. These values describe the review workflow; they do not count methods or establish the prevalence of a predictive algorithm.</p><p>Similarly, the source-specific retrieval counts for Scopus and Web of Science remain separate. Keeping those mappings distinct prevents a database-level retrieval total from being confused with the number of primary studies covered by the retained reviews or with a denominator attached to an LSTM claim.</p><h3>Boundary of interpretation</h3><p>The synthesis can faithfully report that SVM, LSTM, and neural-network categories are prominent in the reviewed literature, that historical closing-price time series and technical indicators are common information sources, and that accuracy, MSE, RMSE, MAPE, and MAE are frequently discussed metrics. These are paper-level observations preserved by the catalog.</p><p>They are not guarantees that SVM or LSTM performs best, not evidence that a particular feature combination improves forecasting, and not results generated by this Python package. Any claim about predictive performance would require a separate, externally supplied model specification, dataset, target, split, preprocessing policy, and evaluation procedure.</p><p>The implementation therefore reproduces the evidence-synthesis method&#8212;source-qualified grouping, denominator auditing, and conflict preservation&#8212;rather than inventing a meta-analysis or a stock-prediction runtime.</p><h2>The explicit contract for underspecified predictive reproductions</h2><p>What should happen when someone wants to turn this review into a stock-prediction model, but the paper does not define the model? The safe answer is to stop before silently choosing one. A reproduction contract acts as a checklist: it records every decision an external predictive study must provide and reports omissions explicitly.</p><h3>Why a contract is necessary</h3><p><strong>Paper fact.</strong> The source discusses classification and regression, historical prices, technical indicators, text, sentiment, macroeconomic information, and many model families, including SVM, LSTM, ANN, CNN, and ARIMA. However, it does not define a target variable, forecast horizon, label threshold, sequence length, data split, normalization policy, architecture, hyperparameters, optimizer, loss, or training schedule. It also supplies no canonical equations.</p><p><strong>Implementation decision.</strong> The generated package therefore implements <code>predictive_reproduction_contract</code>, not a stock forecaster. Its responsibility is to make omitted decisions visible before model code is written. It does not instantiate or train an SVM, LSTM, ANN, CNN, ARIMA, or any other predictor, and it does not calculate accuracy, MSE, RMSE, MAPE, MAE, NMSE, correlation coefficient <code>R</code>, or POCID.</p><p>This boundary is important because a plausible default is still an unsupported assumption. For example, the paper&#8217;s observations about daily data or approximately 1000-day periods describe reviewed studies; they do not establish a universal lookback window. Likewise, the prominence of LSTM in some review findings does not authorize selecting an LSTM architecture for reproduction.</p><h3>Describe inputs without pretending to implement preprocessing</h3><p><code>ModalitySpecification</code> describes one externally defined input source. Its fields are <code>name</code>, <code>shape_description</code>, <code>dtype</code>, <code>frequency</code>, <code>alignment</code>, and <code>leakage_policy</code>. These are contract descriptions rather than loaded array shapes. A numerical time series might therefore be described as <code>(time, features)</code> without inventing a feature count or sequence length.</p><p>The generated implementation validates that these descriptions are non-empty strings:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;67942120-b518-4048-a0df-834dbace5d37&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass
class ModalitySpecification:
    """Describe one externally defined input modality.

    ``shape_description`` is intentionally textual.  For example,
    ``"(time, features)"`` describes a numerical series without asserting a
    sequence length or feature count.  ``leakage_policy`` must state how
    information leakage is prevented; the paper does not define that policy.
    """

    name: str
    shape_description: str
    dtype: str
    frequency: str
    alignment: str
    leakage_policy: str</code></pre></div><p>Here, a <strong>multimodal input</strong> combines different information types, such as numerical prices and timestamped news. <strong>Alignment</strong> specifies how observations from those sources correspond, for example through a timestamp or another key. <strong>Leakage prevention</strong> specifies how information unavailable at prediction time is excluded. The paper recommends combinations of information sources but does not define either procedure.</p><p><code>validate_modalities</code> checks the declarations without loading, reshaping, resampling, or scaling data. <code>validate_temporal_alignment</code> adds a guard for multimodal configurations: declarations must either share an alignment description or explicitly describe an alignment or join procedure. This is a validation rule introduced by the implementation to prevent an ambiguous input contract; it is not a preprocessing algorithm claimed by the paper.</p><h3>Record the decisions a model would require</h3><p><code>PredictiveConfiguration</code> separates task type from the other predictive choices. Its fields are:</p><ul><li><p><code>task_type</code>, which must be externally identified as classification or regression;</p></li><li><p><code>target_definition</code>, describing what value or label is predicted;</p></li><li><p><code>horizon</code>, identifying how far ahead the prediction is made;</p></li><li><p><code>modalities</code>, containing the declared inputs;</p></li><li><p><code>lookback</code>, describing sequence construction or historical context;</p></li><li><p><code>split_policy</code> and <code>scaling_policy</code>, describing data partitioning and normalization;</p></li><li><p><code>model_family</code> and <code>hyperparameters</code>, describing the selected model externally;</p></li><li><p><code>optimizer</code> and <code>loss</code>, describing training choices; and</p></li><li><p><code>metrics</code>, naming the evaluation measures to report.</p></li></ul><p>The class uses <code>None</code> to mean that a decision was not supplied. It does not replace missing values with defaults. This distinction is especially useful for intermediate Python users: an omitted field is not the same as a field whose value happens to be zero, empty, or unknown.</p><p>The following is an intentionally incomplete configuration. It declares one numerical modality but leaves the target, horizon, and other decisions unspecified:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ffa6120c-36c2-4da2-abe1-298ebe45c2b6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from review_synthesis.predictive_contract import (
    ModalitySpecification,
    PredictiveConfiguration,
    missing_predictive_fields,
)

numerical = ModalitySpecification(
    name="numerical_time_series",
    shape_description="(time, features)",
    dtype="float32",
    frequency="daily",
    alignment="timestamp",
    leakage_policy="external time-aware policy",
)

config = PredictiveConfiguration(
    modalities=[numerical],
)

missing = missing_predictive_fields(config)</code></pre></div><p>The descriptive shape <code>(time, features)</code> does not imply a particular tensor rank beyond the text supplied by the external specification, and it does not imply a lookback length. Similarly, <code>frequency="daily"</code> records an intended external description; it does not cause the package to retrieve or resample daily data.</p><h3>Turn omissions into an actionable report</h3><p><code>missing_predictive_fields</code> returns stable field names for decisions that are absent. For the configuration above, the conceptual result includes fields such as <code>task_type</code>, <code>target_definition</code>, <code>horizon</code>, <code>lookback</code>, <code>split_policy</code>, <code>scaling_policy</code>, <code>model_family</code>, <code>hyperparameters</code>, <code>optimizer</code>, <code>loss</code>, and <code>metrics</code>. The implementation does not infer any of them from the review.</p><p><code>build_missing_specification_report</code> groups those omissions into categories such as <code>task_and_target</code>, <code>temporal_construction</code>, <code>data_preprocessing</code>, <code>model_and_optimization</code>, and <code>evaluation</code>. <code>explain_missing_fields</code> adds plain-language explanations. For instance, it explains that the target must include its value or label construction, while the horizon must include its units. These explanations are derived implementation guidance, not additional paper facts.</p><p><code>contract_report</code> packages the result into a JSON-safe dictionary. Its output distinguishes whether the configuration is validated, which fields are missing, which are supplied, and what the implementation boundary is. The generated function explicitly records that no predictive runtime, equations, or training procedure is supplied by the paper.</p><p>A strict caller can use <code>assert_faithful_reproduction_possible</code> as a gate:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;31d377c0-f1e4-4196-9515-02f29b7dc574&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def assert_faithful_reproduction_possible(config: PredictiveConfiguration) -&gt; None:
    """Reject a predictive contract that still depends on unspecified external decisions."""
    if not isinstance(config, PredictiveConfiguration):
        raise ValidationError("config must be a PredictiveConfiguration")

    validate_modalities(config.modalities)
    validate_temporal_alignment(config.modalities)
    missing = missing_predictive_fields(config)
    if missing:
        details = "; ".join(_explanation_for_field(field) for field in missing)
        raise IncompleteSpecificationError(
            "faithful predictive reproduction is incomplete; missing fields: "
            + ", ".join(missing)
            + ". "
            + details
        )</code></pre></div><p><code>IncompleteSpecificationError</code> is therefore an explicit contract response. It does not mean that model training failed; no model-training operation is defined here. It means that the supplied external specification is insufficient for a faithful predictive reproduction.</p><h3>What this boundary does&#8212;and does not&#8212;guarantee</h3><p>The contract validator can ensure that required declarations exist and that modality descriptions are structurally coherent. It cannot determine whether a chosen target is scientifically appropriate, whether a split truly prevents leakage, or whether a selected architecture is effective. Those decisions belong to the external predictive study and must be documented there.</p><p>This model-family-agnostic design is deliberate. The paper names incompatible families and does not provide a unified interface for combining them. A future reproduction could supply an SVM with tabular features, an LSTM with sequences, or a multimodal architecture, but each would require details absent from this source. The contract records those details once they are supplied; it does not invent them.</p><h3>Worked example: stopping before accidental invention</h3><p>Suppose an external user supplies only the numerical modality shown above. <code>missing_predictive_fields</code> identifies the absent target and horizon along with the rest of the required decisions. <code>contract_report</code> can present those omissions by category, and <code>explain_missing_fields</code> can state why each one matters. A caller may then choose to raise <code>IncompleteSpecificationError</code> through <code>assert_faithful_reproduction_possible</code> rather than proceeding.</p><p>The important outcome is not a prediction. It is an auditable boundary: no sequence length, normalization rule, loss, optimizer, or LSTM architecture has been smuggled into the reproduction. Multimodal inputs likewise remain declarations until an external specification defines their alignment and leakage policy.</p><p>The planned audit and verification layers were not run under the authoritative policy, so this tutorial makes no claim that the code was executed or that the contract path passed verification. Faithfully reproducing this paper means preserving its review evidence and explicitly exposing its missing predictive details; implementing a predictive model requires a separate, external specification.</p><h2>Serialization, reporting, and deterministic example data</h2><p>How do you turn a provenance-rich review catalog into an artifact that another person can inspect without losing uncertainty, dates, or denominators? Use three deliberately separate layers: serialization converts domain objects into JSON-safe data, reporting formats that data for readers, and local command-line tools connect the workflow without pretending to query bibliographic databases or train a stock predictor.</p><h3>Preserve meaning when converting objects to JSON</h3><p><strong>Implementation decision.</strong> The generated <code>serialization.py</code> module treats serialization as a loss-avoidance step, not as a place to calculate new evidence. <code>record_to_dict</code> converts one supported dataclass, while <code>catalog_to_dict</code> and <code>synthesis_to_dict</code> serialize the two main review artifacts. <code>write_json</code> then writes a deterministic UTF-8 document to a local <code>Path</code>.</p><p>The important conversions are explicit:</p><ul><li><p>Enum members such as <code>Database.SCOPUS</code> are written using their stable string values.</p></li><li><p><code>date</code> and <code>datetime</code> values are written as ISO-8601 strings.</p></li><li><p><code>None</code> remains JSON <code>null</code>, which is essential for an explicitly unspecified included-study count.</p></li><li><p>Tuples and other supported sequences become JSON arrays while preserving order.</p></li><li><p>Provenance, confidence, subset identifiers, and denominators are ordinary serialized fields rather than comments that could be lost.</p></li></ul><p>The focused implementation excerpt below shows the public catalog boundary. Notice that <code>catalog_to_dict</code> validates the catalog before converting it, but does not repair or infer missing values.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;05afab68-1e87-4834-a125-bf92851734e1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def catalog_to_dict(catalog: ReviewCatalog) -&gt; dict[str, object]:
    """Serialize a review catalog while retaining evidence provenance."""
    if not isinstance(catalog, ReviewCatalog):
        raise ValidationError("catalog must be a ReviewCatalog value")
    catalog.validate()
    converted = _to_json_value(catalog, "catalog")
    if not isinstance(converted, dict):
        raise ValidationError("serialized catalog must be a dictionary")
    return converted


def synthesis_to_dict(report: SynthesisReport) -&gt; dict[str, object]:
    """Serialize a source-qualified synthesis report without recomputation."""
    if not isinstance(report, SynthesisReport):
        raise ValidationError("report must be a SynthesisReport value")
    converted = _to_json_value(report, "synthesis_report")
    if not isinstance(converted, dict):
        raise ValidationError("serialized synthesis report must be a dictionary")
    return converted</code></pre></div><p>This boundary matters for the paper because evidence is source-qualified. A serialized <code>RMSE</code> label must remain a reported metric name, not become a computed value. Likewise, a null field must remain visibly absent rather than being replaced by a guessed count or parameter.</p><h3>Use schemas to describe, not invent, data</h3><p>The generated <code>schema.py</code> module provides two JSON-schema-like descriptions. <code>catalog_schema()</code> describes persisted review records, coded evidence, claims, and warnings. <code>predictive_configuration_schema()</code> describes the fields an external predictive reproduction must supply, including task type, target, horizon, modalities, split, scaling, model family, hyperparameters, optimizer, loss, and metrics.</p><p>The distinction is important: the predictive schema is a contract for missing information, not a model specification. It does not provide a default LSTM, sequence length, normalization rule, or loss. <code>validate_payload_shape</code> checks required keys and structural types without coercing values or filling gaps.</p><p>For example, the catalog schema explicitly permits an included-study count to be an integer or null:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5b18b0a4-218d-40ac-8a00-03bd7ff6b231&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">"included_study_count": {"type": ["integer", "null"], "minimum": 0},</code></pre></div><p>That <code>null</code> is meaningful. The paper reports ten retained reviews with the sequence <code>27</code>, unspecified, <code>53</code>, <code>12</code>, <code>34</code>, <code>122</code>, <code>24</code>, <code>57</code>, <code>30</code>, and <code>20</code>. The implementation must preserve the unspecified entry instead of estimating it from the other values.</p><h3>Render evidence without changing it</h3><p>Serialization is for machines; reporting is for readers. The generated <code>reporting.py</code> module offers <code>render_summary</code>, <code>render_markdown_report</code>, and <code>render_warning_section</code>. These functions format existing objects and make important boundaries visible:</p><ul><li><p><code>render_markdown_report</code> labels the result as a review-evidence synthesis.</p></li><li><p>Flow counts remain stage-specific rather than being presented as one total.</p></li><li><p>Each dimension summary displays source review identifiers, reported denominators, and uncertainty counts.</p></li><li><p>The denominator audit displays each claim&#8217;s <code>claim_id</code>, <code>subset_id</code>, denominator, and source location.</p></li><li><p>Warnings are rendered as a visible section rather than being hidden in program logs.</p></li></ul><p>The report renderer itself states the scope clearly:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e68620e0-0b8d-47f7-815c-726e10d623e8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">"## Scope and implementation boundary",
"",
"This report represents a systematic review of systematic reviews. Its entries are reported literature evidence, not primary stock-price datasets and not results from a newly trained model.",
"",
"The rendering preserves coded labels, source-review identifiers, uncertainty counts, and denominator metadata. It does not reconstruct equations, calculate metric formulas, combine incompatible study subsets, or report predictive performance.",</code></pre></div><p>This is presentation behavior, not new analysis. For instance, a report may show that a review catalog contains <code>LSTM</code> or <code>MAE</code>, but it does not claim that the package trained an LSTM or evaluated predictions with MAE.</p><h3>Keep the command-line path local</h3><p>The generated <code>cli.py</code> module provides a local interface with two input modes: deterministic synthetic examples or a user-supplied local JSON file. <code>build_parser</code> exposes <code>--example</code>, <code>--input</code>, <code>--output</code>, and <code>--format</code> options. <code>load_local_records</code> parses local candidates and extraction batches; it does not contact Scopus, Web of Science, Rayyan, Yahoo Finance, or another external service.</p><p>The following invocation is documentation for the available interface, not a claim that it was run here:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;72807db0-939a-42f9-8d70-2f1a3a849ebd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">python -m review_synthesis.cli --example --output report.json</code></pre></div><p>With <code>--format json</code>, the CLI can include the serialized catalog, flow counts, synthesis report, and missing-specification report. With <code>--format markdown</code>, it uses <code>render_markdown_report</code>; with <code>--format text</code>, it uses <code>render_summary</code>. The generated output should be understood as a local workflow artifact, not as a database export or a predictive experiment.</p><h3>Worked example: synthetic records to a labeled artifact</h3><p>The example flow begins with <code>example_candidates()</code> and <code>example_extraction_batches()</code> from <code>src/review_synthesis/example_data.py</code>. These functions intentionally create small, deterministic fixtures. The candidates include records attributed to both databases and a repeated title so that provenance and deduplication can be demonstrated. The extraction batch contains separate review-level fields for methods, information sources, and metrics. <code>example_claims()</code> supplies illustrative claims with a denominator, but its text is explicitly synthetic.</p><p>The conceptual pipeline is:</p><ol><li><p>Create local candidates and extraction batches.</p></li><li><p>Supply screening decisions to <code>run_review_pipeline</code>.</p></li><li><p>Obtain a <code>PipelineResult</code> containing the catalog, screening log, flow counts, synthesis report, and missing-specification report.</p></li><li><p>Pass the catalog to <code>catalog_to_dict</code> and the synthesis report to <code>synthesis_to_dict</code>.</p></li><li><p>Pass the report and flow counts to <code>render_markdown_report</code>.</p></li><li><p>Write the combined artifact with <code>write_json</code>.</p></li></ol><p>The dedicated <code>scripts/build_example_report.py</code> script follows that path and labels its output with <code>artifact_type: "synthetic_review_synthesis_example"</code>. Its explanatory metadata also states that the artifact is not the paper&#8217;s bibliographic dataset and contains no trained-model results. This labeling is necessary because deterministic data is useful for teaching the interfaces, but it is not evidence that the paper&#8217;s original records were recovered.</p><p>A second script, <code>scripts/export_reported_facts.py</code>, serves a different purpose. It exports supplied paper facts such as the article identifier, search metadata, source-specific retrieval counts, stage-specific flow counts, and the ten included-study counts. The unspecified review count is preserved as JSON <code>null</code>, and the export includes the supplied ambiguity and warning lists. It is therefore a facts manifest, not a reconstruction of the original database contents.</p><h3>What this artifact boundary guarantees&#8212;and what it does not</h3><p>The serialization and reporting layer guarantees only the behavior represented by its contracts: supported objects can be converted without dropping explicit nulls, provenance, confidence, ordering, or denominator metadata; reports can expose those fields; and local examples can be labeled as synthetic. The paper itself supplies no canonical equations, so no equation or metric formula appears in these artifacts.</p><p>The implementation also does not claim that the paper prescribes a Python dependency stack. The source mentions tools including Python, TensorFlow, NumPy, Pandas, Keras, Scikit-Learn, TA-Lib, TA4J, and MATLAB, but the generated local interfaces use the package&#8217;s own records and standard local file handling rather than treating any named tool as mandatory.</p><p>No code execution, test generation, static verification, semantic code verification, tutorial-section verification, or final quality review was performed under the authoritative run policy. The example commands and artifact flow above explain the intended interfaces only. Faithful reproduction of this paper ends with preserving its review evidence and limitations; a stock-prediction runtime would require a separate external specification.</p><h2>Static and semantic verification strategy&#8212;and what was not performed</h2><p>How can you tell whether this reproduction preserves the paper&#8217;s meaning without accidentally claiming that a stock-prediction model exists? Treat verification as a set of explicit invariants, not as a performance test. The relevant promises are that records retain provenance, exclusions retain reasons, reported counts remain separated, uncertain evidence remains uncertain, and generated reports do not claim newly trained models.</p><p>This section describes the intended verification strategy and the generated audit hooks. It also states the actual status precisely: verification was skipped under the authoritative run policy. No code execution, test generation, semantic code verification, tutorial-section verification, or final quality review occurred.</p><h3>What static verification would inspect</h3><p><strong>Implementation decision.</strong> Static verification would inspect the Python package without running the research workflow. For this project, that means parsing each planned Python file with the Python abstract syntax tree, checking imports and public symbols, inspecting dataclass fields and annotations, and comparing public signatures with the planned interfaces.</p><p>The same review would inspect the JSON-like structures returned by <code>catalog_schema()</code> and <code>predictive_configuration_schema()</code>. The catalog schema is intended to require review records, coded evidence, claims, and warnings while preserving nullable fields. In particular, <code>None</code> must remain available for an explicitly unspecified study count or denominator rather than being replaced with an estimate. The predictive schema documents required external decisions but does not provide defaults for a target, horizon, architecture, or optimizer.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The generated <code>schema.py</code> module exposes structural validation through <code>validate_payload_shape</code>. Its responsibility is deliberately narrow: check required keys and supported types without coercing missing information.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;77d9e0be-6a76-4ad6-bc4c-b142107b49b1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def validate_payload_shape(payload: Mapping[str, object], schema: Mapping[str, object]) -&gt; None:
    """Validate required keys and supported structural types without coercion."""
    if not isinstance(payload, Mapping):
        raise ValidationError("payload must be a mapping")
    if not isinstance(schema, Mapping):
        raise ValidationError("schema must be a mapping")
    _validate_node(payload, schema, "payload")</code></pre></div><p>Readers should notice the phrase &#8220;without coercion.&#8221; A structural validator may reject an invalid payload, but it must not repair a missing provenance field, infer a denominator, or turn an uncertain Table 5 association into a reported fact. Those are evidence decisions, not harmless formatting operations.</p><p>A future enabled static pass would also check that <code>run_review_pipeline</code> depends on the intended local modules and that <code>PipelineResult</code> contains review artifacts rather than predictions, model weights, losses, or newly computed scores. The orchestration boundary is important: the function consumes externally supplied records and decisions, while live database retrieval and model fitting remain outside this paper-faithful implementation.</p><h3>What semantic verification would inspect</h3><p><strong>Derived explanation.</strong> Semantic verification asks whether the code&#8217;s behavior matches the paper&#8217;s scope and the implementation contracts. It is more than checking that names and types look plausible. It would inspect whether every extracted item has a source review and location, whether each exclusion has a screening stage and reason, whether counts remain tied to their workflow stages, and whether claims preserve their original subset and denominator.</p><p>The generated <code>audit.py</code> module provides three focused audit entry points:</p><ul><li><p><code>audit_catalog</code> inspects cross-record references, provenance, uncertainty, denominator preservation, flow counts, and report scope.</p></li><li><p><code>audit_reported_counts</code> compares supplied flow values with the paper&#8217;s reported stage-specific values without modifying them.</p></li><li><p><code>audit_no_model_claims</code> searches generated qualitative text and warnings for unsupported language suggesting that a model was trained or evaluated.</p></li></ul><p>The last check is a safeguard against scope drift. It is not evidence that a model exists or that the audit found a valid model implementation. Its purpose is to catch wording that would contradict the paper-faithful boundary.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;1115bb9e-e957-41b7-8479-8843ae453ea5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def audit_no_model_claims(report: SynthesisReport) -&gt; list[AuditFinding]:
    """Detect unsupported claims that this package trained or evaluated a model."""

    if not isinstance(report, SynthesisReport):
        return [
            _finding(
                "error",
                "report_type",
                "report must be a SynthesisReport instance",
            )
        ]

    text_parts: list[str] = []
    text_parts.extend(report.qualitative_findings)
    text_parts.extend(report.warnings)
    text = "\n".join(text_parts).casefold()
    unsupported_phrases = (
        "trained a model",
        "model was trained",
        "achieved accuracy",
        "predicted prices",
        "generated predictions",
        "fitted model",
        "test-set performance",
        "validation score",
    )</code></pre></div><p>This audit is especially relevant here because the paper names SVM, LSTM, neural networks, ARIMA, and many other methods. A report that says one of these methods achieved a result would exceed the supplied evidence unless that statement were explicitly attributed to a reviewed study and retained with its source context.</p><h3>A hypothetical audit walkthrough</h3><p>Consider a catalog containing one retained review and one uncertain <code>SVM</code> observation extracted from fragmented Table 5 material. A semantic audit would conceptually follow this checklist:</p><ol><li><p>Confirm that the evidence&#8217;s <code>source_review_id</code> refers to a retained review or an explicitly external source.</p></li><li><p>Confirm that its <code>Provenance.source_location</code> identifies the relevant table or section.</p></li><li><p>Confirm that its confidence remains <code>UNCERTAIN</code> when row alignment was not established.</p></li><li><p>Confirm that any quantitative <code>EvidenceClaim</code> retains its <code>subset_id</code> and denominator.</p></li><li><p>Compare the supplied <code>FlowCounts</code> with the reported values: 69 retrieved, 43 reviews read, 17 remaining after abstract screening, 16 after duplicate or same-author handling, and 10 finally retained.</p></li><li><p>Inspect the synthesis text for unsupported claims about training, predictions, or test-set performance.</p></li></ol><p>The constants used for this comparison are source-qualified facts rather than inferred prevalence estimates. <code>REPORTED_RETRIEVAL_COUNTS</code> keeps Scopus at 40 and Web of Science at 29, while <code>REPORTED_FLOW_COUNTS</code> keeps the PRISMA-style values separate by stage. Neither mapping proves that a supplied local fixture reproduces the original database export.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ccfc7fa7-ca10-4adb-a115-235971aa09e9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">REPORTED_RETRIEVAL_COUNTS: Final[Mapping[str, int]] = MappingProxyType(
    {
        "Scopus": 40,
        "Web of Science": 29,
    }
)</code></pre></div><p>If a local example contains different counts, the intended behavior is to report a difference, not to overwrite the local data or claim reconciliation. Similarly, a mismatch does not by itself show that the paper is incorrect; it may simply mean that the local records are synthetic or incomplete.</p><h3>Actual verification status for this tutorial</h3><p><strong>Run-policy fact.</strong> The authoritative policy disabled local static verification, test generation, code execution, semantic code verification, tutorial-section verification, and final quality review. The supplied verification records therefore have a skipped status. This tutorial does not claim that the generated files parsed successfully, that schemas validated real artifacts, that audits returned no findings, or that any example command ran.</p><p>The audit functions are planned safeguards and generated interfaces, not completed verification results. The same distinction applies to <code>PipelineResult</code>: its type describes what <code>run_review_pipeline</code> would return for supplied local inputs, but no execution outcome is asserted here.</p><p>Important limitations remain even if these checks are enabled later. Table 5 is fragmented, so some method, source, and metric associations require manual source-PDF review. The paper&#8217;s geographic inclusion criterion and some review-subset statements are ambiguous. No canonical equations, target definition, forecast horizon, preprocessing policy, architecture, hyperparameters, loss, optimizer, or training schedule were supplied. The local examples are synthetic and must not be presented as the original Scopus or Web of Science records.</p><p>The faithful boundary is therefore concise: reproducing this paper means reproducing its review-selection, provenance, uncertainty, and evidence-synthesis workflow. Any executable stock-prediction model requires a separate external specification, and its correctness cannot be established from this review alone.</p><p>Use the button or URL below to download the source code.</p>
      <p>
          <a href="/__u/onepagecode.substack.com/p/forecasting-stock-prices-with-lstms">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Quant Trading: Stacked Ensemble Machine Learning & LSTMs for Stock Price Forecasting (Python Guide)]]></title><description><![CDATA[Implementing a multi-model hybrid stacking engine combining Random Forest, XGBoost, SVR, and deep LSTMs.]]></description><link>https://onepagecode.substack.com/p/quant-trading-stacked-ensemble-machine</link><guid isPermaLink="false">https://onepagecode.substack.com/p/quant-trading-stacked-ensemble-machine</guid><dc:creator><![CDATA[Onepagecode]]></dc:creator><pubDate>Wed, 12 Aug 2026 20:25:27 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_W6n!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3078fdad-8e80-40a9-be17-640eb18fa447_1254x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This review surveys conventional machine learning, time-series, deep-learning, and ensemble methods for stock-price forecasting and stock-trend classification. It also reports an implementation comparing SVR, MLPR, KNN, random forest, XG-Boost, LSTM, and a stacked Random Forest + XG-Boost + LSTM ensemble on Yahoo Finance data for TAINIWALCHM and AGROPHOS. The reported ensemble achieved the best RMSE and R2 values in the paper's comparison, although the exact preprocessing, feature construction, stacking procedure, and experimental code are not fully specified.</p><h3>Use the URL at the end of this article to download the source code</h3><h2>Implementation Assumptions</h2><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><ul><li><p>The Section 4/Table 1 80/20 chronological split is the primary reproduction choice; the Section 3 75/25 recommendation is retained as a documented conflict.</p></li><li><p>Yahoo Finance access is optional at runtime; local CSV fixtures and deterministic synthetic data support offline demonstrations.</p></li><li><p>Ticker symbols, selected features, target column, scaler, look-back window, validation protocol, seeds, hyperparameter selections, and ensemble wiring remain explicit configuration choices because the paper does not specify them.</p></li><li><p>The LSTM is implemented with TensorFlow/Keras rather than a hand-written recurrent cell because equations eq<em>8, eq</em>9, and eq_10 are incomplete.</p></li><li><p>XG-Boost is accessed through an optional local Python dependency, with a clear compatibility boundary and no network-dependent behavior.</p></li><li><p>Canonical equation LaTeX is stored only from the supplied records; empty or damaged equation records are not reconstructed.</p></li><li><p>Table 2 values are stored as reported references and are never represented as independently verified results.</p></li><li><p>All data transformations that learn parameters are fitted on training data only.</p></li><li><p>The implementation is stock-specific by default, with separate preprocessing and model fitting for each security.</p></li><li><p>No live API call, model training, or code execution is claimed by this plan.</p></li></ul><h2>Scope, evidence status, and reproduction target</h2><p>What does it mean to reproduce this paper when the paper does not include the original source code? The practical answer is to reproduce the documented experimental structure while making every missing choice visible. This project therefore aims to be auditable and configurable, not to claim an exact numerical reconstruction of the paper's experiment.</p><p>The paper is primarily a systematic review of stock-price forecasting and trend-classification methods. Its main implementation comparison, in Section 4, contains six individual regressors or forecasting models&#8212;SVR, MLPR, KNN, random forest, XG-Boost, and LSTM&#8212;plus a proposed Random Forest + XG-Boost + LSTM ensemble. A regression model predicts a numeric value, here a stock-price target. A baseline is an individual model used as a comparison point for the ensemble.</p><p>The two reported securities are TAINIWALCHM, described as Tainwala Chemicals and Plastics, and AGROPHOS, described as Agro Phos. The paper reports historical data obtained through Yahoo Finance and gives date ranges for these names, but it does not provide the exact Yahoo Finance ticker identifiers needed to make an unambiguous download request.</p><h3>What is specified, and what is not</h3><p>The strongest implementation evidence comes from Table 1. It specifies an 80% training and 20% testing split for the proposed ensemble, MSE loss and Adam for the neural component, a maximum of 50 epochs, batch size 32, two LSTM layers, dropout of <code>0.2</code>, and a dense layer with 25 units. It also lists candidate values for random-forest and XG-Boost hyperparameters.</p><p>This conflicts with the paper's more general pipeline description, which mentions a 75/25 split. The framework uses the Table 1 80/20 split as its primary reproduction choice because it is the more specific statement about the proposed ensemble. The conflict remains recorded rather than silently discarded.</p><p>Several choices needed for an executable experiment are absent: the exact ticker symbols, selected features, target column, scaling method, LSTM look-back window, forecast horizon, validation procedure, random seeds, selected hyperparameters, and tuning algorithm. The paper also calls the combination &#8220;stacked&#8221; without defining its wiring. Stacking usually means fitting a second-stage model on component predictions, whereas blending or weighted averaging combines predictions directly, often with fixed weights. These are different operations, so the implementation exposes the distinction instead of presenting one as recovered fact.</p><p>The generated configuration objects make these decisions explicit. The following excerpt is copied from <code>src/stock_forecasting/config.py</code>; notice that feature columns, target, split fraction, sequence settings, and scaler are all fields rather than hidden constants.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2b3b1171-bb5d-4dfa-8403-4ace767388e4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True, slots=True)
class PreprocessingConfig:
    """Make feature, target, split, sequence, and scaling decisions explicit.

    The paper's Table 1 specifies an 80/20 split for the proposed ensemble,
    while its generic pipeline mentions 75/25.  The default-free configuration
    requires the caller to choose a fraction; a value of ``0.8`` is the
    primary reproduction choice described by the paper.  The target column,
    look-back, horizon, feature set, and scaler are not specified by the paper.
    """

    feature_columns: tuple[str, ...]
    target_column: str
    train_fraction: float
    lookback: int
    horizon: int
    scaler_name: Optional[str]</code></pre></div><p>Here, <code>feature_columns</code> names the input columns, while <code>target_column</code> identifies the value to forecast. The paper mentions possible OHLCV and secondary data but does not establish the final feature set or target. <code>lookback</code> is the number of historical rows supplied to an LSTM sequence, and <code>horizon</code> describes how far ahead its target is aligned. Both are implementation decisions. <code>scaler_name</code> records whether a transformation such as standard or min-max scaling is used; the paper does not specify which one.</p><h3>The manifest as an audit boundary</h3><p>The experiment manifest separates paper facts from implementation decisions and unresolved ambiguities. <code>ExperimentManifest</code> records the stock identity, source configuration, feature and target choices, split policy, sequence settings, tuning description, model configurations, and ensemble combiner. <code>validate_manifest</code> checks that these fields are sufficiently explicit and serializable.</p><p>The generated class exposes this purpose directly through its fields:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d72e9a14-3d2a-43f8-81ad-b84a2cf14663&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True, slots=True)
class ExperimentManifest:
    """Audit record for one stock-specific reproduction configuration.

    The manifest deliberately stores unresolved paper details as explicit fields or
    notes rather than hiding them in defaults.  ``model_configurations`` records
    selected settings when supplied; an empty mapping means that no selected model
    settings were provided by the caller or recovered from the paper.
    """

    paper_id: str
    stock: str
    data_source: Mapping[str, Any]
    feature_columns: tuple[str, ...]
    target_column: str
    train_fraction: float
    split_policy: str
    scaler_name: str | None
    lookback: int
    horizon: int
    seed: int | None
    tuning_method: str
    validation_fraction: float | None
    model_configurations: Mapping[str, Any]
    combiner: str
    ensemble_weights: tuple[float, ...] | None
    meta_features: str
    package_metadata: Mapping[str, Any]</code></pre></div><p>A small worked configuration would therefore identify the stock label separately from the source ticker, choose features such as <code>Open</code>, <code>High</code>, <code>Low</code>, and <code>Volume</code>, choose <code>Close</code> as the target, record an 80/20 chronological split, select a look-back and scaler, and state whether tuning is unspecified or explicitly configured. It would also record a seed and a combiner such as <code>weighted_average</code> or <code>linear_stacking</code>. Those values would document the reproduction run; they would not become claims about what the paper originally did.</p><p>For example, the public configuration contract requires an explicit data request rather than silently selecting a ticker:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;afe48fb1-aa77-4559-8ca5-8520cb89ad0d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True, slots=True)
class DataSourceConfig:
    """Describe one explicit historical-data request.

    ``ticker`` is deliberately required from the caller because the paper does
    not provide the exact Yahoo Finance symbols for TAINIWALCHM or AGROPHOS.
    Date endpoint inclusivity and interval semantics remain properties of the
    selected Yahoo Finance client.
    """

    ticker: str
    start: str
    end: str
    interval: str</code></pre></div><p>The source ticker, date strings, and interval are provenance, not merely convenience arguments. They make it possible to distinguish two runs that use different Yahoo Finance symbols, endpoint behavior, or sampling frequencies.</p><h3>What the reported results mean</h3><p>The paper's Table 2 reports RMSE and R2 for each stock and model. RMSE, or root mean square error, summarizes prediction error in the target's units; smaller values indicate smaller typical squared errors. R2, or the coefficient of determination, compares the predictions with a constant-mean reference. It can be negative for a poor regression model and is not required to remain between zero and one.</p><p>The paper reports the RF + XG-Boost + LSTM ensemble as having the strongest comparison values: RMSE 2.0247 and R2 0.9921 for TAINIWALCHM, and RMSE 1.2658 and R2 0.9897 for AGROPHOS. These are reported references only. The generated code has not been executed, no Yahoo Finance request has been made, and no model training, test run, syntax check, or independent numerical verification has occurred. Consequently, this tutorial must not describe those values as reproduced results.</p><p>One additional fidelity issue is retained rather than corrected: the supplied paper discussion refers to the AGROPHOS date range in one sentence using the TAINIWALCHM name. The framework records the source and stock explicitly so that a reader can resolve this issue deliberately when supplying data.</p><p>In the rest of the tutorial, <code>fit_models_and_generate_predictions</code> will be treated as the training-and-prediction procedure, <code>stacked_rf_xgboost_lstm</code> as the configurable ensemble method, and <code>evaluate_regression_predictions</code> as the held-out evaluation step. A held-out test set is data reserved until model fitting and model selection are complete. Keeping it separate is a derived implementation safeguard, not evidence that the original paper used exactly the same procedure.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/p/quant-trading-stacked-ensemble-machine?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/p/quant-trading-stacked-ensemble-machine?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/onepagecode.substack.com/p/quant-trading-stacked-ensemble-machine?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><p>The central reproduction principle is therefore simple: preserve the paper's stated settings, expose every missing setting, and label reported metrics as references until an independently specified and executed experiment supports a comparison.</p><h2>Project layout and public contracts</h2><p>How can one keep a stock-forecasting reproduction understandable when it contains data cleaning, several model families, an LSTM, an ensemble, and multiple reporting steps? The practical answer is to give each stage one responsibility and connect stages with explicit contracts. In this project, data modules obtain and validate rows, preprocessing modules construct examples, model adapters fit estimators, ensemble modules align and combine predictions, and metric modules evaluate the final outputs.</p><p>This separation is an implementation design derived from the paper's pipeline description. The paper says that data acquisition, preprocessing, model training, evaluation, and tuning are distinct stages, but it does not prescribe a Python package layout. The generated package makes that progression visible through modules under <code>data</code>, <code>models</code>, <code>ensemble</code>, <code>training</code>, <code>metrics</code>, and <code>experiments</code>.</p><h3>A map of the package</h3><p>The lightweight package entry point is <code>src/stock_forecasting/__init__.py</code>. It exports configuration types but does not import TensorFlow, XG-Boost, or Yahoo Finance integrations eagerly. This matters because configuration and local data work should remain available even when optional modeling or network dependencies are not installed.</p><p>The dependency boundary is declared in <code>pyproject.toml</code>. The core installation contains NumPy, pandas, and scikit-learn, while deep learning, Yahoo Finance access, XG-Boost, plotting, and development tools are optional extras:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;toml&quot;,&quot;nodeId&quot;:&quot;2889ef6c-0fac-4c0d-aa33-151647909cf9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-toml">[project.optional-dependencies]
# TensorFlow supplies the Keras LSTM implementation described in Table 1.
deep-learning = [
    "tensorflow&gt;=2.13",
]
# Yahoo Finance access is opt-in; local CSV and synthetic workflows remain offline.
yahoo = [
    "yfinance&gt;=0.2.30",
]
# XG-Boost is optional because the paper does not specify its package or version.
xgboost = [
    "xgboost&gt;=1.7",
]
plotting = [
    "matplotlib&gt;=3.7",
]</code></pre></div><p>This is not a claim about the paper's original package versions. It is a reproducibility decision for the generated framework. In particular, a local CSV or synthetic-data demonstration should not require a network client, and importing configuration should not fail merely because TensorFlow is unavailable.</p><p>The main hand-off records are defined in <code>src/stock_forecasting/types.py</code>. They carry arrays together with timestamps and validate their basic shape and finiteness requirements. This is important for the methods <code>preprocess_time_series</code>, <code>fit_models_and_generate_predictions</code>, <code>stacked_rf_xgboost_lstm</code>, and <code>evaluate_regression_predictions</code>: each method depends on receiving data with the intended meaning, not merely an object that happens to be an array.</p><h3>Two representations of the same forecasting problem</h3><p>Conventional regressors use a tabular feature matrix, which we will call <code>X_tab</code>. Its shape is rank 2:</p><ul><li><p><code>(n_samples, n_features)</code> means one independent feature row per training or evaluation example.</p></li><li><p><code>y</code> is the aligned numeric target, usually represented as <code>(n_samples,)</code> or <code>(n_samples, 1)</code>.</p></li></ul><p>The LSTM instead uses a sequence tensor, which we will call <code>X_seq</code>. Its shape is rank 3:</p><ul><li><p><code>(n_sequences, lookback, n_features)</code> means each example contains an ordered window of <code>lookback</code> observations.</p></li><li><p>The corresponding <code>y</code> contains one future target for each sequence.</p></li></ul><p>The distinction is semantic. A tabular row says, in effect, &#8220;use these features for this example.&#8221; A sequence window says, &#8220;use this ordered history to predict the associated future target.&#8221; The generated <code>TabularSplit</code> and <code>SequenceSplit</code> records enforce these different contracts rather than treating every input as a generic two-dimensional matrix.</p><p>A focused portion of <code>TabularSplit</code> shows the tabular contract and its alignment checks:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;91676ee9-aab7-4b17-882b-d1e1d73e6f3f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass
class TabularSplit:
    """Chronological train/test arrays for conventional regressors.

    ``X_train`` and ``X_test`` have shape ``(n_samples, n_features)``. Targets
    may have shape ``(n_samples,)`` or ``(n_samples, 1)`` and must align with
    their corresponding timestamp indices.
    """

    X_train: np.ndarray
    X_test: np.ndarray
    y_train: np.ndarray
    y_test: np.ndarray
    train_index: pd.Index
    test_index: pd.Index

    def __post_init__(self) -&gt; None:
        x_train = _as_numeric_array(self.X_train, "X_train")
        x_test = _as_numeric_array(self.X_test, "X_test")
        if x_train.ndim != 2 or x_test.ndim != 2:
            raise ValueError(
                "X_train and X_test must both have rank 2 "
                "(n_samples, n_features)"
            )
        if x_train.shape[1] != x_test.shape[1]:
            raise ValueError("X_train and X_test must have the same feature count")</code></pre></div><p>The record also checks that both partitions are nonempty, that target lengths match sample counts, that timestamp indices have the expected lengths, and that training and test indices do not overlap. Those checks protect the chronological split required by the reproduction plan. They do not determine the split ratio; the surrounding configuration makes the Table 1 choice of 80/20 explicit while documenting the paper's conflicting generic 75/25 recommendation.</p><p>The sequence counterpart uses the same idea with one additional axis:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a776796c-d1c4-4562-9885-05514e2690de&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass
class SequenceSplit:
    """Chronological train/test windows for recurrent models.

    Each feature tensor has shape ``(n_sequences, lookback, n_features)``. The
    target timestamp denotes the future observation associated with its window;
    it therefore has one label per sequence rather than one label per timestep.
    """

    X_train: np.ndarray
    X_test: np.ndarray
    y_train: np.ndarray
    y_test: np.ndarray
    train_index: pd.Index
    test_index: pd.Index

    def __post_init__(self) -&gt; None:
        x_train = _as_numeric_array(self.X_train, "X_train")
        x_test = _as_numeric_array(self.X_test, "X_test")
        if x_train.ndim != 3 or x_test.ndim != 3:
            raise ValueError(
                "X_train and X_test must both have rank 3 "
                "(n_sequences, lookback, n_features)"
            )
        if x_train.shape[1:] != x_test.shape[1:]:
            raise ValueError(
                "train and test sequence tensors must share lookback and feature dimensions"
            )</code></pre></div><p>Here <code>x_train.shape[1:]</code> represents <code>(lookback, n_features)</code>. The two partitions must agree on those dimensions even though they contain different numbers of sequences. The sequence indices identify target timestamps, not every timestamp inside each window. This distinction becomes essential later when RF and XG-Boost predictions are aligned with LSTM predictions: forming windows can remove early rows, so component predictions cannot be combined by position alone.</p><h3>Prediction and preprocessing provenance</h3><p><code>PreprocessingState</code> stores fitted transformation information such as the feature scaler, target scaler, scaler name, feature count, and number of rows used for fitting. The paper mentions scaling but does not specify the scaler type, fit scope, or reporting scale. The generated record therefore preserves those choices instead of hiding them inside a model adapter. A training-only fit scope is a leakage-control requirement: a scaler must not learn distribution information from the held-out period.</p><p><code>PredictionBundle</code> couples a model name, a pandas index, and one prediction per index label. Its contract accepts values shaped <code>(n,)</code> or <code>(n, 1)</code>, but requires the number of values to equal the number of timestamps and rejects non-finite values. This is more than defensive programming. A prediction without its timestamp is not enough for the ensemble method <code>stacked_rf_xgboost_lstm</code>, because the RF, XG-Boost, and LSTM outputs must refer to the same target observations before they are combined.</p><p><code>MetricRecord</code> performs the final identity bookkeeping: it stores the stock, model, RMSE, R2, and reporting scale. The metric value alone is insufficient for an experiment comparing two securities and several models. A record must say what was measured and on which scale.</p><h3>A common tabular model interface</h3><p>The conventional models share a small protocol in <code>src/stock_forecasting/models/base.py</code>. <code>RegressorProtocol</code> requires <code>fit</code> and <code>predict</code>; <code>validate_regression_arrays</code> checks the rank, numeric type, finiteness, and sample alignment of tabular data:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ccf37f98-ea5a-4839-ad37-83a77c42cdb5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@runtime_checkable
class RegressorProtocol(Protocol):
    """Protocol implemented by tabular regression model adapters.

    Implementations consume a rank-2 feature matrix with shape
    ``(n_samples, n_features)`` and a one-dimensional target vector with shape
    ``(n_samples,)``.  ``fit`` returns the fitted estimator to support the
    conventional ``model.fit(...).predict(...)`` usage.
    """

    def fit(self, X: np.ndarray, y: np.ndarray) -&gt; RegressorProtocol:
        """Fit the regressor and return the fitted model."""

    def predict(self, X: np.ndarray) -&gt; np.ndarray:
        """Return one continuous prediction for each input row."""</code></pre></div><p>This contract allows the SVR, MLPR, KNN, random-forest, and XG-Boost adapters to be orchestrated consistently while keeping their library-specific settings local. It does not make the models equivalent: their kernels, neural architectures, neighbor rules, tree construction, and boosting behavior remain different. It simply ensures that the training runner can request a fit and then obtain one prediction per evaluation row.</p><h3>Worked example: tracing two batches</h3><p>Suppose preprocessing produces 240 tabular training rows with four selected features. The tabular contract is <code>X_train.shape == (240, 4)</code> and <code>y_train.shape == (240,)</code>. A conventional model can fit this representation directly and return, for example, one prediction for every held-out timestamp.</p><p>Suppose the same feature stream is converted into windows with a look-back of 20. The recurrent representation is no longer <code>(240, 4)</code>. It is shaped like <code>(n_sequences, 20, 4)</code>, and each sequence target has a timestamp after its input window. The exact <code>n_sequences</code> depends on the configured horizon and the rows available after splitting; the paper does not specify the look-back or horizon, so these remain explicit implementation decisions.</p><p>In both cases, the index is part of the contract. A tabular prediction bundle might contain one value for each test timestamp. A sequence prediction bundle may begin later because its first valid target requires a complete history. Before ensemble combination, the alignment layer must intersect and order these timestamp sets. The resulting component matrix has three columns&#8212;one each for random forest, XG-Boost, and LSTM&#8212;and one row per common target timestamp.</p><p>The planned tests describe these shape, time-order, and alignment invariants, but they are not evidence of completed checks. Under the run policy, local static verification, semantic code verification, test generation, and code execution were disabled. Thus the package layout and contracts presented here explain the intended implementation boundaries; they do not claim that the generated code ran or that any Table 2 result was reproduced.</p><h2>Data acquisition and provenance</h2><p>How can a stock-forecasting experiment be reproduced if the paper names securities but does not provide the exact download identifiers? Treat the data request itself as part of the experiment. A model input is not merely &#8220;TAINIWALCHM data&#8221;; it is a specific <code>ticker</code>, <code>start</code> date, <code>end</code> date, <code>interval</code>, returned column set, and timestamp policy.</p><p>The paper identifies Yahoo Finance as the source for the two reported securities. It describes TAINIWALCHM as covering 2014&#8211;2023 and AGROPHOS as covering 2018&#8211;2023, subject to ticker availability and API behavior. However, it does not supply the exact Yahoo Finance ticker symbols, endpoint inclusivity, interval semantics, client version, or complete returned schema. One passage also incorrectly associates the AGROPHOS period with TAINIWALCHM. The implementation records these facts as unresolved configuration rather than silently guessing.</p><h3>Make the request explicit</h3><p>The generated <code>DataSourceConfig</code> groups the four essential request fields. Here, <code>ticker</code> identifies the security, <code>start</code> and <code>end</code> delimit the requested historical period, and <code>interval</code> selects the sampling frequency, such as daily data. These values are validated before a client is allowed to make a request.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;0de0bfb8-456f-4725-b5e0-bae5c0d8b691&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True, slots=True)
class DataSourceConfig:
    """Describe one explicit historical-data request.

    ``ticker`` is deliberately required from the caller because the paper does
    not provide the exact Yahoo Finance symbols for TAINIWALCHM or AGROPHOS.
    Date endpoint inclusivity and interval semantics remain properties of the
    selected Yahoo Finance client.
    """

    ticker: str
    start: str
    end: str
    interval: str</code></pre></div><p>This is an implementation contract, not an equation or a setting recovered from the paper. It prevents a convenience default from selecting an unintended security. The configuration also rejects an empty ticker, malformed ISO date strings, and an end value that is not later than the start value. The exact meaning of the end boundary still belongs to the selected client and should be recorded with the snapshot.</p><h3>Yahoo Finance as an opt-in boundary</h3><p><code>YahooFinanceClient.download_historical_data</code> is responsible for one explicit request. It imports the optional <code>yfinance</code> dependency only when the method is called, downloads one ticker, rejects an unavailable or malformed response, normalizes the timestamp index, and attaches provenance. The public convenience function <code>download_historical_data</code> delegates to this client; neither function substitutes a ticker.</p><p>A focused excerpt shows the request boundary and several deliberate choices:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;fdb92284-ac0d-40ab-9d56-56fae5b68923&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">downloaded = yf.download(
    tickers=config.ticker,
    start=config.start,
    end=config.end,
    interval=config.interval,
    auto_adjust=False,
    progress=False,
    threads=False,
)</code></pre></div><p>The paper specifies Yahoo Finance as the source but does not specify <code>auto_adjust</code>, progress behavior, threading, or the client API. Consequently, <code>auto_adjust=False</code> and the other arguments are implementation decisions that must be preserved in the provenance record if the resulting data is used. They are not claims about the original experiment.</p><p>The client then stores the request metadata on the returned timestamp-indexed <code>DataFrame</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;1ecd90c3-2d6d-4cff-b72a-10d7896bdba2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">provenance: dict[str, Any] = {
    "source": "Yahoo Finance",
    "client": "yfinance",
    "ticker": config.ticker,
    "start": config.start,
    "end": config.end,
    "interval": config.interval,
    "auto_adjust": False,
}
frame.attrs["yahoo_finance_provenance"] = provenance
return frame</code></pre></div><p>A timestamp-indexed <code>DataFrame</code> has one observation per row and a <code>DatetimeIndex</code> identifying when that observation occurred. Chronological order is an invariant: the returned index must be increasing, and duplicate timestamps are rejected rather than merged with an undocumented rule. This matters because later train/test splitting and sequence construction depend on temporal order. A successful API response is not automatically valid experimental input; an empty result, ambiguous multi-ticker response, malformed timestamp, or duplicate timestamp raises <code>YahooFinanceAcquisitionError</code>.</p><p>The acquisition wrapper also does not claim to resolve market-data subtleties. Missing trading days may be normal rather than errors, while corporate-action adjustments, timezone conversion, endpoint inclusivity, and Yahoo Finance client versions can change the resulting table. Those choices belong in the data snapshot and its manifest.</p><h3>Offline alternatives: CSV and synthetic data</h3><p>The repository provides two offline sources. <code>load_stock_csv</code> reads a user-supplied snapshot, parses a named date column, rejects unparseable or duplicate timestamps, and returns a sorted <code>DataFrame</code>. This is useful when a Yahoo Finance snapshot has already been obtained, but CSV loading is an implementation addition; it is not the paper's reported acquisition procedure.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;126d5791-a7ec-42a5-9d0d-7bd8823c0a38&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from pathlib import Path

from stock_forecasting.data.csv import load_stock_csv

raw = load_stock_csv(Path("data/sample_stock.csv"), "Date")</code></pre></div><p>The CSV loader deliberately stops at structural loading. Feature selection, null handling, zero-value policy, and target selection remain downstream decisions. This keeps the source adapter from silently changing the observations before the configured preprocessing stage sees them.</p><p>For a fully offline demonstration, <code>make_synthetic_stock_data</code> creates deterministic OHLCV-shaped data. It returns the columns <code>Open</code>, <code>High</code>, <code>Low</code>, <code>Close</code>, <code>Adj Close</code>, and <code>Volume</code> with a unique chronological business-day index. The generator validates positive finite values and basic OHLC ordering before returning.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ba8d8f06-45c6-4bea-9601-3b61cf2d0ddd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from stock_forecasting.data.synthetic import make_synthetic_stock_data

raw = make_synthetic_stock_data(300, seed=7)</code></pre></div><p>This worked example provides 300 rows with a repeatable random seed. It is suitable for demonstrating data flow, configuration, and shape contracts, but it cannot reproduce the paper's Yahoo Finance observations or Table 2 metrics. The generated price process, start date, business-day frequency, and initial value are all synthetic implementation choices.</p><h3>The same workflow with a real request</h3><p>Once an exact user-supplied ticker is known, the equivalent explicit configuration is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2a946e0b-0d36-4d87-a575-c820c9f2f1ce&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from stock_forecasting.config import DataSourceConfig
from stock_forecasting.data.yahoo import YahooFinanceClient

request = DataSourceConfig(
    ticker="EXPLICIT_TICKER",
    start="2014-01-01",
    end="2023-12-31",
    interval="1d",
)
raw = YahooFinanceClient().download_historical_data(request)</code></pre></div><p><code>EXPLICIT_TICKER</code> is intentionally a placeholder in this tutorial example: the supplied paper context does not identify the correct Yahoo Finance symbol, and substituting one would be an unsupported factual claim. For an AGROPHOS request, the caller would likewise provide the verified ticker and the reported 2018&#8211;2023 period. The data source, request values, returned columns, and row count should be saved alongside the CSV rather than inferred later.</p><p>The repository's <code>scripts/download_yahoo_snapshot.py</code> provides that opt-in command-line boundary. It requires <code>--ticker</code>, <code>--start</code>, and <code>--end</code>, so a download cannot accidentally use a hidden security default:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;fce11959-f1d3-44b9-8076-e8c4adcf23d0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">parser.add_argument(
    "--ticker",
    required=True,
    help="Exact Yahoo Finance ticker identifier; no default is provided.",
)</code></pre></div><p>After acquisition, the script writes a CSV and a JSON sidecar containing the request and retrieval metadata. It checks that the frame is nonempty, chronological, and free of duplicate timestamps before writing. A typical command therefore has to make the unresolved data decision visible:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;1f31cdfd-48df-4ea3-8d90-5f8239c37e96&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">python scripts/download_yahoo_snapshot.py \
    --ticker EXPLICIT_TICKER \
    --start 2014-01-01 \
    --end 2023-12-31 \
    --interval 1d \
    --output data/yahoo_snapshot.csv</code></pre></div><p>This command is an interface example, not a report that a download occurred. Under the current run policy, no API call, local execution, or data validation was performed.</p><h3>What provenance protects</h3><p>Provenance is the source configuration attached to a loaded dataset: at minimum, the source name, client, ticker, dates, interval, and relevant adjustment settings. It allows a later reader to distinguish a changed download from a changed model. It also exposes why exact numerical reproduction is unavailable here: the paper omits the exact ticker symbols and several data-handling details, and no original snapshot is included.</p><p>The acquisition stage therefore establishes only a trustworthy boundary. It does not decide whether the target is <code>Close</code> or <code>Adj Close</code>, which OHLCV columns become features, how nulls or zero values are handled, or how sequences are formed. Those decisions belong to the next preprocessing stage and must remain explicit. Synthetic and CSV inputs make the package usable offline, while the Yahoo Finance path preserves the paper's stated source when the required identifier and optional dependency are supplied.</p><h2>Cleaning, feature selection, splitting, and scaling</h2><p>How do we turn a timestamped stock table into fair training examples without allowing information from the future to influence the past? The key is to separate four responsibilities: clean the raw rows, choose the inputs and target explicitly, split observations chronologically, and fit learned transformations only on the training portion.</p><p>This section implements the <code>preprocess_time_series</code> method through <code>prepare_stock_data</code> in <code>src/stock_forecasting/preprocessing.py</code>. The paper mentions OHLCV data and possible secondary inputs, but it does not state the exact feature set. It also does not identify whether the target is <code>Close</code>, <code>Adjusted Close</code>, or another value. Those choices therefore belong in <code>PreprocessingConfig</code>, rather than being inferred from the review text.</p><h3>1. Validate and clean the time series</h3><p>The generated pipeline expects a pandas DataFrame with a unique, ascending <code>DatetimeIndex</code>. Each row represents one observation, such as one trading day. A timestamp index is important because a positional split is meaningful only when the rows are already in chronological order.</p><p>The paper's general pipeline calls for removing redundancies and null values. The implementation makes those operations deterministic: <code>clean_stock_table</code> sorts the index, keeps the last row for a duplicate timestamp, and removes rows containing null values. It also accepts a <code>remove_zero_values</code> argument. In the composed pipeline, that argument is explicitly set to <code>False</code>, because the paper does not say which zero values are invalid. A zero volume, for example, may have a different interpretation from a zero price.</p><p>The core cleaning excerpt is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;9e266df5-b4d7-401e-9e98-bd2269f22d60&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">cleaned = frame.copy(deep=True).sort_index()
cleaned = cleaned.loc[~cleaned.index.duplicated(keep="last")]
cleaned = cleaned.dropna(axis=0, how="any")

numeric_columns = [
    column for column in cleaned.columns if is_numeric_dtype(cleaned[column].dtype)
]
if remove_zero_values and numeric_columns:
    zero_rows = (cleaned.loc[:, numeric_columns] == 0).any(axis=1)
    cleaned = cleaned.loc[~zero_rows]</code></pre></div><p>This excerpt is from <code>src/stock_forecasting/data/validation.py</code>, in <code>clean_stock_table</code>. Notice that the function does not silently drop a target column and does not decide that every zero is an error. After cleaning, it checks that the result is nonempty, chronologically ordered, unique by timestamp, and finite in its numeric columns.</p><p><code>validate_stock_table</code> performs the stricter downstream check. It verifies that required columns exist, are numeric, contain no nulls, and contain only finite values. Its <code>required_columns</code> argument is normally the union of the configured features and target. A malformed table fails early rather than producing a confusing model error later.</p><h3>2. Select <code>X</code> and <code>y</code> explicitly</h3><p>Let <code>X</code> denote the feature table and <code>y</code> denote the regression target. In the tabular representation, <code>X</code> has shape <code>(n_observations, n_features)</code>, while <code>y</code> has shape <code>(n_observations,)</code>. Their timestamp indices must be identical so that each feature row refers to the same target row.</p><p>The generated <code>select_features_and_target</code> function does not guess columns from names such as &#8220;OHLCV.&#8221; Instead, it receives a <code>PreprocessingConfig</code> containing <code>feature_columns</code> and <code>target_column</code>. This is an implementation decision forced by an omission in the paper: OHLCV and secondary data are discussed, but the Section 4 feature set is not specified.</p><p>A focused part of the function is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;dd90eee1-86ab-4654-8700-646cb1ae1abc&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">required_columns = config.feature_columns + (config.target_column,)
# Boundary validation enforces the timestamp, null, numeric, and finite
# value invariants before the configured projection is constructed.
validate_stock_table(frame, required_columns)

if config.target_column in config.feature_columns:
    raise ValueError("target_column must not also be a feature column")

features = frame.loc[:, list(config.feature_columns)].astype(
    np.float64, copy=True
)
target = frame.loc[:, config.target_column].astype(np.float64, copy=True)</code></pre></div><p>This excerpt comes from <code>src/stock_forecasting/data/features.py</code>. The check preventing the target from also appearing as a feature is a leakage safeguard: including the value being predicted among the inputs would make the supervised task ill-defined.</p><p>For the worked example, choose the following configuration explicitly:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;17e01465-2389-4d92-8659-474ef74be2b5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">feature_columns = ("Open", "High", "Low", "Volume")
target_column = "Close"
lookback = 20
horizon = 1
scaler_name = "standard"
train_fraction = 0.8</code></pre></div><p>These are not recovered facts about the paper's original experiment. They are illustrative reproduction decisions. The target could instead be <code>Adjusted Close</code>, and the paper does not establish which one was used. Likewise, the look-back and horizon are absent from the paper.</p><h3>3. Split chronologically</h3><p>The primary reproduction choice is an 80/20 chronological split because Table 1 specifies 80% training and 20% testing for the proposed ensemble. The paper's generic pipeline separately mentions 75/25, so the code does not hide that conflict: callers pass <code>train_fraction</code> explicitly.</p><p><code>chronological_split</code> first checks that <code>X</code> and <code>y</code> are aligned, nonempty, unique, and already sorted. It then slices by position without shuffling:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4d71a5ea-4291-417d-96e4-a2947f32294f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">boundary = int(len(X) * fraction)
if boundary &lt;= 0 or boundary &gt;= len(X):
    raise ValueError(
        "train_fraction produces an empty training or test partition; "
        "provide more observations or a different fraction"
    )

X_train = X.iloc[:boundary].copy()
X_test = X.iloc[boundary:].copy()
y_train = y.iloc[:boundary].copy()
y_test = y.iloc[boundary:].copy()</code></pre></div><p>This is the implementation in <code>src/stock_forecasting/data/splitting.py</code>. The resulting <code>X_train</code> and <code>y_train</code> contain earlier observations, while <code>X_test</code> and <code>y_test</code> contain later observations. The function also checks that the final training timestamp precedes the first test timestamp.</p><p>A validation split, when tuning is enabled, is made only inside the training portion by <code>chronological_train_validation_split</code>. Its API does not accept test data. That boundary is deliberate: a test target must not influence feature transformation, hyperparameter selection, or model construction.</p><h3>4. Fit scaling state on training rows only</h3><p>Scaling changes the numerical representation of features. With standard scaling, for example, a feature is centered and rescaled using statistics estimated from data. The exact scaler type and metric scale are not specified by the paper, so the generated code supports <code>None</code>, <code>"standard"</code>, and <code>"minmax"</code> as explicit choices.</p><p>The important invariant is not the particular scaler; it is the fitting scope. <code>fit_scalers</code> receives <code>X_train</code> and <code>y_train</code>, fits the feature and target transformations there, and returns a <code>PreprocessingState</code>. The test rows are transformed later using that already-fitted state:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f6754874-251d-4df8-b81f-7f0223f6a66b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">state = fit_scalers(X_train, y_train, config.scaler_name)
X_train_scaled = transform_features(state, X_train)
X_test_scaled = transform_features(state, X_test)
y_train_scaled = transform_target(state, y_train).reshape(-1)
y_test_scaled = transform_target(state, y_test).reshape(-1)</code></pre></div><p>This excerpt is from <code>src/stock_forecasting/preprocessing.py</code>. Fitting a scaler on the complete dataset would allow the distribution of the later test period to affect the representation of earlier training data. That is a form of temporal leakage, even though the target values are not directly copied into the features.</p><p>If predictions are reported in the original price units, <code>inverse_transform_target</code> can undo the target transformation. The code exposes that operation rather than assuming whether the paper's RMSE and <code>R2</code> were computed on scaled or original values. Predictions and targets must always be compared on the same scale.</p><h3>5. Construct LSTM windows</h3><p>Tabular models consume independent-looking rows. An LSTM instead consumes ordered windows. The sequence representation, called <code>X_seq</code> here, has shape <code>(n_sequences, lookback, n_features)</code>. Each sequence contains <code>lookback</code> historical rows, and its associated target is located <code>horizon</code> steps after the end of that window.</p><p>For example, with <code>lookback=20</code> and <code>horizon=1</code>, the first window contains 20 observations and predicts the immediately following target. The initial rows that cannot form a complete window are therefore not represented as sequence examples. This is why sequence predictions generally have fewer timestamps than tabular predictions.</p><p>The generated constructor documents and enforces that relationship:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;106d8f0a-dfe0-40c3-8110-8fa89f089edb&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">sequence_features[sequence_number] = features[
    sequence_number : sequence_number + lookback
]
sequence_targets[sequence_number] = targets[target_position]
target_labels.append(ordered_timestamps[target_position])</code></pre></div><p>This excerpt comes from <code>src/stock_forecasting/data/sequences.py</code>. <code>target_position</code> is computed after the input window, so the window does not contain observations from after its target. The target timestamps are retained in <code>SequenceSplit.train_index</code> and <code>SequenceSplit.test_index</code>; those labels are needed later when RF, XG-Boost, and LSTM predictions are aligned.</p><p>The generated <code>make_sequences</code> function uses the Table 1 80/20 ratio because its public signature has no split-ratio argument. In the composed <code>prepare_stock_data</code> pipeline, the constructed windows are re-partitioned using the configured <code>train_fraction</code>, so the preprocessing configuration remains the authoritative record of the split decision.</p><h3>6. Compose the preparation pipeline</h3><p><code>prepare_stock_data</code> connects the stages in this order:</p><ol><li><p>Clean the timestamp-indexed table while retaining zero values by explicit policy.</p></li><li><p>Select the configured features and target.</p></li><li><p>Split tabular observations chronologically.</p></li><li><p>Fit optional feature and target scalers on training rows only.</p></li><li><p>Transform training and test tabular arrays with that state.</p></li><li><p>Construct ordered sequence windows from the transformed series.</p></li><li><p>Return both representations, their indices, the preprocessing state, and an audit dictionary.</p></li></ol><p>The result is a <code>PreparedStockData</code> record. Its <code>tabular</code> field contains rank-2 arrays for classical regressors, and its <code>sequences</code> field contains rank-3 arrays for the LSTM. The audit dictionary records choices such as feature columns, target column, split fraction, look-back, horizon, scaler, zero policy, and the paper's 80/20 versus 75/25 conflict.</p><p>A complete illustrative setup is therefore conceptually:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;30f691b5-43e9-4ef4-bbc8-c0d5293c2b0c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">config = PreprocessingConfig(
    feature_columns=("Open", "High", "Low", "Volume"),
    target_column="Close",
    train_fraction=0.8,
    lookback=20,
    horizon=1,
    scaler_name="standard",
)
prepared = prepare_stock_data(raw_frame, config)

print(prepared.tabular.X_train.shape)
print(prepared.sequences.X_train.shape)
print(prepared.audit["target_column"])</code></pre></div><p>The exact <code>PreprocessingConfig</code> constructor and <code>PreparedStockData</code> record are defined by the generated package; the excerpt illustrates the public data flow and the explicit decisions it must carry. For a sufficiently long input, the first printed shape is rank 2, while the second is rank 3 with its middle dimension equal to <code>20</code> and its final dimension equal to the number of configured features.</p><p>The main lesson is that preprocessing is part of the model specification. The paper does not provide enough detail to recover its original target, features, scaler, look-back, horizon, or zero policy. This implementation keeps those choices visible, preserves chronological alignment, and prevents learned transformations from using the held-out period. No code execution, API call, test run, or independent numerical verification was performed under the run policy.</p><h2>Classical regression baselines</h2><p>Which models should receive the same stock-price rows before the LSTM and ensemble are considered? The paper's Section 4 comparison uses five tabular regressors: support-vector regression (SVR), multilayer perceptron regression (MLPR), K-nearest-neighbor regression (KNN), random-forest regression, and XG-Boost regression. Each model consumes the same training features and produces one continuous prediction for each evaluation row.</p><p>In the notation used here, <code>X_train</code> is a rank-2 feature matrix with shape <code>(n_train, n_features)</code>, and <code>y_train</code> is a one-dimensional target array with shape <code>(n_train,)</code>. <code>X_eval</code> has shape <code>(n_eval, n_features)</code> and must use the same feature order and preprocessing as <code>X_train</code>. The resulting prediction vector has shape <code>(n_eval,)</code>. This common contract makes the models comparable and allows their predictions to be aligned later for the Random Forest + XG-Boost + LSTM ensemble.</p><h3>What the paper specifies&#8212;and what it does not</h3><p><strong>Paper fact.</strong> The paper includes SVR, MLPR, KNN, random forest, and XG-Boost in its reported comparison. It describes random forest as a bagging ensemble of decision trees: multiple trees are trained and their regression outputs are aggregated. It describes XG-Boost as sequential gradient boosting: weak learners are added iteratively while later learners respond to earlier errors through gradient-based updates.</p><p><strong>Implementation decision.</strong> The generated package wraps established Python estimators behind a small local protocol. The paper does not specify the SVR kernel or regularization, the MLPR architecture, the KNN neighbor count, the selected random-forest settings, the selected XG-Boost settings, or the search procedure. Consequently, these values remain constructor arguments or explicit configuration rather than being presented as recovered paper settings.</p><p>The paper does provide candidate values for the random forest and XG-Boost components. These are search candidates, not selected configurations. A candidate grid says which values may be tried; it does not say which value produced the paper's reported result.</p><h3>One shared fit/predict contract</h3><p>The <code>RegressorProtocol</code> in <code>src/stock_forecasting/models/base.py</code> defines the boundary used by the tabular adapters. Its responsibility is deliberately small: <code>fit</code> receives training arrays and returns a fitted model, while <code>predict</code> receives rank-2 evaluation features and returns one prediction per row. The accompanying <code>validate_regression_arrays</code> function checks numeric dtype, rank, nonempty dimensions, finite values, and sample-count agreement.</p><p>The central contract is visible in this focused excerpt:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;15c1794f-e8d0-4395-ae8e-735bb68effe3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@runtime_checkable
class RegressorProtocol(Protocol):
    """Protocol implemented by tabular regression model adapters.

    Implementations consume a rank-2 feature matrix with shape
    ``(n_samples, n_features)`` and a one-dimensional target vector with shape
    ``(n_samples,)``.  ``fit`` returns the fitted estimator to support the
    conventional ``model.fit(...).predict(...)`` usage.
    """

    def fit(self, X: np.ndarray, y: np.ndarray) -&gt; RegressorProtocol:
        """Fit the regressor and return the fitted model."""

    def predict(self, X: np.ndarray) -&gt; np.ndarray:
        """Return one continuous prediction for each input row."""</code></pre></div><p>This is more than defensive programming. If one adapter silently accepts a differently ordered feature matrix, returns a column matrix instead of a vector, or drops rows, the comparison and eventual ensemble can become invalid while still appearing to run. The protocol does not decide which model is best; it ensures that each model receives and returns data in a known form.</p><h3>The individual adapters</h3><p><code>SVRRegressor</code> wraps <code>sklearn.svm.SVR</code>. Its configurable settings include <code>kernel</code>, <code>C</code>, <code>gamma</code>, <code>epsilon</code>, and related estimator parameters. The adapter intentionally leaves feature scaling external. Scaling may be important for SVR, but the paper does not specify a scaler or its fitting scope, so that choice belongs to the preprocessing configuration rather than this model class. Calling <code>predict</code> before <code>fit</code>, changing the feature count, or producing non-finite outputs raises an explicit error.</p><p><code>MLPRRegressor</code> wraps scikit-learn's multilayer perceptron regressor and exposes <code>hidden_layer_sizes</code>, activation, solver, optimization settings, iteration count, and random state. The paper reports an MLPR result but does not describe its hidden layers or training configuration. The generated adapter therefore does not claim that its ordinary defaults reproduce the paper's MLPR.</p><p>KNN is also configurable. The relevant choices include <code>n_neighbors</code>, the weighting rule, and the distance metric. A neighbor count must be positive and cannot exceed the number of training rows. KNN is particularly sensitive to feature scale because distances are computed in feature space: a high-magnitude column can dominate a lower-magnitude column unless preprocessing addresses that issue. Tree-based models generally respond differently to scale because their splits compare feature values rather than relying directly on geometric distance.</p><p>The paper's KNN distance record is identified as <code>eq_2</code>, but its canonical LaTeX is empty because the extracted expression is OCR-damaged. The supplied symbol record describes <code>D</code> as a nonnegative distance between observations, <code>h_i</code> as a data value or observation, <code>p_r</code> as a predicted or query value, <code>n</code> as the number of dimensions or terms, and <code>l</code> as a summation index. The exact vector components, square-root formatting, and index placement are unclear. Therefore, no equation block is reproduced here and no formula is reconstructed.</p><p>Instead, <code>euclidean_distance_eq_2</code> documents the implementation mapping. It accepts two finite rank-1 vectors with equal feature dimensions and computes a standard vector Euclidean distance. The generated file explicitly labels this as a coding decision rather than recovered LaTeX:</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;108efd29-f941-4973-ba55-78c884f3a8b2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def euclidean_distance_eq_2(left: np.ndarray, right: np.ndarray) -&gt; float:
    """Return the vector Euclidean distance associated conceptually with eq_2.

    The paper's canonical eq_2 record has no LaTeX because its extraction is
    damaged. This vector-norm implementation is therefore an explicit coding
    decision, not a reconstruction of the missing equation.
    """</code></pre></div><h3>Random forest and XG-Boost candidate grids</h3><p>The generated random-forest adapter exposes <code>n_estimators</code>, <code>max_depth</code>, <code>max_features</code>, bootstrap behavior, and random-state settings. Here, <code>n_estimators</code> is the number of trees, <code>max_depth</code> limits tree depth, and <code>max_features</code> controls the feature subset considered by a split. The paper's Table 1 candidate values are returned exactly by <code>random_forest_candidates</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4ee272a2-2deb-4e94-99b8-f68c3d4f6f12&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def random_forest_candidates() -&gt; dict[str, tuple[object, ...]]:
    """Return the random-forest candidates listed in Table 1.

    These are candidate values reported by the paper, not selected values. The
    paper does not specify the search procedure or the configuration ultimately
    used for its reported results.
    """

    return {
        "n_estimators": (50, 100, 200),
        "max_depth": (3, 5, 7),
        "max_features": ("sqrt", "log2"),
    }</code></pre></div><p>For XG-Boost, <code>learning_rate</code> controls the contribution of each boosting step, <code>n_estimators</code> is the number of boosting estimators, and <code>max_depth</code> controls the depth of the tree learners. The paper lists <code>max_depth</code> values of <code>3</code>, <code>4</code>, and <code>5</code>; <code>learning_rate</code> values of <code>0.1</code>, <code>0.01</code>, and <code>0.001</code>; and <code>n_estimators</code> values of <code>50</code>, <code>100</code>, <code>150</code>, <code>500</code>, and <code>1000</code>. The <code>xgboost_candidates</code> function preserves those values, while <code>XGBoostRegressor</code> keeps the package dependency, objective, random state, and job count explicit implementation choices.</p><p>The combined grid interface makes the distinction auditable:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;62121604-ce83-4f31-8438-202dca73d5ed&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def paper_candidate_grids() -&gt; dict[str, dict[str, tuple[object, ...]]]:
    """Return the Random Forest and XG-Boost candidate grids from Table 1.

    Returns
    -------
    dict[str, dict[str, tuple[object, ...]]]
        Fresh mappings containing the paper-listed candidate values under the
        keys ``"random_forest"`` and ``"xgboost"``.

    No selected hyperparameters are inferred, and no tuning procedure is
    performed here. The values are obtained from the owning model modules so
    that the candidate definitions remain centralized.
    """</code></pre></div><p><code>paper_candidate_grids</code> is therefore a fidelity utility, not a claim that tuning has happened. The separate tuning layer can choose a search and validation policy, but those choices are absent from the paper and must be recorded if used.</p><h3>Constructing models without hiding their identity</h3><p><code>build_regressor</code> accepts canonical names such as <code>"svr"</code>, <code>"mlpr"</code>, <code>"knn"</code>, <code>"random_forest"</code>, and <code>"xgboost"</code>, along with selected aliases such as <code>"XG-Boost"</code> and <code>"random forest"</code>. It passes the supplied configuration directly to the corresponding adapter. It rejects LSTM and ensemble requests because those require rank-3 sequence inputs and separate orchestration, respectively.</p><p>A small comparison setup can therefore be written as follows:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d177fe17-2687-470e-a306-23741e964f21&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from stock_forecasting.models.factory import build_regressor

model_configs = {
    "svr": {"kernel": "rbf", "C": 1.0},
    "mlpr": {"hidden_layer_sizes": (32,), "random_state": 7},
    "knn": {"n_neighbors": 5, "weights": "distance"},
    "random_forest": {
        "n_estimators": 100,
        "max_depth": 5,
        "max_features": "sqrt",
        "random_state": 7,
    },
    "xgboost": {
        "max_depth": 3,
        "learning_rate": 0.1,
        "n_estimators": 100,
        "random_state": 7,
    },
}

models = {
    name: build_regressor(name, config)
    for name, config in model_configs.items()
}

predictions = {}
for name, model in models.items():
    model.fit(X_train, y_train)
    predictions[name] = model.predict(X_eval)</code></pre></div><p>In this worked example, <code>X_train</code> and <code>y_train</code> must come from the earlier chronological preprocessing stage, and <code>X_eval</code> must use the identical feature order and learned preprocessing state. Each entry in <code>predictions</code> should contain one value for every row in <code>X_eval</code>. The particular settings above are illustrative implementation choices; they are not reported as the paper's selected hyperparameters.</p><h3>Bagging versus boosting</h3><p>Random forest and XG-Boost are both ensembles, but they combine trees differently. Random forest uses bagging: multiple trees are trained with randomized data or feature choices, and their predictions are aggregated to reduce dependence on any one tree. XG-Boost uses sequential boosting: later learners are fitted in response to the current ensemble's errors. This distinction matters when interpreting the models and when documenting tuning parameters. A larger forest changes the number of independently randomized trees, while a larger boosted estimator count extends a sequence of corrective learners.</p><p>The <code>classical_regressors_forward</code> and <code>fit_models_and_generate_predictions</code> method contracts keep these baseline predictions separate until evaluation or explicit ensemble construction. That separation prevents the tabular adapters from silently deciding how the paper's later RF + XG-Boost + LSTM combination should work. The paper calls that model stacked, but its actual wiring remains unspecified; alignment and combination are handled in the following ensemble stage.</p><p>Finally, the generated files and planned tests were not executed or independently verified under this run policy. The candidate grids and adapter contracts described here are implementation artifacts intended for static and semantic review, not evidence that the paper's Table 2 values have been reproduced.</p><h2>The two-layer LSTM regressor</h2><p>How is an LSTM different from the tabular regressors described earlier? A tabular model receives one row of features at a time. An LSTM receives an ordered window of rows, so it can model relationships across time. In this implementation, one input has rank 3 and shape <code>(batch, sequence_length, n_features)</code>: <code>batch</code> is the number of sequences processed together, <code>sequence_length</code> is the look-back window, and <code>n_features</code> is the number of selected input columns.</p><p>The paper specifies two LSTM layers, dropout of <code>0.2</code>, a dense layer with 25 units, the Adam optimizer, MSE loss, batch size 32, and a maximum of 50 epochs. It does not specify the look-back length, recurrent-unit count, exact return-sequence settings, validation behavior, or final output-layer details. The generated implementation therefore records those missing values as implementation choices. It uses 32 recurrent units, returns sequences from the first layer, uses the second layer's final representation, and ends with a linear single-output layer for regression.</p><h3>What the extracted gate equations tell us</h3><p>An LSTM cell uses gates to regulate information. The input gate controls information entering the memory, the forget gate controls information retained from the previous state, and the output gate controls information exposed as the current representation. The paper supplies one displayed expression for each of these gate activations, but not a complete recurrence.</p><p>The first supplied expression represents the input-gate activation. It is reproduced exactly from canonical record <code>eq_8</code>.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;iga = &#963; Wip [ht&#8722;1, Xc] + bi&quot;,&quot;id&quot;:&quot;D54DB8BFA2&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here, <code>iga</code> is the input-gate activation, represented as a vector or tensor. <code>&#963;</code> is the elementwise sigmoid activation. <code>Wip</code> is the input-gate weight matrix, although its dimensions are not supplied. <code>ht&#8722;1</code> is the previous hidden state, and <code>Xc</code> is the current input as printed in the paper; both are described as vectors or tensors. <code>bi</code> is the input-gate bias vector. In practical terms, the gate produces one value per gate unit, conventionally between 0 and 1 because of the sigmoid.</p><p>The helper <code>input_gate_eq_8</code> in <code>src/stock_forecasting/models/lstm.py</code> gives this record a named implementation boundary. It does not claim that the extracted expression is a complete cell equation. The helper validates rank-2 step tensors, concatenates the previous hidden state and current input, applies a matrix multiplication and bias, and then applies a sigmoid through TensorFlow.</p><p>The forget-gate record is <code>eq_9</code>.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;fga = &#963; Wf g [ht&#8722;1, Xc] + b f&quot;,&quot;id&quot;:&quot;3A2D9D7CDB&quot;}" data-component-name="LatexBlockToDOM"></div><p><code>fga</code> denotes the forget-gate activation. <code>&#963;</code> again denotes the elementwise sigmoid. <code>Wf g</code> is the printed forget-gate weight matrix, while <code>ht&#8722;1</code> and <code>Xc</code> denote the previous hidden state and current input. <code>b f</code> is the forget-gate bias vector. As with the input gate, the expression describes a gate activation rather than the complete update of the LSTM cell state.</p><p>The generated function <code>forget_gate_eq_9</code> maps this paper record to the same framework-level affine-and-sigmoid pattern as the input gate. The separate function name is useful for traceability: readers can connect the implementation to <code>eq_9</code> without treating the OCR-damaged notation as a complete executable specification.</p><p>The output-gate record is <code>eq_10</code>.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;Opg = &#963; Wop [ht&#8722;1, Xc] + bo&quot;,&quot;id&quot;:&quot;645F5C68C7&quot;}" data-component-name="LatexBlockToDOM"></div><p><code>Opg</code> is the output-gate activation, <code>&#963;</code> is the sigmoid, <code>Wop</code> is the output-gate weight matrix, and <code>bo</code> is the output-gate bias. The same <code>ht&#8722;1</code> and <code>Xc</code> symbols represent the previous hidden state and current input. The supplied record gives no matrix dimensions, candidate-cell expression, cell-state update, or hidden-state update.</p><p>The corresponding <code>output_gate_eq_10</code> helper therefore exposes only the gate-level mapping. The actual recurrent computation is delegated to Keras rather than reconstructed from incomplete extracted equations. This is important: implementing the missing equations from memory would create a new mathematical specification rather than reproduce the supplied record.</p><h3>Building the two-layer network</h3><p>The core network construction is concentrated in <code>build_lstm_regressor</code>. The excerpt below is copied from <code>src/stock_forecasting/models/lstm.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;226d35af-4fed-42de-a40a-3c6be71b7eb3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    inputs = keras.Input(shape=input_shape, name="stock_sequence")
    # Table 1: first LSTM layer, followed by dropout 0.2 in the paper configuration.
    recurrent = keras.layers.LSTM(units, return_sequences=True, name="lstm_1")(inputs)
    recurrent = keras.layers.Dropout(float(dropout), name="dropout_1")(recurrent)
    # Table 1: second LSTM layer; returning its final state is an explicit wiring choice.
    recurrent = keras.layers.LSTM(units, return_sequences=False, name="lstm_2")(recurrent)
    recurrent = keras.layers.Dropout(float(dropout), name="dropout_2")(recurrent)
    dense = keras.layers.Dense(dense_units, activation="relu", name="dense_25")(recurrent)
    outputs = keras.layers.Dense(1, activation="linear", name="price_output")(dense)
    return keras.Model(inputs=inputs, outputs=outputs, name="two_layer_lstm_regressor")</code></pre></div><p>The input shape excludes the batch dimension, so <code>input_shape</code> is <code>(sequence_length, n_features)</code>. Keras adds the batch dimension at runtime, producing rank-3 inputs. The first LSTM uses <code>return_sequences=True</code>, which preserves one recurrent output for every position in the window and supplies a sequence to the second LSTM. The second layer uses <code>return_sequences=False</code>, so it returns one final representation per input sequence. The generated file explicitly labels this return-sequence wiring as a choice because the paper does not state it.</p><p>Dropout is applied after each LSTM layer using the configured probability. The paper specifies <code>0.2</code>; the code keeps the value configurable so that the unresolved choice is visible rather than hidden. The dense layer has 25 units, matching Table 1, and uses a ReLU activation. The final dense layer has one unit and a linear activation. That final linear output is a reasonable regression decision, but the paper does not explicitly provide the output-layer activation.</p><p>The recurrent-unit count is another implementation decision. The generated <code>LSTMRegressor</code> defaults to 32 units per LSTM layer, but this number should not be described as recovered from the paper. The paper specifies the number of layers and the dense width, not the number of recurrent units.</p><h3>Training with MSE and Adam</h3><p>The LSTM adapter enforces the sequence contract before training. Its <code>X</code> input must be a numeric rank-3 array with shape <code>(batch, sequence_length, n_features)</code>. Its target <code>y</code> may have shape <code>(batch,)</code> or <code>(batch, 1)</code>, but it must contain exactly one target per sequence. The adapter reshapes targets to <code>(batch, 1)</code> before passing them to Keras.</p><p>The training configuration is visible in this excerpt from <code>LSTMRegressor.fit</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;746f8bd1-cf71-47e3-aab6-b25aaa9f91a9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        self.model_.compile(optimizer="adam", loss="mse")
        fit_kwargs: dict[str, Any] = {
            "x": X_array,
            "y": y_array,
            "epochs": self.max_epochs,
            "batch_size": self.batch_size,
            "verbose": self.verbose,
            "shuffle": False,
        }
        if self.validation_fraction is not None:
            fit_kwargs["validation_split"] = float(self.validation_fraction)
        self.history_ = self.model_.fit(**fit_kwargs)</code></pre></div><p>The optimizer and loss match the paper's stated settings. The adapter caps <code>max_epochs</code> at 50 and defaults <code>batch_size</code> to 32. Shuffling is disabled as an implementation safeguard for ordered sequences. A validation split is optional and is not a paper-specified value. If used, it must be recorded in the experiment manifest rather than mistaken for a setting recovered from Section 4.</p><p>MSE is the paper-specified neural loss, but the supplied equation records contain no canonical MSE LaTeX. Consequently, <code>src/stock_forecasting/losses.py</code> does not display or invent a formula. <code>mean_squared_error_loss</code> uses TensorFlow's squared-difference operation followed by mean reduction, while <code>validate_loss_inputs</code> checks that prediction and target tensors have matching shapes and finite floating-point values.</p><h3>Worked shape example</h3><p>Suppose preprocessing selects four features, chooses a look-back of 20 observations, and creates a batch of 32 sequences. The LSTM input then has shape <code>(32, 20, 4)</code>. Each sequence corresponds to one future target, so the target array has shape <code>(32, 1)</code> during model training. The network returns <code>(32, 1)</code>, meaning one continuous forecast for each sequence in the batch.</p><p>The look-back value of 20 in this example is not supplied by the paper; it is an explicit reproduction decision. The sequence builder in <code>src/stock_forecasting/data/sequences.py</code> preserves target timestamps and ensures that each window precedes its target. Initial observations are lost because a complete look-back window must exist before the first sequence can be formed. Those shortened sequence indices must later be aligned with the random-forest and XG-Boost predictions before ensemble construction.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/p/quant-trading-stacked-ensemble-machine?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/p/quant-trading-stacked-ensemble-machine?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/onepagecode.substack.com/p/quant-trading-stacked-ensemble-machine?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><h3>Why the framework implementation matters</h3><p>A hand-written LSTM cell would need the candidate-cell computation, cell-state update, and hidden-state update in addition to the three gate expressions. Those pieces are absent from <code>eq_8</code>, <code>eq_9</code>, and <code>eq_10</code>, and the extracted notation itself contains unresolved formatting such as <code>Xc</code> and spaced subscripts. The generated code therefore uses TensorFlow/Keras for the complete recurrent mechanism and keeps the three named helpers as traceability mappings, not as a claim of a newly reconstructed cell.</p><p>This choice also preserves the boundary between paper facts and implementation decisions. The two LSTM layers, dropout <code>0.2</code>, dense width 25, Adam, MSE, batch size 32, and maximum 50 epochs come from the paper's Table 1 description. The recurrent units, look-back, final linear output, return-sequence behavior, dropout placement, validation split, and scaling behavior are choices required to make a runnable framework. The generated files document those choices, but no LSTM training, code execution, syntax check, or independent verification was performed under the run policy.</p><h2>Training orchestration and leakage-aware tuning</h2><p>How can model selection remain fair when stock observations arrive in time order? Treat the test set as a final exam: it may be used to measure the finished models, but not to choose their settings. The generated implementation separates three roles:</p><ul><li><p><code>X_train</code> and <code>y_train</code> are the features and targets available for fitting.</p></li><li><p><code>X_validation</code> and <code>y_validation</code> are a later portion of the training data used to compare candidate configurations.</p></li><li><p><code>X_test</code> and <code>y_test</code> are the held-out observations reserved for final evaluation.</p></li></ul><p>The paper describes hyperparameter tuning and lists candidate values for random forest and XG-Boost, but it does not specify the search algorithm, validation fraction, scoring rule, random seed, or repeat count. Those are therefore implementation decisions, not recovered paper facts. The primary reproduction choice still follows Table 1's chronological 80/20 split; the paper's generic pipeline separately mentions 75/25, so the split policy must remain documented.</p><h3>Training settings are explicit configuration</h3><p>The <code>TrainingConfig</code> class records the settings that affect fitting. The paper specifies batch size 32 and a maximum of 50 LSTM epochs. The seed and validation fraction are not supplied by the paper and are consequently caller choices.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2fd844ca-3232-4b5d-997c-57cb91ddebfe&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True, slots=True)
class TrainingConfig:
    """Record reproducibility and neural-training settings.

    Batch size 32 and a maximum of 50 epochs come from Table 1.  The seed and
    validation fraction are implementation decisions because the paper does
    not specify them.  ``max_epochs`` is capped at 50 so this configuration
    cannot silently exceed the reported neural-training limit.
    """

    seed: Optional[int]
    batch_size: int = 32
    max_epochs: int = 50
    validation_fraction: Optional[float] = None</code></pre></div><p>A fitted model is an estimator that has already learned parameters from permitted training rows and can respond to <code>predict</code>. A <code>PredictionBundle</code> is different: it associates prediction values with their timestamps. That distinction matters because a fitted estimator has model state, whereas a prediction bundle is an auditable evaluation artifact. The runner returns both through <code>ExperimentPredictions</code>.</p><p>For an illustrative run, a caller can choose a nonnegative seed, retain the paper-aligned batch size and epoch cap, and provide a training-only validation fraction such as <code>0.2</code>. That choice does not reproduce an unstated paper protocol; it makes the validation policy inspectable. If <code>validation_fraction</code> remains <code>None</code>, the generated runner can fit configured models without invoking the optional tabular search interface.</p><h3>Candidate search uses training data only</h3><p>The generated <code>search_regressor</code> function makes exhaustive Cartesian-grid search an explicit implementation decision. It first splits the supplied training arrays chronologically, evaluates each candidate on the validation suffix, selects the candidate with the smallest supplied score, and refits the selected estimator on all supplied training rows. The function does not accept test data.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;fb4f5996-fa11-4e25-891f-fc1f0a55e498&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def search_regressor(
    factory: Callable[[Mapping[str, Any]], RegressorProtocol],
    grid: Mapping[str, Sequence[Any]],
    X_train: np.ndarray,
    y_train: np.ndarray,
    validation_fraction: float,
    scorer: Callable[..., float],
) -&gt; SearchResult:
    """Select and refit a regressor using a chronological training split.

    Candidates are evaluated on the final ``validation_fraction`` of the
    supplied training rows, while the earlier rows are used for candidate
    fitting.  The candidate with the smallest score is selected, which is an
    explicit implementation decision appropriate for RMSE and other
    loss-style scorers.  The paper lists candidate values but does not specify
    this search algorithm, score direction, or validation protocol.

    The function intentionally has no test-data parameter.  Consequently,
    held-out test targets cannot enter candidate selection through this API.
    After selection, a fresh estimator is fitted on all ``X_train``/``y_train``
    rows so callers receive a model trained on the complete permitted sample.
    """</code></pre></div><p>Here, <code>validation_fraction</code> determines how much of the supplied training suffix is reserved for candidate comparison. The <code>scorer</code> is a callable that receives true and predicted validation values; the generated runner supplies an RMSE scorer, for which lower is better. <code>factory</code> constructs a fresh estimator for each parameter combination, preventing one candidate's fitted state from carrying into another candidate.</p><p>The returned <code>SearchResult</code> records <code>selected_params</code>, <code>validation_score</code>, and the refitted <code>estimator</code>. In a worked audit, readers should inspect <code>selected_params</code> and the score used for selection, but should not interpret that score as a held-out test result. The paper supplies random-forest candidates for estimators, depth, and feature selection, and XG-Boost candidates for depth, learning rate, and estimator count; it does not state which values won.</p><h3>The runner coordinates model families</h3><p><code>fit_models_and_generate_predictions</code> is the orchestration method for one prepared stock dataset. It fits tabular models on rank-2 arrays, optionally fits the sequence model on rank-3 arrays, and creates prediction bundles for the appropriate test timestamps. Stock-specific fitting is the default: the caller supplies one prepared dataset per security rather than combining TAINIWALCHM and AGROPHOS into a single training operation.</p><p>The central isolation rule is visible in the runner's public contract:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;9ee79449-0e2a-4c07-9f78-2d0f18a5564f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def fit_models_and_generate_predictions(
    prepared: Any,
    model_configs: Mapping[str, Any],
    training_config: TrainingConfig,
) -&gt; ExperimentPredictions:
    """Fit comparison models using training data and produce held-out predictions.

    The test partition is passed only to ``predict``.  Optional searches use a
    chronological validation suffix of the training partition, as required by the
    leakage-aware search interface.  Stock-specific fitting is represented by the
    caller supplying one prepared dataset per stock.
    """</code></pre></div><p>The runner obtains <code>tabular_split</code> and <code>sequence_split</code> from <code>prepared</code>. It validates that tabular features are rank 2, that their row counts match their targets, and that predictions have one value per evaluation row. For each configured classical model, <code>_fit_tabular_model</code> either fits directly or invokes <code>search_regressor</code>; the test arrays are not passed into either fitting or tuning.</p><p>The LSTM receives <code>sequence_X_train</code> and <code>sequence_X_test</code>, not the tabular arrays. Its settings are passed from <code>TrainingConfig</code>, including <code>max_epochs</code>, <code>batch_size</code>, <code>validation_fraction</code>, and <code>seed</code>. Thus the paper's maximum of 50 epochs and batch size 32 remain visible in the training record. The recurrent architecture still contains choices that the paper does not specify, including the look-back length and recurrent unit count established elsewhere in preprocessing and model configuration.</p><h3>Sequence predictions require additional alignment</h3><p>A look-back window means that the first observations may not have enough historical rows to form an LSTM input. Consequently, the sequence test partition can contain fewer timestamps than the tabular test partition. The runner intersects the tabular and sequence timestamps before fitting the ensemble and before producing its combined evaluation output.</p><p>This is not merely bookkeeping. If <code>X_test</code> contains one row for a timestamp that the LSTM cannot represent, combining its prediction with a neighboring LSTM prediction would compare different target times. The runner therefore creates common training rows and common evaluation rows by timestamp, checks that aligned tabular and sequence targets agree, and preserves chronological order.</p><p>The final <code>ExperimentPredictions</code> object contains fitted <code>models</code>, named <code>predictions</code>, a <code>target</code> bundle, <code>selected_configs</code>, and <code>training_metadata</code>. The metadata records such details as the chronological split policy, validation fraction, seed, epoch cap, batch size, and the fact that test targets were not used for fitting. This is the practical audit trail for the worked example: inspect the selected-configuration records without reporting any numerical value as verified.</p><h3>Manifest and reproducibility limits</h3><p><code>ExperimentManifest</code> complements the runner output by recording the data source, feature columns, target column, split policy, scaler, look-back, horizon, seed, tuning method, model configurations, and ensemble choice. Its purpose is to prevent a later reader from mistaking a default for a paper fact. In particular, the manifest can state that Table 1 is the primary reproduction target while preserving the conflicting 75/25 recommendation from the generic pipeline.</p><p>The paper does not specify early stopping, so the generated configuration should not be described as reproducing an early-stopping protocol. It also does not specify whether repeated trials were averaged, whether validation was chronological, or how ties between candidates were handled. The generated exhaustive chronological search and lower-is-better scoring rule are deliberate implementation policies. A recorded seed improves auditability, but it cannot guarantee stable model selection across libraries, hardware, dependency versions, or data snapshots.</p><p>Chronological validation is preferable here to ordinary random cross-validation because random folds can place later observations in the fitting portion while earlier observations appear in validation. That ordering would not represent the intended forecasting direction. Even chronological validation remains an approximation: market regimes can change, and a single validation suffix may not represent all future conditions.</p><p>Finally, no code execution, API request, training run, test run, syntax check, or semantic code verification occurred in this pipeline. The search and runner behavior described here is the generated design and contract, not an observed execution result. The paper's Table 2 metrics therefore remain reported reference values rather than independently verified outcomes.</p><h2>Aligning and combining the ensemble</h2><p>How can three different models contribute to one stock-price forecast without accidentally combining predictions for different dates? The practical rule is simple: first align every prediction by timestamp, then combine the aligned values. This matters especially for the Random Forest + XG-Boost + LSTM ensemble because the tabular models can predict every evaluation row, while the LSTM loses observations when it creates look-back windows.</p><p>The paper's proposed method combines random forest, XG-Boost, and LSTM predictions, but it does not specify the actual wiring. It calls the model &#8220;stacked&#8221; and also discusses assigning higher weights to models with lower loss. These descriptions do not identify one reproducible operation. A weighted average is a blending operation: fixed or externally supplied weights combine component outputs directly. Stacking usually means that a second-stage model learns how to combine component predictions. The generated code supports both alternatives and records which one was selected.</p><h3>Prediction bundles preserve identity and time</h3><p>A <code>PredictionBundle</code> connects a model name, a pandas index, and a prediction array. Its values may have shape <code>(n,)</code> or <code>(n, 1)</code>, but they always represent one prediction per timestamp. The index must not contain duplicates, and the number of labels must equal the number of prediction rows. This is more than defensive programming: without the index, a numerically valid vector can still be paired with the wrong dates.</p><p>For example, the three ensemble components are named <code>random_forest</code>, <code>xgboost</code>, and <code>lstm</code>. Their predictions should represent the same target column and target scale. The LSTM bundle may have fewer timestamps because its first predictions require a complete look-back window.</p><p>The alignment module represents the result with <code>AlignedPredictions</code>. Its <code>predictions</code> field is a rank-2 matrix with shape <code>(n_aligned_samples, n_components)</code>. For this ensemble, <code>n_components</code> is three, with columns ordered as random forest, XG-Boost, and LSTM. Its <code>target</code> field has shape <code>(n_aligned_samples,)</code>, and its index contains the common timestamps in target order.</p><p>The core implementation performs an inner intersection while preserving the target's order:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;faece91f-cca1-4ae1-985f-ad8ee9f20073&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def align_prediction_bundles(
    bundles: Sequence[PredictionBundle], target: PredictionBundle
) -&gt; AlignedPredictions:
    """Inner-align component predictions and targets by timestamp.

    Args:
        bundles: Component prediction bundles, typically RF, XG-Boost, and
            LSTM predictions. Every bundle must have unique timestamps.
        target: Target values indexed by their evaluation timestamps.

    Returns:
        An :class:`AlignedPredictions` record whose rows refer to exactly the
        same timestamps across every component and the target.
    """
    ...</code></pre></div><p>The excerpt shows the public responsibility without reproducing the entire module. In the generated implementation, the function rejects empty component lists, repeated component names, duplicate indices, invalid numeric values, and an empty timestamp intersection. It obtains common timestamps from the target's original order rather than independently sorting each model's output. This preserves the temporal order used for evaluation.</p><h3>Worked example: one shorter LSTM stream</h3><p>Consider four evaluation dates: <code>2024-01-02</code>, <code>2024-01-03</code>, <code>2024-01-04</code>, and <code>2024-01-05</code>. Suppose random forest and XG-Boost produce four predictions, but the LSTM produces only three because the initial date was consumed by sequence construction. The following excerpt constructs that situation. It is an illustrative example of the public interfaces; it is not an execution result.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7fc1f82c-85d3-4415-bd34-6d818a2c88b6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import numpy as np
import pandas as pd

from stock_forecasting.ensemble.alignment import align_prediction_bundles
from stock_forecasting.types import PredictionBundle

index = pd.to_datetime([
    "2024-01-02", "2024-01-03", "2024-01-04", "2024-01-05"
])
lstm_index = index[1:]

target = PredictionBundle(
    "close",
    index,
    np.array([10.0, 11.0, 12.0, 13.0]),
)
rf = PredictionBundle(
    "random_forest",
    index,
    np.array([10.2, 10.8, 12.1, 12.7]),
)
xgb = PredictionBundle(
    "xgboost",
    index,
    np.array([10.1, 11.1, 11.9, 13.1]),
)
lstm = PredictionBundle(
    "lstm",
    lstm_index,
    np.array([10.9, 12.2, 12.9]),
)

aligned = align_prediction_bundles([rf, xgb, lstm], target)</code></pre></div><p>The common index is the final three dates. Conceptually, the resulting matrix is ordered by those dates and has three columns:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;11af4c9e-a5fa-4acb-a195-790cee88cec7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">             random_forest  xgboost  lstm
2024-01-03           10.8     11.1  10.9
2024-01-04           12.1     11.9  12.2
2024-01-05           12.7     13.1  12.9</code></pre></div><p>The important invariant is that each row refers to one date across all three columns. The first random-forest and XG-Boost predictions are discarded from the ensemble view because there is no corresponding LSTM prediction. The target is reduced to the same dates. This is preferable to padding the LSTM output or silently shifting its predictions.</p><h3>Direct weighted blending</h3><p>The generated <code>weighted_average</code> function implements one explicit interpretation of the paper's weight discussion. It requires a prediction matrix with shape <code>(n_samples, 3)</code> and a weight vector with shape <code>(3,)</code>. The weights must be finite and nonnegative, and the implementation normalizes them to sum to one. The returned prediction vector has shape <code>(n_samples,)</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4ff9a18e-8ec4-469f-a784-728cb717421a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def weighted_average(predictions: np.ndarray, weights: np.ndarray) -&gt; np.ndarray:
    """Combine aligned component predictions using normalized nonnegative weights.

    The paper mentions weighting lower-loss models but does not define weights or
    a fitting procedure.  This implementation therefore makes the combination
    explicit: the three supplied weights are required to be finite and
    nonnegative, then normalized to sum to one.  Columns must be ordered as
    random forest, XG-Boost, and LSTM, and the return value has shape ``(n,)``.
    """
    component_values = _validate_component_predictions(predictions)
    component_weights = _as_finite_float_array(weights, "weights")
    if component_weights.ndim != 1 or component_weights.shape[0] != _COMPONENT_COUNT:
        raise ValueError("weights must have shape (3,)")
    if np.any(component_weights &lt; 0):
        raise ValueError("weights must be nonnegative")
    total = float(np.sum(component_weights))
    if total &lt;= 0.0:
        raise ValueError("at least one weight must be positive")

    normalized_weights = component_weights / total
    return np.asarray(component_values @ normalized_weights, dtype=float)</code></pre></div><p>For equal weighting, a caller can provide <code>(1/3, 1/3, 1/3)</code>. That choice is not recovered from the paper; it is an implementation decision. Likewise, weights derived from validation losses would require a documented rule for converting losses into weights. The code deliberately does not infer such a rule.</p><h3>Learned linear stacking</h3><p>The alternative is <code>LinearStacker</code>, a second-stage linear regressor. It receives the three component predictions as meta-features and learns coefficients and an intercept from training-only component outputs. Here, &#8220;meta-features&#8221; means features supplied to the combiner rather than the original stock columns.</p><p>The fitting function makes the data boundary explicit:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;328d1b9a-50c5-4f2b-b806-ef03e5ca038b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def fit_linear_stacker(
    train_predictions: np.ndarray,
    y_train: np.ndarray,
) -&gt; LinearStacker:
    """Fit a training-only ordinary-least-squares linear stacker.

    The paper calls the ensemble stacked but does not specify a meta-learner.
    Ordinary least squares with an intercept is an explicit implementation
    choice.  ``train_predictions`` must contain exactly the three aligned
    component predictions, and ``y_train`` must contain the corresponding
    training targets; test targets must not be passed here.
    """
    ...</code></pre></div><p>The generated implementation validates the three-column input, checks that <code>y_train</code> has a matching number of rows, adds an intercept column, and solves the linear least-squares problem. Its <code>predict</code> method accepts another <code>(n_samples, 3)</code> matrix and returns one value per row. The important leakage constraint is not the particular linear model; it is that the stacker's fitting targets come only from permitted training data.</p><p>A rigorous stack often uses out-of-fold component predictions when fitting the meta-model. Otherwise, a stacker may learn from component predictions generated on the same rows used to fit those components, which can make the component outputs look unrealistically accurate. The paper does not specify out-of-fold generation, a validation arrangement, or even the meta-learner type. The generated framework therefore exposes ordinary least-squares stacking as a documented implementation choice rather than presenting it as the paper's recovered procedure.</p><h3>Orchestrating the three components</h3><p><code>StackedRFXGBoostLSTM</code> owns one random-forest regressor, one XG-Boost regressor, and one LSTM regressor. Its tabular inputs have rank 2, while its LSTM inputs have rank 3. During fitting, the caller must supply pre-aligned training rows to all three components. The class then creates a three-column training prediction matrix. If the configured combiner is <code>linear_stacking</code>, it fits the meta-model on that matrix and the training targets.</p><p>During prediction, the class generates component outputs, associates the shorter LSTM output with the final evaluation timestamps, and calls <code>align_prediction_bundles</code>. The final values are then produced either by <code>weighted_average</code> or by the fitted <code>LinearStacker</code>. The resulting <code>PredictionBundle</code> is named <code>rf_xgboost_lstm_ensemble</code> and carries the aligned timestamps.</p><p>A configuration makes the ambiguity visible at the call site:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a4f8f987-50de-4345-8cb1-aa1ac3022a18&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from stock_forecasting.config import EnsembleConfig
from stock_forecasting.ensemble.stacked import build_ensemble

ensemble_config = EnsembleConfig(
    combiner="weighted_average",
    weights=(1 / 3, 1 / 3, 1 / 3),
    meta_features="predictions",
)
ensemble = build_ensemble(ensemble_config)</code></pre></div><p>Switching to <code>combiner="linear_stacking"</code> changes the operation from direct blending to a learned second stage. The <code>combiner_metadata</code> property records the selected combiner, weights, meta-feature description, component order, and the fact that the paper's wiring was not specified. This provenance is essential: a future reader can distinguish a paper fact from a reproduction decision.</p><p>The generated implementation has not been executed, and neither its alignment behavior nor its numerical output has been independently verified under the run policy. The intended contracts are nevertheless explicit: common timestamps, matching target semantics and scale, finite predictions, three fixed component columns, and no held-out test targets used to fit a learned combiner.</p><h2>RMSE, R2, and auditable reporting</h2><p>How do we know whether a stock-price prediction is useful? Compare predictions with the correct held-out targets, using the same timestamps and the same value scale. The generated metrics layer implements this final step through <code>evaluate_regression_predictions</code>, while <code>build_metric_report</code> adds stock and model identity for repeatable reporting.</p><p>This section concerns regression evaluation. Let <code>y_true</code> denote the observed target values from the held-out test period, and let <code>y_pred</code> denote the corresponding model predictions. Both are numeric arrays with shape <code>(n_test,)</code> or <code>(n_test, 1)</code>. &#8220;Held-out&#8221; means that these observations were not used to fit the model or select its settings.</p><p>RMSE, or root mean square error, summarizes prediction error in the units of the target. If the target is an original stock price, RMSE is expressed in price units. If the target remains standardized, RMSE is expressed in standardized units instead. The implementation does not silently convert between these scales.</p><p>R2, or R-squared, is the coefficient of determination. It compares the model's residual error with the variation around a constant prediction based on the target mean. R2 can be negative when predictions are worse than that constant baseline; it is therefore incorrect to assume that every valid R2 must lie between zero and one.</p><h3>Alignment is part of the metric contract</h3><p>A numerically valid prediction is not necessarily a valid evaluation. The prediction must refer to the same observations as the target, in the same order. The metric function checks numeric dtype, finiteness, supported shapes, equal lengths, and optional index length and uniqueness. A pandas index can document timestamp alignment, although the caller remains responsible for constructing the two arrays from matching observations.</p><p>The core implementation normalizes one-dimensional and single-column inputs before calling standard regression metrics:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;9a5e2e15-b27e-401e-8986-cfdeeca51e36&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def _validate_metric_inputs(
    y_true: np.ndarray, y_pred: np.ndarray
) -&gt; tuple[np.ndarray, np.ndarray]:
    """Validate target/prediction shape, finiteness, and observation count."""
    true_values = _normalise_regression_values(y_true, "y_true")
    predicted_values = _normalise_regression_values(y_pred, "y_pred")
    if true_values.shape != predicted_values.shape:
        raise ValueError(
            "y_true and y_pred must have the same number of aligned observations; "
            f"got {true_values.shape[0]} and {predicted_values.shape[0]}"
        )
    return true_values, predicted_values</code></pre></div><p>The private helper <code>_validate_metric_inputs</code> converts accepted target shapes to compatible vectors and rejects mismatched lengths. It does not inspect the semantic content of timestamps, so the surrounding pipeline must preserve timestamp order when producing <code>y_true</code> and <code>y_pred</code>.</p><h3>Computing and labeling the metrics</h3><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/p/quant-trading-stacked-ensemble-machine?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/p/quant-trading-stacked-ensemble-machine?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/onepagecode.substack.com/p/quant-trading-stacked-ensemble-machine?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><p>The public <code>rmse</code> function returns a finite, nonnegative scalar. The public <code>r2_score</code> function requires at least two observations because the standard coefficient of determination is undefined for a one-observation target. <code>evaluate_regression_predictions</code> performs the common checks and returns a dictionary containing <code>rmse</code> and <code>r2</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a9748659-9319-47f9-95a0-0b76e9ed6b8f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def evaluate_regression_predictions(
    y_true: np.ndarray,
    y_pred: np.ndarray,
    index: pd.Index | None = None,
    scale: str = "original",
) -&gt; dict[str, float]:
    """Return RMSE and R2 after checking held-out prediction alignment."""
    if not isinstance(scale, str) or not scale.strip():
        raise ValueError("scale must be a nonempty string")

    true_values, predicted_values = _validate_metric_inputs(y_true, y_pred)
    _validate_index(index, true_values.shape[0])

    # Reuse the public metric functions after the common alignment checks.
    metrics: dict[str, float] = {
        "rmse": rmse(true_values, predicted_values),
        "r2": r2_score(true_values, predicted_values),
    }
    if not all(np.isfinite(value) for value in metrics.values()):
        raise ValueError("regression metrics must be finite")
    return metrics</code></pre></div><p>The <code>scale</code> argument is a reporting label, not a transformation instruction. Passing <code>scale="original"</code> means that the caller has already supplied original-scale values. Passing <code>scale="standardized"</code> records a different convention, but the function does not inverse-transform the arrays. If a scaler was used during preprocessing, the caller must apply the documented inverse target transformation before requesting original-price metrics.</p><h3>Worked example: two aligned prediction bundles</h3><p>A <code>PredictionBundle</code> associates a model name, a pandas index, and one prediction per index value. The following example creates a small target bundle and two model outputs. It is an interface example only; it was not executed in this pipeline and does not represent a paper result.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c0f07359-7eee-475d-a474-f3ebbd31a562&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import numpy as np
import pandas as pd

from stock_forecasting.metrics.regression import (
    evaluate_regression_predictions,
)
from stock_forecasting.types import PredictionBundle

index = pd.Index(
    pd.to_datetime(["2023-01-03", "2023-01-04", "2023-01-05"]),
    name="Date",
)
target = PredictionBundle(
    "observed",
    index,
    np.array([10.0, 11.0, 12.0]),
)
svr_prediction = PredictionBundle(
    "svr",
    index,
    np.array([10.2, 10.8, 12.1]),
)
ensemble_prediction = PredictionBundle(
    "ensemble",
    index,
    np.array([10.1, 11.1, 11.9]),
)

svr_metrics = evaluate_regression_predictions(
    target.values,
    svr_prediction.values,
    index=target.index,
    scale="original",
)
ensemble_metrics = evaluate_regression_predictions(
    target.values,
    ensemble_prediction.values,
    index=target.index,
    scale="original",
)</code></pre></div><p>Here, both prediction arrays have three values and use exactly the target's index. The <code>scale="original"</code> label states that these illustrative values are being interpreted in the target's original scale. No numerical output is asserted because execution was disabled.</p><h3>Building auditable records</h3><p>A dictionary of metrics is useful for one call, but a multi-stock comparison also needs to identify the security and model. <code>build_metric_report</code> accepts a mapping from model names to <code>PredictionBundle</code> objects and aligns each prediction to the target index in target order. It returns immutable <code>MetricRecord</code> objects containing <code>stock</code>, <code>model</code>, <code>rmse</code>, <code>r2</code>, and <code>scale</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;677a6092-4a6f-4f09-b24e-caf6d6765c91&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def build_metric_report(
    stock: str,
    prediction_map: Mapping[str, PredictionBundle],
    target: PredictionBundle,
    scale: str,
) -&gt; list[MetricRecord]:
    """Evaluate every named prediction bundle against one aligned target.

    The caller supplies the reporting scale explicitly, for example ``"original"``
    or ``"standardized"``.  This function does not perform inverse scaling.  It
    also does not infer a train/test split or configuration provenance because
    those fields are not part of :class:`MetricRecord`; callers should preserve
    the corresponding experiment manifest separately.
    """</code></pre></div><p>The alignment step uses the intersection of target and prediction timestamps. This is important when, for example, an LSTM has fewer evaluable rows because its look-back windows omit initial observations. An empty timestamp intersection is treated as an error rather than producing misleading metrics.</p><p>A report can then be serialized with <code>save_metric_report</code> to either CSV or JSON. Records are sorted deterministically by stock, model, and scale. The serializer intentionally does not invent split, feature, target, or preprocessing information; those details belong in the experiment manifest.</p><h3>Interpreting the paper's reported values</h3><p>The paper's Table 2 reports RMSE and R2 for TAINIWALCHM and AGROPHOS across SVR, MLPR, KNN, LSTM, random forest, XG-Boost, and the ensemble. Those values are reference claims from the paper, not outputs independently reproduced by this generated implementation. The paper reports the ensemble as having the strongest comparison values, but the exact data preparation, target scale, and ensemble wiring are not fully specified.</p><p>In particular, the reported random-forest RMSE values are much larger than the other listed RMSE values while the corresponding R2 values are high. That combination can occur when target variance or scale differs, but the supplied evidence does not establish the cause. The implementation therefore preserves the reported values as references and does not correct, reinterpret, or claim to explain them.</p><p>The meaningful comparison rule is consistent evaluation: every model for a stock should use the same held-out period, aligned target timestamps, and explicitly recorded scale. A model's RMSE and R2 should not be compared across differently transformed targets without qualification. No API call, model training, metric calculation, test run, or independent verification was performed for this article section.</p><h2>End-to-end offline workflow</h2><p>How can you understand the complete reproduction workflow without immediately depending on live Yahoo Finance data? Start with a deterministic synthetic table, pass it through the same public pipeline used for a local snapshot, and inspect the resulting configuration and provenance. This demonstrates the plumbing without suggesting that synthetic data reproduces the paper's market-data experiment.</p><p>The paper reports Yahoo Finance data for TAINIWALCHM and AGROPHOS, but it does not provide exact ticker identifiers, a data snapshot, all preprocessing choices, or the full ensemble wiring. Consequently, a successful local run is illustrative unless those missing details are recovered. The generated code is designed to make that limitation visible.</p><h3>One orchestration object for one experiment</h3><p>The end-to-end entry point is <code>run_experiment</code> in <code>src/stock_forecasting/experiments/pipeline.py</code>. An <code>ExperimentConfig</code> groups the data-source description, preprocessing policy, training settings, ensemble choice, model configurations, tuning description, and package metadata. An <code>ExperimentResult</code> then carries the prepared data, prediction bundles, metric records, and an <code>ExperimentManifest</code>.</p><p>This separation distinguishes the paper's reported facts from implementation decisions. For example, the Table 1 configuration motivates the chronological 80/20 split, MSE, Adam, batch size 32, and a maximum of 50 LSTM epochs. In contrast, the feature set, target column, look-back window, scaler, random seed, selected tree parameters, and combiner remain explicit choices because the paper does not specify them.</p><p>The central orchestration function is deliberately small because each stage belongs to another module:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f767c0c3-e92f-4953-a531-ac9cc463f798&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def run_experiment(
    frame: pd.DataFrame,
    stock_name: str,
    configs: ExperimentConfig,
) -&gt; ExperimentResult:
    """Prepare, fit, evaluate, and document one stock experiment.

    The supplied frame is assumed to be a local snapshot. This function does
    not make network requests. Preprocessing and model fitting are delegated to
    their owning modules; the runner's held-out predictions are aligned by the
    reporting layer before RMSE and R2 are computed. If optional TensorFlow or
    XG-Boost dependencies are unavailable, the owning model adapter's explicit
    import error is allowed to propagate, preserving the dependency boundary.

    No numerical reproduction of Table 2 is implied: the paper omits several
    choices required for an exact run, and this function performs no comparison
    against the paper's reported values.
    """

    _validate_frame(frame, stock_name)
    if not isinstance(configs, ExperimentConfig):
        raise TypeError("configs must be an ExperimentConfig")

    prepared = prepare_stock_data(frame, configs.preprocessing)
    fitted = fit_models_and_generate_predictions(
        prepared,
        configs.model_configs,
        configs.training,
    )
    reporting_predictions = _convert_to_reporting_scale(fitted, prepared)
    metrics = build_metric_report(
        stock_name,
        reporting_predictions.predictions,
        reporting_predictions.target,
        "original",
    )
    manifest = _build_manifest(stock_name, configs, reporting_predictions)
    return ExperimentResult(
        stock=stock_name,
        prepared=prepared,
        predictions=reporting_predictions,
        metrics=metrics,
        manifest=manifest,
    )</code></pre></div><p>Notice the order: the supplied table is validated, preprocessing is performed, models generate held-out predictions, predictions are converted to the configured reporting scale, metrics are built, and the manifest records the experiment. The function does not download data and does not compare results with Table 2. <code>ExperimentResult</code> is therefore an audit container, not evidence that the paper's reported numbers were reproduced.</p><h3>Worked example: deterministic synthetic input</h3><p>The generated <code>make_synthetic_stock_data</code> function creates a chronological OHLCV-shaped DataFrame using a local NumPy random generator. Its columns are <code>Open</code>, <code>High</code>, <code>Low</code>, <code>Close</code>, <code>Adj Close</code>, and <code>Volume</code>. The dates, stochastic process, starting price, and seed are implementation choices; they are not properties of the paper's Yahoo Finance datasets.</p><p>The offline demonstration calls the generator as follows:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2faf7a12-c382-4386-8cef-7f5877fdeee6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">
def run_demo() -&gt; None:
    """Generate synthetic data, run the configured experiment, and print results."""

    # The synthetic generator is the offline replacement for the paper's
    # Yahoo Finance acquisition step; no network request is made here.
    frame = make_synthetic_stock_data(n_rows=320, seed=7)
    configs = _demo_configuration()
    result = run_experiment(frame, stock_name="SYNTHETIC_DEMO", configs=configs)
    _print_result(result)</code></pre></div><p>This function illustrates the complete data flow: acquisition is replaced by a deterministic local fixture, <code>_demo_configuration()</code> makes the missing modeling choices explicit, and <code>run_experiment</code> handles preparation, fitting, ensemble construction, and reporting. The output is labeled illustrative in the script. It cannot reproduce Yahoo Finance results or the paper's Table 2 values.</p><p>The demonstration configuration chooses <code>Open</code>, <code>High</code>, <code>Low</code>, and <code>Volume</code> as features, <code>Close</code> as the target, a 20-observation look-back, a one-step horizon, standard scaling, and equal-weight prediction blending. These choices are useful for showing the interfaces, but they must not be described as recovered experimental settings. The Table 1 80/20 split and the stated LSTM training settings are the more direct paper-aligned choices; the generic pipeline's 75/25 recommendation remains a documented conflict.</p><p>A conceptual invocation from the repository root is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;bb8680d7-f00c-4441-a018-eac1d393a312&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">python scripts/run_offline_demo.py</code></pre></div><p>This command is an example of intended usage, not a command executed in this pipeline. The full comparison requires the relevant installed modeling dependencies, including TensorFlow and XG-Boost. Synthetic data generation itself uses NumPy and pandas, but that does not mean the complete experiment can run without the optional model packages.</p><h3>The command-line interface and explicit choices</h3><p><code>build_argument_parser</code> in <code>src/stock_forecasting/cli.py</code> exposes both synthetic and CSV pathways. The mutually exclusive source options are visible in this excerpt:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;352b44f1-5fde-46a2-8b45-9b271fbacc5c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    source = parser.add_mutually_exclusive_group(required=True)
    source.add_argument(
        "--csv",
        type=Path,
        help="Path to a local stock CSV snapshot.",
    )
    source.add_argument(
        "--synthetic",
        action="/__u/onepagecode.substack.com/store_true",
        help="Use deterministic synthetic OHLCV-like data for an offline demo.",
    )</code></pre></div><p>The CLI also exposes the unresolved choices rather than concealing them. Its defaults select <code>Close</code> as the target, use <code>Open</code>, <code>High</code>, <code>Low</code>, and <code>Volume</code> as features, choose an 80% training fraction, use a 20-row look-back, a one-step horizon, and standard scaling. The parser help explicitly identifies the look-back and scaler as choices absent from the paper. The ensemble option accepts <code>weighted_average</code> or <code>linear_stacking</code>, because the paper calls the model stacked while also discussing higher weights for lower-loss models without defining the actual operation.</p><p>For a local CSV snapshot, an intended command could be written as:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;2f0726de-9ce6-41b6-9e45-b1c5733739c6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">python -m stock_forecasting.cli \
  --csv data/stock_snapshot.csv \
  --date-column Date \
  --ticker EXPLICIT_TICKER \
  --features Open,High,Low,Volume \
  --target Close \
  --train-fraction 0.8 \
  --lookback 20 \
  --horizon 1 \
  --scaler standard \
  --combiner weighted_average \
  --weights 0.3333333333,0.3333333333,0.3333333333</code></pre></div><p>The exact ticker is intentionally supplied by the user. The paper's names and date ranges are not enough to guarantee a valid Yahoo Finance identifier. Likewise, <code>Close</code> is a configuration choice here, not a paper-established target. Before using real data, revisit the target, feature columns, endpoint dates, interval, scaling policy, sequence alignment, model settings, and ensemble combiner.</p><h3>Loading a local snapshot</h3><p><code>load_stock_csv</code> is an offline implementation addition. It does not claim that the paper used CSV files; it provides a way to work from a saved snapshot rather than relying on a network call. The loader parses the requested date column, rejects malformed or duplicate timestamps, and returns a deterministically sorted <code>DatetimeIndex</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d966d305-4000-4e1d-a098-b7edd8924cef&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    parsed_dates = pd.to_datetime(frame[date_column], errors="coerce")
    invalid_mask = parsed_dates.isna()
    if bool(invalid_mask.any()):
        invalid_rows = [int(position) for position in invalid_mask[invalid_mask].index[:5]]
        suffix = "" if int(invalid_mask.sum()) &lt;= 5 else " ..."
        raise ValueError(
            f"Date column {date_column!r} contains "
            f"{int(invalid_mask.sum())} unparseable value(s); "
            f"example row labels: {invalid_rows}{suffix}"
        )

    result = frame.drop(columns=[date_column]).copy()
    result.index = pd.DatetimeIndex(parsed_dates, name=date_column)
    result = result.sort_index(kind="mergesort")

    if result.index.has_duplicates:
        duplicate_count = int(result.index.duplicated(keep=False).sum())
        raise ValueError(
            f"CSV snapshot contains {duplicate_count} row(s) with duplicate "
            f"timestamps; resolve duplicates before loading: {csv_path}"
        )</code></pre></div><p>The important contract is provenance and ordering. The date column becomes the index, and duplicate timestamps are an explicit failure rather than being silently aggregated or discarded. Feature selection, null handling, zero-value policy, scaling, and target construction remain responsibilities of the preprocessing pipeline.</p><h3>What the output means</h3><p>The demonstration prints metric records and a manifest. A metric record identifies the stock, model, RMSE, R2, and reporting scale. The manifest records the choices needed to interpret those values, including the data-source label, feature and target configuration, split, sequence settings, model configurations, tuning description, and combiner.</p><p>This is particularly important for the paper's Table 2. The paper reports RMSE and R2 for TAINIWALCHM and AGROPHOS, including an ensemble result, but the supplied implementation has not been executed and the missing experimental details prevent an exact reconstruction. A locally produced metric is therefore an output of the chosen local configuration, not an independently verified Table 2 result. Even using real Yahoo Finance data would not remove the uncertainty about the original ticker symbols, selected features, scaler, look-back, validation protocol, selected hyperparameters, or ensemble wiring.</p><p>The generated CLI includes a warning in its JSON payload that says the outputs do not independently reproduce Table 2. That warning is part of the implementation's audit posture: it prevents an illustrative offline workflow from being mistaken for a verified paper replication.</p><h3>Practical dependency boundary</h3><p>The CSV and synthetic loaders are designed for offline data handling. The Yahoo Finance client is isolated in its own module, so the CLI does not make an implicit network request. However, the full model comparison includes a TensorFlow/Keras LSTM and an optional XG-Boost adapter. Those dependencies may be unavailable in an offline environment. A missing dependency should produce an explicit adapter-level error rather than silently replacing the model with another algorithm.</p><p>No API call, training run, test run, syntax check, or semantic code verification occurred for this tutorial section. The workflow and commands describe the generated interfaces and intended behavior. They should be treated as a reproducible starting point whose outputs require local execution, matching data, and careful interpretation before they can be compared with the paper's reported references.</p><h2>Fidelity ledger, limitations, and responsible interpretation</h2><p>How can a codebase be reproducible without being an exact reproduction? The distinction is evidence. A <strong>paper fact</strong> is a setting or result explicitly supplied by the paper. An <strong>implementation decision</strong> is a choice made because the paper leaves a detail open. An <strong>ambiguity</strong> is an unresolved detail that prevents a unique reconstruction. A <strong>reported reference</strong> is a value transcribed from the paper, not a result independently produced by this project.</p><p>For this paper, the Section 4 comparison is the main reproduction target: SVR, MLPR, KNN, random forest, XG-Boost, LSTM, and the Random Forest + XG-Boost + LSTM ensemble. The generated framework preserves that scope, but it does not claim exact numerical reproduction. The paper does not provide source code, a data snapshot, complete preprocessing instructions, or the ensemble's actual wiring.</p><h3>What is known and what remains a decision</h3><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><strong>Paper facts.</strong> Table 1 specifies an 80/20 training/test split for the proposed ensemble, MSE and Adam for the neural component, a maximum of 50 epochs, batch size 32, two LSTM layers, dropout <code>0.2</code>, and a dense layer with 25 units. It also lists candidate values for random-forest and XG-Boost parameters. The paper reports Yahoo Finance data for TAINIWALCHM and AGROPHOS and evaluates models using RMSE and R2.</p><p><strong>Derived implementation rationale.</strong> The framework uses a chronological split because later observations should represent the held-out period. It fits learned preprocessing and any learned combiner using training data only. It aligns predictions by timestamp because LSTM sequence construction can remove initial rows. These are safeguards and design explanations, not evidence that the original experiment used precisely the same procedures.</p><p><strong>Implementation decisions and ambiguities.</strong> The exact Yahoo Finance ticker identifiers, selected feature columns, target column, scaler, look-back window, forecast horizon, validation procedure, random seed, selected hyperparameters, and ensemble combiner are not specified. The code therefore exposes them through configuration and records them in <code>ExperimentManifest</code> instead of hiding them in defaults. The generic pipeline also mentions a 75/25 split, which conflicts with Table 1's more specific 80/20 setting; this implementation treats 80/20 as the primary reproduction choice and records the conflict.</p><p>The paper's dataset description contains another fidelity issue: AGROPHOS is associated with the 2018&#8211;2023 period, but one sentence incorrectly uses the TAINIWALCHM name for that description. The framework does not silently repair the source text. The stock label and the API ticker remain explicit fields so a user can document the chosen interpretation.</p><h3>The manifest as an audit boundary</h3><p>The manifest separates a stock label from the data-source ticker and stores preprocessing, training, tuning, and ensemble choices. A compact excerpt from <code>src/stock_forecasting/experiments/schema.py</code> shows the fields that prevent an apparently reproducible run from becoming an undocumented guess:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f1ba5ddc-afb2-4db2-97d5-a05dde502a7d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True, slots=True)
class ExperimentManifest:
    """Audit record for one stock-specific reproduction configuration.

    The manifest deliberately stores unresolved paper details as explicit fields or
    notes rather than hiding them in defaults.  ``model_configurations`` records
    selected settings when supplied; an empty mapping means that no selected model
    settings were provided by the caller or recovered from the paper.
    """

    paper_id: str
    stock: str
    data_source: Mapping[str, Any]
    feature_columns: tuple[str, ...]
    target_column: str
    train_fraction: float
    split_policy: str
    scaler_name: str | None
    lookback: int
    horizon: int
    seed: int | None
    tuning_method: str
    validation_fraction: float | None
    model_configurations: Mapping[str, Any]
    combiner: str
    ensemble_weights: tuple[float, ...] | None
    meta_features: str</code></pre></div><p>The important point is not the dataclass syntax itself. <code>feature_columns</code> and <code>target_column</code> make the supervised task explicit; <code>lookback</code> and <code>horizon</code> identify the temporal construction; <code>scaler_name</code> and <code>train_fraction</code> expose preprocessing choices; and <code>tuning_method</code>, <code>model_configurations</code>, <code>combiner</code>, and <code>ensemble_weights</code> record choices the paper does not recover. <code>validate_manifest</code> is responsible for rejecting incomplete or invalid records, such as an empty target name, an invalid fraction, duplicate features, or an ensemble weight vector with the wrong length.</p><p>A compact fidelity ledger for one hypothetical run could therefore read as follows:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;97664212-68c0-445d-8abd-8c5e2bbb903e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">Paper fact: Table 1 specifies an 80/20 split and a 25-unit dense layer.
Implementation decision: use Close as target and Open, High, Low, Volume as features.
Unresolved ambiguity: the paper does not define whether the ensemble is weighted blending or learned stacking.
Reported reference: Table 2 lists the ensemble's AGROPHOS RMSE as 1.2658 and R2 as 0.9897; this value is not independently verified here.</code></pre></div><p>This record keeps a chosen target separate from a paper-specified training ratio, and it keeps a transcribed number separate from an observed output of the generated code.</p><h3>Damaged and incomplete equations</h3><p>The equation registry in <code>src/stock_forecasting/equations.py</code> is deliberately conservative. For KNN, <code>eq_2</code> has an empty canonical LaTeX string because the extracted Euclidean-distance expression is damaged. The code may implement a vector distance for the KNN adapter, but that is an implementation decision and must not be presented as reconstructed paper notation. The registry preserves the empty value:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a9f57f7c-0573-4be5-980e-4ae19e2adf4c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">"eq_2": CanonicalEquation(
    equation_id="eq_2",
    latex="",
    meaning="Euclidean-distance expression for KNN.",
),</code></pre></div><p>The <code>CanonicalEquation</code> record stores an identifier, the supplied LaTeX text, and its meaning. <code>get_equation</code> retrieves that immutable record without normalizing OCR artifacts. This makes <code>eq_2</code> traceable while respecting the rule that missing mathematical text must not be invented.</p><p>The paper's LSTM gate records are intact enough to preserve as supplied, but they do not define a complete LSTM recurrence. They omit the candidate-cell computation and the cell-state and hidden-state updates. The framework therefore uses a Keras LSTM for the complete trainable model and treats the following records as conceptual gate mappings only.</p><p>The input-gate record is introduced here to identify the information-control operation printed by the paper:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;iga = &#963; Wip [ht&#8722;1, Xc] + bi&quot;,&quot;id&quot;:&quot;82C1C6A970&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here, <code>iga</code> is the input-gate activation, <code>&#963;</code> is an elementwise sigmoid, <code>Wip</code> is the printed input-gate weight matrix, <code>ht&#8722;1</code> is the previous hidden state, <code>Xc</code> is the current input as printed, and <code>bi</code> is the input-gate bias. The paper does not specify matrix dimensions, so these symbols represent compatible vectors, tensors, and parameters rather than a recoverable complete shape contract. In code, <code>input_gate_eq_8</code> names the mapping to the framework LSTM concept; it does not claim to hand-implement the missing cell equations.</p><p>The forget-gate record describes the corresponding gate that controls retained information:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;fga = &#963; Wf g [ht&#8722;1, Xc] + b f&quot;,&quot;id&quot;:&quot;F5EEA18B11&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here, <code>fga</code> is the forget-gate activation, <code>&#963;</code> is the sigmoid function, <code>Wf g</code> is the printed forget-gate weight matrix, <code>ht&#8722;1</code> is the previous hidden state, <code>Xc</code> is the current input as printed, and <code>b f</code> is the forget-gate bias. The extraction has damaged subscript and multiplication formatting, and the equation still omits the subsequent cell-state update. The generated <code>forget_gate_eq_9</code> function therefore provides traceability to the paper record rather than a claim of full mathematical reconstruction.</p><p>The output-gate record describes the gate controlling the exposed hidden representation:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;Opg = &#963; Wop [ht&#8722;1, Xc] + bo&quot;,&quot;id&quot;:&quot;179EEBA567&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here, <code>Opg</code> is the output-gate activation, <code>&#963;</code> is sigmoid, <code>Wop</code> is the output-gate weight matrix, <code>ht&#8722;1</code> is the previous hidden state, <code>Xc</code> is the current input as printed, and <code>bo</code> is the output-gate bias. As with the other two records, dimensions and the remaining recurrent updates are not supplied. The generated <code>output_gate_eq_10</code> function names this correspondence while the actual model is delegated to the Keras implementation.</p><p>The important fidelity rule is to preserve these three strings exactly, including their extracted notation, while explaining that they are incomplete. Extending them into a standard textbook LSTM formula would change the supplied record and would falsely imply that the paper specified the missing steps.</p><h3>Ensemble terminology and reporting limits</h3><p>The paper calls the Random Forest + XG-Boost + LSTM model &#8220;stacked,&#8221; but it also discusses assigning greater weight to lower-loss models. These are not automatically the same operation. <strong>Weighted averaging</strong>, also called a form of blending, combines component predictions using explicit weights. <strong>Stacking</strong> normally trains a second-stage model, or meta-model, on component predictions. The generated framework exposes both alternatives and records the selected <code>combiner</code>; neither is presented as the recovered paper implementation.</p><p>The same caution applies to the reported metrics. RMSE is an error measure in the target's units, while R2 is the coefficient of determination. They are meaningful only when predictions and targets refer to the same timestamps and scale. Table 2 values are <strong>reported references</strong>. In particular, the unusually large random-forest RMSE values are preserved rather than corrected: the supplied material does not establish whether they reflect scaling, evaluation inconsistency, or a reporting issue.</p><p>The paper's conclusions do not establish a universal forecasting method. Stock relationships can change across securities, market regimes, date ranges, data revisions, and preprocessing choices. Accordingly, a result from one period should not be treated as a general guarantee. The paper also frames AI predictions as support for investment decisions rather than their sole determinant; that limitation remains relevant even when a model produces favorable retrospective metrics.</p><h3>What would improve fidelity</h3><p>Exact reproduction would require the original Yahoo Finance ticker identifiers and a fixed data snapshot, including interval and endpoint behavior. It would also require the selected feature and target columns, null and zero handling, scaler and inverse-transform policy, sequence look-back and horizon, validation and tuning procedure, selected model parameters, random seeds, package versions, complete LSTM settings, and the actual ensemble wiring. Source code or an original experiment manifest would resolve much of this uncertainty.</p><p>Under the current run policy, no API call, training run, test execution, syntax check, or semantic code verification was performed. The generated tests and ledger are planned safeguards, not evidence that the implementation matches Table 2. The responsible interpretation is therefore precise: the framework reproduces the paper's documented model families and stated configuration elements while making its unresolved choices explicit; it does not establish independently verified financial-forecasting results.</p><h2>Static and semantic verification plan</h2><p>How can we tell whether this reproduction framework respects the paper's assumptions without pretending that the experiment has already run? Separate four activities: <strong>static verification</strong>, <strong>semantic verification</strong>, <strong>testing</strong>, and <strong>execution</strong>. Static verification inspects structure such as syntax, imports, names, and configuration values. Semantic verification reviews whether the code's behavior is consistent with the intended method, including time alignment and leakage boundaries. Tests encode repeatable checks for those contracts. Execution actually runs the package, trains models, calls data sources, or produces metrics.</p><p>For this project, execution was disabled by the run policy. Local static verification and code semantic verification were also skipped. Therefore, the checks described below are planned invariants and test specifications, not observed passes. In particular, no test, syntax check, API call, model-training run, or numerical reproduction of Table 2 was performed.</p><h3>What the planned checks protect</h3><p>The generated tests focus on six categories:</p><ol><li><p><strong>Rank and shape.</strong> Tabular features must remain rank 2, sequence inputs rank 3, and predictions must have one value per aligned observation.</p></li><li><p><strong>Time order.</strong> Cleaning must produce ordered, unique timestamps; chronological splits must not overlap; and test observations must follow training observations.</p></li><li><p><strong>Sequence alignment.</strong> Every look-back window must contain observations before its target timestamp. Initial rows lost during window construction must not be silently paired with another model's predictions.</p></li><li><p><strong>Leakage control.</strong> Learned preprocessing state, candidate selection, and any ensemble meta-model must use training data only.</p></li><li><p><strong>Finite outputs and metric alignment.</strong> Predictions and targets must have equal lengths, matching timestamps, and finite numeric values before RMSE or R2 is computed.</p></li><li><p><strong>Paper traceability.</strong> Canonical equation text and Table 1 candidate grids must be preserved exactly, while damaged equations must remain unreconstructed.</p></li></ol><p>These are implementation constraints derived from the method cards, not additional results reported by the paper.</p><h3>Data and temporal invariants</h3><p>The data tests in <code>tests/test_data_invariants.py</code> express the earliest contracts. For example, the cleaning test checks ordering, uniqueness, duplicate handling, and the retained final duplicate value:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f53c04ba-271d-4ee1-b8d3-8dd387d359e7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def test_clean_table_orders_and_deduplicates() -&gt; None:
    """Cleaning sorts timestamps and keeps the final occurrence of duplicates."""
    timestamps = pd.to_datetime(
        ["2024-01-03", "2024-01-01", "2024-01-02", "2024-01-02"]
    )
    raw = pd.DataFrame(
        {
            "Open": [3.0, 1.0, 2.0, 20.0],
            "Close": [3.5, 1.5, 2.5, 25.0],
            "Volume": [30.0, 10.0, 20.0, 200.0],
        },
        index=pd.DatetimeIndex(timestamps, name="Date"),
    )

    cleaned = clean_stock_table(raw, remove_zero_values=False)

    assert cleaned.index.is_monotonic_increasing
    assert cleaned.index.is_unique
    assert list(cleaned.index) == list(pd.to_datetime(["2024-01-01", "2024-01-02", "2024-01-03"]))
    assert cleaned.loc[pd.Timestamp("2024-01-02"), "Close"] == pytest.approx(25.0)
    assert cleaned.shape == (3, 3)</code></pre></div><p>The important point is not the particular fixture dates. It is that <code>clean_stock_table</code> must make duplicate and ordering behavior explicit. The paper says to remove redundancies and nulls, but it does not prescribe this exact duplicate policy; the test documents the generated implementation's chosen contract.</p><p>The chronological split test checks the primary 80/20 reproduction decision:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;0df7a763-2cbe-4137-aee7-7f63a1bb1ebd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def test_chronological_split_has_no_overlap() -&gt; None:
    """The split preserves order and places every test observation after training."""
    index = pd.date_range("2024-01-01", periods=10, freq="D", name="Date")
    X = pd.DataFrame({"feature": np.arange(10, dtype=float)}, index=index)
    y = pd.Series(np.arange(10, dtype=float) + 100.0, index=index, name="target")

    X_train, X_test, y_train, y_test = chronological_split(X, y, train_fraction=0.8)

    assert len(X_train) == len(y_train) == 8
    assert len(X_test) == len(y_test) == 2
    assert X_train.index.equals(y_train.index)
    assert X_test.index.equals(y_test.index)
    assert set(X_train.index).isdisjoint(set(X_test.index))
    assert X_train.index[-1] &lt; X_test.index[0]</code></pre></div><p>This protects both shape and chronology. It does not establish that 80/20 is the only correct split: the paper's generic pipeline mentions 75/25, while Table 1 specifies 80/20 for the proposed ensemble. The test simply records which policy the reproduction framework has chosen as its primary target.</p><h3>Sequence and scaling invariants</h3><p>An LSTM consumes a rank-3 array with shape <code>(batch, sequence_length, n_features)</code>. The sequence test checks that a look-back window precedes its future target and that the target index is retained:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;89169f58-7712-419b-8c1a-8fdcdbe51b74&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def test_sequences_preserve_target_alignment() -&gt; None:
    """Each sequence contains only observations before its future target."""
    index = pd.date_range("2024-01-01", periods=8, freq="D", name="Date")
    X = np.arange(8, dtype=float).reshape(-1, 1)
    y = (100.0 + np.arange(8, dtype=float)).reshape(-1, 1)

    result = make_sequences(X, y, index, lookback=3, horizon=2)

    # The first target is position lookback + horizon - 1 = 4.
    assert result.X_train.ndim == 3
    assert result.X_train.shape[1:] == (3, 1)
    assert result.X_train[0, :, 0].tolist() == [0.0, 1.0, 2.0]
    assert result.y_train[0] == pytest.approx(104.0)
    assert result.train_index[0] == index[4]</code></pre></div><p>The paper does not supply the look-back or horizon. Those are implementation decisions, so the test verifies the chosen sequence-construction semantics rather than a paper-specified numerical setting. This invariant is essential to <code>preprocess_time_series</code> and to <code>stacked_rf_xgboost_lstm</code>: tree models may produce predictions for rows that the LSTM cannot represent until enough history exists.</p><p>Scaling has a separate leakage constraint. A scaler is a learned transformation, so it must be fitted on training rows only. The planned test checks that the stored means come from <code>X_train</code> and <code>y_train</code>, not from the deliberately extreme test values:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b92cb3af-a184-4401-ba02-7898573762b0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def test_scaler_fit_scope_is_training_only() -&gt; None:
    """The fitted preprocessing state records only training rows and excludes test values."""
    X_train = np.array([[1.0], [3.0], [5.0]], dtype=float)
    y_train = np.array([10.0, 20.0, 30.0], dtype=float)
    X_test = np.array([[1000.0]], dtype=float)
    y_test = np.array([10000.0], dtype=float)

    state = fit_scalers(X_train, y_train, scaler_name="standard")

    assert state.fit_rows == X_train.shape[0]
    assert state.feature_scaler.mean_.tolist() == pytest.approx([3.0])
    assert state.target_scaler.mean_.tolist() == pytest.approx([20.0])</code></pre></div><p>The paper does not specify the scaler type, whether targets are scaled, or whether reported metrics use original prices. The invariant therefore concerns fit scope, not a particular scaler choice. Any inverse transformation used before reporting must preserve target alignment and scale consistency.</p><h3>Equation traceability without reconstruction</h3><p>Equation traceability has a special rule in this project: the supplied canonical equation records are authoritative. Equation <code>eq_2</code>, the KNN distance expression, has empty canonical LaTeX because the source extraction is damaged. The implementation may provide the named <code>euclidean_distance_eq_2</code> adapter as an explicit vector-distance decision, but it must not present reconstructed LaTeX as if it came from the paper.</p><p>The same rule applies to the LSTM gate records. The supplied canonical text for <code>eq_8</code> describes the input-gate activation as printed:</p><p>The input-gate record is preserved exactly below.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;iga = &#963; Wip [ht&#8722;1, Xc] + bi&quot;,&quot;id&quot;:&quot;B112F81F9B&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here <code>iga</code> is the input-gate activation, <code>&#963;</code> is the elementwise sigmoid function, <code>Wip</code> is the printed input-gate weight matrix, <code>ht&#8722;1</code> is the previous hidden state, <code>Xc</code> is the current input notation used in the extracted equation, and <code>bi</code> is the input-gate bias. These are vector or tensor quantities in an implementation, although the paper does not provide dimensions. The generated <code>input_gate_eq_8</code> function is a traceable helper name, not evidence that the incomplete printed expression defines the entire recurrent cell.</p><p>The forget-gate record is preserved exactly as <code>eq_9</code>.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;fga = &#963; Wf g [ht&#8722;1, Xc] + b f&quot;,&quot;id&quot;:&quot;BAD81294D1&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here <code>fga</code> is the forget-gate activation, <code>&#963;</code> is again applied elementwise, <code>Wf g</code> is the printed forget-gate weight matrix, <code>ht&#8722;1</code> is the previous hidden state, <code>Xc</code> is the current input, and <code>b f</code> is the forget-gate bias. The corresponding code symbol is <code>forget_gate_eq_9</code>.</p><p>The output-gate record is preserved exactly as <code>eq_10</code>.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;Opg = &#963; Wop [ht&#8722;1, Xc] + bo&quot;,&quot;id&quot;:&quot;1571328EEA&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here <code>Opg</code> is the output-gate activation, <code>Wop</code> is the output-gate weight matrix, <code>ht&#8722;1</code> is the previous hidden state, <code>Xc</code> is the current input, and <code>bo</code> is the output-gate bias. The generated mapping is named <code>output_gate_eq_10</code>.</p><p>The three records describe individual gate activations only. They omit the candidate-cell computation, cell-state update, and hidden-state update. The framework-level Keras LSTM therefore supplies the complete recurrent behavior; the code does not hand-reconstruct unsupported equations. The planned traceability test checks exact strings and function names:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6b332364-b574-444e-8273-753647a2db2c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def test_damaged_equations_remain_empty() -&gt; None:
    """Ensure damaged source records are not silently reconstructed in the registry."""
    for equation_id in ("eq_2", "eq_4", "eq_6"):
        assert get_equation(equation_id).latex == ""


def test_lstm_methods_reference_eq_8_eq_9_eq_10() -&gt; None:
    """Check that the three conceptual gate adapters are exposed as named callables."""
    gate_functions = {
        "eq_8": input_gate_eq_8,
        "eq_9": forget_gate_eq_9,
        "eq_10": output_gate_eq_10,
    }

    for equation_id, function in gate_functions.items():
        assert callable(function)
        assert equation_id in function.__name__</code></pre></div><p>These checks are planned evidence-preservation checks. They do not numerically validate an LSTM and do not turn incomplete equations into a complete mathematical specification.</p><h3>Model-contract and paper-fidelity checks</h3><p>The model-contract tests target the interfaces used by <code>fit_models_and_generate_predictions</code>. The LSTM contract is rank 3 in and one regression output per sequence:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;75a54ac8-1094-46a7-8e24-255edef80e21&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def test_lstm_uses_rank_three_inputs_and_single_output() -&gt; None:
    """The configured neural model has (batch, sequence, feature) input and (batch, 1) output."""
    pytest.importorskip("tensorflow")

    model = build_lstm_regressor(
        input_shape=(6, 4),
        units=5,
        dropout=0.2,
        dense_units=25,
    )

    assert model.input_shape == (None, 6, 4)
    assert model.output_shape == (None, 1)
    assert sum(layer.__class__.__name__ == "LSTM" for layer in model.layers) == 2</code></pre></div><p>This checks the stated two-layer architecture, dropout setting, dense width, and shape contract when TensorFlow is available. It does not check training quality or assert that the unspecified recurrent-unit count is the paper's count. Similarly, <code>test_paper_candidate_grids_match_table_1</code> is intended to compare the random-forest and XG-Boost candidate values with Table 1. Candidate values are not selected hyperparameters, and this check does not imply that the paper used the same search procedure as the generated code.</p><h3>Ensemble alignment and leakage checks</h3><p>The ensemble is the area where semantic review matters most. Matching array shapes alone cannot prove that random-forest, XG-Boost, and LSTM predictions refer to the same timestamps. The alignment test deliberately gives each component different order and coverage, then requires an intersection ordered by the target:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d6549d57-2391-48ad-a306-f5409e3d3056&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def test_component_predictions_align_by_timestamp() -&gt; None:
    """Component predictions use the target order and only common timestamps."""
    target_index = _index(0, 1, 2, 3)
    target = PredictionBundle("target", target_index, np.array([10.0, 20.0, 30.0, 40.0]))

    random_forest = PredictionBundle(
        "random_forest",
        _index(3, 1, 0),
        np.array([39.0, 19.0, 9.0]),
    )
    xgboost = PredictionBundle(
        "xgboost",
        _index(2, 0, 3),
        np.array([29.0, 11.0, 41.0]),
    )
    lstm = PredictionBundle("lstm", _index(0, 2, 3), np.array([10.5, 30.5, 40.5]))

    aligned = align_prediction_bundles((random_forest, xgboost, lstm), target)

    expected_index = _index(0, 3)
    np.testing.assert_array_equal(aligned.index, expected_index)
    np.testing.assert_allclose(
        aligned.predictions,
        np.array([[9.0, 11.0, 10.5], [39.0, 41.0, 40.5]]),
    )</code></pre></div><p>The resulting component matrix has shape <code>(n_aligned_samples, 3)</code>, with columns for random forest, XG-Boost, and LSTM. A second planned check verifies that a linear stacker is fit from component predictions and training targets only. This is especially important because the paper calls the ensemble &#8220;stacked&#8221; but also discusses higher weights for lower-loss models. Weighted averaging and learned stacking are different operations, and neither wiring is recovered from the paper.</p><h3>Worked verification checklist</h3><p>For one prepared dataset and one ensemble output, the intended review checklist is:</p><p>| Invariant | Planned question | Evidence location | |---|---|---| | Rank | Is tabular <code>X</code> rank 2, sequence <code>X</code> rank 3, and each prediction vector aligned? | <code>tests/test_model_contracts.py</code> | | Time order | Are timestamps unique, chronological, and non-overlapping across train and test? | <code>tests/test_data_invariants.py</code> | | Window safety | Does every LSTM window precede its target? | <code>test_sequences_preserve_target_alignment</code> | | Scaling leakage | Was preprocessing state fitted only on training rows? | <code>test_scaler_fit_scope_is_training_only</code> | | Equation fidelity | Are intact strings unchanged and <code>eq_2</code> still empty? | <code>tests/test_equation_traceability.py</code> | | Table 1 fidelity | Do RF and XG-Boost candidate grids match the supplied values? | <code>test_paper_candidate_grids_match_table_1</code> | | Ensemble alignment | Do all three components share the same target timestamps? | <code>test_component_predictions_align_by_timestamp</code> | | Meta-model leakage | Was a learned combiner fit without held-out test targets? | <code>test_stacker_fit_excludes_test_targets</code> | | Metric inputs | Do <code>y_true</code> and <code>y_pred</code> have equal aligned lengths? | <code>test_metric_inputs_must_align</code> |</p><p>This checklist labels intended obligations, not completed results. A future verification run could use it to review one synthetic or local-CSV experiment before considering a real Yahoo Finance reproduction.</p><h3>Verification boundary and final caution</h3><p>The planned tests are contract tests, not evidence that the paper's reported RMSE and R2 values have been reproduced. Synthetic data and local CSV data can exercise the pipeline, but they are not substitutes for the paper's unavailable Yahoo Finance snapshot and unspecified preprocessing and stacking decisions. Likewise, a successful model fit would not establish that Table 2 was independently verified.</p><p><strong>Advanced detail.</strong> A learned stacker can introduce a second form of leakage if its training predictions were generated on the same rows used to fit the component models. A robust stacking design often uses out-of-fold component predictions, but the paper does not specify such a procedure. The generated framework therefore records the combiner choice and its training data boundary rather than claiming that the paper recovered this detail.</p><p>The honest conclusion is narrow: the repository defines useful invariants for <code>preprocess_time_series</code>, <code>fit_models_and_generate_predictions</code>, <code>stacked_rf_xgboost_lstm</code>, and <code>evaluate_regression_predictions</code>, but those invariants were not executed or independently reviewed in this pipeline. Treat them as a verification plan to apply before using any resulting metrics, especially for financial decisions.</p><p>Use the button or URL below to download the source code.</p>
      <p>
          <a href="/__u/onepagecode.substack.com/p/quant-trading-stacked-ensemble-machine">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Quant Trading: Kyle’s Price Impact & Order Flow Alpha Generation (Python Guide)]]></title><description><![CDATA[Estimating microstructure price impact, signed order flow, and Amihud illiquidity for stock return forecasting.]]></description><link>https://onepagecode.substack.com/p/quant-trading-kyles-price-impact</link><guid isPermaLink="false">https://onepagecode.substack.com/p/quant-trading-kyles-price-impact</guid><dc:creator><![CDATA[Onepagecode]]></dc:creator><pubDate>Tue, 11 Aug 2026 20:15:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!H7nv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460a6c39-6bbb-41db-95d4-0bdccd3fdf89_1428x482.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Download the Source Code Using the URL At the End of This Article</h2><p>The paper estimates Kyle's price-impact coefficient from CRSP daily equity data and tests whether order-flow-based measures forecast contemporaneous and subsequent monthly stock returns. It constructs signed order flow, total volume, volume volatility, a within-month price-impact regression estimator, and an Amihud-style illiquidity estimator. The theoretical interpretation is that adverse selection and temporary price impact, rather than only risk compensation, can generate an illiquidity premium.</p><h2>Implementation Assumptions</h2><ul><li><p>Required CRSP, H.15, WRDS, factor, and delisting datasets are supplied locally as files; no network access or credentials are used.</p></li><li><p>PERMNO and calendar-month keys are the canonical identifiers throughout the empirical pipeline.</p></li><li><p>The default sample period is 2020&#8211;2025 when the available data support it.</p></li><li><p>The default expanding-window interpretation is used because the procedural description explicitly specifies expansion, despite occasional use of the word rolling.</p></li><li><p>The default empirical lambda regression is uncentered because Equation 12 displays no intercept; this choice is exposed in configuration because the paper is ambiguous.</p></li><li><p>Zero daily price changes receive sign zero.</p></li><li><p>Firm-month winsorization scope is configurable and never silently assumed.</p></li><li><p>Newey&#8211;West lag length, covariance treatment, portfolio weighting, and missing-data rules are explicit configuration fields because they are unspecified in the supplied paper extraction.</p></li><li><p>OCR-damaged equations eq<em>3, eq</em>6, eq<em>8, and eq</em>16 are not reconstructed beyond the supplied meanings and prose.</p></li><li><p>Reported paper coefficients and summary statistics are stored as comparison targets, not as independently reproduced results.</p></li><li><p>No code execution, numerical verification, or semantic validation is claimed under the supplied run policy.</p></li></ul><h2>Scope, evidence boundaries, and pipeline architecture</h2><p>What would it take to turn the paper&#8217;s liquidity and return-predictability tests into a reproducible Python workflow? The practical answer is not one model call. It is a chain of contracts: local daily records are cleaned, daily activity is summarized by firm and calendar month, estimators are aligned with subsequent returns, and each analysis produces labeled outputs with provenance.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The paper studies whether observable trading activity contains information about contemporaneous and future stock returns. Its structural motivation comes from Kyle-style price impact: informed and noise trading combine into signed aggregate order flow, and a market maker adjusts prices in response. The empirical data do not reveal those latent theoretical orders separately. Instead, the reproduction uses signed daily volume as an observable proxy, alongside unsigned volume, volume volatility, and two firm-month illiquidity or price-impact estimators.</p><p>This distinction is central. The theoretical aggregate flow is a model quantity; the empirical <code>signedflow</code> column is constructed from CRSP observations. Likewise, <code>lambda</code> is not one universally defined field. Its meaning depends on whether it was produced by the within-month price-impact regression or by the Amihud-style average.</p><h3>The evidence boundary</h3><p>The supplied paper describes a principal sample covering 2020&#8211;2025 and reports coverage and regression results, including firm and firm-month counts. Those values are paper-reported comparison targets, not independently reproduced results in this workflow. A full numerical reproduction requires local CRSP data and, for particular analyses, H.15 risk-free rates, WRDS book-to-market and factor data, and possibly delisting-return and point-in-time membership files.</p><p>The implementation therefore follows three categories:</p><ul><li><p><strong>Paper facts:</strong> the stated sample period, exchange filter, minimum 15 nonzero-volume-day firm-month rule, daily-to-monthly feature definitions, and the planned regression and forecast families.</p></li><li><p><strong>Implementation decisions:</strong> the exact adjustment-field convention, unchanged-price sign handling, winsorization scope, covariance treatment, forecast pooling, Newey&#8211;West lag, and portfolio weighting.</p></li><li><p><strong>Unavailable evidence:</strong> missing source data, incomplete adjustment instructions, and OCR-damaged or underspecified parts of the paper.</p></li></ul><p>The code keeps the second category in configuration rather than silently choosing hidden behavior. It also avoids filling missing data with paper-reported values. The package is local-only: it reads files supplied by the user and does not access a network or credentials.</p><h3>The pipeline as an accounting ledger</h3><p>The canonical identifiers are <code>PERMNO</code>, the CRSP firm identifier, and <code>month</code>, a normalized calendar-month key. The intended data flow is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;31e34331-971c-4156-9be6-4e867c43dd6a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">daily CRSP rows
    -&gt; filtered and adjusted daily rows
    -&gt; daily dollar volume, signed flow, and Amihud ratios
    -&gt; one row per (PERMNO, month)
    -&gt; lambda estimates, returns, and controls
    -&gt; contemporaneous and next-month model inputs
    -&gt; regressions, forecasts, portfolios, and reports</code></pre></div><p>Every transition should answer four questions:</p><ol><li><p>What source columns were used?</p></li><li><p>What is the resulting row shape and key?</p></li><li><p>Which time period does each value describe?</p></li><li><p>Which convention controls missing or ambiguous values?</p></li></ol><p>The generated <code>ReproductionConfig</code> records many of these choices. Its defaults are explicitly described as conventions rather than uniquely paper-determined facts:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4f498a82-0e59-4e0f-b120-b40c6cdc56b6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class ReproductionConfig:
    """Validated settings for local data preparation and empirical analyses.

    Defaults reflect implementation decisions made where the supplied paper
    extraction is ambiguous: a 2020--2025 sample, exchanges 1/2/3, a
    15-nonzero-volume-day filter, an uncentered lambda estimator, zero sign for
    unchanged prices, monthly winsorization, equal-weighted deciles, and no
    assumed covariance correction.  These defaults are not claimed to be the
    paper's uniquely specified conventions.
    """

    sample_start: DateLike = "2020-01-01"
    sample_end: DateLike = "2025-12-31"</code></pre></div><p><code>validate_config</code> checks ranges and supported options, but it does not establish that a local file exists or that its contents match the paper&#8217;s sample. That separation matters: configuration validation is not empirical verification.</p><h3>What each orchestration layer does</h3><p><code>build_empirical_panel(config)</code> is the main preprocessing boundary. It loads the configured CRSP daily and monthly files, applies exchange and date filters, constructs adjusted prices and daily features, aggregates firm-month variables, estimates both lambda columns, attaches available controls, applies configured winsorization, and creates same-firm next-month targets. Its output should contain at most one row for each <code>PERMNO</code>-<code>month</code> pair.</p><p><code>run_empirical_reproduction(panel, config)</code> receives that normalized panel rather than loading data itself. It isolates model-specific inputs and delegates to the empirical modules:</p><ul><li><p>pooled contemporaneous and one-month-ahead return regressions;</p></li><li><p>lambda-return regressions for both estimators and intercept variants;</p></li><li><p>firm-level expanding-window forecasts and pooled forecast evaluation;</p></li><li><p>monthly Fama&#8211;MacBeth regressions and signal-sorted decile portfolios.</p></li></ul><p>The reporting layer, <code>build_reproduction_report(panel, results, config)</code>, assembles coverage, descriptive, regression, forecast, and portfolio tables. Its design keeps computed outputs separate from paper targets. A target annotation can describe a comparison value, but it must not overwrite a locally computed coefficient or observation count.</p><p>The two command-line entry points reflect this separation. <code>scripts/build_panel.py</code> constructs auditable daily and firm-month artifacts. <code>scripts/run_reproduction.py</code> builds or loads a panel, runs the model pipeline, and writes report tables. Neither command is evidence that execution occurred here; they are local interfaces for a future run with suitable data.</p><h3>A small worked example: one firm through the ledger</h3><p>Consider a synthetic firm with <code>PERMNO</code> equal to <code>10001</code> and three daily observations in January. The rows might contain the following conceptual fields</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!H7nv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460a6c39-6bbb-41db-95d4-0bdccd3fdf89_1428x482.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!H7nv!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460a6c39-6bbb-41db-95d4-0bdccd3fdf89_1428x482.png 424w, /__u/substackcdn.com/image/fetch/$s_!H7nv!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460a6c39-6bbb-41db-95d4-0bdccd3fdf89_1428x482.png 848w, /__u/substackcdn.com/image/fetch/$s_!H7nv!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460a6c39-6bbb-41db-95d4-0bdccd3fdf89_1428x482.png 1272w, /__u/substackcdn.com/image/fetch/$s_!H7nv!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460a6c39-6bbb-41db-95d4-0bdccd3fdf89_1428x482.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!H7nv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460a6c39-6bbb-41db-95d4-0bdccd3fdf89_1428x482.png" width="1428" height="482" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/460a6c39-6bbb-41db-95d4-0bdccd3fdf89_1428x482.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:482,&quot;width&quot;:1428,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:58290,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/210720812?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460a6c39-6bbb-41db-95d4-0bdccd3fdf89_1428x482.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!H7nv!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460a6c39-6bbb-41db-95d4-0bdccd3fdf89_1428x482.png 424w, /__u/substackcdn.com/image/fetch/$s_!H7nv!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460a6c39-6bbb-41db-95d4-0bdccd3fdf89_1428x482.png 848w, /__u/substackcdn.com/image/fetch/$s_!H7nv!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460a6c39-6bbb-41db-95d4-0bdccd3fdf89_1428x482.png 1272w, /__u/substackcdn.com/image/fetch/$s_!H7nv!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460a6c39-6bbb-41db-95d4-0bdccd3fdf89_1428x482.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The adjustment module first produces an auditable <code>adj_price</code> and then a firm-local <code>price_change</code>. The daily-feature functions use those rows to create <code>dollar_volume</code>, <code>signed_flow</code>, and <code>amihud_ratio</code>. Under the configured zero-change convention, the unchanged observation contributes zero signed flow rather than being forced into a buy or sell direction. The upward and downward observations contribute signed volume with opposite directions.</p><p>All three rows receive the same calendar <code>month</code> key. If the firm has at least the configured minimum number of nonzero-volume days in the actual input, <code>build_firm_month_features</code> reduces the daily rows to one firm-month row. That row receives the sum of unsigned volume, the sample standard deviation of daily volume, and the sum of daily signed flow. The Amihud estimator uses only daily rows with valid returns and strictly positive dollar volume, so its valid-day count can differ from the count used for other features.</p><p>Later modules consume these columns differently. The panel regressions use the current firm-month volume features to explain either the current return or the same firm&#8217;s next calendar-month return. The lambda-return regressions use one current-month lambda column at a time. Forecasting uses the available current information inside an expanding chronological window. Portfolio formation ranks on a configured current-information signal and only then associates the resulting portfolios with subsequent returns.</p><p>This example explains data shapes and timing; it is not a paper calibration, an executed calculation, or a reproduced result.</p><h3>Evidence-preserving interpretation</h3><p>The paper contains unresolved contradictions that the pipeline should preserve rather than harmonize silently. Its theoretical discussion gives inconsistent comparative-static descriptions involving noise-trading variance and lambda. Its narrative expectations for signed-flow coefficients do not always agree with reported one-month-ahead coefficient signs. Section and table references are also inconsistent in the supplied extraction.</p><p>These issues affect interpretation, but they do not justify changing observed or reported signs in code. Result tables should retain specification labels, estimator names, sample counts, and provenance so that differences can be inspected directly.</p><p>The same principle applies to damaged equations. No formula is reconstructed when the supplied canonical record is incomplete or OCR-damaged. The implementation can document a prose-defined behavior or expose an explicit configuration choice, but it should not present an inferred formula as though it came from the paper.</p><p>Finally, all verification flags are disabled under the run policy. Static verification, semantic code verification, execution, test generation, and final quality review were not performed. The planned invariants&#8212;unique firm-month keys, same-firm target alignment, finite denominators, nondecreasing forecast windows, and separation of paper targets from computed outputs&#8212;are useful design requirements, not checks claimed to have passed.</p><h2>Local data contracts and preprocessing decisions</h2><p>Before estimating liquidity or return predictability, which observations should count as comparable trading records? The reproduction answers that question with a <strong>data contract</strong>: an explicit agreement about column names, types, keys, filters, missing values, and external inputs. Configuration keeps uncertain choices visible instead of burying them inside preprocessing functions.</p><p>The paper&#8217;s empirical sample is based on CRSP daily equity data, with a stated main period of 2020&#8211;2025 and a default exchange filter for CRSP exchange codes 1, 2, and 3. The pipeline then creates monthly firm observations. These are paper-supported design elements. By contrast, the exact application of CRSP adjustment factors, return conventions, delisting-return treatment, winsorization scope, and inference settings is not fully specified. Those choices must remain configurable.</p><h3>Required local inputs</h3><p>The normalized daily table has one row for one <code>PERMNO</code> on one trading date. <code>PERMNO</code> is the CRSP firm identifier, and the daily uniqueness key is the pair <code>(PERMNO, date)</code>. The required source fields are:</p><ul><li><p><code>PERMNO</code>: firm identifier.</p></li><li><p><code>date</code>: trading date.</p></li><li><p><code>EXCHCD</code>: CRSP exchange code.</p></li><li><p><code>DlyPrc</code>: daily price.</p></li><li><p><code>DlyVol</code>: daily share volume.</p></li><li><p><code>DlyRet</code>: daily stock return.</p></li><li><p><code>ShrOut</code>: shares outstanding.</p></li><li><p><code>DisFacPr</code>: CRSP price-adjustment field.</p></li><li><p><code>DisFacShr</code>: CRSP share-adjustment field.</p></li></ul><p><code>DlyBid</code> and <code>DlyAsk</code> may also be retained when available, but the core volume and lambda constructions do not require them.</p><p>The monthly inputs use a different key. Firm-level monthly returns and controls use <code>(PERMNO, month)</code>, where <code>month</code> is normalized to a calendar-month timestamp. Aggregate risk-free data use <code>month</code> alone. The paper also identifies H.15 Treasury-bill data, WRDS book-to-market and factor data, and possible delisting and point-in-time membership information for particular analyses. These files are external requirements; they cannot be inferred from CRSP daily rows or fabricated by the loader.</p><p>The daily loader is local-only. Its public responsibility is to read a supported file, normalize types and dates, preserve source fields, sort chronologically, and reject duplicate firm-date observations. A focused excerpt shows the boundary check:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;591d2cf1-5843-4b2c-a57f-5d568b56e245&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def validate_crsp_daily_columns(
    frame: pd.DataFrame, columns: DataColumnConfig
) -&gt; None:
    """Validate the normalized CRSP daily schema and key invariants.

    The required schema contains PERMNO, date, EXCHCD, DlyPrc, DlyVol, DlyRet,
    ShrOut, DisFacPr, and DisFacShr under the names in ``columns``.  Bid and ask
    are deliberately optional because the paper identifies them as optional
    inputs and the core feature pipeline does not require them.
    """
    required = columns.required_daily_columns()
    missing = [name for name in required if name not in frame.columns]
    if missing:
        raise ValueError(f"CRSP daily frame is missing required columns: {missing}")

    if not isinstance(frame.index, pd.RangeIndex):
        # Index type is not part of the data contract; reset-indexed input is
        # normalized by the loader, while direct callers receive a clear error.
        if frame.index.has_duplicates:
            raise ValueError("CRSP daily frame index must not contain duplicates")

    if frame[columns.permno].isna().any():
        raise ValueError("PERMNO contains missing values")
    if frame[columns.date].isna().any():
        raise ValueError("date contains missing values")

    permno = pd.to_numeric(frame[columns.permno], errors="coerce")
    if permno.isna().any() or ((permno % 1) != 0).any():
        raise ValueError("PERMNO must contain integer-like values")
    if not pd.api.types.is_datetime64_any_dtype(frame[columns.date]):
        raise TypeError("date must be datetime-like after normalization")

    duplicate_keys = frame.duplicated(
        subset=[columns.permno, columns.date], keep=False
    )
    if bool(duplicate_keys.any()):
        raise ValueError("CRSP daily input contains duplicate PERMNO-date observations")</code></pre></div><p><code>validate_crsp_daily_columns</code> receives a pandas <code>DataFrame</code> and a <code>DataColumnConfig</code>. Its output is <code>None</code> when the schema is acceptable; malformed schemas raise an exception rather than silently dropping or repairing fields. The loader itself returns a table with one normalized row per firm-date observation. It does not establish that the source data have the same point-in-time or corporate-action treatment as the paper.</p><h3>Configuration makes ambiguity explicit</h3><p><code>DataColumnConfig</code> maps source names to the fields used downstream. <code>ReproductionConfig</code> stores sample dates, local paths, filters, adjustment conventions, winsorization, covariance, forecasting, and portfolio choices. The following is a compact local configuration example copied from the generated tutorial:</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f35ccb91-83ce-4629-b26d-8165f0043c8f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from pathlib import Path

from liquidity_premium.config import ReproductionConfig

config = ReproductionConfig(
    sample_start="2020-01-01",
    sample_end="2025-12-31",
    daily_path=Path("data/crsp_daily.parquet"),
    monthly_returns_path=Path("data/monthly_returns.parquet"),
    risk_free_path=Path("data/risk_free_rates.csv"),
    auxiliary_monthly_path=Path("data/auxiliary_monthly.parquet"),
    output_dir=Path("outputs"),
)</code></pre></div><p>The defaults additionally specify exchange codes <code>(1, 2, 3)</code>, a minimum of 15 nonzero-volume days, the <code>price_only</code> adjustment convention, monthly winsorization, an uncentered lambda estimator, equal-weighted deciles, and no fixed Newey&#8211;West lag unless one is supplied. These are not all paper facts. The period, exchange codes, and 15-day rule are supported by the supplied paper description; the adjustment, winsorization, covariance, and portfolio defaults resolve underspecified choices in the implementation.</p><p>For example, a run can make those choices more visible:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;df0664c7-a7d6-4e06-b5c0-77d4bbd88688&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">config = ReproductionConfig(
    sample_start="2020-01-01",
    sample_end="2025-12-31",
    daily_path=Path("data/crsp_daily.parquet"),
    monthly_returns_path=Path("data/monthly_returns.parquet"),
    risk_free_path=Path("data/risk_free_rates.csv"),
    auxiliary_monthly_path=Path("data/auxiliary_monthly.parquet"),
    adjustment_convention="price_only",
    winsorization_scope="month",
    newey_west_lags=4,
    portfolio_weighting="equal",
)</code></pre></div><p>The value <code>4</code> here is an implementation example, not a lag length specified by the paper. The paper does not provide an exact Newey&#8211;West lag. Similarly, <code>price_only</code> is a selected convention, not a uniquely established interpretation of <code>DisFacPr</code> and <code>DisFacShr</code>.</p><h3>Filter order and the 15-day rule</h3><p>The intended order is to load and normalize the daily data, restrict the sample period and exchanges, construct the calendar-month key, and apply the minimum-volume-day rule before aggregation. The paper&#8217;s default exchange filter retains codes 1, 2, and 3. The minimum-volume filter retains a firm-month when it contains at least 15 finite, strictly positive <code>DlyVol</code> observations.</p><p>The generated filter is deliberately a row filter: it returns the daily observations belonging to qualifying groups, leaving aggregation to a later module.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;da18bf74-515e-484c-a3c5-7ae0d14eb647&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def filter_minimum_volume_days(
    frame: pd.DataFrame,
    minimum_days: int = 15,
) -&gt; pd.DataFrame:
    """Retain firm-month groups with enough nonzero-volume trading days.

    A qualifying day has finite, strictly positive ``DlyVol``.  The count is
    computed independently for each ``PERMNO`` and calendar month, matching
    the paper's minimum-15-nonzero-volume-day filter.  This function is
    deliberately a row filter: aggregation and feature construction occur in
    the downstream aggregation module.  The original columns and row order
    are preserved for surviving observations.
    """
    _require_columns(frame, ("PERMNO", "date", "DlyVol"))
    if not isinstance(minimum_days, int) or isinstance(minimum_days, bool):
        raise TypeError("minimum_days must be an integer")
    if minimum_days &lt; 1:
        raise ValueError("minimum_days must be positive")

    dates = _coerce_dates(frame["date"], "date")
    volume = pd.to_numeric(frame["DlyVol"], errors="coerce")
    valid_volume = volume.notna() &amp; volume.gt(0) &amp; volume.ne(float("inf")) &amp; volume.ne(float("-inf"))

    month_key = dates.dt.to_period("M")
    group_keys = pd.DataFrame(
        {"PERMNO": frame["PERMNO"], "_month": month_key},
        index=frame.index,
    )
    counts = valid_volume.groupby([group_keys["PERMNO"], group_keys["_month"]], dropna=False).sum()
    row_counts = pd.MultiIndex.from_arrays(
        [group_keys["PERMNO"], group_keys["_month"]],
        names=["PERMNO", "_month"],
    )
    qualifying = counts.ge(minimum_days)
    mask = qualifying.reindex(row_counts, fill_value=False).to_numpy()
    return frame.loc[mask].copy()</code></pre></div><p>The function accepts a daily <code>DataFrame</code> and returns a same-column <code>DataFrame</code> containing only rows from qualifying firm-month groups. It does not create a monthly row or calculate a return. A group with missing, infinite, or zero volume does not count that observation toward the threshold. The resulting aggregation must later enforce uniqueness of <code>(PERMNO, month)</code>.</p><h3>Adjustment fields and return conventions</h3><p>The paper requires split-adjusted prices and names <code>DisFacPr</code> and <code>DisFacShr</code>, but the supplied extraction does not give an unambiguous formula for applying both fields. The generated <code>compute_adjusted_price</code> function therefore exposes three conventions: <code>raw</code>, <code>price_only</code>, and <code>price_and_shares</code>. It preserves the raw adjustment columns instead of overwriting them.</p><p>This distinction matters because adjusted prices feed later price changes, signed-flow directions, and potentially dollar-volume calculations. The selected convention can change derived variables, so it belongs in configuration and provenance.</p><p><code>validate_adjustment_factors</code> checks that the adjustment fields exist and are finite and strictly positive under a factor-based convention. <code>compute_adjusted_price</code> returns a float-compatible series aligned with the input rows. Invalid adjusted prices raise an error rather than becoming silently usable values. These checks constrain the input but do not prove that a chosen convention matches CRSP documentation or the paper&#8217;s intended processing.</p><p>Monthly returns require a separate decision. <code>load_monthly_returns</code> normalizes the firm-month key and exposes a recognized return field as <code>StockRet</code>, but it does not automatically combine delisting returns or convert raw returns to excess returns. <code>load_risk_free_rates</code> normalizes aggregate monthly rates as <code>risk_free_rate</code>, but does not convert annualized yields to monthly returns. Those transformations depend on source units and conventions that the paper does not state consistently across all analyses.</p><p>This separation prevents a raw return, a delisting-adjusted return, and an excess return from being treated as interchangeable. Each model input should identify which convention it uses.</p><h3>Monthly auxiliary data and duplicate protection</h3><p>The monthly loaders normalize date-like values to month-start timestamps. Thus, dates within the same calendar month map to the same <code>month</code> key. Firm-month auxiliary data must not contain duplicate <code>(PERMNO, month)</code> rows, while aggregate risk-free data must not contain duplicate <code>month</code> rows. Duplicate keys are rejected rather than silently aggregated, because silent aggregation could conceal a data join error.</p><p>The loaders preserve auxiliary columns such as book-to-market or factor returns. They do not invent unavailable values. If WRDS factors, H.15 rates, point-in-time membership, or delisting returns are absent, the corresponding analysis must either be omitted or run with an explicitly documented reduced input set.</p><h3>Provenance without fabricated evidence</h3><p>A provenance record documents the configuration and local input paths. It does not claim that the files exist, that they are complete, or that they reproduce the paper&#8217;s reported coverage. The generated export function makes this boundary explicit:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f801e59c-dcd0-4f03-bd0c-473cfcaf9e0e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def write_provenance(
    path: Path,
    config: ReproductionConfig,
    inputs: list[Path],
) -&gt; None:
    """Write configuration and local input-path metadata for a reproduction run.

    The provenance record documents choices that affect reproducibility, including
    adjustment, winsorization, covariance, forecast, and portfolio conventions.
    It records paths only; it does not assert that external CRSP, H.15, WRDS, or
    delisting data are present, complete, or sufficient to reproduce paper results.
    """
    if not isinstance(config, ReproductionConfig):
        raise TypeError("config must be a ReproductionConfig")
    if not isinstance(inputs, list):
        raise TypeError("inputs must be a list of pathlib.Path values")
    if any(not isinstance(input_path, Path) for input_path in inputs):
        raise TypeError("each input must be a pathlib.Path")

    payload: dict[str, object] = {
        "schema_version": "1.0",
        "record_type": "liquidity_premium_reproduction_provenance",
        "paper_id": "2607.01377v1",
        "paper_title": "Liquidity Premium and Investment Horizons",
        "configuration": asdict(config),
        "input_paths": [str(input_path) for input_path in inputs],
        "notes": [
            "This record describes a local reproduction configuration and does not contain computed results.",
            "External data availability, preprocessing conventions, and unspecified paper choices affect reproducibility.",
            "Paper-reported targets are not represented as independently reproduced values.",
        ],
    }
    write_json(payload, path)</code></pre></div><p><code>write_provenance</code> receives a configuration and a list of local <code>Path</code> objects and writes metadata through <code>write_json</code>. The important invariant is evidentiary: a paper target, such as the reported firm or firm-month count, must remain a comparison target and must never be inserted into a computed report merely because it is known from the paper.</p><h3>Worked example: one firm and two keys</h3><p>Consider a conceptual daily input for one firm, <code>PERMNO=10001</code>, on several dates in January 2020. Each row has <code>DlyPrc</code>, <code>DlyVol</code>, <code>DlyRet</code>, <code>ShrOut</code>, <code>DisFacPr</code>, and <code>DisFacShr</code>, plus <code>EXCHCD=1</code>. After date normalization, all rows receive the same calendar month key, <code>2020-01-01</code>. The daily key is therefore <code>(10001, date)</code>, while the eventual aggregate key is <code>(10001, 2020-01-01)</code>.</p><p>If one row has missing volume, it does not count toward the 15-day threshold. If a row has zero volume, it also does not count. If the firm has fewer than 15 finite, strictly positive-volume days in January, every January row is removed by <code>filter_minimum_volume_days</code>; no partial firm-month is created. With at least 15 qualifying days, the later aggregation can produce exactly one January row for <code>PERMNO=10001</code>.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The later pipeline assigns responsibilities as follows:</p><ol><li><p><code>load_crsp_daily</code> validates and sorts the source rows.</p></li><li><p><code>filter_exchange_codes</code> removes observations outside codes 1, 2, and 3.</p></li><li><p><code>filter_sample_period</code> applies inclusive sample-date bounds.</p></li><li><p><code>compute_adjusted_price</code> creates an audit-preserving adjusted-price series under the configured convention.</p></li><li><p>Daily feature functions use the adjusted prices, volumes, and returns to create derived columns.</p></li><li><p>The aggregation module converts qualifying daily rows into one firm-month row.</p></li><li><p>Monthly loaders and alignment functions attach returns and controls without crossing firm boundaries.</p></li></ol><p>This example explains keys, timing, and responsibilities only. It is not a CRSP result, and no generated code or numerical output is claimed to have been executed.</p><h3>What a full reproduction still needs</h3><p>A complete numerical reproduction depends on local CRSP data and, for relevant specifications, H.15, WRDS, factor, point-in-time, and delisting inputs. It also requires explicit choices for:</p><ul><li><p>application of <code>DisFacPr</code> and <code>DisFacShr</code>;</p></li><li><p>raw, delisting-adjusted, or excess-return conventions;</p></li><li><p>winsorization scope and timing;</p></li><li><p>covariance and standard-error treatment;</p></li><li><p>handling of missing calendar months and external controls;</p></li><li><p>portfolio and risk-free-rate conventions.</p></li></ul><p>The paper&#8217;s reported coverage and summary values remain paper claims until a run with suitable data produces comparable outputs. Under the authoritative run policy, execution, static verification, semantic code verification, and numerical validation were not performed. The implementation therefore documents a reproducible boundary and a set of explicit decisions, not a completed empirical match.</p><h2>3. From daily prices and volume to firm-month features</h2><p>How does a table of daily CRSP observations become the monthly predictors used in the paper? The pipeline first creates row-level quantities&#8212;adjusted price, price change, dollar volume, signed flow, and an Amihud ratio&#8212;and then compresses those rows into one record per <code>PERMNO</code> and calendar month.</p><p>Unsigned volume measures how many shares traded. Signed flow preserves a direction: the same volume contributes positively after an upward price change and negatively after a downward change. An unchanged price contributes zero under the implementation convention. The Amihud-style ratio instead measures absolute return per dollar traded, so a zero-dollar-volume observation has no valid ratio.</p><h3>Daily inputs and shape flow</h3><p>The normalized daily frame contains one row per security-date observation. The relevant columns are <code>PERMNO</code>, <code>date</code>, <code>DlyPrc</code>, <code>DlyVol</code>, and <code>DlyRet</code>; <code>DisFacPr</code> and <code>DisFacShr</code> provide the adjustment fields named by the paper. The daily feature functions preserve the input index and add derived columns without changing the number of daily rows.</p><p>The intended shape transition is:</p><ul><li><p>daily source table: <code>(n_daily_rows, n_columns)</code>;</p></li><li><p>daily feature table: <code>(n_daily_rows, n_columns + derived_columns)</code>;</p></li><li><p>firm-month table: one row for each retained <code>(PERMNO, month)</code> key.</p></li></ul><p><code>make_month_key</code> converts each parseable date to a pandas monthly <code>Period</code>. The aggregation key is therefore the pair <code>PERMNO</code> and <code>month</code>, not merely the calendar month. This prevents observations from different firms being combined.</p><h3>Adjusted prices and daily dollar volume</h3><p>The paper names CRSP adjustment fields but does not uniquely specify how <code>DisFacPr</code> and <code>DisFacShr</code> should be applied. <code>compute_adjusted_price</code> therefore exposes explicit conventions such as <code>raw</code>, <code>price_only</code>, and <code>price_and_shares</code>. The selected convention is an implementation decision and should be recorded in configuration and provenance; it is not a paper fact. Raw source columns remain available for auditability.</p><p><code>compute_daily_price_change</code> calculates chronological differences within each <code>PERMNO</code>. A missing prior price produces a missing change. The daily dollar-volume function then prefers the derived <code>adj_price</code> column and multiplies it by <code>DlyVol</code>. Share volume must be finite and nonnegative. Missing inputs remain missing rather than being converted into artificial values.</p><p>A focused excerpt shows the dollar-volume boundary:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;55622678-ad4f-4f2d-b5de-f1c9656a673d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def compute_daily_dollar_volume(frame: pd.DataFrame) -&gt; pd.Series:
    """Compute daily dollar volume as price times share volume.

    The function uses ``adj_price`` when it is already present, because the
    preprocessing pipeline constructs price-based features from the selected
    split-adjusted price.  Otherwise it falls back to ``DlyPrc`` and uses its
    absolute value, matching CRSP's convention that a negative price can mark
    a bid/ask-related observation.  This fallback is an implementation
    decision; the supplied paper does not fully specify the adjustment-factor
    convention.
    """
    if not isinstance(frame, pd.DataFrame):
        raise TypeError("frame must be a pandas DataFrame")

    price_column = (
        _ADJUSTED_PRICE_COLUMN
        if _ADJUSTED_PRICE_COLUMN in frame.columns
        else _DEFAULT_PRICE_COLUMN
    )
    if price_column not in frame.columns:
        raise KeyError(
            f"Missing price column: expected {price_column!r} or "
            f"{_ADJUSTED_PRICE_COLUMN!r}"
        )
    if _DEFAULT_VOLUME_COLUMN not in frame.columns:
        raise KeyError(f"Missing share-volume column: {_DEFAULT_VOLUME_COLUMN!r}")

    price = _numeric_series(frame[price_column], price_column).abs()
    volume = _numeric_series(frame[_DEFAULT_VOLUME_COLUMN], _DEFAULT_VOLUME_COLUMN)
    invalid_volume = volume.notna() &amp; (~np.isfinite(volume) | (volume &lt; 0))
    if bool(invalid_volume.any()):
        raise ValueError("Share volume must be finite and nonnegative when present")

    dollar_volume = (price * volume).astype("float64")
    dollar_volume.name = "dollar_volume"
    finite = dollar_volume.notna() &amp; np.isfinite(dollar_volume)
    dollar_volume.loc[dollar_volume.notna() &amp; ~finite] = np.nan
    return dollar_volume.reindex(frame.index)</code></pre></div><p>The important implementation invariant is that the returned <code>dollar_volume</code> is aligned row-for-row with the input. It is nonnegative when present, but it can be missing when price or share volume is unavailable.</p><h3>Signed order flow</h3><p>The paper&#8217;s monthly signed-flow construction begins with daily share volume and the direction of the split-adjusted price change. Its purpose is to create an observable proxy for directional order flow before aggregation. The supplied canonical expression is:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\mathrm{signedflow}_{it} = \\sum_{\\tau \\in t} \\mathrm{Volume}_{i\\tau} \\times \\operatorname{sign}(\\Delta P_{i\\tau})&quot;,&quot;id&quot;:&quot;256C9C43DF&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here, <code>i</code> identifies the firm and <code>t</code> identifies the month. <code>&#964;</code> indexes trading days within that month. <code>Volume_i&#964;</code> is daily share volume, while <code>&#916;P_i&#964;</code> is the split-adjusted daily price change. The sign function maps a positive change to <code>1</code>, a negative change to <code>-1</code>, and an unchanged change to <code>0</code> under this implementation convention. The result, <code>signedflow_it</code>, is one signed scalar for the firm-month.</p><p>The code separates this operation into a daily function and a monthly aggregation. The daily function implements the direction step:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b3e65460-934b-4f4f-9abe-0e2d0d269f6d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def compute_daily_signed_flow(
    volume: pd.Series, price_change: pd.Series
) -&gt; pd.Series:
    """Compute daily signed order flow from share volume and price direction.

    This implements Eq. 9: daily volume is multiplied by the sign of the
    split-adjusted daily price change.  An unchanged price has sign zero, so
    its signed flow is exactly zero.  Missing volume or price change remains
    missing rather than being treated as a trading direction.
    """
    volume_values = _numeric_series(volume, "volume")
    change_values = _numeric_series(price_change, "price_change")
    if not volume_values.index.equals(change_values.index):
        raise ValueError("volume and price_change must have identical indexes")

    invalid_volume = volume_values.notna() &amp; (
        ~np.isfinite(volume_values) | (volume_values &lt; 0)
    )
    if bool(invalid_volume.any()):
        raise ValueError("volume must be finite and nonnegative when present")

    # Eq. 9: signed flow uses volume times the sign of the price change.
    direction = np.sign(change_values)
    signed_flow = (volume_values * direction).astype("float64")
    signed_flow.name = "signed_flow"
    signed_flow.loc[change_values.isna() | volume_values.isna()] = np.nan
    finite = signed_flow.notna() &amp; np.isfinite(signed_flow)
    if bool((signed_flow.notna() &amp; ~finite).any()):
        raise ValueError("signed flow contains non-finite values")
    return signed_flow.reindex(volume.index)</code></pre></div><p>The function accepts two one-dimensional, identically indexed series. It returns another one-dimensional series of the same length. A missing price change is not interpreted as a sell direction; the corresponding signed flow is missing. This preserves a distinction between no directional movement, represented by zero, and unavailable information, represented by missingness.</p><h3>Monthly total volume and dispersion</h3><p>The paper defines monthly unsigned share volume as follows:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\mathrm{sumvolume}_{it} = \\sum_{\\tau \\in t} \\mathrm{Volume}_{i\\tau}&quot;,&quot;id&quot;:&quot;3051979249&quot;}" data-component-name="LatexBlockToDOM"></div><p><code>sumvolume_it</code> is the nonnegative total share volume for firm <code>i</code> in month <code>t</code>. In <code>aggregate_monthly_volume</code>, it is computed from valid daily <code>DlyVol</code> values and returned as the <code>sumvolume</code> field. The function also returns <code>nonzero_volume_days</code>, which records the activity count used by the minimum-activity filter.</p><p>The paper uses within-month volume standard deviation as a proxy for noise-trading variation. The canonical LaTeX for equation 8 is unavailable in the supplied extraction, so it is not displayed or reconstructed here. The implementation follows the paper&#8217;s prose description and uses the sample standard deviation, equivalent to pandas <code>ddof=1</code>, over valid daily volume observations. This is a prose-derived implementation detail, not a transcription of missing equation LaTeX.</p><p>The minimum-volume rule retains only firm-month groups with at least 15 strictly positive-volume trading days. The filter is applied at the group level before aggregation, so all daily rows from a nonqualifying group are removed. The 15-day rule normally supplies enough observations for a numeric sample standard deviation, although missing values are still handled explicitly.</p><p>The aggregation function then produces one record per firm-month:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;0d0c2754-077d-4058-927c-5ed6c23eed9f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># Eq. 7: monthly total unsigned share volume.
total_volume = float(valid_volume.sum())

# Eq. 8: sample standard deviation of daily volume; canonical LaTeX is
# unavailable, so the documented ddof=1 convention is used.
stdvolume = float(valid_volume.std(ddof=1)) if len(valid_volume) &gt;= 2 else np.nan

# Eq. 9: monthly signed order flow is the sum of daily signed flow.
signedflow = float(signed_flow.sum(min_count=1)) if signed_flow.notna().any() else np.nan</code></pre></div><p>The resulting columns are <code>sumvolume</code>, <code>nonzero_volume_days</code>, <code>stdvolume</code>, and <code>signedflow</code>, with <code>amihud_lambda</code> added when daily Amihud ratios are supplied. <code>build_firm_month_features</code> sorts the output and rejects duplicate <code>PERMNO</code>-month keys.</p><h3>Amihud-style illiquidity</h3><p>The daily Amihud quantity measures absolute return per dollar traded. The paper&#8217;s monthly estimator is the mean of valid daily ratios:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\hat{\\lambda}^{\\mathrm{Amihud}}_{it} = \\frac{1}{n} \\sum_{\\tau \\in t} \\frac{|r_{i\\tau}|}{\\mathrm{DollarVolume}_{i\\tau}}&quot;,&quot;id&quot;:&quot;374CFE303E&quot;}" data-component-name="LatexBlockToDOM"></div><p>In this expression, <code>r_i&#964;</code> is the daily stock return, <code>DollarVolume_i&#964;</code> is daily dollar trading volume, and <code>n</code> is the number of valid daily observations used in the average. <code>i</code> and <code>t</code> retain their firm and month meanings. The estimate is nonnegative when valid.</p><p>The daily implementation excludes rows with missing or non-finite returns and rows whose dollar-volume denominator is zero. It never turns division by zero into infinity. The valid daily ratios are then averaged within <code>aggregate_monthly_volume</code> as <code>amihud_lambda</code>.</p><p>Equation 15 gives the same estimator under the Method A label:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\hat{\\lambda}^{A}_{it} = \\frac{1}{n} \\sum_{\\tau \\in t} \\frac{|r_{i\\tau}|}{\\mathrm{DollarVolume}_{i\\tau}}.&quot;,&quot;id&quot;:&quot;17FBC2BE1E&quot;}" data-component-name="LatexBlockToDOM"></div><p>The superscript <code>A</code> identifies Method A, the Amihud-style level estimator. The symbols have the same meanings as in the preceding equation: <code>n</code> is the valid-day count, <code>r_i&#964;</code> is daily return, and <code>DollarVolume_i&#964;</code> is a positive daily denominator. In code, both equation records map to <code>compute_daily_amihud_ratio</code> followed by the within-month mean. The two equation labels describe the same calculation rather than two different daily algorithms.</p><p>A focused denominator check from the generated implementation is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;9eaf9ec2-4bbf-4f08-929c-7c201149ddbf&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">valid = (
    returns.notna()
    &amp; np.isfinite(returns)
    &amp; dollars.notna()
    &amp; np.isfinite(dollars)
    &amp; (dollars &gt; 0)
)
ratio = pd.Series(np.nan, index=returns.index, dtype="float64", name="amihud_ratio")
ratio.loc[valid] = (returns.loc[valid].abs() / dollars.loc[valid]).astype("float64")
if bool((ratio.notna() &amp; ~np.isfinite(ratio)).any()):
    raise ValueError("Amihud ratios must be finite when present")
return ratio</code></pre></div><p>Thus, a zero-dollar-volume day does not contribute to the monthly mean. This can make the number of valid Amihud days smaller than the number of days used for volume aggregation.</p><h3>Worked example: three daily observations</h3><p>Consider one firm with three already adjusted daily observations. Suppose the share volumes are <code>100</code>, <code>200</code>, and <code>300</code>, and the corresponding price changes are positive, zero, and negative. The signed contributions are therefore <code>+100</code>, <code>0</code>, and <code>-300</code>; the unchanged price contributes exactly zero. The unsigned total volume is <code>600</code>.</p><p>If the three valid daily volumes are used for the sample dispersion, the calculation is the sample standard deviation of <code>100</code>, <code>200</code>, and <code>300</code>, using denominator <code>n - 1</code> rather than <code>n</code>. This example is conceptual: the paper&#8217;s minimum filter requires at least 15 positive-volume days for a retained firm-month, so three rows alone would not survive the actual aggregation filter.</p><p>For the Amihud calculation, suppose the daily returns are <code>0.01</code>, <code>0.02</code>, and <code>-0.03</code>, and the corresponding positive dollar volumes are <code>1,000</code>, <code>2,000</code>, and <code>0</code>. The first two ratios are valid. The third is excluded because its denominator is zero; it does not become an infinite ratio and is not included in <code>n</code>. The monthly Amihud value is consequently the mean of the valid ratios only.</p><p>The same fields would then flow into later stages as follows: <code>sumvolume</code>, <code>stdvolume</code>, <code>signedflow</code>, and <code>amihud_lambda</code> become firm-month predictors; <code>StockRet</code> is merged from monthly data; and the alignment stage creates a same-firm next-month target. No numerical result from this illustrative example is a paper result or a verified execution outcome.</p><h3>Orchestration and evidence boundary</h3><p><code>build_daily_audit_table</code> exposes the cleaned daily frame with derived fields for inspection. <code>build_empirical_panel</code> calls the daily stage, aggregates firm-month features, constructs controls, estimates lambda values, merges monthly returns, and aligns next-month returns. The generated pipeline does not access a network; CRSP and auxiliary files must already exist locally.</p><p>The implementation enforces key structural constraints such as nonnegative volume, finite valid Amihud ratios, at least 15 positive-volume days in retained groups, and unique <code>PERMNO</code>-month rows. However, the supplied verification records state that static verification and semantic code verification were skipped under the run policy. These constraints describe intended implementation behavior, not checks claimed to have passed. The adjustment convention, exact corporate-action treatment, return and delisting conventions, and winsorization policy remain configuration-dependent.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/p/quant-trading-kyles-price-impact?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/p/quant-trading-kyles-price-impact?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/onepagecode.substack.com/p/quant-trading-kyles-price-impact?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><h2>The theoretical Kyle relations and their empirical interpretation</h2><p>How can a price-impact coefficient connect trading activity to prices without being confused with a directly observed quantity? The paper&#8217;s Kyle-style model provides the intuition: an informed trader responds to a fundamental-value deviation, noise traders add orders that are not separately observed, and a market maker adjusts the transaction price according to total signed order flow. The empirical pipeline then uses observed signed volume as a proxy for this latent flow; it does not observe the theoretical informed and noise orders individually.</p><p>This section therefore has two layers. The first is a scalar theoretical API in <code>src/liquidity_premium/models/kyle.py</code>. The second is the empirical estimation workflow described elsewhere, where daily signed flow and price changes are used to estimate firm-month price impact. The theoretical functions accept scalar values and return scalar values. They are not fitted to CRSP data by the generated module.</p><h3>Informed demand responds to a value discrepancy</h3><p>In the single-period model, <code>v</code> is the asset&#8217;s liquidation value, while <code>p_0</code> is its initial or prior price. Their difference represents the fundamental-value discrepancy available to the informed trader. The positive scalar <code>beta</code> measures informed-trading intensity. The paper expresses informed demand as follows:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;x = \\beta (v - p_0), \\beta > 0&quot;,&quot;id&quot;:&quot;3241A68059&quot;}" data-component-name="LatexBlockToDOM"></div><p>This is <code>eq_1</code>. The symbol <code>x</code> is the informed trader&#8217;s scalar market order. If <code>v</code> exceeds <code>p_0</code>, the order is positive under this convention; if <code>v</code> is below <code>p_0</code>, it is negative. The restriction <code>beta &gt; 0</code> means that the intensity scales the direction implied by the value discrepancy rather than reversing it.</p><p>The generated function <code>informed_order(v, p0, beta)</code> implements this relation. It validates that the inputs are finite scalars and that <code>beta</code> is strictly positive, then returns the scalar order quantity. The validation is an implementation safeguard corresponding to the model&#8217;s stated positivity restriction; it does not estimate <code>beta</code> or infer <code>v</code> from market data.</p><p>A focused excerpt from <code>src/liquidity_premium/models/kyle.py</code> shows the mapping:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;042c0095-27be-466c-99cc-a5184798d07e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def informed_order(v: float, p0: float, beta: float) -&gt; float:
    """Compute informed demand from the supplied single-period Kyle relation.

    Parameters
    ----------
    v:
        Asset liquidation value.
    p0:
        Initial or prior asset price.
    beta:
        Positive informed-trading intensity.

    Returns
    -------
    float
        Informed order quantity ``x``.  The sign follows ``v - p0``.

    Notes
    -----
    Implements eq. 1: ``x = beta (v - p0)``.  This is a theoretical
    calculation, not an estimator of empirical order flow.
    """
    liquidation_value = _validate_finite_scalar(v, "v")
    initial_price = _validate_finite_scalar(p0, "p0")
    intensity = _validate_positive_scalar(beta, "beta")

    # Eq. 1: informed demand is proportional to the liquidation-value deviation.
    return intensity * (liquidation_value - initial_price)</code></pre></div><p>The final comment is important for reproduction fidelity: the function computes <code>x</code> from supplied scalar inputs, but it does not produce the empirical <code>signedflow</code> feature.</p><h3>Aggregate flow is latent in the model</h3><p>The theoretical informed order is only one component of total order flow. Let <code>u</code> denote a noise-trader order and let <code>y</code> denote aggregate signed order flow. In the model, aggregate flow is formed by adding the informed and noise components. The helper <code>aggregate_order_flow(informed, noise)</code> performs that addition and preserves the signs of both scalar inputs.</p><p>This distinction matters when reading the empirical pipeline. The model&#8217;s <code>y</code> is latent: the data do not separately identify <code>x</code> and <code>u</code>. By contrast, empirical <code>signedflow</code> is constructed from daily share volume and the sign of the observed split-adjusted price change. It is a proxy motivated by the theory, not an observed measurement of <code>y</code>, <code>x</code>, or <code>u</code>.</p><h3>The market maker maps flow into price</h3><p>The paper&#8217;s market-maker rule says that the transaction price equals the prior price plus a flow-dependent price adjustment. The coefficient <code>lambda</code> is the Kyle price-impact coefficient: it measures the price displacement associated with one unit of aggregate signed flow. The paper writes:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;p = p_0 + \\lambda y, \\lambda > 0&quot;,&quot;id&quot;:&quot;87831548B9&quot;}" data-component-name="LatexBlockToDOM"></div><p>This is <code>eq_2</code>. Here, <code>p</code> is the transaction price, <code>p_0</code> is the initial or prior price, <code>y</code> is aggregate theoretical flow, and <code>lambda</code> is a positive scalar. Because <code>lambda</code> is positive, positive aggregate flow raises the transaction price relative to <code>p_0</code>, while negative aggregate flow lowers it.</p><p>The generated <code>market_maker_price(p0, lam, aggregate_flow)</code> function implements this pricing rule. It checks that <code>p0</code>, <code>lam</code>, and the aggregate flow are finite and that <code>lam</code> is strictly positive. The function is a theoretical pricing helper, not an empirical regression. In particular, passing an observed signed-flow value to it would be an API illustration, not proof that the observed value is the model&#8217;s latent <code>y</code>.</p><p>The relevant implementation excerpt is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d38e10bd-8af0-449c-8887-1aca34fb3498&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def market_maker_price(p0: float, lam: float, aggregate_flow: float) -&gt; float:
    """Apply the linear market-maker pricing rule to aggregate flow.

    Parameters
    ----------
    p0:
        Initial or prior asset price.
    lam:
        Positive Kyle price-impact coefficient.
    aggregate_flow:
        Aggregate signed order flow ``y``.

    Returns
    -------
    float
        Transaction price ``p``.

    Notes
    -----
    Implements eq. 2: ``p = p0 + lambda y``.  The positive-lambda validation
    ensures that positive (negative) flow produces a positive (negative) price
    displacement relative to ``p0``.
    """
    initial_price = _validate_finite_scalar(p0, "p0")
    price_impact = _validate_positive_scalar(lam, "lam")
    flow = _validate_finite_scalar(aggregate_flow, "aggregate_flow")

    # Eq. 2: the market maker maps aggregate signed flow into price impact.
    return initial_price + price_impact * flow</code></pre></div><p>The empirical within-month regression uses the same conceptual direction&#8212;price change related to signed flow&#8212;but estimates a firm-month coefficient from daily observations. That later coefficient is an empirical construction whose exact scale and interpretation depend on the data adjustments, regression convention, and valid observations.</p><h3>Sequential price changes</h3><p>The sequential extension indexes trading rounds by <code>n</code>. In round <code>n</code>, <code>Delta y_n</code> is the order-flow innovation and <code>lambda_n</code> is the price-impact coefficient for that round. The resulting price innovation is:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\Delta p_n = \\lambda_n \\Delta y_n&quot;,&quot;id&quot;:&quot;85DC017B47&quot;}" data-component-name="LatexBlockToDOM"></div><p>This is <code>eq_5</code>. The scalar <code>Delta p_n</code> is the price change in round <code>n</code>; <code>lambda_n</code> is a positive, round-specific price-impact coefficient; and <code>Delta y_n</code> is the round&#8217;s signed order-flow innovation. The equation does not specify a full dynamic path for <code>lambda_n</code>; it only states how a given round&#8217;s coefficient maps that round&#8217;s flow innovation into a price change.</p><p>The generated <code>sequential_price_change(lam_n, delta_y_n)</code> function implements this multiplication. It validates positive <code>lam_n</code> and finite <code>delta_y_n</code>, then returns a scalar. Its sign invariant follows directly from the positivity restriction: a positive flow innovation produces a positive price change, a negative innovation produces a negative price change, and a zero innovation produces zero change.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;3ff834fc-3914-443b-b682-0a1e9e429ae7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def sequential_price_change(lam_n: float, delta_y_n: float) -&gt; float:
    """Compute a sequential-auction price innovation.

    Parameters
    ----------
    lam_n:
        Positive price-impact coefficient for round ``n``.
    delta_y_n:
        Order-flow innovation in round ``n``.

    Returns
    -------
    float
        Round-specific price change ``Delta p_n``.  Its sign equals the sign of
        ``delta_y_n`` unless the innovation is zero.

    Notes
    -----
    Implements eq. 5: ``Delta p_n = lambda_n Delta y_n``.  No dynamic path for
    ``lambda_n`` is inferred or simulated here.
    """
    round_price_impact = _validate_positive_scalar(lam_n, "lam_n")
    flow_innovation = _validate_finite_scalar(delta_y_n, "delta_y_n")

    # Eq. 5: each round's price innovation scales that round's flow innovation.
    return round_price_impact * flow_innovation</code></pre></div><h3>Worked API illustration</h3><p>Consider a hypothetical asset with <code>v = 101.0</code>, <code>p0 = 100.0</code>, and <code>beta = 2.0</code>. The informed-order function represents a value deviation of one price unit scaled by an intensity of two, so the conceptual informed order is positive. Suppose a noise trader submits a negative order; <code>aggregate_order_flow</code> adds that noise order to the informed order. A positive <code>lam</code> then maps the resulting signed flow into a transaction-price displacement.</p><p>For a sequential illustration, set <code>lam_n</code> to a positive value and <code>delta_y_n</code> to either a positive or negative flow innovation. The returned price change follows the innovation&#8217;s sign. These are hand-worked API examples for understanding scalar inputs, scalar outputs, and sign behavior. They are not CRSP observations, paper calibrations, or executed numerical results.</p><p>The planned tests in <code>tests/test_theoretical_kyle.py</code> mirror these direct relations for <code>eq_1</code>, <code>eq_2</code>, and <code>eq_5</code>. Under the run policy, those tests were not executed or verified; their presence describes intended coverage rather than a passed check.</p><h3>What is deliberately excluded</h3><p>The paper extraction does not provide safe canonical formulas for every theoretical statement. <code>eq_3</code>, the closed-form equilibrium values for <code>beta</code> and <code>lambda</code>, is OCR-damaged and is not reconstructed. <code>eq_4</code> is a qualitative sequential comparative-static claim whose conditioning and distributional assumptions are underspecified, so it is not simulated. <code>eq_6</code>, concerning the continuous-time path of <code>lambda(t)</code>, also has damaged canonical LaTeX and is not implemented.</p><p>These exclusions are deliberate evidence boundaries. The generated module implements only the supplied scalar relations <code>eq_1</code>, <code>eq_2</code>, and <code>eq_5</code>; it does not fill gaps by importing a familiar Kyle formula or by choosing an unstated dynamic model. The paper&#8217;s broader model description and its formal single-period presentation are also not fully reconciled, and its statements about how noise-trading variance affects price impact are internally inconsistent. Those comparative statics should remain documented ambiguities rather than being resolved silently in code.</p><p>Finally, the theoretical <code>lambda</code> should not be treated as identical to either empirical firm-month estimator. The empirical pipeline estimates price impact or illiquidity from observed daily price, return, volume, and signed-flow proxies. The theoretical functions explain why such a coefficient is economically meaningful; they do not identify it from the supplied scalar inputs or replace the empirical estimator.</p><h2>Estimating firm-month lambda</h2><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/p/quant-trading-kyles-price-impact?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/p/quant-trading-kyles-price-impact?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/onepagecode.substack.com/p/quant-trading-kyles-price-impact?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><p>How can several daily observations become one monthly measure of price impact? For each <code>PERMNO</code> and calendar month, the reproduction computes two separate scalars. Method B estimates how daily price changes respond to signed daily flow. Method A computes an Amihud-style average of absolute return per dollar traded. They address related notions of illiquidity, but they are not interchangeable columns.</p><p>The daily inputs for one firm-month are one-dimensional vectors with shape <code>(n_days,)</code>: <code>price_change</code> contains split-adjusted daily price changes, <code>signed_flow</code> contains signed daily volume, <code>return_</code> contains daily returns, and <code>dollar_volume</code> contains daily dollar trading volume. The grouped function reduces each valid group to one row containing <code>regression_lambda</code>, <code>amihud_lambda</code>, and estimator-specific valid-day counts.</p><h3>Method B: price change on signed flow</h3><p>The paper&#8217;s within-month regression relates daily price changes to daily signed order flow. Its displayed specification has no intercept, so the default implementation is an uncentered one-predictor regression.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\Delta P_{i\\tau} = \\hat{\\lambda}_{it} \\cdot OF_{i\\tau} + \\eta_{i\\tau}, \\quad \\tau \\in t&quot;,&quot;id&quot;:&quot;8F6C69151E&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here, <code>\Delta P_{i\tau}</code> is the daily split-adjusted price change for firm <code>i</code> on trading day <code>\tau</code>; <code>OF_{i\tau}</code> is that day&#8217;s signed-flow proxy; and <code>\hat{\lambda}_{it}</code> is the estimated scalar price-impact slope for firm <code>i</code> in month <code>t</code>. The residual <code>\eta_{i\tau}</code> captures the part of the daily price change not explained by the signed-flow regressor. The regression is performed separately within each <code>PERMNO</code>-month group, so its input has shape <code>(n_valid_days, 1)</code> for the design matrix and <code>(n_valid_days,)</code> for the response. The returned lambda is one scalar.</p><p>The default <code>include_intercept=False</code> follows the displayed equation. The generated function exposes <code>include_intercept=True</code> as an explicit alternative because the paper does not clearly resolve whether the intramonth estimator itself should include an intercept. This choice is separate from whether a later return regression includes an intercept.</p><p>The core implementation delegates the one-predictor calculation to <code>fit_ols</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;542ec276-069a-434c-b798-78f587dceaca&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    try:
        result = fit_ols(
            flow.reshape(-1, 1),
            prices,
            include_intercept=bool(include_intercept),
        )
    except ValueError:
        # A group-level estimator should preserve missingness for groups that
        # cannot support the requested OLS specification rather than aborting
        # an otherwise valid firm-month batch.
        return None

    coefficient_index = 1 if include_intercept else 0
    estimate = float(result.coefficients[coefficient_index])
    return estimate if np.isfinite(estimate) else None</code></pre></div><p>The <code>flow.reshape(-1, 1)</code> expression makes the design explicitly two-dimensional, while <code>prices</code> remains a one-dimensional response. With an intercept, coefficient index <code>1</code> is the flow slope because the intercept occupies index <code>0</code>; without one, index <code>0</code> is the only coefficient. <code>fit_ols</code> validates finite values, row counts, rank, and residual degrees of freedom. At the group level, an unsupported specification becomes <code>None</code>, preserving estimator-specific missingness rather than stopping all firm-month processing.</p><p>A group also returns <code>None</code> when its signed-flow vector has no usable magnitude or when there are too few observations for the selected design. This is important for a no-intercept regression: a zero-flow group cannot identify a meaningful price-impact slope. The paper&#8217;s broader sample rule retains firm-months with at least 15 nonzero-volume trading days, but it does not state a separate minimum specifically for the intramonth regression. The upstream filter and these estimator-level degeneracy checks therefore serve different purposes.</p><h3>Method A: Amihud-style illiquidity</h3><p>The second estimator converts each valid daily observation into an absolute-return-per-dollar-volume ratio and averages those ratios within the firm-month. The paper records this construction as follows.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\hat{\\lambda}^{\\mathrm{Amihud}}_{it} = \\frac{1}{n} \\sum_{\\tau \\in t} \\frac{|r_{i\\tau}|}{\\mathrm{DollarVolume}_{i\\tau}}&quot;,&quot;id&quot;:&quot;C9045B8AFF&quot;}" data-component-name="LatexBlockToDOM"></div><p>In this equation, <code>\hat{\lambda}^{\mathrm{Amihud}}_{it}</code> is the nonnegative Amihud-style estimate for firm <code>i</code> and month <code>t</code>. The integer <code>n</code> is the number of valid daily observations used in the mean. <code>r_{i\tau}</code> is the daily stock return, and <code>\mathrm{DollarVolume}_{i\tau}</code> is daily dollar trading volume. The absolute-value operator makes the numerator nonnegative, while the denominator must be strictly positive.</p><p>The paper also labels the same construction Method A:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\hat{\\lambda}^{A}_{it} = \\frac{1}{n} \\sum_{\\tau \\in t} \\frac{|r_{i\\tau}|}{\\mathrm{DollarVolume}_{i\\tau}}.&quot;,&quot;id&quot;:&quot;97987BEB9C&quot;}" data-component-name="LatexBlockToDOM"></div><p>Equations <code>eq_13</code> and <code>eq_15</code> therefore map to the same generated function, <code>estimate_amihud_lambda</code>. The function excludes rows with nonpositive dollar volume, so a zero denominator does not create an infinite ratio. Negative dollar-volume inputs are rejected as invalid. A valid result is nonnegative provided at least one daily observation remains.</p><p>The focused implementation is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7a273366-2941-45da-9ad3-f8a4474a4699&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def estimate_amihud_lambda(
    return_: np.ndarray,
    dollar_volume: np.ndarray,
) -&gt; float | None:
    """Estimate the within-firm-month Amihud-style illiquidity measure.

    The inputs are paired daily vectors with shape ``(n_days,)``.  Observations
    with non-finite returns or non-positive dollar volume are excluded, since
    the denominator in eq. 13 and eq. 15 must be strictly positive.  The
    resulting estimate is nonnegative whenever at least one valid observation
    remains.
    """
    returns = _as_float_vector(return_, "return_")
    dollars = _as_float_vector(dollar_volume, "dollar_volume")
    if returns.size != dollars.size:
        raise ValueError(
            "return_ and dollar_volume must have equal lengths: "
            f"{returns.size} != {dollars.size}."
        )
    if bool((dollars &lt; 0).any()):
        raise ValueError("dollar_volume must be nonnegative.")

    valid = np.isfinite(returns) &amp; np.isfinite(dollars) &amp; (dollars &gt; 0.0)
    if not bool(valid.any()):
        return None

    # Eq. 13 and Eq. 15: lambda_hat = mean(|r| / DollarVolume).
    ratios = np.abs(returns[valid]) / dollars[valid]
    if not np.isfinite(ratios).all():
        return None
    estimate = float(np.mean(ratios, dtype=np.float64))
    if estimate &lt; 0.0 or not np.isfinite(estimate):
        return None
    return estimate</code></pre></div><p>Notice the distinction between invalid input and unavailable estimation. A negative dollar volume raises an error because it violates the data contract. A group with no positive dollar-volume observations returns <code>None</code>, because no valid Amihud mean can be formed. The function uses float64 accumulation for the within-group mean and returns a scalar float when the calculation is defined.</p><h3>Grouped firm-month computation</h3><p><code>estimate_firm_month_lambdas</code> applies both estimators independently to every <code>(PERMNO, month)</code> group. It requires <code>price_change</code> and <code>signed_flow</code>, plus a daily return column named <code>DlyRet</code>, <code>return</code>, or <code>return_</code>, and a <code>dollar_volume</code> column. It sorts by firm and month, converts the daily columns to numeric arrays, computes separate validity masks, and records the two estimates and their valid-day counts.</p><p>A compact excerpt shows the separate paths:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6d34fc46-e44b-4e20-920f-c47f817a1956&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        regression_valid = np.isfinite(price_values) &amp; np.isfinite(flow_values)
        amihud_valid = (
            np.isfinite(return_values)
            &amp; np.isfinite(dollar_values)
            &amp; (dollar_values &gt; 0.0)
        )
        regression_lambda = (
            estimate_regression_lambda(
                price_values[regression_valid],
                flow_values[regression_valid],
                include_intercept=bool(include_intercept),
            )
            if bool(regression_valid.any())
            else None
        )
        amihud_lambda = (
            estimate_amihud_lambda(
                return_values[amihud_valid],
                dollar_values[amihud_valid],
            )
            if bool(amihud_valid.any())
            else None
        )</code></pre></div><p>The two masks need not select the same days. For example, a daily price change and signed flow may be available while the daily return is missing, or a return may be available while dollar volume is zero. Consequently, <code>regression_valid_days</code> and <code>amihud_valid_days</code> are reported separately. The output contract is one row per firm-month key, with missingness retained independently for <code>regression_lambda</code> and <code>amihud_lambda</code>.</p><h3>Lambda and the subsequent return target</h3><p>Estimating lambda is only the first step. The current-month estimate must be aligned with the same firm&#8217;s subsequent calendar-month return. The paper writes the return relationship as:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\mathrm{StockRet}_{i,t+1} = \\alpha_i + \\beta_i \\hat{\\lambda}_{it} + \\varepsilon_{i,t+1}.&quot;,&quot;id&quot;:&quot;235887C0C6&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here, <code>\mathrm{StockRet}_{i,t+1}</code> is firm <code>i</code>&#8217;s return in the month after the estimate; <code>\hat{\lambda}_{it}</code> is the current-month lambda; <code>\alpha_i</code> and <code>\beta_i</code> are regression coefficients in the paper&#8217;s notation; and <code>\varepsilon_{i,t+1}</code> is the residual. This equation describes a later return regression, not the daily lambda estimator itself.</p><p>The supplied paper uses firm subscripts on the coefficients but also describes pooled observations and tables. The generated reproduction therefore treats pooled versus firm-specific estimation as an explicit modeling issue rather than assuming that the notation settles it. The alignment module uses an explicit <code>PERMNO</code> and month-plus-one join, so a missing calendar month does not silently become an adjacent observation.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;598642ac-4626-422b-bafc-f5a754011a98&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    result["next_month"] = result[_MONTH] + pd.DateOffset(months=1)

    lookup = result[[_PERMNO, _MONTH, return_column]].rename(
        columns={_MONTH: "_target_month", return_column: "next_month_return"}
    )
    aligned = result.merge(
        lookup,
        left_on=[_PERMNO, "next_month"],
        right_on=[_PERMNO, "_target_month"],
        how="left",
        sort=False,
        validate="one_to_one",
    ).drop(columns=["_target_month"])</code></pre></div><p>The join preserves the same firm identifier and requests exactly one calendar month later. The resulting <code>next_month_return</code> is then available to a later pooled regression; it is not used when computing the current-month lambda.</p><h3>Worked example: vector contracts</h3><p>Consider a synthetic firm-month with four daily observations. The intended inputs have these shapes:</p><ul><li><p><code>price_change</code>: <code>(4,)</code></p></li><li><p><code>signed_flow</code>: <code>(4,)</code></p></li><li><p><code>return_</code>: <code>(4,)</code></p></li><li><p><code>dollar_volume</code>: <code>(4,)</code></p></li></ul><p>A caller can request both estimators through the public functions:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;0e4a5fb3-8126-4ea4-a7ff-c7bd4276e328&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import importlib
import numpy as np

lambda_module = importlib.import_module(
    "liquidity_premium.estimators.lambda"
)

price_change = np.array([0.10, -0.04, 0.08, 0.02], dtype=float)
signed_flow = np.array([100.0, -50.0, 75.0, 25.0], dtype=float)
daily_return = np.array([0.010, -0.004, 0.008, 0.002], dtype=float)
dollar_volume = np.array(
    [10_000.0, 8_000.0, 12_000.0, 9_000.0],
    dtype=float,
)

regression_lambda = lambda_module.estimate_regression_lambda(
    price_change,
    signed_flow,
    include_intercept=False,
)
amihud_lambda = lambda_module.estimate_amihud_lambda(
    daily_return,
    dollar_volume,
)</code></pre></div><p>This excerpt documents the API and shape transition only; it was not executed under the run policy, so no numerical output is claimed. The first result is a scalar regression slope, and the second is a scalar mean ratio. If the same firm-month supplied a zero vector for <code>signed_flow</code>, the regression estimator would return <code>None</code> because the no-intercept slope is degenerate. The Amihud estimator could still be valid if its return and positive dollar-volume vectors remained usable.</p><h3>Method B&#8217;s incomplete source equation</h3><p>The paper describes Method B as the slope from regressing daily price changes on daily volume multiplied by the sign of the daily price change. The generated code maps that prose to the same signed-flow regression path used for <code>eq_12</code>. However, the canonical LaTeX for <code>eq_16</code> is OCR-truncated. No missing formula, intercept convention, or additional term is reconstructed here. The implementation follows only the supplied prose and the complete displayed relationship in <code>eq_12</code>.</p><p>Finally, do not conflate two independent choices: <code>include_intercept</code> on <code>estimate_regression_lambda</code> controls the intramonth price-impact estimator, while <code>include_intercept</code> on the later lambda-return regression controls the return model. The paper&#8217;s estimator ambiguity is therefore isolated in the API rather than silently propagated to every downstream specification.</p><h2>Contemporaneous, predictive, and lambda-return regressions</h2><p>How should a monthly activity measure be connected to a stock return without accidentally using the wrong firm or the wrong month? The reproduction treats every regression row as a keyed observation: <code>PERMNO</code> identifies the firm, <code>month</code> identifies the predictor month, and the target is either that same row&#8217;s return or the same firm&#8217;s return in the immediately following calendar month.</p><p>This section covers three related specifications. The first asks whether monthly volume features are associated with contemporaneous returns. The second uses those features to forecast the next month&#8217;s return. The third replaces the volume features with one firm-month lambda estimate at a time. All three are pooled ordinary least squares (OLS) regressions: observations from eligible firms and months enter one design matrix. The implementation does not constrain coefficient signs.</p><h3>The common panel contract</h3><p>The panel functions expect one row per <code>PERMNO</code>-<code>month</code> key. The base predictor columns are:</p><ul><li><p><code>sumvolume</code>: total unsigned share volume during the month;</p></li><li><p><code>stdvolume</code>: within-month sample standard deviation of daily volume;</p></li><li><p><code>signedflow</code>: monthly signed-flow proxy.</p></li></ul><p>The target column is <code>StockRet</code> for the contemporaneous specification. For a one-month-ahead specification, the code creates <code>next_month_return</code> by matching the current firm and the next calendar month. Optional controls are appended by name. The paper discusses controls such as log market capitalization, book-to-market, momentum, and Amihud illiquidity, but unavailable external controls must remain missing rather than being fabricated.</p><p>The design-building function returns a predictor matrix with shape <code>(n_used_observations, p)</code>, a target vector with shape <code>(n_used_observations,)</code>, and labels preserving predictor order. The shared <code>fit_ols</code> function then optionally prepends an intercept column. With three base predictors and an intercept, the coefficient vector has shape <code>(4,)</code>. The residual vector has length <code>n_used_observations</code>.</p><h3>Equation 10: contemporaneous activity and returns</h3><p>The contemporaneous specification asks whether firm-month activity and the return in that same firm-month are related. The paper records it as follows.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\mathrm{StockRet}_{it} = \\alpha + \\beta_1 \\mathrm{sumvolume}_{it} + \\beta_2 \\mathrm{stdvolume}_{it} + \\beta_3 \\mathrm{signedflow}_{it} + \\varepsilon_{it}.&quot;,&quot;id&quot;:&quot;7015A1024B&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here, <code>StockRet_it</code> is the monthly stock return for firm <code>i</code> in month <code>t</code>. The three predictors are total volume <code>sumvolume_it</code>, volume volatility <code>stdvolume_it</code>, and signed flow <code>signedflow_it</code>. <code>alpha</code> is the pooled regression intercept; <code>beta_1</code>, <code>beta_2</code>, and <code>beta_3</code> are the corresponding slopes; and <code>epsilon_it</code> is the residual for that firm-month observation.</p><p>In the generated implementation, <code>run_panel_return_regression</code> selects this equation with <code>horizon="contemporaneous"</code>. It delegates row construction to <code>build_panel_dataset</code>, creates the ordered predictor list, and passes the resulting arrays to <code>build_panel_design</code> and then <code>fit_ols</code> with an intercept.</p><p>A focused call is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;1a16e3d8-6c29-4cfe-9c40-cb43b6c4648c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from liquidity_premium.regressions.panel import run_panel_return_regression

result = run_panel_return_regression(
    panel,
    horizon="contemporaneous",
    controls=None,
)</code></pre></div><p>The no-control design has columns in this order:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;ca1ab596-084d-437a-bfc3-e4a32df5c6ac&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">sumvolume, stdvolume, signedflow</code></pre></div><p>The returned labels add the specification context, including <code>contemporaneous.no_controls.intercept</code>. Rows with a missing target or missing required predictor are removed before fitting. The OLS primitive requires finite values, more observations than fitted parameters, and a full-rank design. These are numerical input conditions, not economic assumptions.</p><p>The paper does not fully specify fixed effects, weighting, clustering, or panel covariance treatment. The generated panel module therefore exposes a covariance argument but currently uses the shared classical homoskedastic OLS inference. That implementation choice must be recorded in result metadata and must not be presented as uniquely determined by the paper.</p><h3>Equation 11: exact one-month-ahead alignment</h3><p>The predictive version uses month-<code>t</code> activity to explain the same firm&#8217;s return in month <code>t+1</code>.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\mathrm{StockRet}_{i,t+1} = \\alpha + \\beta_1 \\mathrm{sumvolume}_{it} + \\beta_2 \\mathrm{stdvolume}_{it} + \\beta_3 \\mathrm{signedflow}_{it} + \\varepsilon_{i,t+1}.&quot;,&quot;id&quot;:&quot;3613071F77&quot;}" data-component-name="LatexBlockToDOM"></div><p>The notation changes only the target timing: <code>StockRet_i,t+1</code> is the next calendar month&#8217;s return for the same firm. The predictors remain <code>sumvolume_it</code>, <code>stdvolume_it</code>, and <code>signedflow_it</code>, all measured in month <code>t</code>. <code>alpha</code> and the three <code>beta</code> coefficients retain their regression roles, while <code>epsilon_i,t+1</code> is the forecast-regression residual.</p><p>The important implementation detail is that &#8220;next month&#8221; means an explicit calendar key, not merely the next available row for a firm. <code>align_next_month_return</code> normalizes month values, sorts by <code>PERMNO</code> and month, creates a month-plus-one key, and joins the return lookup on both <code>PERMNO</code> and target month. Thus, a January observation with no February row does not receive a March return as its target.</p><p>The core alignment operation is represented by this generated code excerpt:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f81ca492-3531-4681-be71-e04c3d7c1f32&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">result["next_month"] = result[_MONTH] + pd.DateOffset(months=1)

lookup = result[[_PERMNO, _MONTH, return_column]].rename(
    columns={_MONTH: "_target_month", return_column: "next_month_return"}
)
aligned = result.merge(
    lookup,
    left_on=[_PERMNO, "next_month"],
    right_on=[_PERMNO, "_target_month"],
    how="left",
    sort=False,
    validate="one_to_one",
).drop(columns=["_target_month"])</code></pre></div><p><code>_PERMNO</code> and <code>_MONTH</code> are the firm and calendar-month key columns. The lookup contains one return per firm-month. The <code>one_to_one</code> validation expresses the uniqueness contract, while the two-key merge prevents cross-firm leakage. Missing next-month observations remain missing and are later removed when <code>build_panel_dataset</code> prepares complete regression rows.</p><p>The predictive call changes only the horizon argument:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c31c0913-ba47-4dcf-be02-b5c6b9928944&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">next_month_result = run_panel_return_regression(
    panel,
    horizon="next_month",
    controls=None,
)</code></pre></div><p>With three base predictors and an intercept, this again produces four coefficients, but the retained observation count can be smaller because firms&#8217; final available months have no following calendar-month target. The generated tests describe this behavior for synthetic panels, but those tests were not executed under the current run policy.</p><h3>Adding controls without hiding missing data</h3><p>Controls are appended after the three base predictors. For example, a fully populated controlled specification can be requested with:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b16419ab-443e-4204-b6b8-12d2d20f88cb&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">controlled_result = run_panel_return_regression(
    panel,
    horizon="next_month",
    controls=[
        "log_size",
        "book_to_market",
        "momentum",
        "amihud_lambda",
    ],
)</code></pre></div><p>If all four columns are available and numeric, the design contains seven economic predictors plus an intercept. The names must already exist in <code>panel</code>; the regression function raises an error for a missing requested column rather than inventing a substitute. This matters especially for book-to-market and factor-related inputs, which depend on external WRDS or factor files not supplied in the paper context.</p><p>The implementation keeps predictor units unstandardized. The paper does not specify a scaling convention, so raw economic units are retained for coefficient reporting. Every result should identify whether it uses no controls or controls and whether its return is raw, delisting-adjusted, or excess, because the supplied paper uses return conventions in different contexts without fully resolving them.</p><h3>Equation 14: lambda as the sole predictor</h3><p>The lambda-return specification tests whether a current firm-month illiquidity estimate predicts the same firm&#8217;s next-month return.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\mathrm{StockRet}_{i,t+1} = \\alpha_i + \\beta_i \\hat{\\lambda}_{it} + \\varepsilon_{i,t+1}.&quot;,&quot;id&quot;:&quot;F065190037&quot;}" data-component-name="LatexBlockToDOM"></div><p><code>StockRet_i,t+1</code> is the next-month return for firm <code>i</code>. The predictor <code>hat(lambda)_it</code> is the lambda estimate formed from firm <code>i</code>&#8217;s daily observations in month <code>t</code>; it may be the regression-based estimator or the Amihud-style estimator. <code>alpha_i</code> and <code>beta_i</code> are the intercept and slope notation used in the paper, and <code>epsilon_i,t+1</code> is the residual.</p><p>The generated implementation uses a pooled interpretation by default. Although the equation uses firm-specific subscripts on the coefficients, the supplied descriptions and reported pooled observation structure support fitting one regression over the eligible firm-month rows rather than silently fitting a separate regression per firm.</p><p>The lambda-return function recomputes same-firm next-month alignment through <code>align_next_month_return</code>, selects exactly one estimator column, removes rows with missing or non-finite estimator or target values, and calls <code>fit_ols</code>. A focused example is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7a2972c3-1758-4a5b-b0c0-8f9e88390be4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from liquidity_premium.regressions.lambda_return import (
    run_lambda_return_regression,
)

regression_lambda_with_intercept = run_lambda_return_regression(
    panel,
    estimator="regression_lambda",
    include_intercept=True,
)

amihud_lambda_uncentered = run_lambda_return_regression(
    panel,
    estimator="amihud_lambda",
    include_intercept=False,
)</code></pre></div><p>The <code>include_intercept</code> argument belongs to this return regression. It is separate from the convention used earlier when estimating the firm-month regression lambda from daily price changes and signed flow. The generated default for that daily estimator is uncentered because the displayed price-impact equation omits an intercept; the return regression deliberately exposes both alternatives.</p><p>With an intercept, the lambda predictor matrix supplied to <code>fit_ols</code> has shape <code>(n_used_observations, 1)</code>, and the fitted coefficient vector has shape <code>(2,)</code>: intercept and lambda slope. Without an intercept, the coefficient vector has shape <code>(1,)</code> and contains only the lambda slope. The generated test specifies these structural dimensions without asserting any expected economic sign.</p><h3>Comparing estimator and intercept specifications</h3><p>The paper emphasizes that direct lambda-return relationships can be sensitive to both estimator construction and intercept treatment. <code>run_lambda_specification_grid</code> makes those comparisons explicit:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;1a8fcd64-6e4f-46ba-9062-8f8faf51d1e4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from liquidity_premium.regressions.lambda_return import (
    run_lambda_specification_grid,
)

lambda_grid = run_lambda_specification_grid(panel)</code></pre></div><p>The grid contains four combinations:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;b7d0dce9-f3bb-4a0f-b3ff-2b4e50b6e1fe&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">regression_lambda, with intercept
regression_lambda, uncentered
amihud_lambda, with intercept
amihud_lambda, uncentered</code></pre></div><p>Each row stores estimator and intercept labels, coefficient arrays, standard errors, t-statistics, R-squared, and observation count. The two lambda columns are never combined into one regression by this helper, so estimator differences remain visible.</p><h3>A small timing example</h3><p>Consider two firms, <code>A</code> and <code>B</code>, observed in January, February, and March. The January row for <code>A</code> receives <code>A</code>&#8217;s February return as its next-month target; the January row for <code>B</code> receives <code>B</code>&#8217;s February return. February rows similarly receive March returns. March rows have no target unless April is present. A January row for <code>A</code> must never receive <code>B</code>&#8217;s February return, and a January row for a firm with no February record must not receive that firm&#8217;s March return.</p><p>For the contemporaneous model, all six firm-month rows can be eligible if their current-month returns and predictors are present. For the next-month model, only rows with an exact same-firm calendar-next-month return are eligible. The resulting base design has three columns and shape <code>(n_used_observations, 3)</code> before the intercept is added. An uncentered lambda-return design has one column and shape <code>(n_used_observations, 1)</code>.</p><p>This example explains key and shape behavior only. It is not a numerical result and has not been executed.</p><h3>Interpreting signs and inference boundaries</h3><p>The paper&#8217;s narrative sometimes suggests a positive predictive role for signed flow, while the supplied results report negative signed-flow coefficients in the one-month-ahead specifications. The reproduction must preserve those reported signs rather than flip them to match a hypothesis. A coefficient sign is an output of a particular data, return, alignment, winsorization, and inference specification; it is not an input restriction.</p><p>Similarly, classical OLS standard errors and t-statistics are available from the generated <code>fit_ols</code> implementation, but the paper does not uniquely specify panel clustering, fixed effects, weighting, or covariance treatment. The generated code was not executed, and local static verification and code semantic verification were skipped under the authoritative run policy. Consequently, this section documents the equation-to-code mapping and structural contracts, not verified numerical agreement with the paper&#8217;s tables.</p><h2>Expanding-window forecasts and out-of-sample evaluation</h2><p>How can a return forecast imitate an investor making predictions in real time? Start with an initial history, estimate a model using only that history, predict the next available month, then enlarge the history and repeat. This is the intuition behind an <strong>expanding-window forecast</strong>. Unlike a rolling window, which drops old observations, an expanding window retains all eligible historical observations as new information arrives.</p><p>The paper describes this procedure using the first 30 percent of each firm&#8217;s chronological observations as the initial training period. The source alternates between the words &#8220;rolling&#8221; and &#8220;expanding,&#8221; but its explicit procedural description supports the expanding interpretation used here. The generated implementation also makes the integer-rounding rule explicit: the initial length is the ceiling of 30 percent of the firm&#8217;s observations. Because the paper does not specify a rounding rule, this is an implementation decision rather than a uniquely paper-determined fact.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/p/quant-trading-kyles-price-impact?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/p/quant-trading-kyles-price-impact?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/onepagecode.substack.com/p/quant-trading-kyles-price-impact?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><h3>Inputs and timing contract</h3><p><code>generate_expanding_forecasts()</code> expects a firm-month <code>pandas.DataFrame</code> containing one row per <code>PERMNO</code> and calendar <code>month</code>. The required return column is <code>StockRet</code>. The function accepts a risk-free-rate column named <code>rf</code>, <code>risk_free_rate</code>, or <code>r_f</code>, and a lambda column named <code>lambda</code>, <code>lambda_estimate</code>, <code>regression_lambda</code>, or <code>amihud_lambda</code>.</p><p>The shape flow is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;ab0def75-d74d-4e05-9f4e-5b2ac96a8f3f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">firm-month panel:        (n_rows, p)
firm-local training:     (n_training_rows, p)
forecast records:        (n_oos_observations, q)
pooled evaluation sample: (n_oos_observations, 2 predictor/target fields)</code></pre></div><p>For each firm, the implementation sorts rows by calendar month and treats adjacent calendar months as a valid forecast transition. If the next row is not exactly one month later, that forecast origin is skipped rather than treated as an artificial adjacent observation. Missing predictors, missing realized returns, and insufficient training observations are also skipped; the code does not fabricate values.</p><p>At forecast origin <code>j</code>, the current eligible observation is included in the training sample. The subsequent calendar month supplies the predictor values used for the forecast and the realized return used later for evaluation. This ordering is the central no-look-ahead rule: the target return must not enter the model that predicts it.</p><h3>Equation 17: the recursive forecasting model</h3><p>The paper&#8217;s recursive forecasting relation is:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\mathrm{ActualReturn}_t = a + b_1 r_{f,t} + b_2 \\hat{\\lambda}_{it} + \\epsilon_t&quot;,&quot;id&quot;:&quot;28CAA4DE1E&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here, <code>ActualReturn_t</code> is the realized return at time <code>t</code>; <code>a</code> is the intercept; <code>b_1</code> is the coefficient on the risk-free rate <code>r_f,t</code>; <code>b_2</code> is the coefficient on the current firm-month lambda estimate <code>hat(lambda)_it</code>; and <code>epsilon_t</code> is the regression residual. In the generated implementation, the model is fit with an intercept and two predictors in the order risk-free rate, then lambda.</p><p><code>fit_recursive_forecast_model()</code> in <code>src/liquidity_premium/forecast/expanding.py</code> receives a training <code>DataFrame</code>, removes rows missing any required variable, constructs a design matrix with shape <code>(n_training_rows, 2)</code>, and calls <code>fit_ols()</code> from <code>src/liquidity_premium/estimators/ols_core.py</code>. With an intercept, <code>fit_ols()</code> adds a column of ones, so the fitted coefficient vector has shape <code>(3,)</code>: intercept, risk-free-rate coefficient, and lambda coefficient.</p><p>A focused excerpt shows the equation-to-code mapping:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c32349c0-f459-4e66-b75a-80157ebba1a4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    x = complete[[risk_free_column, lambda_column]].to_numpy(dtype=np.float64)
    y = complete[_RETURN].to_numpy(dtype=np.float64)
    if complete.shape[0] &lt;= x.shape[1] + 1:
        raise ValueError("training must contain more observations than fitted parameters.")

    # Eq. 17: ActualReturn_t = a + b_1 r_{f,t} + b_2 lambda_it + epsilon_t.
    result = fit_ols(x, y, include_intercept=True)
    result.labels = ("intercept", "risk_free_rate", "lambda")</code></pre></div><p>The recursive model is estimated separately for each firm in the generated workflow. This follows the firm-level method card, although the paper&#8217;s notation and discussion do not fully resolve whether the forecasting relation should instead be pooled, firm-specific, or hierarchical. That ambiguity should remain visible in provenance metadata.</p><h3>Building the expanding forecasts</h3><p>The public function <code>generate_expanding_forecasts(panel, initial_fraction=0.30)</code> performs the chronological loop. Its algorithm is:</p><ol><li><p>Normalize and sort the panel by <code>PERMNO</code> and month.</p></li><li><p>For each firm, compute <code>ceil(0.30 * n_rows)</code> as the initial window length.</p></li><li><p>Use the current origin and all earlier rows as the candidate training history.</p></li><li><p>Remove incomplete training rows and fit Equation 17 when enough observations remain.</p></li><li><p>Use the next calendar month&#8217;s risk-free rate and lambda as predictors.</p></li><li><p>Store the predicted return and the subsequent realized return.</p></li><li><p>Continue with a nondecreasing training endpoint.</p></li></ol><p>The output records include the firm identifier, origin month, forecast month, actual and predicted returns, training start and end dates, training-observation count, estimator label, and frequency label. The intended output shape is one row per eligible firm-month forecast origin.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c98e1c0a-8361-48e5-a110-60ebdabf3142&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from liquidity_premium.forecast.expanding import generate_expanding_forecasts

forecasts = generate_expanding_forecasts(
    panel,
    initial_fraction=0.30,
)</code></pre></div><p>The first forecast for a firm with <code>n_rows</code> observations uses row <code>ceil(0.30 * n_rows) - 1</code> as its origin, because the current origin is included in training. The next row must be the immediately following calendar month. For example, with 10 ordered monthly rows, the ceiling-based initial length is 3. The first origin is therefore the third row, and the first target is the fourth month. This example explains the indexing convention only; it is synthetic and has not been executed.</p><p>The generated implementation validates several timing invariants through <code>validate_forecast_timing()</code>. It requires unique firm-forecast-month keys, a training endpoint earlier than the forecast month, a positive training-observation count, finite actual and predicted returns, and nondecreasing training endpoints within each firm.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ce63cc07-fd8a-4909-b989-c7f5a8ceaec7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from liquidity_premium.forecast.expanding import validate_forecast_timing

validate_forecast_timing(forecasts)</code></pre></div><p>Calling this function is part of the intended workflow, not evidence that the workflow has been run. Under the authoritative run policy, no code execution, test execution, static verification, or semantic code verification was performed.</p><h3>Equation 18: calibration of stored forecasts</h3><p>Forecast generation and forecast evaluation answer different questions. Generation asks what prediction would have been available at each historical forecast origin. Calibration evaluates the relationship between those stored predictions and their realized outcomes. The paper specifies that pooled evaluation with:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\mathrm{ActualReturn} = \\alpha + \\beta \\times \\mathrm{PredictedReturn} + \\varepsilon&quot;,&quot;id&quot;:&quot;70B0BA6352&quot;}" data-component-name="LatexBlockToDOM"></div><p>In this equation, <code>ActualReturn</code> is the observed out-of-sample return, <code>PredictedReturn</code> is the previously stored forecast, <code>alpha</code> is the calibration intercept, <code>beta</code> measures the calibration slope, and <code>epsilon</code> is the evaluation residual. A calibration regression is not a replacement for the recursive forecasting step; it is applied only after forecasts have been generated.</p><p><code>evaluate_oos_predictions()</code> in <code>src/liquidity_premium/forecast/evaluation.py</code> maps <code>actual_return</code> to the response and <code>predicted_return</code> to the single predictor. It requires an explicit boolean <code>is_out_of_sample</code> or <code>out_of_sample</code> column and filters to rows marked true before fitting Equation 18.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8a181b5c-01aa-4f36-a747-06a965b60951&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from liquidity_premium.forecast.evaluation import evaluate_oos_predictions

oos_evaluation = evaluate_oos_predictions(forecasts)</code></pre></div><p>The core implementation is deliberately narrow:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;1e2b3dbb-0b70-4da3-bbe5-1920047d7e0c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    sample = _out_of_sample_rows(forecasts)
    x = sample[["predicted_return"]].to_numpy(dtype=float)
    y = sample["actual_return"].to_numpy(dtype=float)

    # Eq. 18: ActualReturn = alpha + beta * PredictedReturn + epsilon.
    result = fit_ols(x, y, include_intercept=True)</code></pre></div><p>The resulting <code>RegressionResult</code> has two coefficient labels, <code>intercept</code> and <code>predicted_return</code>, and uses only the explicitly marked out-of-sample rows. It also checks that actual and predicted values are numeric and finite. If the marker is absent, the function raises an error instead of guessing which rows are out of sample.</p><p>An important interface issue should be recorded rather than hidden: the generated <code>generate_expanding_forecasts()</code> excerpt does not itself add an <code>is_out_of_sample</code> column, while <code>evaluate_oos_predictions()</code> requires one. A caller must therefore attach an explicit marker or adapt the output contract before evaluation. The supplied test fixture also uses lowercase names such as <code>permno</code> and <code>training_end_month</code>, whereas the generated forecast function emits <code>PERMNO</code> and <code>training_end</code>. These are unresolved generated-file interface mismatches, not results of a completed verification pass.</p><p>For separate estimator and frequency summaries, <code>evaluate_by_method_and_frequency()</code> groups the explicitly marked rows by <code>estimator</code> and <code>frequency</code>, then applies Equation 18 to each group:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;70a3090d-95fe-46fd-9d24-32a800c97944&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from liquidity_premium.forecast.evaluation import (
    evaluate_by_method_and_frequency,
)

method_evaluations = evaluate_by_method_and_frequency(forecasts)</code></pre></div><p>The paper does not specify a unique covariance or clustering treatment for this pooled evaluation. Consequently, inference metadata from the generated OLS result represents an implementation output and should not be presented as the paper&#8217;s uniquely mandated standard-error procedure.</p><h3>Advanced detail: conventions that affect reproducibility</h3><p>The paper&#8217;s risk-free-rate timing is not fully specified, and return conventions differ across parts of the empirical discussion. The workflow should record whether <code>StockRet</code> is a raw return, a delisting-adjusted return, or an excess return, and which month&#8217;s risk-free rate is used. These choices belong in configuration and provenance rather than being inferred silently.</p><p>The paper also leaves the exact pooled-versus-firm-specific interpretation of the recursive model unresolved. The generated implementation uses firm-local recursive estimation, with a separate model fit for each <code>PERMNO</code>. That choice preserves the firm-level expanding-window procedure in the method card, but it should be described as an implementation interpretation.</p><h3>Conceptual worked example</h3><p>Consider one firm with ordered month rows from January through October. A 30 percent initial fraction produces an initial length of three rows under the generated ceiling rule. January through March form the initial history, March is the first forecast origin, and April is the first target month. The model uses the information available through March, including March&#8217;s eligible return, risk-free rate, and lambda. It produces a prediction for April. After that forecast is stored, the training endpoint can move forward for the next origin.</p><p>The important objects are:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;cdd9becb-43d1-4976-ac57-84396cd68a0c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">initial training rows: 3
first origin:          March
first target:          April
forecast record:       firm, origin month, target month,
                       actual return, predicted return,
                       training endpoint and metadata</code></pre></div><p>If a target month is missing&#8212;for example, March is followed by May&#8212;the implementation does not silently treat May as the next month. It skips that transition because the forecast target is not the immediately subsequent calendar month. This preserves the stated one-month horizon and prevents a calendar gap from being mistaken for a valid forecast.</p><p>This example is conceptual and unexecuted. It demonstrates timing, shapes, and responsibilities rather than a numerical forecast or a reproduced paper result.</p><p>The resulting boundary is straightforward: recursive forecasts must be generated from information available at each origin, and Equation 18 must be fit only to stored out-of-sample pairs. The paper&#8217;s reported forecast and Fama&#8211;MacBeth values remain comparison targets; no numerical agreement is claimed here.</p><h2>Fama&#8211;MacBeth inference and signal-sorted portfolios</h2><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>How can we tell whether a return signal works across firms consistently, rather than only in one pooled regression? The paper uses two related procedures. First, a <strong>Fama&#8211;MacBeth regression</strong> fits a separate cross-sectional regression for each month. Second, a signal-sorted portfolio procedure ranks firms within each month and compares the highest-signal group with the lowest-signal group. Both procedures must use information available at portfolio-formation time; next-month returns are outcomes, not ranking inputs.</p><p>The generated implementation treats the firm-month panel as the central input. Each row identifies a firm with <code>PERMNO</code> and a calendar <code>month</code>, contains a current-information signal such as <code>predicted_return</code> or a lambda estimate, and contains a realized next-month return. The resulting monthly coefficient series and portfolio-return series are separate outputs. They should not be replaced by paper-reported values, especially because the supplied run policy disables execution and verification.</p><h3>Monthly cross-sectional regression</h3><p>Fama&#8211;MacBeth estimation separates the cross-sectional question from the time-series inference question. In a given month, the cross-sectional question is whether firms with different signals have different subsequent returns. The time-series question is whether the monthly slope estimates are consistently different from zero across the usable months.</p><p>The paper&#8217;s monthly cross-sectional specification is:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;r_{i,t+1} = \\alpha_t + \\beta_t \\hat{r}^{\\mathrm{pred}}_{i,t+1} + \\varepsilon_{i,t+1}&quot;,&quot;id&quot;:&quot;CC3548EE34&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here, <code>r_{i,t+1}</code> is firm <code>i</code>&#8217;s realized return in the month after the signal is formed. <code>alpha_t</code> is the intercept estimated separately for month <code>t</code>, and <code>beta_t</code> is that month&#8217;s cross-sectional slope. The signal <code>hat(r)^pred_{i,t+1}</code> is the model-implied predicted return associated with firm <code>i</code>; the supplied paper also discusses lambda-based signals, so the code accepts an explicitly selected signal column. The residual <code>epsilon_{i,t+1}</code> is the unexplained part of the firm&#8217;s subsequent return.</p><p>The function <code>run_monthly_cross_section()</code> in <code>src/liquidity_premium/inference/fama_macbeth.py</code> implements one month of this calculation. Its input is a <code>pandas.DataFrame</code> containing exactly one month&#8217;s observations, a numeric signal column, and a realized-return target. It removes rows with nonnumeric or missing signal or target values, requires at least three usable observations, and calls the shared <code>fit_ols()</code> primitive with an intercept. The returned coefficient vector has shape <code>(2,)</code>: the intercept followed by the slope. The function labels these coefficients <code>("intercept", "slope")</code> and records <code>eq_19</code> in the result metadata.</p><p>A focused call looks like this:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;85529b8a-7ab7-42fa-9268-bd61efba6208&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from liquidity_premium.inference.fama_macbeth import run_monthly_cross_section

january = panel.loc[panel["month"] == pd.Timestamp("2020-01-01")]
monthly_result = run_monthly_cross_section(
    january,
    signal="predicted_return",
    target="next_month_return",
)</code></pre></div><p>The important detail is month isolation: <code>january</code> is fitted independently of February or March. The function does not pool all firms and months into one cross-sectional coefficient. Its classical OLS fields describe that one month; the Fama&#8211;MacBeth aggregation happens afterward.</p><h3>From monthly slopes to Fama&#8211;MacBeth inference</h3><p><code>run_fama_macbeth()</code> loops over the <code>month</code> groups in chronological order. For every usable group, it stores one intercept, one slope, the within-month <code>R-squared</code>, and the number of firms used. If no usable months remain, or if the requested HAC lag is incompatible with the number of months, it raises an error rather than silently producing a summary.</p><p>A complete call is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e64e887c-a7e4-49a8-929b-6c7e355a1605&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from liquidity_premium.inference.fama_macbeth import run_fama_macbeth

fm_results = run_fama_macbeth(
    panel,
    signal="predicted_return",
    nw_lags=4,
)</code></pre></div><p>The generated implementation recognizes the first available target among <code>next_month_return</code>, <code>StockRet_next</code>, and <code>target</code>. This makes the target-column contract explicit, but it does not determine which return convention the paper intended. The signal is always supplied by the caller because the paper extraction does not fully resolve whether portfolio and Fama&#8211;MacBeth rankings should use lambda, predicted return, or another model-implied value.</p><p>The resulting DataFrame has one row per usable month. Its monthly columns include <code>intercept</code>, <code>slope</code>, <code>r_squared</code>, and <code>n_observations</code>; summary fields record the mean coefficients, HAC standard errors, HAC t-statistics, selected signal, target, lag, and usable-month count. The mean slope is an average of monthly slopes, not an average over individual firm-level rows. This distinction is the defining aggregation step in the method.</p><h3>Newey&#8211;West inference operates on monthly coefficients</h3><p>Newey&#8211;West, also called HAC or heteroskedasticity-and-autocorrelation-consistent inference, adjusts uncertainty estimates when a time series may have changing variance or serial dependence. Here, the input is not the original firm-month panel. It is the chronological vector of monthly intercepts or slopes, with shape <code>(n_usable_months,)</code>.</p><p>The helper <code>newey_west_mean_inference()</code> in <code>src/liquidity_premium/inference/newey_west.py</code> requires the lag explicitly:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7efbb43d-21c7-4f9f-8738-df9899aa38ce&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from liquidity_premium.inference.newey_west import (
    newey_west_mean_inference,
)

slope_statistics = newey_west_mean_inference(
    monthly_slopes,
    lags=4,
)</code></pre></div><p>The function validates a one-dimensional finite array and requires a nonnegative lag smaller than the number of observations. It applies Bartlett weights to autocovariances and returns a dictionary containing <code>mean</code>, <code>standard_error</code>, <code>t_statistic</code>, <code>n_observations</code>, and <code>lags</code>. The paper states that Newey&#8211;West adjustment is used but does not specify the lag length. Therefore, <code>4</code> is only an implementation example; it is not a paper-imposed value.</p><p>This ordering matters:</p><ol><li><p>Fit one intercept-and-slope regression per month.</p></li><li><p>Sort those monthly estimates chronologically.</p></li><li><p>Average the monthly estimates.</p></li><li><p>Apply HAC inference to each coefficient time series.</p></li></ol><p>Applying HAC directly to all firm-level rows would answer a different statistical question. The generated test fixture in <code>tests/test_fama_macbeth_portfolios.py</code> is designed to check month-local regression isolation and the propagation of the configured lag, but the test suite was not executed under the authoritative run policy.</p><h3>Signal-sorted deciles without look-ahead</h3><p>A portfolio sort translates a continuous signal into groups. For each month, <code>assign_monthly_deciles()</code> ranks firms from low to high using only the selected current-month signal. It assigns integer labels from <code>1</code> through <code>10</code>, represented in the return table as <code>D1</code> through <code>D10</code>. The next-month return is deliberately not read during assignment.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;9b95cb07-4ad6-48b0-99c7-5c049f0a31cd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from liquidity_premium.portfolios.deciles import assign_monthly_deciles

ranked_panel = assign_monthly_deciles(
    panel,
    signal="predicted_return",
    n_deciles=10,
)</code></pre></div><p>The generated function leaves missing or nonfinite signals unassigned. Ties use pandas&#8217; <code>first</code> ranking policy, so tied observations are resolved by their input order. If a month contains fewer than ten eligible firms, the implementation retains the available firms in lower-numbered deciles and leaves unavailable higher deciles missing. These are explicit implementation decisions because the supplied paper does not specify tie handling or small-cross-section behavior.</p><p>The paper also does not uniquely specify portfolio weighting. The generated <code>compute_decile_returns()</code> function supports equal weighting and value weighting. Equal weighting computes the arithmetic mean of valid realized returns within each month and decile. Value weighting requires a positive, finite <code>market_cap</code> column.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;1d9a95f0-3c4c-4087-9e55-a70d6815259c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from liquidity_premium.portfolios.deciles import compute_decile_returns

decile_returns = compute_decile_returns(
    ranked_panel,
    return_column="next_month_return",
    weighting="equal",
)</code></pre></div><p>For value weighting, the corresponding call is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a0fc487f-fb13-4440-865f-bf019df15e96&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">value_weighted_deciles = compute_decile_returns(
    ranked_panel,
    return_column="next_month_return",
    weighting="value",
)</code></pre></div><p>The output is a DataFrame indexed by month. Its normal shape is <code>(n_months, 10)</code>, with columns <code>D1</code> through <code>D10</code>. Months without a valid return for a particular decile retain a missing value rather than being silently removed. Portfolio formation still depends only on the current signal; <code>return_column</code> is used only after the decile labels exist.</p><h3>The D10&#8211;D1 spread and Sharpe ratios</h3><p>The long-short portfolio is formed by subtracting the lowest-signal portfolio from the highest-signal portfolio in the same month. The generated function performs this row-wise operation:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e4e53f7f-b1f3-452e-be31-42b89c8fec77&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from liquidity_premium.portfolios.deciles import compute_long_short_returns

long_short = compute_long_short_returns(decile_returns)</code></pre></div><p><code>compute_long_short_returns()</code> requires unique monthly indexes and both <code>D1</code> and <code>D10</code> columns. Its output has shape <code>(n_months,)</code> and is named <code>D10_minus_D1</code>. It does not rerank firms, filter on returns, or average D10 and D1 over different months. The spread is a same-month subtraction, preserving the portfolio method&#8217;s timing invariant.</p><p><code>compute_sharpe_ratio()</code> in <code>src/liquidity_premium/portfolios/performance.py</code> summarizes a monthly return series. The generated implementation uses the sample standard deviation and multiplies the monthly ratio by <code>sqrt(12)</code>. If a risk-free series is supplied, it must have exactly the same monthly index and is subtracted without conversion. The paper does not uniquely specify annualization or the conversion of an H.15 yield into a monthly return, so these are implementation decisions rather than paper facts.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;085db9d3-ab43-47f8-b00e-312374c250f0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from liquidity_premium.portfolios.performance import compute_sharpe_ratio

sharpe = compute_sharpe_ratio(d10_returns)</code></pre></div><p><code>summarize_portfolio_performance()</code> collects the decile table, the D10&#8211;D1 series, and Sharpe statistics into a <code>PortfolioResult</code>. It records that weighting was determined upstream and that no risk-free series was used when none was supplied. No paper-reported Sharpe ratio is inserted into this object.</p><h3>Worked example: conceptual month and coefficient series</h3><p>Consider one conceptual month with a small number of firms whose current signals have already been computed. Suppose the signals are ordered from low to high. The portfolio procedure assigns the lowest eligible observations to <code>D1</code> and the highest eligible observations to <code>D10</code>; with fewer than ten firms, not every decile can be populated. Their realized next-month returns are then averaged within the assigned groups.</p><p>For example, after assignment, the monthly return table might have the conceptual structure below. This is an explanatory shape, not an executed or paper-reproduced result:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;535408bf-3e45-4206-a8c0-a9670c85300d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">month       D1      ...     D10
month_t     r_D1    ...     r_D10</code></pre></div><p>The same-month spread is then <code>r_D10 - r_D1</code>. Importantly, the values in the <code>D1</code> and <code>D10</code> columns are outcomes observed after formation; they do not determine membership.</p><p>For Fama&#8211;MacBeth inference, imagine repeating the cross-sectional regression for months <code>t</code>, <code>t+1</code>, and <code>t+2</code>. The output is not one pooled slope but a chronological series:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;00b45a24-6383-4a4e-9081-492034a99a66&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">monthly slopes: [beta_t, beta_t+1, beta_t+2]</code></pre></div><p><code>newey_west_mean_inference()</code> receives that one-dimensional series and a configured lag. The resulting mean and t-statistic summarize the time series of monthly relationships. Neither this conceptual example nor the code excerpts claim that numerical values were computed in the current run.</p><h3>Evidence boundaries and unresolved choices</h3><p>The supplied paper reports Fama&#8211;MacBeth results over a stated 58 months, including negative mean slopes and Newey&#8211;West t-statistics for the Amihud and regression lambda specifications. Those are paper claims and comparison targets, not independently reproduced values. The extraction also refers inconsistently to Table 9 and Table 10, so this section uses neutral method names rather than treating either table reference as definitive.</p><p>Several choices remain visible in the implementation because the paper does not settle them:</p><ul><li><p>the exact ranking signal may be lambda or predicted return;</p></li><li><p>portfolio weighting may be equal or value weighted;</p></li><li><p>tie handling and incomplete deciles are not specified;</p></li><li><p>the Newey&#8211;West lag length is not stated;</p></li><li><p>the exact usable-month range is unclear;</p></li><li><p>Sharpe-ratio annualization and risk-free conversion are implementation conventions.</p></li></ul><p>The generated files encode these choices as arguments or documented behavior rather than silently harmonizing them. Under the run policy, static verification, semantic code verification, test execution, and numerical reproduction were skipped. The code and examples therefore explain the intended responsibilities and invariants, but they do not establish that the implementation passed those checks or matched the paper&#8217;s reported numbers.</p><h2>Reporting, provenance, limitations, and verification boundaries</h2><p>How do you know whether a reproduction table contains a computed result, a paper target, or merely a planned output? Treat every reported value as part of an audit record. The record should identify the input sample, filtering and transformation conventions, model specification, observation count, and verification status. A number without that context is not yet evidence of reproduction.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share Onepagecode&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/onepagecode.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share Onepagecode</span></a></p><p>The reporting layer in this implementation is deliberately downstream. It consumes a cleaned panel and already-created result objects; it does not download data, fit missing models, or substitute values from the paper. Its central distinction is:</p><ul><li><p>a <strong>computed result</strong> is derived from the local inputs passed to the reporting function;</p></li><li><p>a <strong>paper target</strong> is a value reported by the source paper and supplied only for comparison;</p></li><li><p><strong>provenance</strong> records the configuration and inputs that determine how a computed result was produced; and</p></li><li><p><strong>verification status</strong> states whether execution, tests, or independent checks actually occurred.</p></li></ul><p>Under this run policy, generated code was not executed, statically checked, semantically verified, or subjected to final quality review. The reporting design therefore explains how a future run should label outputs; it does not claim that any table matches the paper.</p><h3>Coverage is derived from the supplied frames</h3><p><code>src/liquidity_premium/reporting/coverage.py</code> provides <code>summarize_coverage(daily, panel)</code>. Its inputs are two <code>pandas.DataFrame</code> objects: a daily frame containing a firm identifier and date, and a firm-month panel containing a firm identifier and month-like key. The function normalizes identifiers and timestamps, counts unique firms and distinct firm-month keys, and returns a table with <code>metric</code>, <code>value</code>, and <code>source</code> columns.</p><p>The function does not hard-code the paper&#8217;s reported 9,893 firms or 448,393 firm-month observations. It counts the rows and normalized keys in the frames it receives. This matters because coverage can change with exchange filters, date bounds, point-in-time membership, delisting treatment, missing-data rules, and adjustment conventions.</p><p>A focused excerpt shows the contract and the provenance label attached to computed rows:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;0439a46a-b3ec-4932-95e0-6d06a43f54b3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">rows = [
    _coverage_row("daily_unique_permnos", int(daily_keys["permno"].nunique()), "computed_from_daily_input"),
    _coverage_row("daily_unique_dates", int(daily_keys["date"].nunique()), "computed_from_daily_input"),
    _coverage_row("daily_min_date", daily_start, "computed_from_daily_input"),
    _coverage_row("daily_max_date", daily_end, "computed_from_daily_input"),
    _coverage_row("daily_firm_month_observations", int(daily_keys.assign(month=daily_keys["date"].dt.to_period("M")).drop_duplicates(["permno", "month"]).shape[0]), "computed_from_daily_input"),
    _coverage_row("panel_unique_permnos", int(panel_keys["permno"].nunique()), "computed_from_panel_input"),
    _coverage_row("panel_firm_month_observations", int(panel_keys.drop_duplicates(["permno", "month"]).shape[0]), "computed_from_panel_input"),
    _coverage_row("panel_min_month", panel_start, "computed_from_panel_input"),
    _coverage_row("panel_max_month", panel_end, "computed_from_panel_input"),
]</code></pre></div><p>The output is suitable for a coverage table, but it is not a claim that the input universe is equivalent to the paper&#8217;s universe. <code>summarize_coverage</code> also fails explicitly when required identifier or time columns are missing, or when <code>PERMNO</code> or dates contain invalid values. Those failures protect the count from silently using an unintended column.</p><p>To compare a computed report with a paper value, use <code>compare_coverage_to_paper(report, targets)</code>. The <code>targets</code> mapping is caller-supplied. The resulting comparison retains <code>computed_value</code>, <code>paper_target</code>, <code>difference</code>, and <code>status</code>, plus a note that no verification is claimed. A difference is a diagnostic, not proof that either implementation is correct.</p><h3>Descriptive statistics retain sample meaning</h3><p><code>src/liquidity_premium/reporting/descriptive.py</code> contains <code>summarize_variables(frame, variables, kurtosis_method)</code>. It accepts a panel and an explicit list of columns, excludes missing and non-finite values separately for each variable, and returns counts, means, sample standard deviations, kurtosis, extrema, and selected quantiles. The count is the number of finite numeric observations actually used for that variable, not automatically the total number of panel rows.</p><p>The helper <code>build_table_2_to_4_summaries(panel, kurtosis_method="pandas")</code> organizes available return, volume-feature, and lambda columns into paper-aligned groups. It uses only columns present in the local panel. Missing external variables are not fabricated, and an absent group may produce an empty summary.</p><p>Raw and winsorized samples must remain distinguishable. <strong>Winsorization</strong> clips extreme observations to configured percentile bounds; it does not remove rows. The paper calls for 1st and 99th percentile winsorization but does not fully specify its scope or timing. A report should therefore identify whether a statistic uses raw columns or explicitly named winsorized columns, and should record whether bounds were computed across all observations, by month, or under another configured scope.</p><p>The kurtosis convention is also configurable. The implementation accepts <code>pandas</code>, <code>fisher</code>, or <code>pearson</code>; the source material does not uniquely determine which convention should be used. Paper-reported descriptive values, such as the stated monthly-return mean and extreme maximum, remain comparison targets rather than independently reproduced statistics.</p><h3>Regression tables separate estimates from paper targets</h3><p>A <code>RegressionResult</code> represents an already-fitted result. <code>format_regression_result(result)</code> converts it into one row per uniquely labeled coefficient. Each row carries the computed coefficient, optional standard error and t-statistic, observation count, R-squared, covariance-method metadata, and specification metadata. <code>build_regression_report(results)</code> combines multiple <code>RegressionResult</code> objects while preserving their labels.</p><p>This design prevents a common reporting error: copying a paper coefficient into a local result column when the local model has not been run or has used a different sample. The annotation function instead adds separate comparison fields. The following excerpt is copied from <code>src/liquidity_premium/reporting/regression_tables.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;71493c92-9a32-4e05-bb20-8ebafe31c9ca&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">annotated = table.copy()
annotated["paper_target"] = np.nan
annotated["paper_target_source"] = pd.NA
unmatched: list[dict[str, Any]] = []

for specification, label, value, source in _target_entries(targets):
    mask = annotated["coefficient_label"].eq(label)
    if specification is not None:
        mask &amp;= annotated["specification"].eq(specification)
    if not bool(mask.any()):
        unmatched.append(
            {
                "specification": specification,
                "coefficient_label": label,
                "paper_target": value,
            }
        )
        continue
    annotated.loc[mask, "paper_target"] = value
    annotated.loc[mask, "paper_target_source"] = (
        source or "paper-reported target; not independently verified"
    )</code></pre></div><p><code>attach_paper_target_annotations(table, targets)</code> matches targets by specification and coefficient label when those fields are supplied. Label-only targets can match every row with that label. Unmatched targets are stored in the table attributes rather than discarded. Crucially, the original <code>coefficient</code> column is not overwritten.</p><p>Observation counts must describe the rows used by that particular regression after missing-value filtering. They must not be copied from the broader panel or from a paper table. Similarly, covariance metadata should describe what the fitted result actually supplied. The reporting layer does not infer clustering, fixed effects, weighting, or standard-error treatment that the paper leaves unspecified.</p><h3>Orchestration and local export</h3><p><code>build_reproduction_report(panel, results, config)</code> in <code>src/liquidity_premium/pipeline/run_reporting.py</code> assembles coverage, descriptive summaries, regression results, and already-structured forecast or portfolio tables. Its <code>panel</code> argument is the computed firm-month input; its <code>results</code> dictionary contains model outputs; and <code>config</code> supplies reporting settings and optional paper targets. It preserves table attributes such as paper identity, schema version, sample provenance, and the fact that paper targets are comparison-only.</p><p>The function can use a retained daily frame from <code>results</code> for daily coverage. If no daily frame is available, it falls back to the panel and marks that fallback in report metadata. This is transparent but weaker than reporting coverage from the actual cleaned daily input.</p><p><code>write_reproduction_report(report, output_dir)</code> writes each report table as a local CSV and creates <code>report_manifest.json</code>. <code>write_dataframe</code> in <code>src/liquidity_premium/reporting/export.py</code> accepts <code>.csv</code> and <code>.parquet</code> paths; <code>write_json</code> accepts <code>.json</code> paths and converts common scientific-Python values into JSON-compatible representations. These writers create parent directories when writing, but the existence of an output file would only show that a writer was called. It would not establish that the underlying data or estimates are correct.</p><p>The provenance writer records configuration and input paths without asserting that external data are present or sufficient:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;88157e10-4164-4441-b39e-ae35ac5d4768&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">payload: dict[str, object] = {
    "schema_version": "1.0",
    "record_type": "liquidity_premium_reproduction_provenance",
    "paper_id": "2607.01377v1",
    "paper_title": "Liquidity Premium and Investment Horizons",
    "configuration": asdict(config),
    "input_paths": [str(input_path) for input_path in inputs],
    "notes": [
        "This record describes a local reproduction configuration and does not contain computed results.",
        "External data availability, preprocessing conventions, and unspecified paper choices affect reproducibility.",
        "Paper-reported targets are not represented as independently reproduced values.",
    ],
}</code></pre></div><p>The command-line entry point <code>scripts/run_reproduction.py</code> loads a local JSON or YAML configuration, builds or loads a panel, calls <code>run_empirical_reproduction</code>, assembles the report, and writes local artifacts. It performs no network access. A typical planned invocation is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;e48539b1-61da-470e-889f-51cbe997d7c4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">python scripts/run_reproduction.py --config config/reproduction.yaml</code></pre></div><p>This command is an execution example, not evidence that execution occurred in this tutorial.</p><h3>Worked example: one report row, two kinds of evidence</h3><p>Consider a conceptual regression table row with the following fields:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;005913f0-59fc-4371-8036-7029189c0340&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">specification:             lambda_amihud_with_intercept
coefficient_label:        lambda
coefficient:               computed local estimate
paper_target:              paper-reported comparison value
n_observations:            rows used by this local regression
r_squared:                 local regression R-squared
sample_provenance:         input panel and missing-value policy
paper_target_source:       paper-reported target; not independently verified
verification_status:       not executed or verified</code></pre></div><p>The <code>coefficient</code> and <code>paper_target</code> fields answer different questions. The first would describe a local fitted estimate if the model had been run; the second records what the paper reported. Keeping them in separate columns prevents a target from overwriting a computed result and makes disagreement visible without implying which value is authoritative.</p><p>The same principle applies to coverage and descriptive statistics. A computed row should identify its input frame and sample scope. A paper target should identify its source and remain explicitly unverified. A report with no computed model result should contain no invented coefficient merely because a paper table contains one.</p><h3>What remains unresolved</h3><p>A full numerical reproduction requires local CRSP daily data and, depending on the analysis, monthly returns, delisting returns, H.15 risk-free rates, WRDS book-to-market and factor data, and potentially point-in-time membership data. The paper extraction does not fully specify several conventions that can change every downstream table:</p><ul><li><p>how <code>DisFacPr</code> and <code>DisFacShr</code> are applied;</p></li><li><p>how raw, delisting-adjusted, and excess returns are defined in each specification;</p></li><li><p>how missing prices, unchanged prices, and zero dollar volume are handled;</p></li><li><p>the scope and timing of winsorization;</p></li><li><p>panel covariance, clustering, fixed effects, and weighting choices;</p></li><li><p>whether recursive forecasts are pooled or estimated separately by firm;</p></li><li><p>the Newey&#8211;West lag length;</p></li><li><p>the portfolio signal, tie handling, weighting, risk-free alignment, and annualization; and</p></li><li><p>the availability and point-in-time treatment of external datasets.</p></li></ul><p>The source also contains OCR-damaged or incomplete equation records for the closed-form theoretical equilibrium, the continuous-time lambda statement, the volume standard-deviation expression, and Method B. Those formulas are not reconstructed here. The implementation can use the prose-supported sample standard deviation and Method B mapping, but the missing canonical LaTeX remains an evidence limitation.</p><p>Reported signs and narratives must also remain as reported. In particular, the paper&#8217;s narrative expectations for signed-flow predictability conflict with some negative one-month-ahead coefficients, and its theoretical discussion contains inconsistent statements about noise-trading variance and lambda. Reporting should preserve those distinctions rather than silently harmonize them.</p><h3>Verification boundary</h3><p>The planned verification strategy would inspect equation-to-function traceability, daily-row alignment, unique <code>PERMNO</code>-month keys, coefficient shapes, same-firm temporal alignment, forecast chronology, look-ahead leakage, monthly cross-sectional isolation, and separation of paper targets from computed values. It would also confirm that damaged equations are not reconstructed and that provenance accompanies exported tables.</p><p>Those checks were not performed here. Local static verification was disabled, code semantic verification was skipped, test generation and execution were disabled, and no final quality review occurred. Consequently, this tutorial can explain intended responsibilities and invariants but cannot claim that the generated files pass them. Semantic plausibility is not empirical validation: only a completed local run with appropriate data and independent checks could support numerical reproduction claims.</p><p>Use the button or URL below to download the source code.</p>
      <p>
          <a href="/__u/onepagecode.substack.com/p/quant-trading-kyles-price-impact">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Quant Trading: Feature-Wise Compositional RNNs & Grey Wolf Optimization for Stock Prediction (PyTorch Guide)]]></title><description><![CDATA[Building an end-to-end multi-stream LSTM, GRU, and SRU neural architecture with metaheuristic GWO hyperparameter optimization.]]></description><link>https://onepagecode.substack.com/p/quant-trading-feature-wise-compositional-084</link><guid isPermaLink="false">https://onepagecode.substack.com/p/quant-trading-feature-wise-compositional-084</guid><dc:creator><![CDATA[Onepagecode]]></dc:creator><pubDate>Sun, 09 Aug 2026 20:00:39 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_W6n!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3078fdad-8e80-40a9-be17-640eb18fa447_1254x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What This Article Builds</h2><p>The paper proposes a feature-wise compositional recurrent neural network approach for multivariate stock-price forecasting. The five OHLCV inputs are modeled separately with stacked LSTM, GRU, or SRU recurrent layers, regularized with dropout, fused by concatenation, optionally processed by additional recurrent layers and dense layers, and optimized using either Random Search (RS) or Grey Wolf Optimizer (GWO). The study evaluates 54 configurations on daily Hang Seng Index data from Yahoo Finance. The reported best result is an LSTM-GWO configuration, followed by GRU-GWO and SRU-GWO configurations.</p><h2>Implementation Assumptions</h2><ul><li><p>The five input channels are ordered Open, High, Low, Close, Volume.</p></li><li><p>The 20-day and 40-day values are treated as input window lengths, not forecast horizons.</p></li><li><p>Because the paper does not identify the target or horizon, the implementation exposes target<em>column and forecast</em>horizon as configuration choices; the tutorial must label these as implementation decisions.</p></li><li><p>Scaling defaults to training-only fitting over [0, 0.95] to reduce leakage, while documenting that the paper does not specify the fitting partition.</p></li><li><p>The paper's opaque configuration labels such as LSTM-GWO (1-1-0-1) are retained as benchmark labels and are not decoded into architecture parameters.</p></li><li><p>The training loss, gradient optimizer, search spaces, GWO budget, and metric formulas are configurable implementation decisions rather than asserted paper facts.</p></li><li><p>SRU is represented behind a local interface because the paper does not specify an SRU variant or API.</p></li><li><p>Yahoo Finance acquisition is optional for offline reproducibility; local OHLCV files and deterministic synthetic data are supported.</p></li><li><p>No reported benchmark is treated as verified reproduction output.</p></li><li><p>No canonical equations were supplied, so no equation-level LaTeX or equation_id mapping is asserted.</p></li><li><p>Verification is limited to planned static and semantic checks; execution, generated tests, and local verification are disabled by policy.</p></li></ul><h2>Scope, evidence status, and reproduction target</h2><p>What should happen when five market measurements describe the same trading day? The paper&#8217;s central idea is to avoid sending all of them immediately into one undifferentiated recurrent pathway. Instead, Open, High, Low, Close, and Volume each receive their own temporal pathway. The resulting representations are then joined before the final forecast. A useful intuition is to treat the five pathways as specialists: each studies one signal over time, and a prediction head combines their reports.</p><p>Here, <code>OHLCV</code> means the ordered features Open, High, Low, Close, and Volume. The generated model expects tensors in <code>batch &#215; time &#215; feature</code> layout. For a 20-day input window, for example, one batch has shape <code>batch &#215; 20 &#215; 5</code>; for the alternative window, the shape is <code>batch &#215; 40 &#215; 5</code>. The final dimension must retain the canonical OHLCV order.</p><h3>What the paper establishes</h3><p>The supplied paper context supports a feature-wise compositional recurrent method. Each OHLCV stream is processed independently by a stacked LSTM, GRU, or SRU family. Dropout is used for regularization, the five representations are fused through concatenation, and optional recurrent and dense layers produce the stock-price prediction. The paper also compares two outer hyperparameter strategies:</p><ul><li><p><strong>Random Search (RS)</strong> samples candidate configurations stochastically.</p></li><li><p><strong>Grey Wolf Optimizer (GWO)</strong> maintains a population of candidates and updates them according to their relative objective values.</p></li></ul><p>The paper reports 54 evaluated configurations and identifies an LSTM-GWO configuration as its best reported result. It also reports GRU-GWO and SRU-GWO benchmark configurations. These are useful comparison targets, but they do not fully describe how to construct the corresponding models.</p><p>The phrase &#8220;univariate encoding&#8221; in the extracted paper should be read carefully. The implementation target is not one single-variable forecasting model. It is feature-wise encoding: five separate one-channel sequences are modeled and then fused into a multivariate representation.</p><h3>What remains underdetermined</h3><p>An exact reproduction cannot be recovered from the supplied context alone. Several decisions needed by executable code are absent or ambiguous:</p><ul><li><p>The Yahoo Finance symbol, date range, adjustment policy, and missing-row treatment are not uniquely specified.</p></li><li><p>The target column and forecast horizon are not identified. The paper names stock-price prediction, but does not establish that the target is <code>Close</code> or that the horizon is one day.</p></li><li><p>The recurrent widths, exact layer counts, hidden-state extraction rule, dense widths, activations, and SRU implementation are missing.</p></li><li><p>The training loss, gradient optimizer, learning-rate schedule, stopping rule, and checkpoint policy are unspecified.</p></li><li><p>RS bounds, sampling distributions, trial counts, and validation details are absent.</p></li><li><p>GWO population size, iteration budget, bounds, initialization, discrete-parameter encoding, and update specification are absent.</p></li><li><p>Metric formulas and conventions, including percentage-error zero handling and evaluation scale, are not supplied.</p></li></ul><p>The labels <code>LSTM-GWO (1-1-0-1)</code>, <code>GRU-GWO (2-1-1-1)</code>, and <code>SRU-GWO (2-2-0-0)</code> therefore remain opaque benchmark labels. They must not be decoded into presumed layer counts, dropout switches, or dense-layer structures.</p><h3>How the generated package responds</h3><p>The generated package treats the paper as a method specification plus a set of reported targets, not as a complete executable recipe. <code>DataConfig</code> exposes the target, horizon, window length, split, and scaling range. <code>ModelConfig</code> exposes the recurrent family and architecture choices. <code>TrainingConfig</code>, <code>SearchConfig</code>, and <code>EvaluationConfig</code> make training, search, and metric conventions explicit. <code>validate_reproduction_config</code> provides a common boundary for checking these records.</p><p>The model assembly is represented by <code>CompositionalRNN</code> and <code>build_compositional_model</code>. Their intended responsibilities are to create five independent recurrent encoders, concatenate their representations, apply the configured fusion treatment, and produce the configured output dimension. RS and GWO remain separate outer-search interfaces, but both are intended to evaluate candidates through a shared training and validation objective.</p><p>This separation is important. A derived implementation decision&#8212;such as using <code>Close</code> as the target, a one-step horizon, or a particular loss&#8212;can be changed when better evidence becomes available without changing the paper&#8217;s feature-wise architecture or the public configuration interfaces. It also prevents a convenient default from being mistaken for a fact reported by the paper.</p><h3>Worked example: an explicit tutorial configuration</h3><p>For a small paper-oriented setup, choose <code>window_length=20</code>, <code>target_column='Close'</code>, and <code>forecast_horizon=1</code> in <code>DataConfig</code>. These values are tutorial choices: the paper supports 20-day input windows, but it does not specify <code>Close</code> as the target or a one-step forecast horizon. A corresponding <code>ModelConfig</code> can select <code>LSTM</code> and an explicitly documented number of stream layers and units, while a <code>TrainingConfig</code> records the chosen loss, optimizer, batch size, epochs, and seed.</p><p>The resulting input contract is <code>batch &#215; 20 &#215; 5</code>, and the configured output dimension determines the prediction shape. That shape contract describes the scaffold; it does not demonstrate that the configuration reproduces the paper&#8217;s reported LSTM-GWO result.</p><h3>Evidence status</h3><p>This section and the generated package distinguish three kinds of statements:</p><ol><li><p><strong>Paper facts:</strong> feature-wise OHLCV processing, recurrent-family comparison, concatenation fusion, RS and GWO comparison, and the reported benchmark labels and values.</p></li><li><p><strong>Derived explanations:</strong> the &#8220;five specialists&#8221; intuition and the interpretation of separate pathways as feature-wise rather than single-variable modeling.</p></li><li><p><strong>Implementation decisions:</strong> target and horizon defaults, exact architecture settings, preprocessing policies, training choices, search budgets, and metric conventions.</p></li></ol><p>The run policy disabled execution, local static verification, semantic code verification, generated test execution, tutorial verification, and final quality review. Consequently, this package should be described as a configurable reproduction scaffold, not as a verified implementation or a successful reproduction of the published results.</p><h2>Data contract: HSI OHLCV acquisition and provenance</h2><p>How can a forecasting model learn from market data if the source rows, feature order, and cleaning decisions are unclear? The data contract answers that question before any recurrent layer is built. It specifies which five measurements enter the model, preserves their chronological order, and records choices that the paper does not disclose.</p><h3>What the paper specifies&#8212;and what it does not</h3><p>The paper identifies daily Hang Seng Index history obtained through Yahoo Finance/YFinance and names the raw Open, High, Low, Close, and Volume fields as its OHLCV inputs. In the generated package, the canonical order is fixed as <code>Open</code>, <code>High</code>, <code>Low</code>, <code>Close</code>, <code>Volume</code>. A raw table therefore has shape <code>num_days &#215; 5</code>: each row represents one observation and each column represents one input channel.</p><p>The paper does not provide a unique Yahoo Finance ticker, date range, download options, adjusted-price setting, missing-value policy, or duplicate-row policy. These are not minor details. Different tickers or date ranges change the observations, while adjusted and unadjusted prices can produce different price histories. The generated implementation consequently makes these choices explicit and stores them in <code>AcquisitionMetadata</code> rather than presenting a particular choice as recovered paper methodology.</p><p>The reproduction scope also stays deliberately narrow. It retains the five raw OHLCV fields and does not add technical indicators or other engineered predictors. Technical indicators are mentioned by the paper as possible future work, not as part of the evaluated method.</p><h3>Canonical columns and provenance</h3><p><code>CANONICAL_OHLCV_COLUMNS</code> is the package-level contract used by acquisition and model-facing code:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;74402b89-e812-4503-99bd-59233d70db58&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># Canonical feature order required by the paper-oriented data and model APIs.
OHLCVColumns: TypeAlias = tuple[str, str, str, str, str]
CANONICAL_OHLCV_COLUMNS: OHLCVColumns = (
    "Open",
    "High",
    "Low",
    "Close",
    "Volume",
)</code></pre></div><p>This excerpt is copied from <code>src/compositional_rnn_stock/types.py</code>. <code>OHLCVColumns</code> describes a five-name tuple, while <code>CANONICAL_OHLCV_COLUMNS</code> supplies the concrete order. Preserving this order matters because later model code interprets the final input axis as five channels. Reordering columns&#8212;for example, placing Volume before Close&#8212;would change the meaning of the model input without changing its shape.</p><p>The acquisition record captures the choices surrounding those columns:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;142157c1-e384-4501-95c3-cd5e0d2d783f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class AcquisitionMetadata:
    """Provenance and cleaning-policy record for an OHLCV acquisition.

    The paper does not specify the Yahoo Finance symbol, date range, adjustment
    mode, missing-value policy, or duplicate-row policy.  These fields preserve
    the choices made by a local reproduction instead of treating them as paper
    facts.
    """

    source: str
    symbol: str | None
    start: str | None
    end: str | None
    adjusted: bool | None
    frequency: str
    requested_columns: OHLCVColumns
    missing_value_policy: str
    duplicate_policy: str
    ordering_policy: str</code></pre></div><p>This is an implementation record, not an additional claim about the paper. For a network acquisition, <code>source</code>, <code>symbol</code>, <code>start</code>, <code>end</code>, and <code>adjusted</code> identify the request. <code>requested_columns</code> records the five retained fields, and the remaining policy fields describe how the returned table was accepted. For an offline archive, some fields such as <code>symbol</code> may be <code>None</code>, but the artifact still records that the source was a local CSV.</p><h3>Two acquisition paths</h3><p>The generated adapter supports a network path through <code>load_hsi_ohlcv_from_yfinance</code> and an offline path through <code>load_ohlcv_csv</code>. The dispatcher requires exactly one source, preventing an ambiguous call that supplies both a symbol and a local file:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8bd47048-e70a-4219-bf03-8e9af725aae5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def acquire_hsi_ohlcv_data(
    *,
    symbol: str | None = None,
    local_path: Path | None = None,
    start: str | None = None,
    end: str | None = None,
    adjusted: bool = False,
) -&gt; tuple[pd.DataFrame, AcquisitionMetadata]:
    """Dispatch to an explicit local-file or Yahoo Finance acquisition mode."""

    if (symbol is None) == (local_path is None):
        raise ValueError("provide exactly one of symbol or local_path")
    if local_path is not None:
        return load_ohlcv_csv(local_path, CANONICAL_OHLCV_COLUMNS)
    return load_hsi_ohlcv_from_yfinance(symbol=symbol, start=start, end=end, adjusted=adjusted)</code></pre></div><p>This excerpt is copied from <code>src/compositional_rnn_stock/data/acquisition.py</code>. The function returns a pandas table plus its provenance record. In local mode, it delegates to <code>load_ohlcv_csv</code>; in Yahoo Finance mode, it delegates to <code>load_hsi_ohlcv_from_yfinance</code>. The optional <code>yfinance</code> import is isolated inside the network function, so an offline workflow does not need that package merely to load an archived file.</p><p>For a local archive, a typical call is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;1ed485ea-24b2-4a74-b25a-2dc19aedab41&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">frame, metadata = load_ohlcv_csv(Path("data/hsi_ohlcv.csv"), CANONICAL_OHLCV_COLUMNS)</code></pre></div><p>The generated CSV loader looks for a conventional <code>Date</code>, <code>date</code>, <code>Timestamp</code>, or <code>timestamp</code> column and uses it as the index when present. It then selects and orders only the canonical OHLCV fields. An archive should therefore preserve its date column whenever the identity of trading days matters. If no recognized date column exists, the loader retains the file's row index; that is a documented limitation of the local input rather than evidence that dates were absent from the original paper dataset.</p><p>For Yahoo Finance, the caller must supply the symbol and may supply explicit date bounds and an adjustment choice. The paper does not identify these values, so a tutorial command should not silently substitute a supposedly canonical ticker or time range. The command-line entry point instead requires explicit source and target-related settings, while the acquisition function records the selected source parameters.</p><h3>Validation and cleaning</h3><p>Acquisition first validates the structural contract. <code>validate_ohlcv_table</code> requires a nonempty pandas table, all five columns, numeric finite values, an increasing index, and no duplicate index values. It does not infer missing observations or create rows. The relevant responsibility is summarized by its public interface:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;95333448-680f-4774-966b-2bd0b684fb58&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def validate_ohlcv_table(frame: pd.DataFrame) -&gt; None:
    """Validate a chronological table containing exactly the five OHLCV fields.

    Validation deliberately rejects missing values and duplicates rather than
    fabricating observations.  The paper leaves those policies unspecified, so
    callers must clean or otherwise resolve such records before acquisition
    results are accepted.
    """</code></pre></div><p>The implementation's cleaning layer makes the missing and duplicate policies explicit:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5e47d94e-3664-4c38-b8c0-d334e89e8cdd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">cleaned = clean_ohlcv(frame, missing_policy="drop", duplicate_policy="reject")</code></pre></div><p>This call is the focused usage pattern exposed by <code>src/compositional_rnn_stock/data/cleaning.py</code>. The `</p><h2>Scaling, target definition, and sliding windows</h2><p>How does an ordered table of daily market observations become training data without allowing future information to leak backward? The preprocessing pipeline answers this in three stages: scale the five OHLCV channels, construct overlapping time windows, and split the resulting supervised samples chronologically.</p><p>The paper states that the features are normalized to the interval <code>[0, 0.95]</code> and that the model uses 20-day and 40-day temporal windows. It does not supply a canonical scaling equation, identify the target column, or specify the forecast horizon. The generated implementation therefore makes those choices explicit instead of presenting them as recovered paper facts.</p><h3>Scaling five channels independently</h3><p><code>FeaturewiseMinMaxScaler</code> treats Open, High, Low, Close, and Volume as five separate numeric channels. It stores one minimum and maximum for each channel, so the much larger numerical scale of Volume does not determine the scaling of price features. Its final feature axis must have length five, while arbitrary leading dimensions are allowed. Thus it can process both a raw table shaped <code>num_days &#215; 5</code> and model inputs shaped <code>samples &#215; window_length &#215; 5</code>.</p><p>The conventional feature-wise min-max mapping used by the generated scaler is an implementation decision because no canonical equation was supplied. The implementation also chooses to fit statistics on training observations only. This is a leakage-prevention policy: test observations are transformed using training metadata, rather than helping define that metadata. The paper states the target interval but does not state which partition fits the normalization statistics, so this choice must remain visible in provenance.</p><p>A focused part of the generated scaler shows the fitting contract:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;fa546f31-7a9e-4181-8d3c-c377a9ff2da9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">scaler = FeaturewiseMinMaxScaler(feature_range=feature_range)
scaler.fit(training_values)
return scaler</code></pre></div><p>Here, <code>training_values</code> is expected to contain only the earlier chronological observations selected by the caller. The scaler records <code>feature_min_</code>, <code>feature_max_</code>, and feature-specific scale metadata. Constant training features are mapped to the lower bound and restored to their fitted constant during inverse transformation, avoiding division by zero. Values outside the fitted extrema are not clipped by <code>transform</code>; the generated code documents this as an intentional choice so extrapolation remains visible.</p><p>The <code>fit</code>, <code>transform</code>, and <code>inverse_transform</code> methods preserve the input shape. That matters later when a prediction is converted back from normalized units to original price units. If evaluation is performed on original prices, the target column must be identified so that the corresponding feature scaling metadata can be applied.</p><h3>Target and horizon are explicit choices</h3><p>The paper describes stock-price prediction and lists OHLCV predictors, but it does not say whether the target is Close, adjusted Close, another price field, or a multivariate output. It also does not specify how far into the future the target lies. The generated windowing API therefore requires both <code>target_index</code> and <code>forecast_horizon</code>.</p><p><code>window_length</code> describes how many past observations enter one model input. For paper-oriented runs, it is treated as either 20 or 40 trading days. <code>forecast_horizon</code> is different: it describes the offset from the end of the input window to the target observation. The 20-day and 40-day values should not be interpreted as prediction horizons merely because they are temporal quantities.</p><p>For a start position <code>s</code>, the generated implementation uses the input rows from <code>s</code> through the end of the selected window, then takes the target at the configured future position. With <code>window_length=3</code> and <code>forecast_horizon=1</code>, the first input contains rows 0, 1, and 2, and its target is row 3. A horizon of 2 would instead select row 4 for that same first input.</p><p>The key implementation fragment is copied from <code>make_sliding_windows</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;993cae5d-2cea-4bdf-bb7c-9dea89a70d4d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">sample_count = day_count - window_length - forecast_horizon + 1
if sample_count &lt;= 0:
    raise ValueError(
        "insufficient observations for the requested window and horizon: "
        f"num_days={day_count}, window_length={window_length}, "
        f"forecast_horizon={forecast_horizon}"
    )

inputs = np.empty((sample_count, window_length, _FEATURE_COUNT), dtype=np.float64)
targets = np.empty(sample_count, dtype=np.float64)

target_timestamps = None if timestamp_array is None else np.empty(sample_count, dtype=timestamp_array.dtype)

for start in range(sample_count):
    end = start + window_length
    target_position = end + forecast_horizon - 1
    inputs[start] = array[start:end]
    targets[start] = array[target_position, target_index]
    if target_timestamps is not None:
        target_timestamps[start] = timestamp_array[target_position]</code></pre></div><p>The resulting <code>WindowedDataset</code> records <code>inputs</code> with shape <code>samples &#215; window_length &#215; 5</code> and scalar <code>targets</code> with shape <code>samples</code>. It also records the selected target column, horizon, and optional target timestamps. The five channels remain in canonical Open, High, Low, Close, Volume order; window construction does not reorder or mix them.</p><p>The following focused example is the same alignment pattern used by the generated contract tests. It uses synthetic arrays only and does not represent Hang Seng Index data:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6bf4f2b0-13ca-438d-b16e-85ce6d6f9222&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">dataset = make_sliding_windows(
    values,
    target_index=3,
    window_length=3,
    forecast_horizon=2,
    timestamps=timestamps,
)

assert dataset.inputs.shape == (4, 3, 5)
assert dataset.targets.shape == (4,)
np.testing.assert_array_equal(dataset.inputs[0], values[0:3])
np.testing.assert_array_equal(dataset.targets, values[4:8, 3])</code></pre></div><p>This example uses target index 3, which corresponds to Close in the canonical order. That index is an example configuration choice, not evidence that the paper definitively forecasts Close.</p><h3>Chronological splitting</h3><p>After windows and targets have been aligned, <code>chronological_split</code> divides complete supervised samples without shuffling. Its boundary is the floor of <code>sample_count * train_fraction</code>. The first partition contains earlier target timestamps, and the second contains later ones. The function rejects a split that would leave either partition empty.</p><p>The generated test expresses the central ordering invariant:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4834397e-6fc6-4363-a411-5ea0e0a44d48&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">train, test = chronological_split(dataset, train_fraction=0.6)

expected_train_count = int(dataset.sample_count * 0.6)
assert train.sample_count == expected_train_count
assert test.sample_count == dataset.sample_count - expected_train_count
assert train.sample_count &gt; 0
assert test.sample_count &gt; 0
assert train.timestamps is not None
assert test.timestamps is not None
assert int(train.timestamps[-1]) &lt; int(test.timestamps[0])</code></pre></div><p>These are contract tests supplied in the generated repository, but they were not executed under the authoritative run policy. They describe intended behavior rather than reporting a completed verification.</p><p>The paper reports an 80%/20% chronological partition and gives sample counts of 4,678 training samples and 1,170 test samples. Those counts are useful reproduction targets, but they do not uniquely reveal the raw date range, the number of overlapping windows, or whether the 20-day and 40-day datasets were constructed separately.</p><h3>How the prepared-data pipeline orders its work</h3><p><code>prepare_data</code> records the selected target, horizon, window length, split ratio, scaling range, and cleaning policies. Its relevant flow is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6574f66e-9308-470b-b4e6-306a0b33b159&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">observation_split = _observation_split_index(
    raw_values.shape[0], data_config.split_ratio
)

scaler = fit_scaler_on_training_observations(
    raw_values=raw_values,
    split_index=observation_split,
    feature_range=data_config.scaling_range,
)
scaled_values = scaler.transform(raw_values)
dataset = make_sliding_windows(
    values=scaled_values,
    target_index=_target_index(data_config.target_column),
    window_length=data_config.window_length,
    forecast_horizon=data_config.forecast_horizon,
    timestamps=timestamps,
)
train_dataset, test_dataset = chronological_split(
    dataset,
    train_fraction=data_config.split_ratio,
)</code></pre></div><p>Notice the distinction between the observation boundary used to fit the scaler and the later sample split. The generated pipeline first determines an observation-level boundary, fits the scaler on the leading observations, transforms the ordered series, constructs windows, and then partitions the completed samples. This is the package's documented leakage-avoidance design, not a uniquely specified procedure from the paper. Because overlapping windows can span an observation boundary, a stronger reproduction should confirm the intended split order from the original experiment details.</p><p>The pipeline also records <code>scaler_fit_policy</code> as <code>training_observations_only</code>, along with the target column, horizon, observation counts, window sample counts, and scaling range. These records make it possible to replace an assumption later without changing the public interfaces.</p><h3>Returning to original units</h3><p>Normalized inputs are useful for model training, but reported price errors may need to be calculated in original units. <code>FeaturewiseMinMaxScaler.inverse_transform</code> restores all five channels, while a target-specific evaluation adapter can select the configured target feature. This distinction is important because the paper does not specify whether its metrics were calculated on normalized values or inverse-transformed prices.</p><p>The generated offline demonstration follows the same policy on synthetic data: it fits the scaler on an earlier chronological portion, transforms the complete ordered sequence, and constructs 20-day windows. The demonstration is useful for checking shape contracts conceptually, but it is not HSI data and cannot establish the paper's reported performance.</p><p>In summary, the preprocessing contract is precise about shapes and ordering but deliberately honest about unresolved semantics. The implementation preserves five feature channels, treats 20 and 40 as input lengths, aligns each window with an explicitly configured future target, and prevents test observations from fitting the scaler. The target column, forecast horizon, split order, and normalization convention remain choices that must be documented for any reproduction run.</p><h2>The feature-wise compositional recurrent model</h2><p>How can five related market signals contribute to one forecast without being mixed too early? The paper&#8217;s answer is feature-wise composition: Open, High, Low, Close, and Volume each travel through an independent recurrent pathway. Their learned summaries are then placed side by side and passed to a prediction head. This is not five separate forecasting models. The final forecast is multivariate because the five representations are fused before prediction.</p><p>The paper supports this overall structure, but it does not uniquely specify the recurrent widths, exact layer counts, representation extraction rule, dense-layer widths, or activation functions. The generated package therefore implements these details through <code>ModelConfig</code>. The code is a configurable reproduction scaffold, not a verified reconstruction of every hidden architecture choice.</p><h3>Tensor flow through five independent streams</h3><p>The public model input, <code>x</code>, has shape <code>batch &#215; time_steps &#215; 5</code>. The final dimension follows the canonical order <code>Open</code>, <code>High</code>, <code>Low</code>, <code>Close</code>, <code>Volume</code>. A 20-day run therefore uses <code>batch &#215; 20 &#215; 5</code>; a 40-day run uses <code>batch &#215; 40 &#215; 5</code>.</p><p><code>FeatureWiseEncoder</code> splits that final dimension into exactly five tensors. Each slice has shape <code>batch &#215; time_steps &#215; 1</code>, so one recurrent stack receives one feature channel rather than the complete OHLCV vector. The five encoders are separate module instances, which means their parameters are not shared.</p><p>The following excerpt shows the stream-level contract. Notice that the recurrent stack is configured with <code>return_sequences=False</code>; this local choice selects the final time-step representation for fusion. The paper does not say whether the final hidden state, a complete sequence, or another aggregation was used.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d39fbac2-1f9b-419c-ad03-76b5e599b452&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class FeatureStreamEncoder(nn.Module):
    """Encode one OHLCV channel with its own recurrent stack.

    The input contract is ``(batch, time_steps, 1)`` and the output contract is
    ``(batch, hidden_size)``.  A separate instance is created for every OHLCV
    channel, so parameters are not shared between feature streams.
    """

    def __init__(
        self,
        recurrent_family: RecurrentFamily | str,
        hidden_size: int,
        layers: int,
        dropout_rate: float = 0.0,
        feature_name: str | None = None,
    ) -&gt; None:
        super().__init__()
        if isinstance(hidden_size, bool) or not isinstance(hidden_size, int) or hidden_size &lt;= 0:
            raise ValueError(f"hidden_size must be a positive integer; received {hidden_size!r}")
        if isinstance(layers, bool) or not isinstance(layers, int) or layers &lt;= 0:
            raise ValueError(f"layers must be a positive integer; received {layers!r}")
        if feature_name is not None and feature_name not in CANONICAL_OHLCV_COLUMNS:
            raise ValueError(
                f"feature_name must be one of {CANONICAL_OHLCV_COLUMNS!r}; received {feature_name!r}"
            )

        self.feature_name = feature_name
        self.input_size = 1
        self.hidden_size = hidden_size
        self.layer_count = layers
        self.recurrent_family = RecurrentFamily.coerce(recurrent_family)
        self.recurrent = build_recurrent_stack(
            family=self.recurrent_family,
            input_size=self.input_size,
            hidden_size=hidden_size,
            layers=layers,
            return_sequences=False,
        )
        self.dropout = DualDropout(stream_rate=dropout_rate, fusion_rate=0.0)</code></pre></div><p>The important invariant is <code>input_size=1</code>: every stream receives one feature channel. After recurrent processing, each stream produces a fixed-width tensor with shape <code>batch &#215; hidden_size</code>. Stream dropout is then applied. Dropout is a training-time regularizer: it masks values while the model is training and is disabled when the module is in evaluation mode.</p><p><code>FeatureWiseEncoder.forward</code> preserves the channel order with a one-element split along the final axis:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;311d58d0-5c28-4120-afa3-4ecf2e4805fd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        # Each slice retains its one-channel axis: (batch, time_steps, 1).
        feature_slices = torch.split(x, split_size_or_sections=1, dim=-1)
        if len(feature_slices) != 5:
            raise RuntimeError(f"expected five feature slices, received {len(feature_slices)}")

        encoded_streams = tuple(
            encoder(feature_slice)
            for encoder, feature_slice in zip(self.encoders, feature_slices)
        )</code></pre></div><p>This code does not perform early averaging, summation, or mixing. The five encoded outputs remain separate until the fusion module receives them.</p><h3>One interface for LSTM, GRU, and SRU</h3><p>A recurrent layer processes a sequence while maintaining a learned state that carries information across time steps. LSTM and GRU are the two standard recurrent families exposed by the generated wrapper. The paper also evaluates SRU, but the supplied paper context does not identify an SRU variant or Python API.</p><p><code>RecurrentFamily</code> provides a common selector, and <code>build_recurrent_stack</code> returns a <code>RecurrentStack</code> with the same input and output contract for each family. For a stream, the input is <code>batch &#215; time_steps &#215; 1</code>. With final-time-step reduction, the output is <code>batch &#215; hidden_size</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6ec89a4f-e254-425a-8b3f-c4a697feee12&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class RecurrentFamily(str, Enum):
    """Supported recurrent-family selectors.

    ``SRU`` is included because it is one of the families evaluated by the
    paper.  The paper does not identify an SRU variant or Python API, so the
    local adapter below is an explicit implementation choice rather than a
    claim of exact SRU reproduction.
    """

    LSTM = "LSTM"
    GRU = "GRU"
    SRU = "SRU"</code></pre></div><p>For LSTM and GRU, the wrapper constructs batch-first PyTorch stacks. The local <code>SRUAdapter</code> is different: it is a dependency-free compatibility adapter supplied by the generated implementation. Its documentation explicitly says that it is not a canonical SRU implementation. Consequently, selecting <code>SRU</code> makes the interface available for experimentation, but it must not be described as proof of paper-equivalent SRU behavior.</p><p>The wrapper also validates rank, feature width, positive time steps, and floating-point input. Invalid inputs fail explicitly rather than being silently reshaped. This matters because changing <code>batch &#215; time_steps &#215; features</code> to another ordering would change the meaning of the recurrent computation.</p><h3>Dropout and concatenation fusion</h3><p>The paper describes dual dropout: one stage after feature-specific recurrent processing and another stage around feature fusion. The first stage is clear enough to place after each stream encoder. The second stage is not: &#8220;around fusion&#8221; does not establish whether it is before concatenation, after concatenation, or both. <code>DualDropout</code> and <code>ConcatenationFusion</code> expose this as <code>fusion_placement</code>.</p><p>Concatenation means placing the representations end to end along the final dimension. If each of the five streams has width <code>hidden_size</code>, the fused representation has width <code>5 &#215; hidden_size</code>. The operation preserves the batch dimension and does not combine channels by summation.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;128a3857-95cc-40f1-8835-a9c386e4236a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        # The paper's stated fusion operation is concatenation along the
        # representation dimension; stream representations are never summed.
        fused = torch.cat(processed_streams, dim=-1)

        if self.fusion_placement in {"after_concatenation", "both"}:
            fused = self.fusion_dropout.apply_fusion(fused)
        return fused</code></pre></div><p><code>validate_stream_shapes</code> requires exactly five floating-point tensors, matching leading dimensions, a positive representation width, the same dtype, and the same device. A mismatch such as one stream having a different batch size is rejected before concatenation. The default generated placement applies fusion dropout after concatenation, but that default is an implementation decision rather than a recovered paper setting.</p><h3>Optional post-fusion processing and prediction</h3><p>After fusion, <code>PostFusionHead</code> can either send the fused vector directly to dense layers or apply additional recurrent processing first. The paper allows post-fusion recurrent layers but does not specify how a fused vector should be treated as a sequence. The generated head makes one explicit choice: it adds a single time step, giving a tensor shaped <code>batch &#215; 1 &#215; fused_features</code>, and then uses the final representation from that one-step recurrent stack.</p><p>Dense layers then produce the configured target output. The generated implementation uses ReLU between configurable dense layers, followed by a final linear layer. Both the widths and the activation choice are local implementation decisions. The output shape is <code>batch &#215; output_dimension</code>, where <code>output_dimension</code> must agree with the separately configured target definition.</p><h3>Assembling the complete model</h3><p><code>CompositionalRNN</code> connects the stages in order: validate the input, encode five streams, concatenate them, apply the optional post-fusion head, and validate the prediction shape. Its central forward path is deliberately explicit:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d3dec4a4-308a-4547-a683-a4ee298ce97b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        # Feature-wise recurrent encoding: five tensors of shape
        # (batch, stream_units), preserving the canonical OHLCV order.
        encoded_streams = self.feature_encoder(x)
        if len(encoded_streams) != self.input_feature_count:
            raise RuntimeError(
                f"expected {self.input_feature_count} encoded streams, "
                f"received {len(encoded_streams)}"
            )

        # Paper architecture stage: fuse representations by concatenation.
        fused = self.fusion(encoded_streams)
        expected_fused_shape = (x.shape[0], self.fused_representation_size)
        if tuple(fused.shape) != expected_fused_shape:
            raise RuntimeError(
                "fusion returned an unexpected shape: "
                f"received {tuple(fused.shape)}, expected {expected_fused_shape}"
            )

        # Optional post-fusion recurrence and dense prediction are delegated to
        # the head; its recurrent representation is a configurable choice.
        prediction = self.post_fusion_head(fused)</code></pre></div><p>The model boundary requires exactly five channels and a positive batch and time dimension. It returns <code>batch &#215; output_dimension</code>; an unexpected output shape raises an error. These checks enforce the paper-supported architecture without pretending that unspecified widths or layer counts have been recovered.</p><h3>Focused configuration example</h3><p>The following excerpt is adapted exactly from the generated model-contract test. It demonstrates a small local configuration with one recurrent layer, eight stream units, one dense layer, and one output. Those values are demonstration settings, not the paper&#8217;s reported best architecture. The <code>20</code> in the input shape represents the paper-oriented window choice, while the target semantics remain configured elsewhere.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5405e223-d628-4a8a-9e7a-cb61829bae85&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def _make_model_config(
    family: RecurrentFamily | str = RecurrentFamily.LSTM,
    *,
    stream_dropout: float = 0.0,
    fusion_dropout: float = 0.0,
) -&gt; ModelConfig:
    """Build a small deterministic configuration for contract tests.

    These tests verify structural invariants of the implementation rather than
    the paper's unresolved architecture details or opaque benchmark labels.
    """
    return ModelConfig(
        recurrent_family=family,
        stream_layers=1,
        stream_units=8,
        stream_dropout=stream_dropout,
        fusion_dropout=fusion_dropout,
        fusion_placement="after_concatenation",
        post_fusion_layers=0,
        post_fusion_units=8,
        dense_layers=(8,),
        output_dimension=1,
    )</code></pre></div><p>A caller would construct the model and prepare an input with shape <code>batch &#215; 20 &#215; 5</code> as follows:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;66127dca-712c-4cfb-806e-d93747e412b5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">model = build_compositional_model(_make_model_config())
inputs = _make_inputs(batch_size=4, time_steps=20)

predictions = model(inputs)</code></pre></div><p>The intended contract is <code>predictions</code> with shape <code>4 &#215; 1</code> for this configuration. This excerpt was not executed under the run policy, so it demonstrates the interface and expected shapes only; it does not establish runtime correctness or paper-level reproduction.</p><p>The opaque labels in the reported results&#8212;<code>LSTM-GWO (1-1-0-1)</code>, <code>GRU-GWO (2-1-1-1)</code>, and <code>SRU-GWO (2-2-0-0)</code>&#8212;are deliberately not translated into <code>ModelConfig</code> fields. The supplied paper context does not define what those four positions mean. Keeping them as benchmark labels prevents an attractive but unsupported architecture claim.</p><p>The generated test plan reinforces structural invariants such as five independent streams, additive fusion width, common family dispatch, and training-only dropout. Those tests were not executed: local static verification, semantic verification, and code execution were disabled or skipped by policy.</p><h2>Training one candidate without hiding missing specifications</h2><p>What does it mean to train one candidate model in this reproduction? The candidate receives batches of chronological windows, produces predictions, computes a selected scalar loss, and updates its trainable parameters with gradients. After each epoch, a separate validation pass records loss with dropout disabled and without changing the parameters. This inner training loop is distinct from Random Search and Grey Wolf Optimizer (GWO), which choose different candidate configurations outside the loop.</p><h3>Paper facts and implementation choices</h3><p>The paper says that recurrent models are trained and compared using validation performance, and it names learning rate, batch size, and training epochs among the optimized hyperparameters. However, the supplied paper context does not specify the training loss, base gradient optimizer, initialization, learning-rate schedule, random seed, batch-shuffling policy, early-stopping rule, or checkpoint-selection rule. It also does not clearly define how the reported 80%/20% partition relates to the validation loss discussed in the method.</p><p>The generated code makes these gaps visible. <code>TrainingConfig</code> supplies the local choices, while <code>ForecastLoss</code> supports explicit <code>mse</code>, <code>mae</code>, and <code>huber</code> objectives. These are available implementation options, not claims about which loss reproduced the paper. Likewise, <code>make_optimizer</code> supports explicit gradient optimizers for a candidate; this inner optimizer should not be confused with RS or GWO, which operate at the outer configuration-search level.</p><p>The generated pipeline uses the prepared test partition as the validation loader for a one-candidate run. Its source documentation explicitly calls this an implementation decision because the paper does not define a separate validation construction. Consequently, this scaffold should not be described as recovering the paper&#8217;s exact validation protocol.</p><h3>Validating predictions and targets</h3><p>A forecasting loss is meaningful only when each prediction is paired with its corresponding target. <code>validate_loss_inputs</code> requires PyTorch tensors with floating-point dtypes, identical shapes, at least one element, and finite values. A mismatch, empty batch, integer tensor, or non-finite value raises an error rather than allowing silent broadcasting or an invalid objective.</p><p>The public loss interface is <code>ForecastLoss</code>. Its output is a scalar tensor, so it can participate in backpropagation. The generated implementation uses mean reduction for each supported local objective. The following excerpt shows the explicit dispatch; it is copied from <code>src/compositional_rnn_stock/losses.py</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a9e90480-9144-465c-b3cb-7f0d814c96c5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        if self.name == "mse":
            loss = F.mse_loss(predictions, targets, reduction="mean")
        elif self.name == "mae":
            loss = F.l1_loss(predictions, targets, reduction="mean")
        else:
            loss = F.smooth_l1_loss(
                predictions,
                targets,
                beta=self.huber_delta,
                reduction="mean",
            )</code></pre></div><p>The inputs have the model&#8217;s configured target shape. In the scalar-target pipeline, the model returns <code>batch &#215; 1</code>, while window datasets expose targets as <code>batch</code>. The training loop performs only this unambiguous single-output reshape; other mismatches are rejected. That safeguard matters because accidental broadcasting could produce a finite-looking loss with incorrect semantics.</p><h3>One epoch: update in training mode, evaluate in validation mode</h3><p>During training, <code>train_one_candidate</code> puts the model in training mode, reads a batch shaped <code>batch &#215; time_steps &#215; 5</code>, computes predictions, aligns the targets, backpropagates the scalar loss, and calls the selected optimizer. The five-channel input contract is checked before the model is used. A training loader with no batches or non-finite values is rejected.</p><p>The core update sequence is shown below. This excerpt is copied from <code>src/compositional_rnn_stock/training/loops.py</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ac4c30d0-4f17-4fbe-9f48-d2a43c40de77&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        for batch in train_loader:
            inputs, targets = _prepare_batch(batch, device, "training")
            optimizer.zero_grad(set_to_none=True)
            predictions = model(inputs)
            targets = _align_targets(predictions, targets)
            loss = loss_fn(predictions, targets)
            loss.backward()
            optimizer.step()</code></pre></div><p>After the training batches, <code>evaluate_loss</code> switches the model to evaluation mode and wraps inference in <code>torch.no_grad()</code>. Evaluation mode is important for the compositional model because dropout must be disabled when validation predictions are measured. The function restores the model&#8217;s previous training state afterward, computes a value-weighted mean over the validation targets, and rejects an empty or non-finite result.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;0df7a03e-5319-43fa-aaae-6262455811b7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    was_training = model.training
    model.eval()
    total_loss = 0.0
    total_values = 0
    try:
        with torch.no_grad():
            for batch in loader:
                inputs, targets = _prepare_batch(batch, device, "validation")
                predictions = model(inputs)
                targets = _align_targets(predictions, targets)
                loss = loss_fn(predictions, targets)
                value_count = int(targets.numel())
                total_loss += float(loss.detach().cpu().item()) * value_count
                total_values += value_count
    finally:
        model.train(was_training)</code></pre></div><p>This separation is the main training invariant: validation loss is observed, not optimized directly. The validation loader is never passed to <code>backward()</code> or <code>optimizer.step()</code> by this workflow. The paper&#8217;s exact split and model-selection policy remain unresolved, so the generated code retains the final state after the configured epochs rather than silently inventing a best-checkpoint rule.</p><h3>Optimizer and reproducibility settings</h3><p><code>make_optimizer</code> constructs the inner gradient optimizer from the configured name and learning rate. The generated implementation accepts <code>adam</code>, <code>adamw</code>, and <code>sgd</code> in this function. It validates that the learning rate is positive and finite, that the model has trainable parameters, and that the optimizer name is supported. The paper does not identify which of these, if any, was used.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5d49edbf-e68b-4aed-a0d6-0dbf6cd3e197&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    if normalized == "adam":
        return Adam(parameters, lr=learning_rate)
    if normalized == "adamw":
        return AdamW(parameters, lr=learning_rate)
    return SGD(parameters, lr=learning_rate)</code></pre></div><p><code>set_global_seed</code> seeds supported local Python, NumPy, and PyTorch generators. <code>SeedContext</code> additionally snapshots and restores available random states and can request deterministic PyTorch algorithms. These facilities improve the comparability of local experiments, but the paper does not report its seed or deterministic-backend settings. A seed therefore belongs in provenance, not in a claim that the paper used the same setting.</p><h3>Histories and provenance records</h3><p>Each completed epoch becomes an <code>EpochRecord</code> containing its epoch number, training loss, and validation loss. <code>TrainingHistory</code> preserves insertion order and rejects non-increasing epoch numbers. Its <code>best_epoch</code> helper can identify the earliest record with the smallest available validation loss, but the training loop does not automatically restore that epoch&#8217;s weights. Selecting and restoring a best checkpoint would be an additional explicit implementation decision.</p><p>The outer record is <code>TrainingRun</code>. It links the trained <code>CompositionalRNN</code>, its <code>TrainingHistory</code>, the <code>ModelConfig</code>, the <code>TrainingConfig</code>, and a prepared-data identifier. That identifier records acquisition and preprocessing metadata, window length, forecast horizon, target column, and train/test sample counts. This prevents a local loss curve from becoming detached from choices such as <code>Close</code> versus another target, a 20-day versus 40-day window, or a particular scaling policy.</p><h3>Worked example: defining a local candidate run</h3><p>The following configuration fragment is copied from <code>scripts/train_model.py</code>. It selects <code>Close</code> and a one-step horizon, but those are tutorial-level implementation choices because the paper does not identify the target column or forecast horizon.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5521ead8-53ec-4cae-b6d9-e6845c5d06f2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    data_config = DataConfig(
        symbol=args.symbol,
        local_path=args.local_path,
        start=args.start,
        end=args.end,
        target_column=args.target_column,
        forecast_horizon=args.forecast_horizon,
        window_length=args.window_length,
        split_ratio=args.split_ratio,
        scaling_range=(0.0, 0.95),
    )</code></pre></div><p>A focused call path inside <code>train_one_candidate</code> then creates the selected loss and optimizer and trains for the configured number of epochs. This excerpt is copied from <code>src/compositional_rnn_stock/training/loops.py</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;74cba30e-5aab-4657-9c3f-0a68525e7332&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    loss_fn = _loss_for_config(config)
    optimizer = make_optimizer(model, config)
    device = _model_device(model)
    history = TrainingHistory()</code></pre></div><p>At the pipeline level, <code>run_training</code> validates the prepared data and configurations, builds chronological loaders, constructs the model, checks that loader inputs have shape <code>batch &#215; time_steps &#215; 5</code>, and delegates to <code>train_one_candidate</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a0aa783c-1211-4a96-b32c-c5e65427c8cd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    trained_model, history = train_one_candidate(
        model=model,
        train_loader=train_loader,
        validation_loader=validation_loader,
        config=training_config,
    )</code></pre></div><p>This call defines one local experiment. It does not establish that the loss, optimizer, architecture settings, validation protocol, or final model state match the paper. It also does not execute here: the authoritative run policy disabled code execution and verification.</p><h3>Why this boundary matters</h3><p>The gradient loop answers, &#8220;How do we fit one selected candidate?&#8221; RS and GWO answer a different question, &#8220;Which candidate settings should we try next?&#8221; Keeping those responsibilities separate allows both search methods to use the same loss, loaders, validation objective, and provenance schema. It also makes missing paper specifications inspectable instead of burying them in defaults.</p><p>For this scaffold, the honest output of training is a locally configured model, an epoch-aligned history, and provenance describing the choices. No training result, convergence behavior, or successful reproduction of the paper&#8217;s reported benchmarks is claimed. Verification was skipped under policy, so the code excerpts describe the generated interfaces and intended control flow rather than a checked execution outcome.</p><h2>Random Search and Grey Wolf Optimizer</h2><p>How should a reproduction choose among many possible recurrent-network recipes? Random Search (RS) samples independent candidates, while Grey Wolf Optimizer (GWO) maintains a population of candidates and moves them using the best-ranked candidates as guides. In both cases, a candidate is more than a model family: it can include recurrent-layer count, recurrent units, dropout rates, learning rate, batch size, and training epochs.</p><p>The paper names both RS and GWO and reports that GWO performed better across the LSTM, GRU, and SRU families. However, it does not specify the search bounds, probability distributions, number of trials, wolf population, iteration count, initialization, update settings, invalid-candidate policy, or randomization controls. The generated package therefore treats these as explicit implementation configuration. No displayed setting in this section should be read as the paper's missing experimental setting.</p><h3>One objective shared by both methods</h3><p>The outer optimizer does not train a model directly. Instead, it proposes a typed <code>CandidateConfig</code>. The shared candidate objective builds the corresponding compositional model, creates or receives chronological training and validation loaders, trains the candidate, and records its best validation loss. Validation loss is used here as a local implementation choice because the paper discusses validation performance but does not define the loss or the exact selection rule.</p><p><code>CandidateResult</code> stores the candidate configuration, objective, metrics, provenance, success status, and any failure information. <code>SearchResult</code> stores the selected result and the complete search history. This separation matters: a failed candidate remains visible in the history rather than being silently treated as the best candidate, and the held-out test set is not needed for hyperparameter selection.</p><p>The objective's public protocol is deliberately small. Its candidate has type <code>CandidateConfig</code>, its data bundle must provide training and validation loaders, and its <code>seed</code> is recorded for reproducibility. The generated implementation explicitly records that test data were not used for selection:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e03f7909-591c-4967-9850-be571e916000&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">provenance: dict[str, Any] = {
    "seed": seed,
    "objective": "minimum_validation_loss",
    "selection_basis": "lowest validation loss recorded during training",
    "test_data_used_for_selection": False,
}</code></pre></div><p>This excerpt comes from <code>src/compositional_rnn_stock/search/objectives.py</code>. The objective returns a <code>CandidateResult</code> after <code>train_one_candidate</code> supplies a <code>TrainingHistory</code>; the history must contain a validation loss from which the lowest recorded value can be selected. The code does not claim that this objective is the paper's exact loss or checkpoint policy.</p><p>The same selection helper serves both optimizers. Its <code>minimize</code> argument makes the direction explicit, which is important because validation loss is normally minimized, whereas some metrics are maximized:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;21d134cb-c632-492a-a6ec-cba1a5ef5aa2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def select_best_candidate(
    results: Sequence[CandidateResult],
    minimize: bool,
) -&gt; CandidateResult:
    """Select a successful candidate using an explicit objective direction."""
    if not isinstance(minimize, bool):
        raise TypeError(f"minimize must be boolean; received {type(minimize).__name__}")
    if not isinstance(results, Sequence):
        raise TypeError("results must be a sequence of CandidateResult instances")

    validated: list[CandidateResult] = []
    for index, result in enumerate(results):
        if not isinstance(result, CandidateResult):
            raise TypeError(
                f"results[{index}] must be a CandidateResult; "
                f"received {type(result).__name__}"
            )
        if result.success:
            if result.objective is None:
                raise ValueError(
                    f"successful results[{index}] must contain an objective"
                )
            validated.append(result)

    if not validated:
        failure_count = sum(1 for result in results if not result.success)
        raise ValueError(
            "cannot select a best candidate: no successful candidate results "
            f"were available ({failure_count} recorded failures)"
        )

    if minimize:
        return min(validated, key=lambda result: float(result.objective))
    return max(validated, key=lambda result: float(result.objective))</code></pre></div><p>The function accepts a sequence of results, rejects malformed successful records, ignores unsuccessful records for selection, and raises an error when no successful candidate exists. That failure behavior is a derived implementation safeguard, not a reported paper procedure.</p><h3>Typed search spaces</h3><p>A search space must represent different kinds of values. Recurrent-layer counts, units, batch sizes, and epochs are integer-valued. Dropout rates and learning rates are continuous. A model family is categorical when one search spans LSTM, GRU, and SRU. <code>ParameterDomain</code> gives each parameter one of these meanings and validates its bounds or allowed values.</p><p>The following excerpt is copied from <code>src/compositional_rnn_stock/search/space.py</code> and shows the domain constructors exposed by the generated package:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;57e703b9-c36f-4bd1-899d-11825f5f82ab&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    @classmethod
    def continuous(cls, lower: float, upper: float) -&gt; "ParameterDomain":
        return cls("continuous", lower=lower, upper=upper)

    @classmethod
    def integer(cls, lower: int, upper: int) -&gt; "ParameterDomain":
        return cls("integer", lower=lower, upper=upper)

    @classmethod
    def categorical(cls, values: Sequence[Any]) -&gt; "ParameterDomain":
        return cls("categorical", values=tuple(values))

    @classmethod
    def discrete(cls, values: Sequence[Any]) -&gt; "ParameterDomain":
        return cls("discrete", values=tuple(values))</code></pre></div><p><code>continuous</code> and <code>integer</code> domains use inclusive lower and upper bounds. <code>categorical</code> and <code>discrete</code> domains hold explicit values. The distinction between categorical and discrete values is useful when documenting intent, although both are represented by positions when a vector-based optimizer needs numeric coordinates.</p><p><code>SearchSpace</code> keeps the named domains in a stable order and exposes encoded bounds. <code>CandidateConfig</code> then groups the decoded values into <code>ModelConfig</code>, <code>TrainingConfig</code>, and any explicitly declared extra parameters. <code>candidate_from_mapping</code> validates a named mapping before constructing that typed record. The paper's list of optimized hyperparameters tells us what should be represented, but not the actual ranges; those ranges must be supplied by the caller.</p><h3>Random Search: independent proposals</h3><p>RS draws one value from every domain for each trial. The generated <code>random_search</code> function accepts a <code>SearchSpace</code>, a callable <code>CandidateObjective</code>, and a search configuration containing a seed, trial count, data bundle, and objective direction. Each trial receives a derived candidate seed, and each result is appended to the history, including failures.</p><p>The core loop is implemented as follows in <code>src/compositional_rnn_stock/search/random_search.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e541c4f1-c540-468c-9d09-d6311bbee6a2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    history: list[CandidateResult] = []
    for trial_index in range(trial_count):
        candidate = sample_random_candidate(space, rng)
        candidate_seed = seed + trial_index
        try:
            result = objective(candidate, data_bundle, candidate_seed)
            if not isinstance(result, CandidateResult):
                raise TypeError(
                    "objective must return CandidateResult; "
                    f"received {type(result).__name__}"
                )
        except Exception as exc:
            result = CandidateResult(
                configuration=candidate,
                objective=None,
                metrics={},
                provenance={
                    "method": "random_search",
                    "trial_index": trial_index,
                    "seed": candidate_seed,
                    "failure_type": type(exc).__name__,
                },
                success=False,
                error=f"{type(exc).__name__}: {exc}",
            )
        else:
            result.provenance.update(
                {
                    "method": "random_search",
                    "trial_index": trial_index,
                    "sampler_seed": seed,
                    "candidate_seed": candidate_seed,
                }
            )
        history.append(result)</code></pre></div><p>Notice the two seeds recorded in successful results: the sampler seed identifies the sequence of sampled candidates, while the candidate seed identifies the training run. This is a useful provenance design for comparing methods locally. It does not establish that the paper used the same seeding scheme.</p><p>A tutorial experiment might declare three trials for a quick demonstration, but that would be a tutorial budget only. The paper's statement that 54 configurations were evaluated must not be reverse-engineered into a presumed RS trial count or into a presumed division between model families and optimizers.</p><h3>GWO: population-guided proposals</h3><p>GWO works with a matrix of wolf positions. A position is a numeric vector whose coordinates correspond to the ordered parameters in <code>SearchSpace</code>. The best three successful candidates are called alpha, beta, and delta in the generated implementation. Their encoded vectors guide a new population. After each update, positions are clipped to the declared bounds and decoded back into valid typed candidates.</p><p>The generated update function accepts <code>positions</code> with shape <code>population &#215; encoded_dimension</code>. Each leader vector has shape <code>encoded_dimension</code>. The update is intentionally documented as a local GWO-style choice because no canonical GWO equations or settings were supplied in the paper:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;cdc33e84-bc63-426e-997b-d429069bc2e5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def update_wolf_positions(
    positions: Any,
    alpha: Any,
    beta: Any,
    delta: Any,
    iteration: int,
    total_iterations: int,
    rng: random.Random,
) -&gt; np.ndarray:
    """Return one local GWO-style update for every wolf.

    ``positions`` has shape ``(population, encoded_dimension)`` and each leader
    has shape ``(encoded_dimension,)``.  The implementation uses the usual
    linearly decreasing exploration coefficient and independent random draws.
    These update rules are implementation decisions, not equations supplied by
    the paper.
    """
    current = np.asarray(positions, dtype=float)
    leader_arrays = [np.asarray(value, dtype=float) for value in (alpha, beta, delta)]
    if current.ndim != 2:
        raise ValueError("positions must have shape (population, encoded_dimension)")
    if current.shape[0] &lt; 1:
        raise ValueError("positions must contain at least one wolf")</code></pre></div><p>The remainder of the function validates leader shapes and iteration bounds, generates independent random coefficients, averages the three leader-guided proposals, and returns an array with the same population-by-dimension shape. The linearly decreasing exploration coefficient, random update details, and alpha/beta/delta ranking are implementation decisions rather than recovered paper settings.</p><h3>Decoding mixed parameters safely</h3><p>A continuous vector cannot directly be used as a batch size or a model-family name. <code>CandidateCodec</code> bridges that gap. It encodes typed values into a stable vector ordering and decodes them by clipping to the domain, rounding integer-like coordinates, selecting categorical positions, and constructing validated <code>ModelConfig</code> and <code>TrainingConfig</code> records.</p><p>The public decoding contract is concise:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6e2e6022-31a4-4c3e-98fd-d7137dabb043&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    def decode(self, vector: Sequence[float] | np.ndarray) -&gt; CandidateConfig:
        array = np.asarray(vector, dtype=float)
        if array.ndim != 1 or array.shape[0] != self.dimension:
            raise ValueError(
                f"encoded candidate must have shape ({self.dimension},); received {array.shape}"
            )
        if not np.all(np.isfinite(array)):
            raise ValueError("encoded candidate must contain only finite values")
        values = {
            name: self.space.domains[name].decode_value(float(array[index]))
            for index, name in enumerate(self.space.parameter_names)
        }
        return self._build_candidate(self.space.validate_candidate(values))</code></pre></div><p>For example, a vector coordinate intended for batch sizes <code>(8, 16, 32)</code> is clipped to the valid index range and rounded to one of those positions. That prevents GWO from producing an invalid fractional batch size. It also means that several nearby continuous positions can decode to the same typed candidate; this is an expected consequence of using a continuous optimizer for mixed parameters.</p><h3>Worked configuration example</h3><p>The following excerpt is a small configuration-only example based on the generated search-contract test. It demonstrates integer, continuous, and discrete domains without asserting that these bounds came from the paper:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;24f944ed-4707-417f-8993-9ba3ef37a3fd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">PARAMETER_DOMAINS = {
    "stream_layers": ParameterDomain.integer(1, 3),
    "units": ParameterDomain.integer(8, 32),
    "dropout": ParameterDomain.continuous(0.0, 0.5),
    "learning_rate": ParameterDomain.continuous(1.0e-4, 1.0e-2),
    "batch_size": ParameterDomain.discrete((8, 16, 32)),
    "epochs": ParameterDomain.integer(1, 4),
}

space = SearchSpace(domains=PARAMETER_DOMAINS)
codec = CandidateCodec(space)</code></pre></div><p>Here, <code>stream_layers</code>, <code>units</code>, <code>batch_size</code>, and <code>epochs</code> are discrete choices, while <code>dropout</code> and <code>learning_rate</code> vary continuously. <code>codec</code> can translate a validated <code>CandidateConfig</code> into a numeric vector for GWO and decode a clipped vector back into a typed candidate. The displayed ranges are deliberately modest tutorial settings; they are not paper-reported bounds.</p><p>In a complete local run, the same <code>space</code> and shared objective can be passed to RS or GWO. The pipeline functions <code>run_random_search</code> and <code>run_gwo_search</code> bind both methods to a chronological training/validation split created from the prepared training partition. Their documented contract excludes <code>PreparedData.test</code> from selection, so the held-out test data are reserved for evaluation after a candidate has been chosen.</p><h3>Practical limits of the search reproduction</h3><p>A fair local comparison requires RS and GWO to use the same feature order, scaling policy, candidate-training routine, validation partition, objective, and failure policy. Only the proposal mechanism should differ. The generated pipeline records the validation fraction as a local choice and shares the candidate objective, but this does not recover the paper's undisclosed validation protocol.</p><p>The paper's GWO superiority claim remains a reported paper result, not a result established by this scaffold. In particular, the generated GWO update, population size, iteration count, bounds, initialization, and discrete decoding cannot be labeled as the paper's implementation. The four-part labels in <code>LSTM-GWO (1-1-0-1)</code>, <code>GRU-GWO (2-1-1-1)</code>, and <code>SRU-GWO (2-2-0-0)</code> also remain opaque; the search code does not decode them.</p><p>The related contract tests are intended to check local invariants such as candidate decoding, seeded RS sampling, GWO shape and bound handling, and the common objective interface. They were not executed in this run. Static verification and semantic code verification were also skipped, so the generated files should be reviewed before relying on them operationally. No search, training run, or reproduction result is claimed here.</p><h2>Held-out evaluation, metrics, and residual analysis</h2><p>How do we know whether a forecast is useful? First, each prediction must be paired with the correct held-out observation. Then both arrays must be evaluated under the same scale and metric conventions. The evaluation layer therefore aligns predictions and targets, optionally restores original price units, computes the seven metrics named by the paper, and summarizes residual behavior.</p><p>This section distinguishes three kinds of statements. The paper names the metrics and discusses error distributions, but it does not provide canonical formulas, denominator rules, residual signs, or plotting conventions. The generated implementation supplies explicit conventional choices for those missing details. The resulting values would be local evaluation outputs&#8212;not verified reproductions of the paper's reported results.</p><h3>From model output to aligned evaluation arrays</h3><p>The generated <code>predict_dataset</code> function runs inference in loader order. It switches the model to evaluation mode so dropout is disabled, uses no-gradient inference, collects scalar predictions and targets batch by batch, and restores the model's previous training-mode flag afterward. Its scalar-target contract accepts arrays shaped <code>samples</code> or <code>samples &#215; 1</code>. It does not reorder or silently truncate samples.</p><p>The key alignment operation is exposed separately as <code>align_predictions_and_targets</code>. It rejects unequal sample counts and preserves the order emitted by the loader. This is why held-out loaders should normally use <code>shuffle=False</code>: a shuffled loader can still produce pairs within a batch, but it no longer represents the original chronological output order.</p><p>The following excerpt is copied from <code>predict_dataset</code>. Notice the two independent checks: the model output and the batch targets are normalized to the scalar-target representation, and their batch counts must agree before either is appended.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;706da20e-4ef5-4f19-be0f-0b5543061746&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">                outputs = model(inputs)
                if not isinstance(outputs, Tensor):
                    raise TypeError(
                        f"model output for batch {batch_index} must be a tensor"
                    )
                predictions, _ = _normalise_scalar_targets(outputs, "predictions")
                batch_targets, _ = _normalise_scalar_targets(targets, "targets")
                if predictions.shape[0] != batch_targets.shape[0]:
                    raise ValueError(
                        f"batch {batch_index} prediction and target counts differ: "
                        f"{predictions.shape[0]} != {batch_targets.shape[0]}"
                    )
                prediction_parts.append(predictions)
                target_parts.append(batch_targets)</code></pre></div><p>The function returns predictions in loader order, while the targets are retained internally for alignment validation. The model input remains the package-wide <code>batch &#215; time &#215; 5</code> contract; this evaluation helper specifically supports the configured scalar-output case. A multivariate target would require an explicit extension rather than an implicit interpretation.</p><h3>Choosing the evaluation scale</h3><p>The paper states that inputs are scaled to <code>[0, 0.95]</code>, but it does not say whether its reported metrics were computed in normalized space or after inverse transformation to price units. The generated pipeline exposes both choices through <code>EvaluationConfig</code>.</p><p>In normalized evaluation, <code>actual_targets</code> and <code>predictions</code> remain in scaled space. RMSE then has normalized units. In original-unit evaluation, <code>inverse_transform_target</code> embeds each scalar target into a five-feature row, applies <code>FeaturewiseMinMaxScaler.inverse_transform</code>, and extracts the configured target column. RMSE can then be interpreted in price units. MAPE, RMSPE, and PBIAS remain percentage-valued under the local conventions, while agreement metrics remain unitless.</p><p>The target column itself is also unresolved by the paper. The generated pipeline requires a configured choice such as <code>Close</code>; it does not assume that the paper's phrase &#8220;stock price&#8221; uniquely means Close or adjusted Close. The forecast horizon is likewise recorded separately from the input <code>window_length</code>.</p><p>The following excerpt is copied from <code>inverse_transform_target</code> and shows why the scaler needs a target index even though the prediction contains only one value per sample.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a27e9eed-77a8-41f8-9716-abf6b14bc732&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    embedded = np.zeros((flat_values.shape[0], scaler.n_features), dtype=np.float64)
    embedded[:, target_index] = flat_values
    restored = scaler.inverse_transform(embedded)[:, target_index]

    if len(original_shape) == 1:
        return restored
    return restored.reshape(original_shape)</code></pre></div><p>This operation preserves the prediction shape while using the scaler's five feature-specific statistics. It is an implementation mapping, not a formula recovered from the paper. The scaler must already be fitted, and its fitting policy is recorded as training-only in the generated preprocessing pipeline to reduce leakage.</p><h3>The seven named metrics</h3><p>The paper names <code>R&#178;</code>, <code>RMSE</code>, <code>MAPE</code>, <code>RMSPE</code>, <code>PBIAS</code>, Willmott Index, and <code>NSE</code>. Because no canonical equation records were supplied, the generated <code>metrics.py</code> module documents and implements conventional local definitions. The important practical point is consistency: every candidate must use the same scale, alignment rule, sign convention, and zero-denominator policy.</p><ul><li><p><code>R&#178;</code> summarizes explained variation relative to an actual-value mean baseline. The generated function stores the conventional result on the unit-interval scale by default. A display convention can multiply it by 100, which supports the paper's percentage-style reporting. Thus a paper display of <code>99.2427%</code> is distinct from the underlying unit-interval representation <code>0.992427</code>.</p></li><li><p><code>RMSE</code> measures the typical magnitude of prediction error in the units supplied to the function. It is therefore scale-sensitive: normalized inputs yield normalized-unit RMSE, while inverse-transformed prices yield price-unit RMSE.</p></li><li><p><code>MAPE</code> averages absolute relative errors and returns a percentage. It cannot divide by zero actual values, so <code>EvaluationConfig</code> selects whether such observations raise an error, are ignored, or receive a defined local treatment.</p></li><li><p><code>RMSPE</code> is the root-mean-square version of relative error and uses the same zero-actual policy.</p></li><li><p><code>PBIAS</code> summarizes signed aggregate bias. The generated default residual direction is <code>actual_minus_predicted</code>, so positive PBIAS indicates underprediction under that convention.</p></li><li><p>Willmott Index measures agreement using the generated module's selected standard convention. Its denominator can be undefined for degenerate data, in which case the local implementation returns <code>NaN</code>.</p></li><li><p><code>NSE</code>, or Nash&#8211;Sutcliffe Efficiency, compares squared forecast error with variation around the actual-series mean. Like <code>R&#178;</code>, it is undefined when the actual series has no variation.</p></li></ul><p>The module's report assembler keeps these choices together. The following excerpt is copied from <code>compute_metric_report</code>; it shows that one configuration controls the conventions used by the individual metric functions.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;0a39b119-2caa-4839-b7d4-36dfc141ce25&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    _validate_zero_policy(str(zero_policy))
    return {
        "r_squared": r_squared(actual_vector, predicted_vector, str(r2_convention)),
        "rmse": rmse(actual_vector, predicted_vector),
        "mape": mape(actual_vector, predicted_vector, str(zero_policy)),
        "rmspe": rmspe(actual_vector, predicted_vector, str(zero_policy)),
        "pbias": pbias(actual_vector, predicted_vector, str(pbias_convention)),
        "willmott_index": willmott_index(
            actual_vector,
            predicted_vector,
            str(willmott_convention),
        ),
        "nash_sutcliffe_efficiency": nash_sutcliffe_efficiency(actual_vector, predicted_vector),
    }</code></pre></div><p>Before reaching this block, <code>_aligned_arrays</code> converts supported array-like inputs to finite one-dimensional vectors and rejects mismatched shapes. This failure behavior matters: a metric calculated from mispaired samples can look numerically plausible while describing the wrong forecast errors.</p><h3>A small illustrative interface example</h3><p>The following excerpt is adapted directly from the generated metric test's synthetic setup. It demonstrates the intended interface with short local arrays; it is not HSI data, and no numerical result from this example should be compared with the paper's benchmarks. The test uses <code>EvaluationConfig()</code> to select the generated defaults rather than claiming that those defaults were specified by the paper.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e253d247-f3be-4e6c-a8e5-ec3b51cb4b90&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    actual = np.array([100.0, 105.0, 110.0, 115.0], dtype=float)
    predicted = np.array([99.0, 106.0, 109.0, 116.0], dtype=float)

    # Defaults are deliberately supplied by EvaluationConfig rather than
    # inferred from the paper, whose metric conventions are unspecified.
    report = compute_metric_report(actual, predicted, EvaluationConfig())</code></pre></div><p>A corresponding residual summary can be created from the same aligned arrays. In the generated implementation, <code>compute_residuals(actual, predicted)</code> returns <code>actual - predicted</code> with the same shape, and <code>summarize_residuals</code> flattens the result before calculating its count, mean, population standard deviation, standardized skewness, and excess kurtosis. Fewer than three observations, or zero variance, leave skewness undefined; fewer than four observations, or zero variance, leave excess kurtosis undefined. These estimator and insufficiency rules are implementation decisions because the paper only mentions mean, standard deviation, skewness, and kurtosis without defining their conventions.</p><h3>Residuals and visual analysis</h3><p>Residuals make systematic behavior easier to inspect than a single score. Under the generated <code>actual_minus_predicted</code> convention, positive residuals indicate underprediction and negative residuals indicate overprediction. <code>compare_error_distributions</code> creates a stable table for named residual groups, preserving empty groups with an explicit count and undefined summary values rather than inventing observations.</p><p>The visualization adapters implement the analysis types mentioned by the paper:</p><ul><li><p><code>plot_error_boxplots</code> shows grouped residual spread and central tendency.</p></li><li><p><code>plot_error_violins</code> shows distribution shape.</p></li><li><p><code>plot_metric_comparison</code> gives each metric its own panel so percentage and unitless measures are not placed on one misleading axis.</p></li><li><p><code>plot_pbias_radar</code> displays signed PBIAS values using a locally selected radial convention.</p></li><li><p><code>plot_taylor_summary</code> uses a locally selected correlation-angle and standard-deviation representation.</p></li></ul><p>These functions operate on local predictions and reports. They do not contain the paper's plotted points, styling, axes, or source predictions. Therefore they can support analogous analysis, but they cannot claim exact reproduction of the paper's figures.</p><h3>Evaluate only after model selection</h3><p>The evaluation pipeline enforces the intended experiment order conceptually: train candidates and compare them using the configured validation objective, select a candidate, and only then read the chronological held-out partition for final metrics. <code>evaluate_training_run</code> evaluates a retained trained model. <code>evaluate_search_result</code> requires that a selected trained model be retained; it does not silently retrain from an incomplete search record.</p><p>The resulting <code>EvaluationRun</code> stores predictions, aligned actual targets, residuals, the metric report, and provenance. Provenance includes the evaluation scale, target column, forecast horizon, window length, metric conventions, residual sign, held-out sample count, and preprocessing metadata. This makes a future comparison auditable: a difference in RMSE can be investigated as a possible scale, target, or convention difference rather than treated as unexplained model behavior.</p><p>The paper's seven metrics and its discussion of residual distributions are therefore represented, but not overclaimed. The formulas and conventions are local implementation choices, the target and evaluation scale remain configurable, and no execution or verification occurred in this run. Any future local report must be labeled as a computed result and kept separate from the paper's reported benchmark records.</p><h2>Reported LSTM-GWO, GRU-GWO, and SRU-GWO benchmarks</h2><p>How should you use the paper&#8217;s reported numbers when the original experiment cannot yet be reconstructed exactly? Treat them as reference records, not as expected outputs that automatically validate a new run. The generated package stores the reported values separately from locally computed predictions and metrics, so a later experiment can be compared with the paper without confusing the two.</p><h3>What the paper reports</h3><p>The paper identifies its strongest reported configuration as LSTM-GWO <code>(1-1-0-1)</code>. The four-part label is opaque: the supplied paper context does not define whether its fields represent layer counts, dropout settings, dense layers, or another encoding. The label must therefore remain a benchmark identifier rather than being translated into <code>ModelConfig</code> values.</p><p>The reported LSTM-GWO metrics are:</p><ul><li><p>R&#178;: <code>99.2427%</code></p></li><li><p>RMSE: <code>339.3902</code></p></li><li><p>MAPE: <code>1.1721%</code></p></li><li><p>RMSPE: <code>1.6221%</code></p></li><li><p>PBIAS: <code>0.0523</code></p></li><li><p>Willmott Index: <code>0.9981</code></p></li><li><p>NSE: <code>0.9924</code></p></li></ul><p>The paper also reports the following best labels and associated values:</p><ul><li><p><strong>GRU-GWO `(2-1-1-1)`:</strong> R&#178; <code>99.2322%</code>, RMSE <code>341.7225</code>, MAPE <code>1.1821%</code>, and PBIAS <code>&#8722;0.1357</code>.</p></li><li><p><strong>SRU-GWO `(2-2-0-0)`:</strong> R&#178; <code>99.2009%</code>, RMSE <code>348.6384</code>, and MAPE <code>1.2080%</code>.</p></li></ul><p>The remaining GRU and SRU metric associations are not clear in the supplied extraction. They are intentionally left unavailable rather than inferred from nearby text or reconstructed tables. Likewise, the configuration labels remain undecoded.</p><p>These values also use mixed display conventions. R&#178; is shown by the paper as a percentage, whereas Willmott Index and NSE are shown on a unit-interval scale. A local metrics implementation may store R&#178; internally as <code>0.992427</code> and display it as <code>99.2427%</code>, but that conversion must be recorded explicitly before comparing values.</p><h3>How the generated code preserves the evidence</h3><p><code>PaperBenchmark</code> is an immutable record containing the recurrent family, optimizer, opaque configuration label, metric values, units, and provenance. Its metric schema always contains the seven named metrics, but unavailable values are represented by <code>None</code>. The following excerpt is copied from <code>reported_benchmarks</code> in <code>src/compositional_rnn_stock/evaluation/benchmarks.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;cef51fc2-8d4d-419a-b82f-71a9430bf9ae&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def reported_benchmarks() -&gt; tuple[PaperBenchmark, ...]:
    """Return the three benchmark records supplied in the paper context.

    Missing GRU and SRU associations remain ``None`` rather than being inferred from
    other reported values. R&#178; is stored in the paper's percentage display convention.
    """
    return (
        PaperBenchmark(
            family="LSTM",
            optimizer="GWO",
            configuration_label="1-1-0-1",
            metrics=_benchmark_metrics(
                r_squared=99.2427,
                rmse=339.3902,
                mape=1.1721,
                rmspe=1.6221,
                pbias=0.0523,
                willmott_index=0.9981,
                nash_sutcliffe_efficiency=0.9924,
            ),
            metric_units=dict(_METRIC_UNITS),
        ),</code></pre></div><p>The function returns records marked as supplied paper targets. It does not train a model, generate predictions, or assert that any local candidate achieved these values. The LSTM record is complete, while the later GRU and SRU records retain <code>None</code> for unavailable associations.</p><p>The generated contract test makes this distinction explicit when it retrieves the records:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;44453fcf-df6b-421b-9c5b-ff0e6c7a1415&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">benchmarks = reported_benchmarks()

assert [(item.family, item.optimizer, item.configuration_label) for item in benchmarks] == [
    ("LSTM", "GWO", "1-1-0-1"),
    ("GRU", "GWO", "2-1-1-1"),
    ("SRU", "GWO", "2-2-0-0"),
]</code></pre></div><p>This is a planned test file, not evidence of an executed check. Under the authoritative run policy, code execution and verification were disabled.</p><h3>Comparing a future local report</h3><p>The <code>compare_to_report</code> function accepts a local metric mapping and one <code>PaperBenchmark</code>. It produces rows only when both sides have a value. Differences are expressed as local minus reported in the benchmark&#8217;s displayed units. Thus a missing GRU RMSPE value does not become a guessed comparison row, and a local unit-interval R&#178; can be converted to the paper&#8217;s percentage convention when the local report has not declared another unit.</p><p>A comparison is meaningful only when the upstream experiment is also aligned. The provenance should identify the target column, forecast horizon, scaling range and fitting policy, window length, chronological split, recurrent-family configuration, seed, training settings, search settings, evaluation scale, and metric conventions. Without those fields, a numerical difference cannot reveal whether the models or datasets were actually comparable.</p><p>The evaluation pipeline places these comparisons alongside a local <code>EvaluationRun</code>, whose predictions, aligned targets, residuals, metric report, and provenance are kept together. This connects benchmark comparison to the feature-wise compositional model: the local prediction must first come from the five independent OHLCV streams, concatenation fusion, and the selected recurrent and dense head. A benchmark difference cannot repair an architecture or target choice that the paper leaves unspecified.</p><h3>Reported evidence versus local conclusions</h3><p>The paper states that GWO outperformed Random Search across the evaluated LSTM, GRU, and SRU families. That statement is reported evidence from the paper. It is not a result established by this unexecuted scaffold. A local RS-versus-GWO comparison would require the same preprocessing, validation objective, candidate space, and declared search budget for both methods, followed by actual execution and evaluation.</p><p>The safe interpretation is therefore:</p><ol><li><p><code>reported_benchmarks()</code> preserves the supplied reference records.</p></li><li><p><code>compare_to_report</code> provides a labeled comparison mechanism for future local outputs.</p></li><li><p>Missing metrics and opaque configuration labels remain unresolved.</p></li><li><p>No local metric is claimed to match the paper.</p></li></ol><p>The benchmark records are useful precisely because they retain their provenance and limitations. They provide targets for a future, better-specified reproduction rather than proof that the current configurable implementation reproduces the published experiment.</p><h2>Offline synthetic demonstration and artifact provenance</h2><p>How can you learn the repository&#8217;s data and model interfaces without downloading market data? Use the offline demonstration. It creates clearly labeled synthetic OHLCV rows, applies the same five-channel scaling and windowing contracts, builds a configurable compositional model, and inspects the expected tensor shapes. This is a plumbing demonstration only: the synthetic series is not Hang Seng Index data and cannot validate forecasting performance.</p><h3>Why use an offline demonstration?</h3><p>The paper describes daily Hang Seng Index history obtained from Yahoo Finance, but the supplied context does not identify a unique ticker, date range, adjustment policy, or download configuration. Yahoo Finance data can also change as providers revise historical records. An offline example avoids those dependencies while making every local choice visible.</p><p>The generated example uses a deterministic pseudo-random generator when given a selected <code>seed</code>. It creates five columns in the canonical order <code>Open</code>, <code>High</code>, <code>Low</code>, <code>Close</code>, <code>Volume</code>. The values are synthetic prices and volumes; they are not intended to reproduce the empirical distribution of the HSI.</p><p>The example&#8217;s generator is defined in <code>examples/offline_synthetic_demo.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;abc436cd-5eca-4335-8193-b4e8f5483a68&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def make_synthetic_ohlcv(num_days: int, seed: int) -&gt; pd.DataFrame:
    """Create deterministic, clearly labeled synthetic OHLCV observations.

    The returned table has shape ``(num_days, 5)`` and canonical column order:
    Open, High, Low, Close, Volume. Values are generated only for this
    demonstration; they are not intended to model the empirical HSI series.
    """
    if isinstance(num_days, bool) or not isinstance(num_days, int) or num_days &lt;= 0:
        raise ValueError("num_days must be a positive integer")
    if isinstance(seed, bool) or not isinstance(seed, int) or seed &lt; 0:
        raise ValueError("seed must be a non-negative integer")

    rng = np.random.default_rng(seed)
    daily_returns = rng.normal(loc=0.0004, scale=0.012, size=num_days)
    close = 18_000.0 * np.exp(np.cumsum(daily_returns))
    open_noise = rng.normal(loc=0.0, scale=35.0, size=num_days)
    open_price = close + open_noise
    intraday_spread = np.abs(rng.normal(loc=55.0, scale=18.0, size=num_days))
    high = np.maximum(open_price, close) + intraday_spread
    low = np.minimum(open_price, close) - intraday_spread
    volume = np.maximum(
        1.0,
        1_000_000.0 + rng.normal(loc=0.0, scale=120_000.0, size=num_days),
    )

    values = np.column_stack((open_price, high, low, close, volume))
    return pd.DataFrame(
        values,
        index=pd.date_range("2020-01-01", periods=num_days, freq="D"),
        columns=list(CANONICAL_OHLCV_COLUMNS),
    )</code></pre></div><p>The function accepts a positive number of rows and a non-negative integer seed. Its output is a table with shape <code>num_days &#215; 5</code>. The validation errors are useful failure cases: a boolean is not accepted as an integer, and invalid row counts or seeds are rejected rather than silently corrected.</p><h3>Scaling an earlier chronological portion</h3><p>The paper states the intended feature range <code>[0, 0.95]</code>, but it does not specify exactly which observations fit the scaling statistics. The demonstration chooses an earlier chronological portion for fitting. This is a leakage-prevention decision made by the implementation, not a recovered paper detail. The remaining observations are transformed with the already fitted scaler.</p><p>The relevant portion of <code>main</code> is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4d5edc40-1ffb-44e3-9fa9-b6ab4a043288&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">data_config = DataConfig(
    target_column="Close",
    forecast_horizon=1,
    window_length=20,
    split_ratio=0.8,
    scaling_range=(0.0, 0.95),
)
raw_frame = make_synthetic_ohlcv(num_days=96, seed=7)
raw_values = raw_frame.loc[:, list(CANONICAL_OHLCV_COLUMNS)].to_numpy(
    dtype=np.float64
)

# Fit on an earlier chronological portion to demonstrate a leakage-avoiding
# choice. The paper states the [0, 0.95] range but does not specify fit scope.
scaler_split = int(raw_values.shape[0] * data_config.split_ratio)
scaler = FeaturewiseMinMaxScaler(feature_range=data_config.scaling_range)
scaler.fit(raw_values[:scaler_split])
scaled_values = scaler.transform(raw_values)</code></pre></div><p>Here, <code>target_column='Close'</code>, <code>forecast_horizon=1</code>, and <code>window_length=20</code> are tutorial configuration choices. The paper does not identify the target column or forecast horizon; it only describes 20-day and 40-day input windows. Similarly, the demonstration&#8217;s 96 synthetic rows and seed <code>7</code> are not paper data.</p><p><code>FeaturewiseMinMaxScaler</code> requires a final feature axis of length five. It stores separate scaling metadata for each OHLCV channel, so Volume does not determine the scaling of the price columns. Its <code>inverse_transform</code> method is available when later evaluation should return predictions to original price units. The example does not calculate or present forecast metrics.</p><h3>Building windows and inspecting model shapes</h3><p>After scaling, the example selects the configured target column and constructs overlapping supervised windows:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;deeea9eb-5275-4195-a15d-9504e5544a46&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">target_index = list(CANONICAL_OHLCV_COLUMNS).index(data_config.target_column)
windowed = make_sliding_windows(
    scaled_values,
    target_index=target_index,
    window_length=data_config.window_length,
    forecast_horizon=data_config.forecast_horizon,
    timestamps=raw_frame.index.to_numpy(),
)

# The model follows the paper's feature-wise structure: five independent
# recurrent streams, concatenation fusion, and a one-value prediction head.
model_config = ModelConfig(
    recurrent_family="LSTM",
    stream_layers=1,
    stream_units=16,
    stream_dropout=0.0,
    fusion_dropout=0.0,
    use_post_fusion_recurrent=False,
    dense_layers=(16,),
    output_dimension=1,
)
model = build_compositional_model(model_config)
model.eval()

batch_size = min(4, windowed.sample_count)
batch_inputs = torch.as_tensor(windowed.inputs[:batch_size], dtype=torch.float32)
with torch.no_grad():
    predictions = model(batch_inputs)</code></pre></div><p><code>make_sliding_windows</code> returns inputs with shape <code>samples &#215; window_length &#215; 5</code> and scalar targets with shape <code>samples</code>. With this tutorial configuration, each input contains 20 time steps and five channels. The target is the selected <code>Close</code> value one step after the input window according to the implementation&#8217;s explicit alignment convention.</p><p><code>build_compositional_model</code> then constructs the feature-wise recurrent model. Each batch input has shape <code>batch &#215; 20 &#215; 5</code>; the model separates the five channels, encodes them independently, concatenates their representations, and returns a prediction with shape <code>batch &#215; output_dimension</code>. The model configuration shown here is an educational choice. It does not decode the paper&#8217;s opaque benchmark label <code>LSTM-GWO (1-1-0-1)</code>, nor does it claim to recover the paper&#8217;s exact layer widths or training recipe.</p><p>The source example prints these contracts:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;36e496df-59e8-4656-bc47-0131774f87f9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">print("Synthetic demonstration only; no paper result is being reproduced.")
print(f"raw OHLCV shape: {raw_values.shape}")
print(f"scaled OHLCV range: {data_config.scaling_range}")
print(
    "window input shape (samples, time_steps, features): "
    f"{windowed.inputs.shape}"
)
print(f"window target shape: {windowed.targets.shape}")
print(
    "model batch input shape (batch, time_steps, features): "
    f"{tuple(batch_inputs.shape)}"
)
print(f"model prediction shape (batch, output_dimension): {tuple(predictions.shape)}")</code></pre></div><p>These statements document intended dimensions; this run did not execute the example. In particular, no output numbers, predictions, or paper-comparison result should be inferred from the excerpt.</p><h3>Recording provenance and arrays locally</h3><p>A demonstration becomes more useful when its assumptions travel with its outputs. Artifact provenance should identify at least the data source, cleaning policy, scaling range and fit scope, window length, target column, forecast horizon, model family, architecture settings, seed, and evaluation conventions. For real HSI work, the archived raw file and the Yahoo Finance query parameters should also be retained.</p><p>The generated artifact module provides local persistence through <code>save_json_record</code>, <code>save_array</code>, and <code>load_json_record</code>. The README illustrates the JSON-record call as follows:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;fc0dd15d-26ce-4f9a-baaf-9565fab8c236&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">save_json_record(report, Path('artifacts/report.json'))</code></pre></div><p>A JSON record is appropriate for serializable configuration, acquisition metadata, histories, benchmark labels, and metric reports. Numeric predictions and residuals should be stored with <code>save_array</code>, which preserves their numeric shape and dtype. <code>load_json_record</code> reads a previously saved top-level JSON object. The artifact implementation uses local filesystem operations and rejects implicit overwrites, helping prevent an old experiment from being silently replaced.</p><p>For example, a future evaluation record could contain a local run identifier, <code>window_length</code>, <code>target_column</code>, <code>forecast_horizon</code>, scaler settings, model configuration, seed, metric conventions, and a separate reference to prediction and residual arrays. It should distinguish computed local values from the paper&#8217;s reported benchmark records. The generated interfaces provide the storage operations, while the exact record contents remain an experiment-provenance responsibility.</p><h3>What this demonstration does not establish</h3><p>The synthetic path confirms the intended software contracts conceptually: five ordered channels can be scaled, transformed into windows, and supplied to a feature-wise compositional model. It does not establish that the model forecasts the HSI accurately, that the chosen target and horizon match the paper, or that the reported LSTM-GWO, GRU-GWO, or SRU-GWO values can be reproduced.</p><p>No example execution, artifact creation, test execution, or verification occurred under the run policy. A stronger reproduction would archive the original Yahoo Finance data query and downloaded rows, supply the missing architecture and optimization details, run the configured experiment, and then evaluate predictions under explicitly documented metric conventions. Until then, the offline example is a reproducibility aid&#8212;not evidence of paper-level performance.</p><h2>Verification status, limitations, and responsible reproduction claims</h2><p>How can a carefully organized implementation be useful without overstating what it proves? The answer is to separate <strong>invariants</strong>, <strong>planned checks</strong>, and <strong>verified results</strong>. The generated package records the paper&#8217;s intended data flow, model structure, search interfaces, evaluation conventions, and provenance. However, this run did not execute code or perform verification. Its correct status is therefore <code>verification_skipped</code>: a configurable reproduction scaffold with explicit gaps, not a validated reproduction of the published experiment.</p><h3>What the implementation is designed to preserve</h3><p>Several structural requirements are clear enough to review independently of the paper&#8217;s missing experimental details:</p><ul><li><p>The input feature order is exactly <code>Open</code>, <code>High</code>, <code>Low</code>, <code>Close</code>, <code>Volume</code>.</p></li><li><p>Model inputs use <code>batch &#215; time &#215; 5</code> ordering.</p></li><li><p>Preprocessing constructs chronological windows and keeps training samples before later samples.</p></li><li><p>The scaler targets the stated interval <code>[0, 0.95]</code>. Fitting it on earlier training observations only is a leakage-prevention implementation choice; the paper does not specify its fitting partition.</p></li><li><p>The model creates exactly five independent feature streams.</p></li><li><p>Stream representations are fused by concatenation, rather than by early summation or mixing.</p></li><li><p>Dropout is intended for training mode and should be inactive during evaluation.</p></li><li><p>Predictions and targets must remain sample-aligned before metric calculation.</p></li><li><p>Search candidates must decode to valid typed values for layer counts, units, dropout rates, learning rates, batch sizes, and epochs.</p></li><li><p>RS and GWO should use the same candidate objective, preprocessing, and validation protocol when compared locally.</p></li><li><p>Benchmark labels such as <code>LSTM-GWO (1-1-0-1)</code> remain opaque identifiers and are not translated into architecture settings.</p></li></ul><p>These are implementation contracts and interpretation safeguards. They are not evidence that a particular model reached the paper&#8217;s reported metrics.</p><h3>Planned contract tests are not completed checks</h3><p>The repository contains generated test files intended to make these invariants reviewable later. For data handling, <code>test_scaler_range_and_inverse_round_trip</code>, <code>test_window_target_alignment</code>, and <code>test_chronological_split_has_no_reordering</code> address scaling behavior, window-to-target indexing, and chronological partitioning. The model tests include <code>test_model_accepts_five_feature_input</code>, <code>test_fusion_dimension_is_additive</code>, <code>test_recurrent_family_dispatch_is_consistent</code>, and <code>test_dropout_differs_only_in_training_mode</code>.</p><p>The remaining planned tests cover local metric conventions, candidate encoding, search bounds, and provenance. Examples include <code>test_rmse_zero_for_equal_arrays</code>, <code>test_percentage_metric_zero_policy</code>, <code>test_candidate_codec_validates_discrete_parameters</code>, <code>test_gwo_positions_stay_in_bounds_after_decoding</code>, <code>test_reported_benchmarks_preserve_supplied_values</code>, and <code>test_incomplete_metrics_remain_missing</code>.</p><p>These names describe intended checks, not completed checks. Under the authoritative run policy, test generation and execution were disabled. No test result, passing assertion, successful import, or model output should be inferred from the presence of these files.</p><h3>Static and semantic verification status</h3><p>Static verification and semantic verification address different questions. Static review would inspect syntax, imports, public interfaces, tensor-shape contracts, and configuration propagation without running the experiment. Semantic review would assess whether the implementation&#8217;s behavior matches the intended method, including feature independence, concatenation, dropout mode changes, search-objective consistency, and metric conventions.</p><p>Both forms of verification were skipped here. Local static verification reports the status <code>verification_skipped</code>, and semantic code verification was also skipped. Code execution, test execution, tutorial-section verification, and final quality review were disabled by policy. Consequently, this section must not claim that the generated code is correct, that the tests pass, or that any local result matches the paper.</p><h3>Why exact reproduction remains underdetermined</h3><p>The supplied paper context does not uniquely determine several decisions that materially affect results:</p><ul><li><p>The Yahoo Finance ticker or symbol, date range, download options, adjustment policy, missing-value handling, and duplicate-row handling.</p></li><li><p>The forecast target and horizon. The paper names &#8220;stock price&#8221; and lists OHLCV inputs, but does not identify whether the target is <code>Close</code>, adjusted <code>Close</code>, another price, or a multivariate output.</p></li><li><p>The exact interpretation of the 20-day and 40-day windows beyond treating them as input lengths.</p></li><li><p>The normalization formula, fitting partition, clipping behavior, and inverse-transformation procedure.</p></li><li><p>Recurrent layer counts, widths, activations, hidden-state extraction, dense-layer structure, and exact placement of both dropout stages.</p></li><li><p>The SRU variant and Python implementation or API.</p></li><li><p>The training loss, gradient optimizer, initialization, learning-rate schedule, seed, shuffling policy, stopping rule, and checkpoint-selection rule.</p></li><li><p>RS search domains, distributions, trial budget, and validation protocol.</p></li><li><p>GWO population size, initialization, bounds, iteration count, update schedule, objective, and encoding of discrete parameters.</p></li><li><p>Metric formulas, percentage conventions, zero-denominator handling, bias signs, and the scale used for evaluation.</p></li></ul><p>The paper&#8217;s reported total of 54 configurations does not resolve these omissions. Nor can the opaque four-part labels in the LSTM-GWO, GRU-GWO, and SRU-GWO records be safely decoded from the supplied evidence.</p><h3>A responsible review checklist</h3><p>Before treating a future local run as evidence, review the following items and record each one in the experiment provenance:</p><ol><li><p><strong>Data shape:</strong> confirm that the prepared inputs use <code>batch &#215; time &#215; 5</code> and the canonical OHLCV order.</p></li><li><p><strong>Chronology:</strong> confirm that the split and window-target alignment preserve temporal order and do not use future observations during fitting.</p></li><li><p><strong>Scaling:</strong> record the <code>[0, 0.95]</code> range, the fitting partition, constant-feature policy, and any inverse transformation.</p></li><li><p><strong>Target semantics:</strong> record the target column and forecast horizon instead of assuming that the paper specifies them.</p></li><li><p><strong>Model structure:</strong> record the five independent encoders, recurrent family, representation choice, dropout placement, fusion operation, post-fusion layers, and output dimension.</p></li><li><p><strong>Training:</strong> record the loss, gradient optimizer, learning rate, batch size, epochs, seed, validation protocol, and checkpoint policy.</p></li><li><p><strong>Search:</strong> record the typed domains, objective direction, trial or wolf budget, bounds, discrete decoding, and failure handling.</p></li><li><p><strong>Evaluation:</strong> record whether metrics use normalized or original units, every metric convention, percentage display rules, and zero handling.</p></li><li><p><strong>Benchmarks:</strong> preserve reported values separately from local values, keep missing GRU and SRU associations missing, and leave configuration labels opaque.</p></li><li><p><strong>Verification:</strong> record whether code and tests were actually executed and reviewed; do not substitute planned test names for evidence.</p></li></ol><p>This checklist is a review target, not a statement that the items were checked in this run.</p><h3>What a stronger reproduction would require</h3><p>A stronger claim would require the original data query or an archived copy of the downloaded OHLCV data, the complete architecture and configuration tables, the exact target and forecast horizon, and the full training and validation protocol. It would also require the RS and GWO search spaces and budgets, the selected SRU implementation, canonical metric definitions, and the original prediction or error outputs needed to compare visual analyses.</p><p>After those details were supplied, the implementation could use its provenance records to replace local choices without changing the overall public interfaces. The resulting experiment would still need to be executed, checked, and compared under a documented protocol before it could be described as a reproduction.</p><p>Attention mechanisms, hybrid architectures, and technical indicators are mentioned as future work in the paper. They are outside this reproduction workflow and should not be added to close the evidence gaps.</p><p>The appropriate conclusion at this stage is deliberately modest: the package makes the paper-to-code assumptions visible and organizes the intended pipeline, but neither the generated scaffold nor the supplied benchmark constants establish successful or exact reproduction.</p><p>Use the button or URL below to download the source code.</p>
      <p>
          <a href="/__u/onepagecode.substack.com/p/quant-trading-feature-wise-compositional-084">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Quant Trading: Deep GRU Recurrent Neural Networks for Stock Price Forecasting (Python Guide)]]></title><description><![CDATA[Building an end-to-end multi-layer GRU time-series model for weekly stock return prediction on AAPL and MSFT.]]></description><link>https://onepagecode.substack.com/p/quant-trading-deep-gru-recurrent</link><guid isPermaLink="false">https://onepagecode.substack.com/p/quant-trading-deep-gru-recurrent</guid><dc:creator><![CDATA[Onepagecode]]></dc:creator><pubDate>Thu, 06 Aug 2026 03:55:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_W6n!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3078fdad-8e80-40a9-be17-640eb18fa447_1254x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Use the URL to download the source code!</h2><p>The paper develops and evaluates separate GRU recurrent neural-network models for forecasting weekly adjusted closing prices of Apple Inc. (AAPL) and Microsoft Corporation (MSFT). Each series contains 263 observations from Yahoo Finance. The workflow consists of exploratory analysis, missing-value and outlier checks, decomposition, ADF stationarity testing, autocorrelation analysis, three-period sliding-window transformation, training-only MinMax normalization, chronological splitting, Hyperband-based hyperparameter tuning, GRU training with Adam, early stopping and checkpointing, inverse transformation, and validation forecasting evaluation with MSE, MAE, MAPE, and forecast accuracy. The selected AAPL model has three GRU layers with 240, 16, and 16 units; the selected MSFT model has two GRU layers with 224 and 48 units. Reported MAPE values are 3.37% for AAPL and 2.55% for MSFT.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>Implementation Assumptions</h2><ul><li><p>Use appendix records as the primary local data source when the exact Yahoo Finance retrieval and weekly aggregation procedure cannot be established.</p></li><li><p>Represent each raw asset series as 263 chronological observations and each windowed dataset as 260 samples with shape (260, 3, 1).</p></li><li><p>Use chronological sample partitions of 156 training, 52 testing, and 52 validation samples.</p></li><li><p>Keep AAPL and MSFT data, scalers, models, histories, checkpoints, and metrics separate.</p></li><li><p>Treat exploratory decomposition, ADF testing, autocorrelation, histograms, and boxplots as diagnostics only; they do not modify forecasting inputs.</p></li><li><p>Use one explicit, documented MinMaxScaler convention for reproducibility, while exposing the convention because the paper does not specify whether raw training prices or flattened window values were used.</p></li><li><p>Use the Table 1 configurations as the final reproduction targets: AAPL widths (240, 16, 16), MSFT widths (224, 48), zero dropout, learning rate 0.001, MSE, batch size 8, maximum 100 epochs, and patience 7.</p></li><li><p>Keep the testing subset distinct from validation until an explicit evaluation-subset option is selected.</p></li><li><p>Do not claim exact paper metric reproduction because source versions, tuner space, random seed, callback details, and evaluation protocol are incomplete.</p></li><li><p>No code execution, tests, static verification, semantic verification, tutorial verification, or final review is claimed under the supplied run policy.</p></li></ul><h2>Scope, Paper Facts, and Reproducibility Limits</h2><p>What exactly is being reproduced? The paper forecasts the next weekly adjusted closing price for two stocks: Apple, identified as <code>AAPL</code>, and Microsoft, identified as <code>MSFT</code>. Each stock is treated as its own univariate time series. In other words, the AAPL model does not receive MSFT prices, and neither model uses technical indicators, sentiment, volume, macroeconomic variables, or other explanatory features.</p><p>The model sees a short ordered history of adjusted prices and produces one next-price forecast. Adjusted closing prices, rather than ordinary closing prices, are the paper's retained price field because they account for effects such as dividends and stock splits. This distinction matters: substituting ordinary closing prices would change the input data and would no longer be the same experiment.</p><h3>Three evidence categories</h3><p>It helps to separate what is known from the paper, what follows mathematically from the data dimensions, and what the implementation must choose.</p><p><strong>Paper facts</strong> include 263 weekly adjusted closing-price observations for each asset, separate GRU models, a three-period sliding window, chronological partitioning, training-only MinMax normalization, Hyperband-based tuning, Adam optimization, early stopping, checkpointing, and evaluation with MSE, MAE, MAPE, and forecast accuracy. The final reported AAPL architecture has GRU widths of 240, 16, and 16. The final MSFT architecture has widths of 224 and 48. Both configurations report zero dropout, a learning rate of <code>0.001</code>, MSE loss, batch size <code>8</code>, a maximum of <code>100</code> epochs, and patience <code>7</code>.</p><p><strong>Derived dimensional correction.</strong> Windowing creates supervised examples before splitting. Three preceding prices are needed to predict one following price. Therefore, 263 raw observations produce 260 supervised samples. The paper-consistent chronological partition is consequently 156 training samples, 52 testing samples, and 52 validation samples. Each recurrent input has shape <code>(n, 3, 1)</code>: <code>n</code> examples, three time steps, and one price feature. Each target is standardized in the generated API as shape <code>(n, 1)</code>, one scalar target per example.</p><p><strong>Implementation decisions.</strong> Several details are absent from the supplied paper text. The generated project uses local appendix-style CSV files instead of inventing a Yahoo Finance download and weekly-resampling procedure. It exposes the choice between the testing and validation subsets rather than silently combining them. It also records a scaler convention, diagnostic parameters, Keras layer details, callback behavior, and tuner search ranges as implementation choices rather than paper facts.</p><p>No canonical equation records were supplied. This section therefore explains the tensor contracts, algorithm responsibilities, and data flow in prose and code. It does not reconstruct GRU gate equations, scaling formulas, loss formulas, or metric formulas.</p><h3>The end-to-end reproduction boundary</h3><p>The package is organized as a sequence of isolated per-asset stages:</p><ol><li><p>Load and validate one local adjusted-price series.</p></li><li><p>Run exploratory diagnostics without modifying the forecasting data.</p></li><li><p>Convert 263 prices into 260 three-step supervised windows.</p></li><li><p>Split those windows chronologically into 156, 52, and 52 samples.</p></li><li><p>Fit normalization using training data only and transform the later subsets.</p></li><li><p>Build the selected AAPL or MSFT GRU architecture.</p></li><li><p>Optionally run asset-specific Hyperband tuning, or instantiate the reported final configuration directly.</p></li><li><p>Train with Adam, validation-loss monitoring, early stopping, and checkpointing.</p></li><li><p>Select either the testing or validation subset explicitly.</p></li><li><p>Inverse-transform forecasts and targets, calculate metrics, and compare them with the paper's Table 3 references.</p></li></ol><p>The central orchestration function is <code>run_asset_pipeline</code>. Its inputs are an asset identifier, a local data directory, an output directory, an <code>evaluation_subset</code> value of either <code>"validation"</code> or <code>"test"</code>, and a Boolean <code>tune</code> flag. Its output is an <code>AssetRunResult</code> containing the raw <code>PriceSeries</code>, diagnostic records, the <code>WindowedDataset</code>, scaled <code>DatasetSplits</code>, scaler, model, training result, and forecast result. This retained-artifact design makes the intermediate boundaries inspectable instead of reducing the workflow to a single opaque training call.</p><p>The pipeline rejects unsupported assets, missing data directories, invalid evaluation-subset names, and invalid tuning flags. The lower-level validation stage is responsible for checking the expected 263 rows, chronology, missingness, finite numeric values, and positive adjusted prices. A failure at these boundaries should stop the reproduction rather than silently repair or reinterpret the input.</p><p>Here is the public configuration interface that captures the paper's final model settings and reference values:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;08414873-23d1-4702-af8a-b06a6de9eb6b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from stock_forecasting.config import (
    get_default_split_config,
    get_final_model_config,
    get_reference_metrics,
)

config = get_final_model_config("AAPL")
split = get_default_split_config()
reference = get_reference_metrics("AAPL")</code></pre></div><p><code>get_final_model_config</code> returns an immutable <code>AssetModelConfig</code>. For AAPL, its <code>units</code> are <code>(240, 16, 16)</code>; for MSFT, they are <code>(224, 48)</code>. It also carries the reported dropout, learning rate, loss name, batch size, epoch limit, and patience. <code>get_default_split_config</code> returns the derived 156/52/52 split and rejects configurations that do not total 260 samples. <code>get_reference_metrics</code> returns Table 3 values for comparison only; it does not calculate a forecast and does not claim that those values were achieved.</p><h3>Worked example: trace one AAPL run</h3><p>Assume that <code>data/raw/AAPL.csv</code> contains the normalized local input with columns <code>date</code> and <code>adjusted_close</code>, one asset per file, and 263 chronological rows. A complete AAPL pipeline call is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8d3a090e-3066-46c2-9da9-972eaa1c972b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from pathlib import Path

from stock_forecasting.pipeline import run_asset_pipeline

result = run_asset_pipeline(
    asset="AAPL",
    data_dir=Path("data/raw"),
    output_dir=Path("artifacts"),
    evaluation_subset="validation",
    tune=False,
)</code></pre></div><p>This call expresses several important decisions. <code>asset="AAPL"</code> keeps the model univariate and isolated. <code>evaluation_subset="validation"</code> follows the paper's statement that validation forecasting was reported, but it does not consume or merge the separate testing subset. <code>tune=False</code> selects the reported final Table 1 configuration directly rather than pretending that the incomplete original Hyperband metadata can be recovered.</p><p>The resulting <code>AssetRunResult</code> retains the main stages:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e101d2f2-f805-45ce-9806-387c66edec4d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">raw_prices = result.series.prices
window_inputs = result.windowed.inputs
scaled_training_inputs = result.splits.train_inputs
trained_model = result.model
forecast = result.forecast
comparison = result.diagnostics["reference_comparison"]</code></pre></div><p>The expected raw price array has length 263. The window input array has shape <code>(260, 3, 1)</code>. The scaled training input array has 156 samples, while the testing and validation arrays each have 52. <code>forecast</code> contains date-aligned actuals, predictions, and metrics after inverse transformation. <code>comparison</code> contains computed, reference, and difference fields for the four reported metrics; it is a descriptive comparison, not a verification result.</p><p>The lower-level pipeline ordering is visible in the generated implementation. This excerpt shows the key dimensional transitions and the explicit scaler choice:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;396ad934-811f-49ba-9789-e6e786385916&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    # load_validate_price_series: local appendix records are the offline input.
    series = load_asset_records(data_dir, canonical_asset)
    quality = validate_price_series(series)
    diagnostics = _run_diagnostics(series)
    diagnostics["quality"] = quality

    # create_windowed_supervised_dataset: 263 raw observations become 260 samples.
    windowed = create_windowed_supervised_dataset(series, window_length=3)

    # chronological_split_and_minmax_scale: 156/52/52 is applied after windowing.
    split_config: SplitConfig = get_default_split_config()
    raw_splits = split_windowed_dataset(windowed, split_config)
    scaled_splits, scaler = _scale_splits(
        raw_splits,
        convention=ScalerConvention.RAW_TRAINING_PRICES,
    )</code></pre></div><p>Notice that diagnostics are computed before modeling but are not passed into the GRU as features. Also notice that the split occurs after windowing and that the scaler is fitted through <code>_scale_splits</code> using training data. The generated pipeline records <code>ScalerConvention.RAW_TRAINING_PRICES</code> as its selected decision, while the paper itself leaves open whether the original implementation fitted on raw training prices or on flattened training windows and targets.</p><p>After preprocessing, the pipeline chooses the model configuration, builds the model, trains it, and evaluates the explicitly selected subset:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;085b8398-9d36-416c-a1aa-c2af36743e3f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    model_config = _select_model_config(
        canonical_asset,
        scaled_splits,
        tune=tune,
        output_dir=output_dir,
    )
    diagnostics["model_config"] = model_config
    diagnostics["tuning_requested"] = tune

    model = build_asset_gru_model(model_config, input_shape=(3, 1))
    checkpoint_path = output_dir / canonical_asset / "best.weights.h5"
    checkpoint_path.parent.mkdir(parents=True, exist_ok=True)
    training = train_gru_with_callbacks(
        model=model,
        splits=scaled_splits,
        config=model_config,
        checkpoint_path=checkpoint_path,
    )

    evaluation_inputs, evaluation_targets, evaluation_dates = select_evaluation_subset(
        scaled_splits,
        normalized_subset,
    )
    forecast = evaluate_forecasts(
        model=training.model,
        inputs=evaluation_inputs,
        targets=evaluation_targets,
        dates=evaluation_dates,
        scaler=scaler,
    )</code></pre></div><p>The model consumes scaled inputs with shape <code>(batch, 3, 1)</code> and returns one normalized scalar per sample. <code>evaluate_forecasts</code> converts predictions and targets back to price units using the training-fitted scaler before producing the reported metrics. The <code>evaluation_dates</code> are the dates of the target observations, so plotted forecasts correspond to the prices they are intended to predict.</p><p>For convenience, <code>run_both_assets</code> repeats this workflow for <code>AAPL</code> and <code>MSFT</code> independently. It returns a dictionary keyed by ticker and uses validation evaluation by default. Calling <code>run_asset_pipeline</code> directly is the appropriate way to select the testing subset for a particular asset.</p><h3>Command-line entry point</h3><p>The generated script exposes the unresolved evaluation choice and optional tuning choice explicitly:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;a0506b09-14a7-4b37-8258-2146830ea730&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">python scripts/run_reproduction.py \
  --data-dir data/raw \
  --output-dir artifacts \
  --evaluation-subset validation</code></pre></div><p>To keep the distinct testing subset as the held-out evaluation target, use:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;2fbd93c8-aa11-4914-844d-e67be2d9bfae&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">python scripts/run_reproduction.py \
  --data-dir data/raw \
  --output-dir artifacts \
  --evaluation-subset test</code></pre></div><p>Adding <code>--tune</code> requests the optional asset-specific Hyperband path. The non-tuning path is the clearer reproduction target when the goal is to instantiate the paper's reported Table 1 architectures. The exact original tuner search space, random seed, number of Hyperband iterations, and stopping criteria are unavailable, so a locally selected configuration cannot be described as an exact recovery of the paper's search.</p><h3>What the reference metrics do&#8212;and do not&#8212;mean</h3><p>The paper reports the following Table 3 reference values: AAPL MSE <code>55.80</code>, MAE <code>6.11</code>, MAPE <code>3.37%</code>, and forecast accuracy <code>96.63%</code>; MSFT MSE <code>144.44</code>, MAE <code>9.41</code>, MAPE <code>2.55%</code>, and forecast accuracy <code>97.45%</code>. The generated <code>compare_to_reference</code> function reports differences between a local run and these values. It does not turn the comparison into a pass/fail claim unless a caller supplies a tolerance, and even then the result describes the selected run rather than proving scientific equivalence.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://onepagecode.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">This Substack is reader-supported. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>These values may differ because the exact source-data version, weekly aggregation procedure, scaler fitting convention, random initialization, unspecified Keras defaults, callback behavior, Hyperband search space, and evaluation-subset protocol are incomplete. The lower MSFT MAPE reference also illustrates why relative and absolute errors should not be conflated: a higher price level can produce larger absolute errors while still producing a lower percentage error.</p><h3>Reproducibility and verification limits</h3><p>This is a reproduction scaffold, not a claim that the original experiment has been exactly rerun. The supplied run policy disabled code execution, test generation, local static verification, semantic code verification, tutorial verification, and final quality review. The local verification record therefore reports a skipped status, not a passing result, and semantic code verification was also skipped. No commands, training runs, tests, generated metrics, or reference matches are claimed here.</p><p>The paper itself also limits interpretation. It uses only two technology stocks and 263 weekly observations per asset. A price-history-only GRU may smooth abrupt movements and can be affected by distribution shift. The paper does not supply rolling-origin evaluation, confidence intervals, statistical significance testing, or broad benchmark comparisons. Consequently, forecasts should be treated as decision-support information for analysis and risk management, not as an autonomous trading system.</p><p>The practical scope is therefore precise: preserve the two isolated adjusted-price series, reproduce the 260-sample dimensional interpretation, keep the 156/52/52 chronology and training-only scaling boundary, expose unresolved protocol choices, and label Table 1 and Table 3 as paper references. Any numerical conclusion should wait until execution and the relevant verification stages are actually enabled.</p><h2>Local Adjusted-Price Inputs and Diagnostic Analysis</h2><p>Before training a forecasting model, ask a simpler question: <strong>do we have the right observations in the right order?</strong> For this reproduction, each asset is an independent weekly sequence of adjusted closing prices. The diagnostic stage checks that sequence, summarizes it, and records useful time-series behavior without changing the data later supplied to the GRU.</p><p>The paper uses 263 observations for each of <code>AAPL</code> and <code>MSFT</code>. Its retained price field is the adjusted closing price, not the ordinary close. The implementation keeps dates alongside prices so that later windows and forecasts can remain aligned with their target dates.</p><h3>What is fixed and what is chosen</h3><p>The paper establishes several input facts: the two assets are modeled separately, each series contains 263 observations, and no missing values were reported. It also describes an upward trend, non-stationarity, strong autocorrelation, and no visible outliers.</p><p>Other details are not fully specified. The exact Yahoo Finance retrieval query, weekly aggregation rule, calendar handling, and source-data version are unavailable. Similarly, the paper does not state the decomposition method or period, the ADF regression and lag settings, or the autocorrelation lag range and interval construction. The generated code makes these choices explicit rather than presenting them as recovered paper settings.</p><p>Most importantly, diagnostics are descriptive. They do not remove outliers, replace values, difference the series for forecasting, or add decomposition components as GRU features.</p><h3>Load local appendix-style records</h3><p>The reproduction uses local CSV files instead of silently downloading from Yahoo Finance. This is an implementation decision that makes the input inspectable and avoids inventing an unspecified retrieval and resampling procedure. The expected layout is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;a2e40de3-0f57-4152-9f07-da8bc9d3a86a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">data/raw/
&#9500;&#9472;&#9472; AAPL.csv
&#9492;&#9472;&#9472; MSFT.csv</code></pre></div><p>The loader accepts date and adjusted-price column aliases, but the normalization script writes the stable <code>date</code> and <code>adjusted_close</code> schema. A focused excerpt from <code>src/stock_forecasting/io/appendix_loader.py</code> shows the public loading contract:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;638ee727-d8a1-401e-8845-44069d4d907d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from pathlib import Path

from stock_forecasting.data_types import PriceSeries
from stock_forecasting.exceptions import DataValidationError


def load_appendix_csv(path: Path, asset: str) -&gt; PriceSeries:
    """Load one local appendix-style CSV as a validated price series.

    The local appendix file is deliberately used instead of downloading from Yahoo
    Finance: the paper identifies Yahoo Finance as provenance but does not specify
    the exact retrieval, calendar, or weekly aggregation procedure.  The loader
    therefore accepts already prepared records and preserves the adjusted-close
    field as the sole forecasting variable.
    """</code></pre></div><p><code>load_appendix_csv</code> reads one file and returns a <code>PriceSeries</code> containing <code>dates</code>, <code>prices</code>, and the canonical asset name. It checks optional asset labels so a file intended for AAPL cannot quietly contain MSFT rows. It sorts records chronologically and converts prices to floating-point values. It does not remove rows identified visually as unusual.</p><p>For the standard directory layout, use <code>load_asset_records</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e3e1042d-f7f0-4bd9-ae22-a29ae3f7de6a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from pathlib import Path

from stock_forecasting.io.appendix_loader import load_asset_records

series = load_asset_records(Path("data/raw"), "AAPL")
print(series.asset)
print(series.dates.shape)
print(series.prices.shape)</code></pre></div><p>The generated <code>UnavailableYahooSource</code> is deliberately not a network client. Its responsibility is to make an unsupported source choice fail clearly rather than download a different dataset:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;0ec1fe2b-df3b-4e77-ae3d-e007de540f38&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class UnavailableYahooSource:
    """Explicit offline adapter for the paper's Yahoo Finance provenance."""

    def get_series(self, asset: str) -&gt; PriceSeries:
        """Reject external retrieval and direct the caller to local data."""
        if not isinstance(asset, str) or asset not in _ALLOWED_ASSETS:
            raise DataValidationError(
                "asset must be either 'AAPL' or 'MSFT'; "
                f"received {asset!r}"
            )
        raise DataValidationError(
            f"No offline Yahoo Finance adapter is configured for {asset}. "
            "Provide the paper's local appendix-style records and load them "
            "with load_appendix_csv or load_asset_records."
        )</code></pre></div><p>This failure is intentional. The paper names Yahoo Finance as provenance, but the supplied specification does not identify the exact query or weekly aggregation operation.</p><h3>Validate the 263-observation contract</h3><p>Loading a CSV is not the same as validating an experiment input. <code>validate_price_series</code> checks the paper's structural assumptions and returns a quality summary. Its default expected count is 263, although callers can use another count for isolated fixtures.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;eff13f5e-52b7-4010-9da7-40d56a098726&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from stock_forecasting.data_validation import validate_price_series

quality = validate_price_series(series)
print(quality)</code></pre></div><p>The validation boundary requires all of the following:</p><ul><li><p><code>dates</code> and <code>prices</code> each contain 263 entries;</p></li><li><p>dates are present, unique, and strictly increasing;</p></li><li><p>prices are one-dimensional, finite, and strictly positive;</p></li><li><p>the retained field is adjusted closing price;</p></li><li><p>no observations are removed as outliers.</p></li></ul><p>The lower-level functions make the invariants visible. <code>check_chronological_dates</code> rejects missing, duplicate, or non-increasing dates. <code>check_adjusted_prices</code> rejects nonnumeric, non-finite, multidimensional, or non-positive values. <code>summarize_quality</code> reports counts, extrema, endpoints, chronology, and the outlier policy without modifying the series.</p><p>A concise end-to-end loading example is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;102b09d8-b07f-43ac-8552-7648e62cbf74&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from pathlib import Path

from stock_forecasting.data_validation import validate_price_series
from stock_forecasting.io.appendix_loader import load_asset_records

DATA_DIR = Path("data/raw")
aapl = load_asset_records(DATA_DIR, "AAPL")
quality = validate_price_series(aapl)

print(aapl.asset)       # AAPL
print(aapl.dates.shape) # Expected shape: (263,)
print(aapl.prices.shape) # Expected shape: (263,)
print(quality)</code></pre></div><p>The corresponding MSFT call should be made independently. Do not concatenate the two <code>PriceSeries</code> objects: asset isolation is a modeling requirement, not merely an organizational preference.</p><h3>Prepare files without downloading data</h3><p>When source records are supplied in an appendix-style directory, <code>scripts/prepare_local_data.py</code> validates and normalizes them into the expected raw-data layout. It requires local files and writes only <code>date</code> and <code>adjusted_close</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;f07df259-5b3d-4173-87a5-2306e596ab76&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">python scripts/prepare_local_data.py \
    --input-dir appendix \
    --output-dir data/raw</code></pre></div><p><code>prepare_asset_file</code> loads one asset, invokes <code>validate_price_series</code> with <code>expected_count=263</code>, and writes a normalized CSV. It does not fill missing values, infer missing weeks, or correct outliers. Consequently, a short or malformed input fails before it can enter the modeling pipeline.</p><h3>Descriptive statistics and reference comparisons</h3><p>The paper's descriptive table reports count, mean, standard deviation, minimum, quartiles, median, and maximum for each 263-observation series. The generated <code>compute_descriptive_statistics</code> function returns those fields as a pandas Series:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d2af515f-4a3e-4278-a1b2-f241e528aa4d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from stock_forecasting.eda.descriptive import (
    compare_with_table_2,
    compute_descriptive_statistics,
)

stats = compute_descriptive_statistics(aapl)
print(stats)

# Generated statistic minus the paper's Table 2 reference value.
differences = compare_with_table_2(stats, aapl.asset)
print(differences)</code></pre></div><p>The stored Table 2 values are paper references, not assertions that an unverified local file matches the paper's source exactly. Differences may reflect source-data versions, weekly aggregation, or numeric details. The helper reports signed differences and does not label them as a pass or fail.</p><h3>Plot the distribution and chronology</h3><p>The plotting functions in <code>src/stock_forecasting/eda/plots.py</code> return Matplotlib axes and read the source arrays without mutating them. The histogram describes the price distribution, the boxplot supports visual outlier screening, and the chronological plot reveals trend and fluctuations.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;fd3d2390-93c8-4f78-b814-3473015a815c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import matplotlib.pyplot as plt

from stock_forecasting.eda.plots import (
    plot_boxplot,
    plot_price_distribution,
    plot_price_series,
    save_figure,
)

figure, axes = plt.subplots(3, 1, figsize=(10, 12), constrained_layout=True)
plot_price_distribution(aapl, ax=axes[0])
plot_boxplot(aapl, ax=axes[1])
plot_price_series(aapl, ax=axes[2])
save_figure(figure, Path("artifacts/diagnostics/AAPL_overview.png"))
plt.close(figure)</code></pre></div><p>A boxplot point is evidence for inspection only. The paper says that no visible outliers were found, but it gives no formal threshold or correction rule. Therefore, this implementation retains every validated observation even if a different local file produces a visually unusual point.</p><h3>Decomposition: diagnostic components only</h3><p>The paper asks for decomposition into trend, seasonal, and residual components, but it does not specify how. The generated wrapper requires a <code>DecompositionConfig</code>, making the method and period visible:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5f703fbe-eb12-4bed-8a02-cf52fef0c21a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from stock_forecasting.eda.decomposition import (
    DecompositionConfig,
    decompose_price_series,
)

# Explicit implementation decision for this diagnostic run.
config = DecompositionConfig(method="additive", period=52)
components = decompose_price_series(aapl, config)

for name, component in components.items():
    print(name, component.shape, component.index.equals(aapl.dates))</code></pre></div><p>Here, additive decomposition and a 52-observation period are implementation decisions for a weekly series. They are not parameters recovered from the paper. The returned dictionary contains <code>trend</code>, <code>seasonal</code>, and <code>residual</code> pandas Series, each aligned to the original 263 dates. Boundary component values may be missing because the decomposition needs neighboring observations; they are retained rather than imputed.</p><p>The components are not passed into the GRU. The forecasting input remains the single adjusted-price feature. This separation prevents an exploratory choice from silently changing the proposed model.</p><h3>ADF stationarity diagnostic</h3><p>The paper reports raw-price non-stationarity and gives approximate ADF p-values of <code>0.4594</code> for AAPL and <code>0.9229</code> for MSFT. Those values depend on test settings, and the supplied paper text does not specify the deterministic regression term or lag-selection policy. The generated <code>run_adf_test</code> therefore requires both:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ab195c79-a9bc-412d-a7a4-2dfec0ddb346&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from stock_forecasting.eda.stationarity import interpret_adf, run_adf_test

# Explicit implementation choices, not recovered paper settings.
aapl_adf = run_adf_test(aapl, regression="c", autolag="AIC")
print(aapl_adf)
print(interpret_adf(aapl_adf, alpha=0.05))</code></pre></div><p><code>ADFReport</code> contains the test statistic, p-value, and critical values. <code>interpret_adf</code> uses the paper's stated significance threshold of <code>0.05</code>: a p-value below that threshold is reported as rejection of the unit-root null, while a p-value at or above it is reported as failure to reject.</p><p>Failure to reject is a diagnostic conclusion under one test specification, not proof about every future path. Likewise, the presence of non-stationarity does not mean that a GRU mathematically removes non-stationarity or remains valid under distribution shift. The ADF call does not difference or otherwise alter the price series used later.</p><p>Repeat the explicitly configured diagnostic for MSFT:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8d16144f-7c17-4fff-bc9d-8b231799fd5a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">msft = load_asset_records(Path("data/raw"), "MSFT")
msft_adf = run_adf_test(msft, regression="c", autolag="AIC")
print(interpret_adf(msft_adf, alpha=0.05))</code></pre></div><h3>Autocorrelation diagnostic</h3><p>Autocorrelation measures how values in the ordered series relate to earlier values at selected lags. The paper describes strong autocorrelation but does not specify the lag range or confidence-interval construction. The generated implementation therefore accepts an inclusive <code>max_lag</code> and returns only lag-aligned values:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6701601c-7058-44fa-a29f-862b7674b75f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from stock_forecasting.eda.autocorrelation import compute_autocorrelation

aapl_acf = compute_autocorrelation(aapl, max_lag=40)
print(aapl_acf.lags.shape)   # (41,)
print(aapl_acf.values.shape) # (41,)
print(aapl_acf.values[0])     # 1.0 by contract</code></pre></div><p>Using lags 0 through 40 is an implementation decision. <code>AutocorrelationReport</code> requires contiguous lags beginning at zero and equal-length <code>lags</code> and <code>values</code> arrays. The report does not include confidence intervals because no interval method was supplied.</p><p>The diagnostic can be plotted without adding those values to the forecasting features:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;258aaa32-f436-4830-bb86-a75b67a1a946&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import matplotlib.pyplot as plt

figure, axis = plt.subplots(figsize=(10, 4))
axis.stem(aapl_acf.lags, aapl_acf.values)
axis.set_title("AAPL adjusted-price autocorrelation")
axis.set_xlabel("Lag")
axis.set_ylabel("Autocorrelation")
save_figure(figure, Path("artifacts/diagnostics/AAPL_acf.png"))
plt.close(figure)</code></pre></div><h3>Run both assets from the command line</h3><p>The diagnostic script applies the same explicit parameter choices independently to AAPL and MSFT and writes JSON reports, decomposition CSV files, and figures:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;7750d675-ac1d-4dc0-85b2-5c0f52f30957&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">python scripts/run_diagnostics.py \
    --data-dir data/raw \
    --output-dir artifacts/diagnostics</code></pre></div><p>Its defaults record an additive decomposition with period 52, an ADF constant term with AIC autolag, significance threshold <code>0.05</code>, and autocorrelation through lag 40. Because several of these values are not paper facts, the script stores them as implementation metadata in each asset's diagnostic report.</p><h3>What this stage establishes</h3><p>After this stage, the intended state is a pair of independently validated <code>PriceSeries</code> objects and non-mutating diagnostic artifacts. The qualitative expectations from the paper are an upward trend, non-stationary raw prices, substantial autocorrelation, and no visible outliers. Those are interpretive reference points, not execution results claimed here.</p><p>No code execution, tests, static verification, semantic code verification, tutorial verification, or final quality review was performed under the supplied run policy. The next stage can therefore explain the planned transformation without implying that these diagnostics have already produced verified outputs.</p><p>The following section turns each unchanged 263-value sequence into supervised examples: three preceding adjusted prices become one input sequence, and the following price becomes its target. That transformation remains separate from the diagnostic plots and tests described here.</p><h2>From a Price Sequence to Supervised Windows</h2><p>How can a plain chronological price list become input that a recurrent neural network can learn from? The paper uses a three-period sliding window: take three consecutive adjusted prices, use them as one short sequence, and assign the following adjusted price as the target. Then move forward by one observation and repeat.</p><p>This transformation is performed separately for each asset and before any train, test, or validation split. AAPL and MSFT therefore produce independent <code>WindowedDataset</code> objects; their prices are never combined into one multivariate input.</p><h3>The input and target contract</h3><p>The generated code represents one raw asset as a <code>PriceSeries</code>. Its <code>dates</code> and <code>prices</code> fields are aligned one-to-one, with 263 chronological observations and positive adjusted prices. The <code>create_windowed_supervised_dataset</code> function then returns a <code>WindowedDataset</code> with three coordinated fields:</p><ul><li><p><code>inputs</code>: recurrent input windows with shape <code>(260, 3, 1)</code>;</p></li><li><p><code>targets</code>: one scalar next-price target per window, with shape <code>(260, 1)</code>; and</p></li><li><p><code>target_dates</code>: the dates associated with those target prices, with length 260.</p></li></ul><p>The three dimensions of <code>inputs</code> mean <code>(samples, time steps, features)</code>. There are 260 samples, each sequence contains 3 time steps, and each time step contains 1 feature: the adjusted price. The singleton feature dimension is important because Keras recurrent layers expect a feature axis even for a univariate series.</p><p>The paper supplies the three-period window and 263 raw observations. The count of supervised samples follows from the implementation: the first three observations form an input but do not yet have a preceding window of their own, so each subsequent raw observation can serve as one target. Consequently, 263 raw observations produce 260 samples.</p><h3>Building the windows</h3><p>The public function accepts a validated <code>PriceSeries</code> and defaults to <code>window_length=3</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;11837903-17ee-4e35-ab53-c35abc317614&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from stock_forecasting.supervised.windowing import create_windowed_supervised_dataset

windowed = create_windowed_supervised_dataset(series, window_length=3)

print(windowed.inputs.shape)       # (260, 3, 1)
print(windowed.targets.shape)      # (260, 1)
print(windowed.target_dates.shape) # (260,)</code></pre></div><p>Inside <code>create_windowed_supervised_dataset</code>, the implementation constructs each input from a consecutive slice and adds the feature axis. The following excerpt is the central transformation copied from <code>src/stock_forecasting/supervised/windowing.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c066ca03-4dfd-4e9b-af51-69c158db9c14&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    # Three preceding adjusted prices predict the next adjusted price.
    inputs = np.stack(
        [prices[index : index + int(window_length)] for index in range(sample_count)],
        axis=0,
    )[..., np.newaxis]
    targets = prices[int(window_length) :].reshape(-1, 1)
    target_dates = dates[int(window_length) :]</code></pre></div><p>For sample index 0, the input slice contains raw prices at positions 0, 1, and 2. The target slice begins at position 3, so its first value is the price immediately after that input. The <code>[..., np.newaxis]</code> operation changes a two-dimensional collection of windows into the rank-three recurrent shape <code>(samples, 3, 1)</code>. The target reshape standardizes one scalar per sample as <code>(samples, 1)</code>, matching the intended scalar model output.</p><h3>A worked example with inspectable values</h3><p>A monotonic synthetic series makes the boundary behavior easy to see without depending on external data. This is an illustration only, not a replacement for the paper's AAPL or MSFT records.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;cd4c8080-236c-49eb-99cb-a431bbb19f42&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import numpy as np
import pandas as pd

from stock_forecasting.data_types import PriceSeries
from stock_forecasting.supervised.windowing import (
    create_windowed_supervised_dataset,
)

_RAW_COUNT = 263

raw_dates = pd.date_range("2020-01-03", periods=_RAW_COUNT, freq="W-FRI")
raw_prices = np.arange(1, _RAW_COUNT + 1, dtype=np.float64)

series = PriceSeries(
    dates=raw_dates,
    prices=raw_prices,
    asset="AAPL",
)

windowed = create_windowed_supervised_dataset(series)</code></pre></div><p>The first sample contains prices 1, 2, and 3. Its target is price 4, and its target date is the fourth date in the raw series:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f678bc5e-5ee1-46da-a7d5-3246916d6f55&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">np.testing.assert_allclose(
    windowed.inputs[0, :, 0],
    np.array([1.0, 2.0, 3.0]),
)
np.testing.assert_allclose(windowed.targets[0], np.array([4.0]))
assert windowed.target_dates[0] == raw_dates[3]</code></pre></div><p>The final sample contains the last three prices before the final raw observation. Its target is 263, and its date is the final raw date:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;81b48c26-c4e3-4cc9-8845-971c62cc8d71&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">np.testing.assert_allclose(
    windowed.inputs[-1, :, 0],
    np.array([260.0, 261.0, 262.0]),
)
np.testing.assert_allclose(windowed.targets[-1], np.array([263.0]))
assert windowed.target_dates[-1] == raw_dates[-1]</code></pre></div><p>This boundary example makes the no-future-leakage invariant visible: the target never appears inside its own input window. It also explains why target dates, rather than the final input dates, are retained. A forecast plotted against <code>target_dates</code> is compared with the observation it was intended to predict.</p><h3>Structural validation and failure cases</h3><p><code>WindowedDataset</code> validates the recurrent and target shapes, matching sample counts, finite values, positive prices, and chronological target dates. The companion <code>validate_windowed_dataset</code> function checks the expected relationship between source length, window length, and sample count. It rejects a non-integer or non-three window length in this reproduction, because the generated container deliberately implements the paper-specific <code>(n, 3, 1)</code> contract.</p><p>The function also rejects common data problems before a model sees them. It raises a domain-specific <code>DataValidationError</code> when the input is not a <code>PriceSeries</code>, when the series does not contain exactly 263 prices and dates, when values are non-finite or non-positive, or when the constructed result violates the expected shapes. These checks protect both temporal alignment and the downstream Keras input contract.</p><p>The supplied test file expresses the important boundary checks with deterministic prices:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8008de06-8d8a-4f8b-8473-283e4939953a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def test_recurrent_shape_is_three_by_one() -&gt; None:
    """Verify the recurrent input and scalar-target tensor contracts."""
    dataset = create_windowed_supervised_dataset(_synthetic_price_series())

    assert dataset.inputs.shape == (_EXPECTED_SAMPLE_COUNT, _WINDOW_LENGTH, 1)
    assert dataset.targets.shape == (_EXPECTED_SAMPLE_COUNT, 1)
    assert dataset.target_dates.shape == (_EXPECTED_SAMPLE_COUNT,)
    assert dataset.inputs.dtype == np.float64
    assert dataset.targets.dtype == np.float64</code></pre></div><p>The other generated tests check that the sample count is exactly 260 and that the first and last windows use the correct preceding prices, targets, and dates. Those tests are planned verification artifacts; under the current run policy, they were not executed and no passing result is claimed.</p><h3>Why windowing comes before splitting</h3><p>The paper-consistent dimensional interpretation is to construct all 260 one-step samples first and then partition those samples chronologically into 156 training, 52 testing, and 52 validation samples. Splitting the raw 263 observations first would describe a different boundary procedure and could change which windows are available at each partition boundary.</p><p>After windowing, each sample already has an input, a target, and a target date. The later splitting stage can therefore slice all three arrays together while preserving order. This keeps the temporal direction clear: no future target is included in an earlier sample, and no random shuffle is needed to create the supervised dataset.</p><h3>Advanced detail: target shape conventions</h3><p>Some forecasting libraries store scalar targets as a one-dimensional array with shape <code>(n,)</code>. That representation can be valid, but this implementation intentionally standardizes targets as <code>(n, 1)</code>. The choice makes the target contract match a model whose final output is one value per sample and helps prevent accidental broadcasting during training and evaluation. It is an implementation convention, not an additional feature or a change to the paper's one-step forecasting task.</p><p>No display equation is included here: the supplied paper context contains no canonical equation records. The index-level procedure, array shapes, and code slices above provide the supported mathematical meaning without reconstructing an equation from memory.</p><h2>Chronological Splitting and Training-Only Scaling</h2><p>How do we prevent a forecasting experiment from learning about the future before evaluation? The answer is procedural: create all supervised windows first, place them in chronological order, divide those windows into contiguous partitions, and learn scaling parameters from the training partition only. Later testing and validation values may be transformed with that state, but they must not determine it.</p><p>This section uses the paper-consistent dimensional interpretation. Each asset begins with 263 weekly adjusted closing prices. A three-period one-step window produces 260 supervised samples. Those samples are then divided into 156 training samples, 52 testing samples, and 52 validation samples. AAPL and MSFT go through this process independently.</p><h3>What the paper specifies&#8212;and what it leaves open</h3><p>The paper specifies chronological partitioning and training-only MinMax normalization. This protects the temporal direction of the experiment: earlier samples are used before later samples. It does not specify whether the scaler was fitted to the raw one-dimensional training-price series, to flattened training windows, or to training inputs and targets together.</p><p>That missing detail matters because a MinMax scaler stores data-dependent bounds. Different fitting populations can produce different normalized values, training losses, and eventually price-unit metrics. The implementation therefore exposes two named <code>ScalerConvention</code> choices. This is an implementation decision, not a recovered detail from the paper.</p><p>The paper also keeps testing and validation conceptually distinct, although its reported metrics are described as validation evaluation. Because the exact use of the testing partition is unclear, the reproduction retains both partitions and requires the caller to choose explicitly later.</p><h3>Partition the windowed samples</h3><p><code>split_windowed_dataset</code> receives a <code>WindowedDataset</code> and a <code>SplitConfig</code>. The dataset contains recurrent inputs, scalar targets, and target dates. The configuration contains the three requested partition sizes. The function slices all three fields at the same boundaries and returns a <code>DatasetSplits</code> object with fields such as <code>train_inputs</code>, <code>test_inputs</code>, and <code>validation_inputs</code>.</p><p>The important algorithm is short and deliberately does not shuffle:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;83ff934d-ad48-48a1-98e2-f3e62e9aa4b7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">train_end = config.train_size
test_end = train_end + config.test_size

train_inputs = dataset.inputs[:train_end]
train_targets = dataset.targets[:train_end]
train_dates = dataset.target_dates[:train_end]

test_inputs = dataset.inputs[train_end:test_end]
test_targets = dataset.targets[train_end:test_end]
test_dates = dataset.target_dates[train_end:test_end]

validation_inputs = dataset.inputs[test_end:]
validation_targets = dataset.targets[test_end:]
validation_dates = dataset.target_dates[test_end:]</code></pre></div><p>With the default configuration, <code>train_end</code> is 156 and <code>test_end</code> is 208. Therefore the training indices are 0 through 155, testing indices are 156 through 207, and validation indices are 208 through 259. Every windowed sample is consumed once.</p><p>The configuration is supplied by <code>get_default_split_config</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;53d73da1-9ff3-4a69-bae6-2ec533173c98&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def get_default_split_config() -&gt; SplitConfig:
    """Return the paper-consistent chronological 156/52/52 split.

    The split is a derived dimensional correction: 263 raw observations minus
    a three-period window produce 260 supervised samples.
    """
    return SplitConfig(train_size=156, test_size=52, validation_size=52)</code></pre></div><p>The <code>SplitConfig</code> constructor rejects sizes that do not sum to 260. That check encodes the derived dimensional contract rather than treating 263 raw observations as 263 supervised examples.</p><h3>Worked example: inspect the three boundaries</h3><p>The following example uses the monotonic synthetic <code>WindowedDataset</code> introduced in the windowing discussion. It is useful because the sample indices are easy to inspect; it is not AAPL or MSFT data.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;0563b864-2fc8-4ddc-af25-73c1440a2555&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">split_config = get_default_split_config()
splits = split_windowed_dataset(windowed, split_config)

assert splits.train_inputs.shape == (156, 3, 1)
assert splits.train_targets.shape == (156, 1)
assert len(splits.train_dates) == 156

assert splits.test_inputs.shape == (52, 3, 1)
assert splits.test_targets.shape == (52, 1)
assert len(splits.test_dates) == 52

assert splits.validation_inputs.shape == (52, 3, 1)
assert splits.validation_targets.shape == (52, 1)
assert len(splits.validation_dates) == 52</code></pre></div><p>Notice the shape invariants. Each input remains rank three: sample count, sequence length 3, and one feature. Each target remains a rank-two scalar batch with shape <code>(n, 1)</code>. Dates have one entry per target, not one entry per input timestep.</p><p><code>validate_split_boundaries</code> checks these shapes and also checks finite positive values, chronological order within each partition, and non-overlap between adjacent partitions. This function is a boundary guard: it does not repair an incorrectly ordered dataset or infer a missing split.</p><p>Random splitting would undermine the intended experiment. It could place later price regimes in the training partition and earlier regimes in validation, making the held-out data no longer a clean later period. The generated splitter therefore performs contiguous slicing and never calls a random shuffling operation.</p><h3>Fit the scaler after splitting</h3><p>Scaling occurs after the chronological split so that the fitting function receives training arrays only. The generated API is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;80d3afc3-6d34-44af-a2ca-2a487b0f1a59&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">scaler = fit_training_scaler(
    train_inputs=splits.train_inputs,
    train_targets=splits.train_targets,
    convention=ScalerConvention.RAW_TRAINING_PRICES,
)</code></pre></div><p><code>fit_training_scaler</code> returns a one-feature scikit-learn <code>MinMaxScaler</code>. It validates that inputs have shape <code>(n, 3, 1)</code> and that targets have shape <code>(n,)</code> or <code>(n, 1)</code>. In this reproduction, the selected convention is recorded through the enum value, so experiment metadata can state which policy was requested.</p><p>There is an important implementation limitation to understand precisely. The current helper does not receive the original raw training series separately. Consequently, its <code>RAW_TRAINING_PRICES</code> branch fits from the price values available in the training inputs and targets. The <code>FLATTENED_TRAINING_WINDOWS</code> branch currently uses the same supplied training values after flattening. The two names preserve the unresolved protocol choice in the public API, but they do not yet implement two different fitting populations. This should not be mistaken for evidence that the paper used either convention.</p><p>The fitting function deliberately receives no testing or validation arrays. Applying a fitted scaler later is allowed; refitting it later is not.</p><h3>Transform while preserving shapes</h3><p>Use the same scaler for every partition belonging to the same asset:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;20214da7-57b0-4c5d-8eca-e24b003f8995&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">scaled_train_inputs = transform_windows(scaler, splits.train_inputs)
scaled_test_inputs = transform_windows(scaler, splits.test_inputs)
scaled_validation_inputs = transform_windows(scaler, splits.validation_inputs)

scaled_train_targets = transform_targets(scaler, splits.train_targets)
scaled_test_targets = transform_targets(scaler, splits.test_targets)
scaled_validation_targets = transform_targets(scaler, splits.validation_targets)

assert scaled_train_inputs.shape == (156, 3, 1)
assert scaled_test_inputs.shape == (52, 3, 1)
assert scaled_validation_inputs.shape == (52, 3, 1)
assert scaled_train_targets.shape == (156, 1)
assert scaled_test_targets.shape == (52, 1)
assert scaled_validation_targets.shape == (52, 1)</code></pre></div><p><code>transform_windows</code> flattens the final one-feature view temporarily for the scaler, then reshapes the result back to <code>(n, 3, 1)</code>. <code>transform_targets</code> similarly preserves the target's scalar-batch convention. Neither function refits the scaler, and neither changes sample order or date arrays.</p><p>AAPL and MSFT need separate scaler instances. Their prices have different levels and distributions, and the paper models the assets in isolation. Reusing one scaler across both series would introduce cross-asset information and would not match the stated univariate design.</p><h3>Inverse transformation before evaluation</h3><p>The GRU is trained on normalized values, so its predictions are initially in normalized target units. Before calculating price-unit metrics, convert both predictions and actual targets back with the same training-fitted scaler:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;04102c71-6778-4f5a-b82f-06ce80333895&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">restored_training_targets = inverse_targets(scaler, scaled_train_targets)</code></pre></div><p>The inverse operation preserves the input rank. For a held-out validation batch, an array with shape <code>(52, 1)</code> returns to the same shape and represents adjusted-price values again. The evaluation stage then flattens or otherwise validates the arrays according to its metric contract, aligns them with the 52 validation target dates, and computes metrics in original units.</p><p>This ordering is essential. Evaluating normalized values would produce errors whose units and scale depend on preprocessing rather than on the stock price itself. The intended sequence is therefore: fit on training values, transform all model inputs and targets, predict in normalized space, inverse-transform predictions and actuals, then calculate price-unit metrics.</p><h3>Planned tests and verification status</h3><p>The generated test file contains focused checks for the split and scaler invariants, including the expected 156/52/52 shapes, date contiguity, unchanged scaler state after transforming validation values, shape preservation, and inverse transformation. These tests are planned artifacts, not evidence of completed verification. Under the supplied run policy, code execution, testing, local static verification, semantic code verification, and tutorial verification were disabled. No test result is claimed here.</p><h3>Reproduction checklist</h3><p>For this preprocessing stage, a reviewable run should record the following:</p><ol><li><p>The input asset and its 263 chronological adjusted prices.</p></li><li><p>The three-period windowing step that produced 260 samples.</p></li><li><p>The <code>SplitConfig</code> values <code>156</code>, <code>52</code>, and <code>52</code>.</p></li><li><p>The fact that testing remains separate from validation.</p></li><li><p>The selected <code>ScalerConvention</code> value.</p></li><li><p>The scaler's training-only fitting population.</p></li><li><p>The recurrent and target shapes before and after transformation.</p></li><li><p>The scaler instance retained for inverse transformation.</p></li><li><p>The explicit held-out subset selected for later evaluation.</p></li></ol><p>The result is leakage-aware preprocessing, but not a guarantee of accurate forecasting. Scaling choices can affect the numerical outcome even when the GRU architecture and training settings remain unchanged. Exact reproduction still depends on unresolved source-data, preprocessing, and evaluation details.</p><h2>Asset-Specific GRU Architectures</h2><p>How does the implementation turn a three-price sequence into one next-price forecast? A gated recurrent unit, or GRU, reads the input steps in order and maintains an internal representation of the sequence. Here, each sample contains three prior adjusted prices for one asset, so the model receives a tensor with shape <code>(batch, 3, 1)</code> and returns one normalized scalar for every sample, with shape <code>(batch, 1)</code>.</p><p>The paper uses separate univariate models. The AAPL model never receives MSFT values, and the MSFT model never receives AAPL values. This isolation is important because the paper evaluates asset-specific architectures rather than a shared multivariate network.</p><h3>Table 1 configurations versus implementation choices</h3><p>The reported final configurations are paper facts. AAPL uses three GRU layers with widths <code>(240, 16, 16)</code>, while MSFT uses two layers with widths <code>(224, 48)</code>. Both configurations use dropout <code>0.0</code>, learning rate <code>0.001</code>, MSE loss, batch size <code>8</code>, a maximum of <code>100</code> epochs, and patience <code>7</code>.</p><p>The generated <code>AssetModelConfig</code> stores these values together with the asset identifier:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;35bf9aee-4998-4158-8bef-cf7c6703e3fc&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class AssetModelConfig:
    """Immutable configuration for one isolated asset GRU model.

    ``units`` contains the number of units in each stacked GRU layer.  The
    paper-selected values are AAPL ``(240, 16, 16)`` and MSFT ``(224, 48)``.
    Activation functions, initializers, return-sequence behavior, and the
    output-layer details are not specified by the supplied paper text and are
    intentionally not represented as paper facts here.
    """

    asset: str
    units: tuple[int, ...]
    dropout: float
    learning_rate: float
    loss_name: str
    batch_size: int
    max_epochs: int
    patience: int</code></pre></div><p>The validation in <code>AssetModelConfig</code> rejects unsupported assets, empty or non-positive unit tuples, invalid dropout rates, non-positive learning rates, unsupported losses, and invalid training limits. These checks protect the configuration boundary; they do not add new model behavior described by the paper.</p><p>The factory returns the reported settings directly. It does not pretend to recover the original Hyperband search space or random seed:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;63bacd8a-0a10-4085-8de0-6fc3a9a5dd57&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def get_final_model_config(asset: str) -&gt; AssetModelConfig:
    """Return the supplied Table 1 final configuration for ``asset``.

    The layer widths, zero dropout, learning rate, MSE loss, batch size,
    epoch limit, and patience are paper facts.  They are returned directly
    rather than inferred from an unavailable Hyperband search space.
    """
    if asset == "AAPL":
        units = (240, 16, 16)
    elif asset == "MSFT":
        units = (224, 48)
    else:
        raise ValueError(f"unsupported asset {asset!r}; expected AAPL or MSFT")

    return AssetModelConfig(
        asset=asset,
        units=units,
        dropout=0.0,
        learning_rate=0.001,
        loss_name="mse",
        batch_size=8,
        max_epochs=100,
        patience=7,
    )</code></pre></div><h3>Connecting recurrent layers</h3><p>A stacked GRU passes a sequence from each intermediate recurrent layer to the next one. In the generated builder, every GRU except the last uses <code>return_sequences=True</code>. The terminal GRU returns its final representation, which is passed to <code>Dense(1)</code> to produce one scalar output.</p><p>These settings are implementation decisions, not additional paper facts. The supplied paper text does not specify return-sequence behavior, activation functions, recurrent activations, initializers, or the terminal output-layer type. The generated implementation makes those choices explicit so that the model has a definite tensor contract:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4ab0488e-8e4e-4aaf-87c3-f7de4919a461&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">for index, width in enumerate(units):
    is_terminal = index == len(units) - 1
    representation = keras.layers.GRU(
        width,
        activation="tanh",
        recurrent_activation="sigmoid",
        dropout=float(dropout),
        return_sequences=not is_terminal,
        name=f"gru_{index + 1}",
    )(representation)

outputs = keras.layers.Dense(1, name="next_adjusted_close")(representation)
return keras.Model(inputs=inputs, outputs=outputs, name="stacked_gru_forecaster")</code></pre></div><p>The helper <code>validate_gru_architecture</code> checks that the unit tuple is non-empty and contains positive integers, that dropout lies in the interval <code>[0, 1)</code>, and that the input shape is exactly <code>(3, 1)</code>. Its failure cases are deliberate: a different sequence length or feature count would no longer represent the paper's three prior prices and one univariate feature.</p><p>No GRU gate equations are displayed here. The supplied context contains no canonical equation records, so reconstructing equations from general GRU knowledge would incorrectly present an unavailable formula as part of the paper specification.</p><h3>Building an isolated asset model</h3><p><code>build_asset_gru_model</code> accepts one validated <code>AssetModelConfig</code> and an input shape. It delegates layer construction to <code>build_gru_stack</code>, then checks the resulting model's input and output shapes. TensorFlow is imported lazily by the lower-level builder, allowing data and preprocessing components to remain conceptually separate from the optional neural-network runtime.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8f9b4715-f030-4e66-80da-a49c8de39fc3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from stock_forecasting.config import get_final_model_config
from stock_forecasting.models.gru_model import (
    build_asset_gru_model,
    describe_model_contract,
)

# Table 1 paper configuration: AAPL uses three GRU layers.
aapl_config = get_final_model_config("AAPL")
aapl_model = build_asset_gru_model(aapl_config, input_shape=(3, 1))
print(describe_model_contract(aapl_model))</code></pre></div><p>The corresponding MSFT model is constructed independently:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b9b9cf0e-18a8-4e7a-a828-9b7e33a9921a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from stock_forecasting.config import get_final_model_config
from stock_forecasting.models.gru_model import build_asset_gru_model

# Table 1 paper configuration: MSFT uses two GRU layers.
msft_config = get_final_model_config("MSFT")
msft_model = build_asset_gru_model(msft_config, input_shape=(3, 1))</code></pre></div><p><code>describe_model_contract</code> reports the model name, input and output shapes, recurrent layer widths, dropout values, and <code>return_sequences</code> settings. It rejects models without inspectable shapes, without GRU layers, with the wrong scalar output, or with an invalid sequence-return arrangement. The model-building path also raises a domain-specific <code>ConfigurationError</code> if TensorFlow is unavailable or if the constructed model violates the expected shape contract.</p><h3>Worked example: compare the two architectures</h3><p>The following inspection workflow expresses the intended comparison without training either model:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;103cbbd2-2791-4a12-a9f0-77b6adfb5cf3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from stock_forecasting.config import get_final_model_config
from stock_forecasting.models.gru_model import (
    build_asset_gru_model,
    describe_model_contract,
)

for asset in ("AAPL", "MSFT"):
    config = get_final_model_config(asset)
    model = build_asset_gru_model(config, input_shape=(3, 1))
    contract = describe_model_contract(model)

    print(asset, config.units)
    print(contract["input_shape"], contract["output_shape"])</code></pre></div><p>The expected structural interpretation is that AAPL has three recurrent layers with widths <code>240</code>, <code>16</code>, and <code>16</code>, while MSFT has two with widths <code>224</code> and <code>48</code>. Both accept <code>(None, 3, 1)</code> at the model boundary and emit <code>(None, 1)</code>, where <code>None</code> represents a variable batch size. The configurations also carry training settings such as learning rate, loss name, batch size, maximum epochs, and patience, although those settings are applied later during compilation and training.</p><p>The generated contract tests check these structural invariants without asserting forecast quality:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f04196c0-8c45-426a-858b-0c3a7a1415f1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def test_assets_are_not_combined() -&gt; None:
    """AAPL and MSFT are built as independent models with no shared layers."""
    aapl_model = build_asset_gru_model(get_final_model_config("AAPL"))
    msft_model = build_asset_gru_model(get_final_model_config("MSFT"))

    assert aapl_model is not msft_model
    assert _layer_units(aapl_model) == (240, 16, 16)
    assert _layer_units(msft_model) == (224, 48)

    aapl_layer_ids = {id(layer) for layer in _gru_layers(aapl_model)}
    msft_layer_ids = {id(layer) for layer in _gru_layers(msft_model)}
    assert aapl_layer_ids.isdisjoint(msft_layer_ids)
    assert len(_gru_layers(aapl_model)) != len(_gru_layers(msft_model))</code></pre></div><p>This test describes the intended invariant: separate models have separate layer objects and preserve their asset-specific widths. It does not establish that training succeeds, that predictions are accurate, or that the paper's reported metrics are reproduced.</p><h3>Normalized outputs and the next stage</h3><p>The model operates on scaled inputs and produces a scaled next-price prediction. The output is therefore not yet an adjusted closing price in the original price units. The later evaluation stage must use the same asset-specific training-fitted scaler to inverse-transform predictions and targets before computing price-unit metrics.</p><p>The architecture alone cannot establish the paper's reported MSE, MAE, MAPE, or forecast accuracy. Those values also depend on the local data snapshot, scaling convention, training randomness, callback behavior, evaluation subset, and other unspecified protocol details. Under the supplied run policy, this model construction and its tests were not executed, and neither static verification nor semantic code verification was completed.</p><h2>Hyperband Tuning, Adam Training, and Best Checkpoints</h2><p>How should we choose and train a GRU without allowing later observations to influence the model? The workflow has two distinct stages. First, optional Hyperband tuning compares candidate asset-specific configurations using validation loss. Second, the selected configuration is trained with Adam while a separate chronological validation partition controls early stopping and checkpoint selection.</p><p>The paper describes this procedure, but it does not provide enough tuner metadata to reconstruct the original search exactly. The generated implementation therefore supports both a direct reproduction path using the reported Table 1 configurations and an optional, explicitly documented local Hyperband search.</p><h3>Paper facts and implementation decisions</h3><p>The paper reports separate searches and models for <code>AAPL</code> and <code>MSFT</code>. The final configurations are:</p><ul><li><p><code>AAPL</code>: GRU widths <code>(240, 16, 16)</code>;</p></li><li><p><code>MSFT</code>: GRU widths <code>(224, 48)</code>;</p></li><li><p>dropout <code>0.0</code> for both assets;</p></li><li><p>Adam learning rate <code>0.001</code>;</p></li><li><p>MSE loss;</p></li><li><p>batch size <code>8</code>;</p></li><li><p>maximum <code>100</code> epochs; and</p></li><li><p>early-stopping patience <code>7</code>.</p></li></ul><p>The paper also identifies validation loss as the Hyperband selection objective. Training uses the 156-sample training partition, while the separate 52-sample validation partition is supplied for monitoring. The 52-sample testing partition is not passed to model fitting or callback monitoring.</p><p>Several details remain unspecified: the exact Hyperband search space, bracket and iteration settings, random seed, tuner stopping criteria, Keras validation arrangement, callback filename, monitor mode, and whether best weights were restored automatically. The generated code makes these choices explicit rather than presenting them as recovered paper facts. In particular, it uses <code>shuffle=False</code>, minimizes <code>val_loss</code>, restores best weights, and saves weights to a <code>.weights.h5</code> file.</p><h3>A resource-aware search, without claiming exact tuner recovery</h3><p>Hyperband is a resource-aware search strategy. It begins with multiple candidate configurations and allocates more training resources to candidates that appear promising. In this reproduction, a candidate may vary in recurrent depth, unit counts, dropout, learning rate, and loss name. The search is run separately for each asset, so AAPL and MSFT observations never enter the same tuner.</p><p>The search-space container documents the local choices:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;acdb4ae9-cd6e-46f6-8564-e820e5f33aa4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class TuningSearchSpace:
    """Candidate values used by the optional Hyperband reproduction.

    The paper reports that Hyperband searched layer depth, recurrent widths,
    dropout, learning rate, and loss, but it does not provide the original
    candidate ranges.  Consequently, the values returned by
    :func:`default_search_space` are explicit implementation decisions rather
    than recovered paper facts.
    """

    layer_options: tuple[int, ...]
    unit_options: tuple[int, ...]
    dropout_options: tuple[float, ...]
    learning_rates: tuple[float, ...]
    loss_options: tuple[str, ...]</code></pre></div><p><code>TuningSearchSpace</code> contains finite candidate values and <code>validate_search_space</code> rejects malformed options such as nonpositive widths, invalid dropout values, duplicate entries, or unsupported losses. The generated <code>default_search_space</code> includes the reported widths and learning rate so the Table 1 configurations are representable, but its ranges are implementation decisions.</p><p>The search function receives a <code>DatasetSplits</code> object whose scaled training inputs have shape <code>(156, 3, 1)</code> and targets have shape <code>(156, 1)</code>. Its validation inputs and targets have shapes <code>(52, 3, 1)</code> and <code>(52, 1)</code>. It passes only the training and validation arrays to Keras Tuner:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8d807115-09ea-4b33-93a3-26eee5c5856a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">tuner.search(
    train_data.train_inputs,
    train_data.train_targets,
    validation_data=(
        train_data.validation_inputs,
        train_data.validation_targets,
    ),
    epochs=100,
    batch_size=8,
    shuffle=False,
    verbose=0,
)</code></pre></div><p>The objective is explicitly <code>val_loss</code> with direction <code>"min"</code>. The testing partition remains distinct. If TensorFlow or Keras Tuner is unavailable, the function raises a configuration error rather than silently substituting another algorithm. It also rejects invalid shapes, wrong sample counts, non-finite values, and searches that produce no finite validation score.</p><p>The paper's final architecture can be selected directly when the original tuner metadata is unavailable:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ead7f3da-c3d6-4e31-abea-da613e2ffca8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from stock_forecasting.tuning.hyperband import select_reported_configuration

reported_aapl_config = select_reported_configuration("AAPL")
reported_msft_config = select_reported_configuration("MSFT")</code></pre></div><p><code>select_reported_configuration</code> returns the Table 1 configuration. It does not rerun Hyperband and does not claim that the generated tuner recovered the paper's search. This direct path is the appropriate default for the requested reproduction target.</p><h3>Normalized-space MSE and model compilation</h3><p>During training, predictions and targets remain on the scaler's normalized price scale. The model returns one scalar prediction per sample, so predictions and targets should use compatible scalar-batch shapes, preferably <code>(batch, 1)</code>. The <code>models.losses</code> module validates this contract and rejects unintended broadcasting or mismatched batch lengths.</p><p>The paper selects MSE for both final asset models. No canonical loss equation was supplied in the paper context, so this tutorial does not reconstruct or display one. Instead, <code>resolve_loss("mse")</code> delegates the numerical loss behavior to the configured Keras/TensorFlow version. This keeps the framework's reduction behavior in one place and avoids presenting an invented formula as a paper equation.</p><p>Compilation constructs Adam with the configured learning rate and attaches the resolved loss:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d8e1d04e-8195-4073-a223-91a66e66735a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># The paper-selected MSE loss operates on normalized scalar predictions
# and targets; inverse transformation is intentionally deferred to
# evaluation.
optimizer = Adam(learning_rate=config.learning_rate)
loss = resolve_loss(config.loss_name)

try:
    model.compile(optimizer=optimizer, loss=loss)
except (TypeError, ValueError) as exc:
    raise ConfigurationError(
        f"Keras rejected the model compilation configuration: {exc}"
    ) from exc</code></pre></div><p><code>compile_gru_model</code> accepts an uncompiled Keras-compatible model and an <code>AssetModelConfig</code>, then returns the same model after compilation. It validates the learning rate and supported loss before importing TensorFlow. A missing TensorFlow installation or a model without a callable <code>compile</code> method is treated as a clear configuration failure.</p><p>The validation loss values reported by the paper are not explicitly labeled as normalized-space or original-price-unit values. The generated code treats training and validation loss as model-space quantities and reserves price-unit MSE, MAE, and MAPE for the inverse-transformed evaluation stage. This is an implementation interpretation, not a confirmed explanation of the paper's reported validation-loss units.</p><h3>Training with chronological validation</h3><p>Once a model configuration is selected, <code>train_gru_with_callbacks</code> compiles the model, builds the callbacks, and calls Keras <code>fit</code>. The important data-flow rule is that only the training partition updates weights:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a0c549ab-62c0-4902-a82b-46db538fa011&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">history = compiled_model.fit(
    splits.train_inputs,
    splits.train_targets,
    validation_data=(splits.validation_inputs, splits.validation_targets),
    epochs=config.max_epochs,
    batch_size=config.batch_size,
    shuffle=False,
    callbacks=callbacks,
    verbose=0,
)</code></pre></div><p>Here, <code>splits.train_inputs</code> has shape <code>(156, 3, 1)</code> and <code>splits.train_targets</code> has shape <code>(156, 1)</code>. The validation arrays have corresponding shapes <code>(52, 3, 1)</code> and <code>(52, 1)</code>. <code>shuffle=False</code> is an implementation decision that preserves chronological sample order; the extracted paper does not explicitly state its shuffle setting.</p><p>The configuration supplies the paper's batch size, epoch limit, and patience. The helper checks these shapes and counts before training, rejects non-finite arrays, and verifies that training and validation dates match their expected partition sizes. It never passes <code>test_inputs</code> or <code>test_targets</code> to <code>fit</code>, so the testing data cannot influence parameter updates or callback selection.</p><h3>Early stopping and best checkpoints</h3><p>Early stopping answers a practical question: when validation loss stops improving, should training continue merely because the maximum epoch count has not been reached? The generated callback policy monitors <code>val_loss</code>, minimizes it, tolerates seven non-improving epochs, and restores the best in-memory weights. A matching <code>ModelCheckpoint</code> saves the best weights locally.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d92ba597-602e-4715-9362-ac57d3b54013&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># The paper specifies validation-loss monitoring and best-state retention;
# mode='min' and restore_best_weights=True make those choices explicit.
early_stopping = keras.callbacks.EarlyStopping(
    monitor=_VALIDATION_MONITOR,
    mode="min",
    patience=config.patience,
    restore_best_weights=True,
    verbose=0,
)
checkpoint = keras.callbacks.ModelCheckpoint(
    filepath=str(checkpoint_path),
    monitor=_VALIDATION_MONITOR,
    mode="min",
    save_best_only=True,
    save_weights_only=True,
    verbose=0,
)
return [early_stopping, checkpoint]</code></pre></div><p>The paper specifies patience <code>7</code> and checkpointing, but not the monitor mode, restoration flag, or file format. The generated implementation chooses <code>mode="min"</code> because lower validation loss is preferred, uses <code>restore_best_weights=True</code>, and requires a checkpoint path ending in <code>.weights.h5</code>. <code>build_training_callbacks</code> creates the parent directory and raises an error for an invalid path or missing TensorFlow.</p><p>After fitting, <code>train_gru_with_callbacks</code> requires the checkpoint file to exist and reloads it. The returned <code>TrainingResult</code> contains the model, training history, and checkpoint path. Thus, forecasting uses the best validation-loss state rather than assuming the final epoch is best. This is a runtime contract; no training or checkpoint creation was performed for this article under the run policy.</p><h3>Worked example: direct reproduction versus optional tuning</h3><p>For the direct Table 1 path, keep tuning disabled. The pipeline then obtains the reported asset configuration through <code>get_final_model_config</code>, builds the corresponding GRU, trains it with the callback policy, and retains separate artifacts under the selected asset directory:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b2a00cd7-c20c-4c62-9797-434df1eaa5fe&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from pathlib import Path

from stock_forecasting.pipeline import run_asset_pipeline

result = run_asset_pipeline(
    asset="AAPL",
    data_dir=Path("data/raw"),
    output_dir=Path("artifacts"),
    evaluation_subset="validation",
    tune=False,
)</code></pre></div><p>This call is procedural rather than a claimed result. The pipeline loads local AAPL records, preserves the 156/52/52 split, fits the scaler on training data, constructs the reported AAPL configuration, trains with validation monitoring, and evaluates the explicitly selected validation subset. The returned <code>AssetRunResult</code> retains the raw <code>PriceSeries</code>, windowed dataset, scaled splits, scaler, model, training result, forecast, and diagnostic metadata.</p><p>To request the optional local Hyperband path from the command line, use:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;3366b4b2-cfef-422a-8df9-e1dc3e04f29e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">python scripts/run_reproduction.py \
    --data-dir data/raw \
    --output-dir artifacts \
    --assets AAPL MSFT \
    --evaluation-subset validation \
    --tune</code></pre></div><p>Without <code>--tune</code>, the command uses the reported Table 1 configurations. With <code>--tune</code>, it uses the generated search space and records tuner artifacts separately for each asset. Neither invocation silently merges testing and validation, and neither command's numerical outcome is claimed here.</p><h3>What this stage establishes&#8212;and what it does not</h3><p>This implementation preserves the paper's main training constraints: separate asset models, validation-loss selection, Adam at learning rate <code>0.001</code>, MSE loss, batch size <code>8</code>, at most <code>100</code> epochs, patience <code>7</code>, and best-weight checkpointing. It also makes unresolved choices visible: tuner ranges, random seed, Hyperband details, Keras defaults, callback behavior, and evaluation protocol.</p><p>The paper's reported AAPL and MSFT validation-loss values and forecast metrics remain reference values, not achieved results. Under the supplied run policy, no training, tuning, code execution, tests, static verification, semantic code verification, tutorial verification, or final quality review was performed.</p><h2>Inverse Transformation, Forecast Metrics, and Reference Comparison</h2><p>How do normalized neural-network outputs become interpretable stock-price forecasts? The evaluation stage reverses the training-time scaling, pairs every prediction with the date it forecasts, and computes errors in original adjusted-price units. It also makes one unresolved protocol choice visible: whether to evaluate the testing partition or the validation partition.</p><p>The paper describes its reported evaluation as validation-based, but it also defines separate training, testing, and validation subsets. Because the supplied paper description does not explain whether testing was ignored, evaluated separately, or used in another way, the generated implementation never merges these partitions silently. The caller must choose one explicitly.</p><h3>Paper facts and implementation decisions</h3><p>The paper reports separate results for AAPL and MSFT. It evaluates next-period adjusted closing prices with MSE, MAE, MAPE, and forecast accuracy. The supplied series contain positive prices, so percentage error is defined for these inputs.</p><p>The implementation makes several choices explicit rather than presenting them as recovered paper details:</p><ul><li><p><code>validation</code> is the default command-line evaluation choice, but <code>test</code> is also available.</p></li><li><p>The scaler is the one fitted from training data only.</p></li><li><p>Predictions and targets are inverse-transformed before price-unit metrics are calculated.</p></li><li><p>A nonpositive actual price is rejected for MAPE rather than handled with an invented zero-denominator rule.</p></li><li><p>The paper's Table 3 values are comparison references, not guaranteed outputs from this implementation.</p></li></ul><p>No canonical metric equations were supplied in the source context. Accordingly, this section explains the metric contracts and their code mappings in prose rather than reconstructing formulas.</p><h3>1. Select exactly one held-out partition</h3><p>After windowing and chronological splitting, each asset has 156 training samples, 52 testing samples, and 52 validation samples. The <code>select_evaluation_subset</code> function accepts a <code>DatasetSplits</code> object and returns the inputs, targets, and target dates from exactly one held-out partition.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f9196055-35a2-470e-a8fd-321ba101bf03&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">inputs, scaled_targets, target_dates = select_evaluation_subset(
    splits,
    subset="validation",
)

print(inputs.shape)          # expected: (52, 3, 1)
print(scaled_targets.shape)  # expected: (52, 1)
print(target_dates.shape)    # expected: (52,)</code></pre></div><p>To select the testing partition instead, change only the explicit protocol argument:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;28aa1b60-ec33-48ef-8fc9-c034b5aca151&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">test_inputs, test_targets, test_dates = select_evaluation_subset(
    splits,
    subset="test",
)</code></pre></div><p>The function returns three aligned objects:</p><ul><li><p>input windows with shape <code>(52, 3, 1)</code>;</p></li><li><p>normalized scalar targets with shape <code>(52, 1)</code>; and</p></li><li><p>a chronological <code>DatetimeIndex</code> of 52 target dates.</p></li></ul><p>Its failure case is deliberate: any subset other than <code>test</code> or <code>validation</code> raises <code>EvaluationProtocolError</code>. This prevents a misspelled or implicit evaluation choice from producing an apparently valid report. The function also does not concatenate the two held-out partitions.</p><h3>2. Predict, restore price units, and preserve dates</h3><p>The GRU receives normalized windows and should produce one normalized scalar per window. <code>evaluate_forecasts</code> validates this contract, calls the model's <code>predict</code> method, converts predictions and targets back through the same training-fitted scaler, and then constructs a <code>ForecastResult</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7bc114ce-eb54-4ffb-b2ac-34f9e96de2cf&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">result = evaluate_forecasts(
    model=model,
    inputs=inputs,
    targets=scaled_targets,
    dates=target_dates,
    scaler=scaler,
)

print(result.dates.shape)       # expected: (52,)
print(result.actuals.shape)     # expected: (52,)
print(result.predictions.shape) # expected: (52,)
print(result.metrics)</code></pre></div><p>The returned <code>ForecastResult</code> contains <code>dates</code>, <code>actuals</code>, <code>predictions</code>, and <code>metrics</code>. The actuals and predictions are one-dimensional price arrays. For either default held-out partition, each has length 52.</p><p>The function checks several invariants before asking the model to forecast:</p><ul><li><p><code>inputs</code> must have shape <code>(n, 3, 1)</code>;</p></li><li><p>targets may have shape <code>(n,)</code> or <code>(n, 1)</code>, then are normalized internally;</p></li><li><p>dates must be unique, nonmissing, and chronological;</p></li><li><p>inputs and targets must contain the same number of samples; and</p></li><li><p>all numeric values must be finite.</p></li></ul><p>It also requires a model with a callable <code>predict</code> method. A prediction returned as <code>(n, 1)</code> is converted to <code>(n,)</code>; an incompatible output shape, a failed prediction, or a non-finite prediction raises <code>EvaluationProtocolError</code>.</p><p>The lower-level <code>inverse_transform_forecasts</code> helper performs only the restoration step. It applies the same scaler to normalized predictions and normalized targets, then returns equal-length one-dimensional arrays in adjusted-price units:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a9bc91b9-ec5f-4acd-b1b0-f9eac504b038&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from stock_forecasting.evaluation.forecasting import inverse_transform_forecasts

price_predictions, price_targets = inverse_transform_forecasts(
    scaler,
    normalized_predictions,
    scaled_targets,
)</code></pre></div><p>Using one training-fitted scaler for both arrays is important. Applying a separately fitted evaluation scaler would change the meaning of the predictions and could introduce information from the held-out data. The paper specifies training-only fitting, but it does not resolve whether fitting used the raw training prices or flattened training windows and targets. That convention must therefore be recorded with the result artifact.</p><h3>3. Understand the metric units</h3><p>Once arrays are back in price units, <code>compute_forecast_metrics</code> delegates to the individual metric functions. The public interface is concise:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f5543649-a80d-49d4-a218-1db3df886a7b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from stock_forecasting.evaluation.metrics import compute_forecast_metrics

metrics = compute_forecast_metrics(
    actuals=result.actuals,
    predictions=result.predictions,
)

print(f"MSE: {metrics.mse:.2f}")
print(f"MAE: {metrics.mae:.2f}")
print(f"MAPE: {metrics.mape:.2f}%")
print(f"Forecast accuracy: {metrics.forecast_accuracy:.2f}%")</code></pre></div><p>The fields have different interpretations:</p><ul><li><p><strong>MSE</strong> is mean squared error. Its unit is squared adjusted-price units, so large errors receive disproportionately more weight.</p></li><li><p><strong>MAE</strong> is mean absolute error in adjusted-price units. It is easier to read as a typical price-distance measure than MSE.</p></li><li><p><strong>MAPE</strong> is mean absolute percentage error, reported in percentage points.</p></li><li><p><strong>Forecast accuracy</strong> follows the paper's convention of taking 100 minus MAPE, also in percentage points.</p></li></ul><p>The last quantity is not classification accuracy. It does not measure whether the model correctly predicted an upward or downward movement. It is simply the paper's complement of a percentage error.</p><p>The metric functions validate equal-length, one-dimensional, finite arrays. <code>compute_mape</code> additionally requires every actual price to be strictly positive. This matches the supplied AAPL and MSFT data and avoids silently choosing how to divide by zero or a negative price. The generated implementation raises <code>EvaluationProtocolError</code> when that assumption is violated.</p><h3>4. Worked example: validation forecast and date-aligned plot</h3><p>The following sequence makes the evaluation protocol visible: choose validation data, evaluate it, and plot the returned date-aligned result. The <code>model</code> and <code>scaler</code> are assumed to have been produced by the preceding training pipeline; this excerpt does not claim that training or forecasting was executed.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8e628bcb-bd2c-4d83-9a25-ed82fd172940&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from pathlib import Path

from stock_forecasting.evaluation.forecasting import (
    evaluate_forecasts,
    select_evaluation_subset,
)
from stock_forecasting.eda.plots import plot_forecast, save_figure

inputs, scaled_targets, target_dates = select_evaluation_subset(
    splits,
    subset="validation",
)

result = evaluate_forecasts(
    model=model,
    inputs=inputs,
    targets=scaled_targets,
    dates=target_dates,
    scaler=scaler,
)

axis = plot_forecast(result)
save_figure(
    axis.figure,
    Path("artifacts/forecasts/AAPL_validation_forecast.png"),
)</code></pre></div><p><code>plot_forecast</code> reads the <code>ForecastResult</code> without mutating it. It checks that dates, actuals, and predictions have equal lengths, then draws both series against the target dates. Those dates are not the dates at the beginning of each input window: they identify the following observation that served as the target. This distinction keeps a forecast visually aligned with the price it was intended to predict.</p><p>To inspect testing separately, repeat the same process with <code>subset="test"</code> and write a distinct artifact name. Do not label a validation result as a test result, since the two partitions represent different chronological portions of the data.</p><h3>5. Compare with Table 3 without claiming reproduction</h3><p>The paper reports these reference values:</p><p>| Asset | MSE | MAE | MAPE | Forecast accuracy | |---|---:|---:|---:|---:| | AAPL | 55.80 | 6.11 | 3.37% | 96.63% | | MSFT | 144.44 | 9.41 | 2.55% | 97.45% |</p><p><code>compare_to_reference</code> retrieves the appropriate paper reference, records the computed value, reference value, and signed difference for every metric, and can optionally calculate absolute-tolerance indicators. <code>format_reference_report</code> renders those fields as a readable report.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;db831897-9500-4f2a-b865-ca5e8a849b17&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from stock_forecasting.evaluation.reference import (
    compare_to_reference,
    format_reference_report,
)

comparison = compare_to_reference(
    asset="AAPL",
    metrics=result.metrics,
)

print(format_reference_report({"AAPL": comparison}))</code></pre></div><p>A comparison with an optional tolerance is still descriptive rather than a proof of reproducibility:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2efb150a-97d9-40b7-9669-575f1ef463d9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">comparison = compare_to_reference(
    asset="MSFT",
    metrics=msft_result.metrics,
    tolerance=1.0,
)

print(format_reference_report({"MSFT": comparison}))</code></pre></div><p>The report keeps units explicit: MSE and MAE differences use price-based units, while MAPE and forecast-accuracy differences use percentage points. Without a tolerance, the report states that tolerance was not assessed. With one, it reports whether each absolute difference falls within that numerical threshold; it does not establish that the data, training process, or evaluation protocol matched the paper.</p><h3>6. Offline artifact reporting</h3><p>The generated <code>scripts/report_results.py</code> script reads JSON forecast artifacts and reports their generated metrics separately from Table 3. Its <code>load_result_artifact</code> function requires dates, actuals, predictions, and all four metric fields. It rejects missing fields, invalid dates, unequal array lengths, non-finite values, and nonpositive actuals.</p><p>A typical command shape is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;4f919a8c-6730-4224-80c0-9441f0cdc620&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">python scripts/report_results.py \
    --artifact-dir artifacts/results \
    --assets AAPL MSFT \
    --tolerance 1.0</code></pre></div><p>This reporter is intentionally offline. It does not download data or infer a missing evaluation subset from the JSON. The artifact-producing workflow should retain the asset name, selected subset, scaler convention, model configuration, target dates, and metric units alongside the arrays. That metadata is necessary when a discrepancy with Table 3 could arise from more than the GRU itself.</p><h3>7. Interpreting absolute and relative errors</h3><p>The paper reports larger absolute errors for MSFT&#8212;MSE <code>144.44</code> and MAE <code>9.41</code>&#8212;than for AAPL&#8212;MSE <code>55.80</code> and MAE <code>6.11</code>. MSFT nevertheless has the lower MAPE, <code>2.55%</code> compared with AAPL's <code>3.37%</code>, and therefore the higher reported forecast accuracy, <code>97.45%</code> compared with <code>96.63%</code>.</p><p>These values illustrate why both absolute and relative measures matter. A stock with a higher price level can have a larger dollar error while that error represents a smaller fraction of the actual price. The metrics answer different questions: MAE asks about price-unit distance, whereas MAPE asks about proportional distance.</p><p>These figures remain paper references in this reproduction. Exact numerical agreement is not guaranteed because the supplied context leaves the source-data version, weekly aggregation, scaler-fitting convention, random seed, complete Hyperband search space, Keras defaults, callback details, and reported evaluation-subset procedure incomplete.</p><h3>8. Evaluation checklist</h3><p>Before interpreting a forecast artifact, confirm that:</p><ul><li><p>the asset is recorded as either AAPL or MSFT and was modeled independently;</p></li><li><p>one explicit subset, <code>test</code> or <code>validation</code>, is recorded;</p></li><li><p>inputs have shape <code>(52, 3, 1)</code> for the selected default held-out partition;</p></li><li><p>dates, actuals, and predictions have equal length and chronological alignment;</p></li><li><p>the scaler was fitted on training data only and reused for inverse transformation;</p></li><li><p>metrics were calculated after returning predictions and targets to price units;</p></li><li><p>MSE and MAE are reported with price-based units;</p></li><li><p>MAPE and forecast accuracy are reported as percentages;</p></li><li><p>forecast accuracy is interpreted as 100 minus MAPE, not as classification accuracy; and</p></li><li><p>Table 3 values are labeled as references rather than achieved results.</p></li></ul><p>The generated metric tests describe these invariants, including inverse transformation, positive-price validation, the forecast-accuracy complement, and date preservation. However, the run policy disabled execution, test generation, static verification, semantic code verification, tutorial verification, and final review. No test result or numerical reproduction claim is made here.</p><h2>Limitations, Static and Semantic Verification Status, and Reproduction Checklist</h2><p>What would make this reproduction trustworthy&#8212;and what can it <em>not</em> establish? The workflow is a useful, structured experiment: it keeps AAPL and MSFT separate, preserves chronological order, prevents later values from fitting the scaler, and records the choices needed to turn the paper description into Python. It is not a guarantee that a GRU will forecast future markets reliably.</p><p>The study uses only 263 weekly adjusted closing-price observations for each of two technology stocks. Each model sees only three preceding prices and predicts one next price. That narrow design is faithful to the paper's scope, but it limits generalization. Market conditions can change, and a price-history-only model may smooth abrupt movements or behave differently under distribution shift. The resulting forecasts should therefore support analysis and financial-risk decisions, not operate as an autonomous trading system.</p><h3>A final separation of evidence</h3><p>A responsible reproduction distinguishes three categories:</p><ul><li><p><strong>Paper facts:</strong> the separate AAPL and MSFT models, adjusted-price inputs, three-period windows, chronological partitions, Table 1 architectures, training settings, and Table 3 reference metrics.</p></li><li><p><strong>Derived structure:</strong> 263 raw observations minus a three-period window produces 260 supervised samples, which are partitioned as 156 training, 52 testing, and 52 validation samples.</p></li><li><p><strong>Implementation decisions:</strong> local CSV input instead of live retrieval; the selected scaler convention; decomposition, ADF, and autocorrelation parameters; Keras defaults not specified by the paper; callback details; and the locally defined Hyperband search space.</p></li></ul><p>This distinction is especially important when comparing numbers. The paper's reference values&#8212;AAPL MSE 55.80, MAE 6.11, MAPE 3.37%, and forecast accuracy 96.63%; MSFT MSE 144.44, MAE 9.41, MAPE 2.55%, and forecast accuracy 97.45%&#8212;are targets for comparison. They are not evidence that this generated code achieved those values.</p><h3>What can prevent exact numerical reproduction</h3><p>Even with the same broad workflow, results can differ because the supplied paper description does not fully specify:</p><ul><li><p>the exact Yahoo Finance data version, retrieval query, weekly aggregation rule, and calendar handling;</p></li><li><p>whether the scaler was fitted on raw training prices or flattened training windows and targets;</p></li><li><p>random seeds and framework initialization;</p></li><li><p>GRU activations, initializers, sequence-output settings, and the terminal output construction;</p></li><li><p>the Hyperband search space, iterations, seed, and stopping rules;</p></li><li><p>early-stopping and checkpoint details; and</p></li><li><p>whether the reported evaluation used validation only, testing only, or another unresolved protocol.</p></li></ul><p>The generated pipeline makes these uncertainties visible. It keeps the testing and validation partitions distinct and requires an explicit <code>evaluation_subset</code> choice. It also records the selected scaler convention and diagnostic parameters rather than implying that they were recovered from the paper.</p><h3>Verification status: skipped, not passed</h3><p>The run policy disabled execution, test generation, local static verification, semantic code verification, tutorial verification, and final quality review. The local verification record consequently reports <code>verification_skipped</code>, not a successful check. Semantic code verification was also skipped. No claim is made here that the package imports correctly, that a model trained, that tests passed, or that any metric was reproduced.</p><p>Static verification and semantic verification would answer different questions if they were later enabled. Static checks could inspect Python parsing, imports, dependency order, configuration consistency, tensor shapes, chronology, and leakage boundaries without training a model. Semantic review would examine whether the implementation's behavior actually matches the intended method&#8212;for example, whether only training values fit the scaler, whether testing data stays out of training, and whether dates remain aligned with target observations. Neither kind of check was completed in this run.</p><h3>Worked example: the complete local workflow</h3><p>The intended workflow begins with local appendix-style files. The repository does not download data or pretend to reproduce an unspecified Yahoo Finance retrieval procedure. Prepare <code>AAPL.csv</code> and <code>MSFT.csv</code> under a local input directory, with <code>date</code> and <code>adjusted_close</code> columns:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f205ad76-47a7-4a30-9227-096598e71e9d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">python scripts/prepare_local_data.py \
  --input-dir appendix \
  --output-dir data/raw</code></pre></div><p>The preparation script validates the local records and writes the normalized schema. Its public function preserves the paper's 263-row contract:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d4982f5d-eb25-4276-a4a1-8f22bf0fbe5e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def prepare_asset_file(
    source_path: Path,
    destination_path: Path,
    asset: str,
) -&gt; None:
    """Validate a local appendix CSV and write its normalized raw-data form.

    The input is never downloaded or augmented.  The loader accepts the paper's
    date and adjusted-close fields (including documented column aliases), while
    this function writes the stable schema expected by later pipeline stages:
    ``date``, ``adjusted_close``.  Records are ordered chronologically and must
    contain exactly the paper's 263 observations.
    """</code></pre></div><p>Next, generate non-mutating diagnostics for both assets:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;44abf02b-0d1d-44bf-8ae2-0b480407dc12&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">python scripts/run_diagnostics.py \
  --data-dir data/raw \
  --output-dir artifacts/diagnostics</code></pre></div><p>These reports and figures describe the data; they do not add decomposition components or other diagnostics to the GRU inputs. The generated command records implementation choices such as the decomposition period, ADF settings, and autocorrelation lag range.</p><p>Finally, run the reproduction with an explicit evaluation partition:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a12a9410-425c-40ca-95d1-25d0c363d7bc&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">python scripts/run_reproduction.py \
  --data-dir data/raw \
  --output-dir artifacts \
  --evaluation-subset validation</code></pre></div><p>To keep the testing partition separate but inspect it instead, choose <code>test</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c03d470b-9f6e-4e97-9a89-d6d86c5f3c7a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">python scripts/run_reproduction.py \
  --data-dir data/raw \
  --output-dir artifacts \
  --evaluation-subset test</code></pre></div><p>The <code>--tune</code> flag is another explicit choice. Without it, <code>run_asset_pipeline</code> uses the reported Table 1 configuration directly. With it, the implementation runs its documented, locally selected Hyperband search space; that search is not claimed to be the paper's unrecovered search space.</p><p>At the Python level, one asset follows the same sequence of responsibilities:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;284bab0d-c0dc-4b56-adb1-869a06932fd7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">result = run_asset_pipeline(
    asset=asset,
    data_dir=data_dir,
    output_dir=output_dir / asset,
    evaluation_subset=args.evaluation_subset,
    tune=args.tune,
)</code></pre></div><p>The returned <code>AssetRunResult</code> is intended to retain the raw <code>PriceSeries</code>, diagnostic record, <code>WindowedDataset</code>, scaled <code>DatasetSplits</code>, scaler, model, training artifact, and <code>ForecastResult</code>. This makes the pipeline inspectable: a reviewer can trace an asset from its 263 observations through 260 windows and the 156/52/52 partition to its selected held-out forecast. AAPL and MSFT are processed in separate calls, so their scalers, models, checkpoints, and metrics do not get combined.</p><h3>Practical reproduction checklist</h3><p>Use the following checklist when execution and verification are available:</p><ol><li><p><strong>Prepare inputs.</strong> Supply one local CSV per asset with <code>date</code> and <code>adjusted_close</code>; do not download or infer missing records.</p></li><li><p><strong>Validate raw data.</strong> Confirm 263 rows, chronological dates, finite positive adjusted prices, and no missing values.</p></li><li><p><strong>Run diagnostics.</strong> Save descriptive statistics, plots, decomposition, ADF, and autocorrelation outputs. Treat them as diagnostic artifacts only.</p></li><li><p><strong>Create windows.</strong> Confirm 260 samples with inputs shaped <code>(260, 3, 1)</code>, targets shaped <code>(260, 1)</code>, and dates aligned to target observations.</p></li><li><p><strong>Check partitions.</strong> Confirm contiguous chronological sizes of 156 training, 52 testing, and 52 validation samples.</p></li><li><p><strong>Check scaling.</strong> Confirm that only training values fit the chosen scaler and that later subsets are transformed without refitting.</p></li><li><p><strong>Build models separately.</strong> Confirm AAPL widths <code>(240, 16, 16)</code> and MSFT widths <code>(224, 48)</code>, with zero dropout and one scalar output per sample.</p></li><li><p><strong>Train carefully.</strong> Confirm Adam at learning rate <code>0.001</code>, MSE loss, batch size <code>8</code>, at most <code>100</code> epochs, patience <code>7</code>, validation monitoring, and best-checkpoint use.</p></li><li><p><strong>Select evaluation data explicitly.</strong> Record whether the held-out subset is <code>validation</code> or <code>test</code>; never merge them silently.</p></li><li><p><strong>Evaluate in price units.</strong> Inverse-transform predictions and targets with the training-fitted scaler before computing MSE, MAE, MAPE, and forecast accuracy.</p></li><li><p><strong>Compare transparently.</strong> Use <code>compare_to_reference</code> to report differences from Table 3, while labeling those values as paper references rather than pass/fail verification.</p></li><li><p><strong>Record the environment.</strong> Preserve data provenance, dependency versions, random seeds, scaler convention, tuner settings, callback settings, and the evaluation choice.</p></li></ol><p>The paper's conclusions also suggest what a stronger future study would need: more assets and time periods, rolling-origin evaluation, benchmark models, uncertainty estimates, confidence intervals, statistical significance analysis, and potentially technical, sentiment, volatility, or macroeconomic features. Those are sensible extensions, but they are outside this reproduction and must not be presented as part of the implemented method.</p><p>The central lesson is procedural rather than predictive: document the evidence boundary, preserve time order, expose unresolved choices, and never turn an unexecuted or unverified pipeline into a claimed result.</p><p>Use the button or URL below to download the source code.</p>
      <p>
          <a href="/__u/onepagecode.substack.com/p/quant-trading-deep-gru-recurrent">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Quant Trading: Feature-Wise Compositional RNNs & Grey Wolf Optimization for Stock Prediction (PyTorch Guide)]]></title><description><![CDATA[Building an end-to-end multi-stream LSTM, GRU, and SRU neural architecture with metaheuristic GWO hyperparameter optimization.]]></description><link>https://onepagecode.substack.com/p/quant-trading-feature-wise-compositional</link><guid isPermaLink="false">https://onepagecode.substack.com/p/quant-trading-feature-wise-compositional</guid><dc:creator><![CDATA[Onepagecode]]></dc:creator><pubDate>Tue, 04 Aug 2026 03:34:36 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_W6n!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3078fdad-8e80-40a9-be17-640eb18fa447_1254x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Use the button at the end of this article to download the source code.<br></h2><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.researchgate.net/publication/406514438_Compositional_RNN_Approach_to_Accurate_Stock_Price_Forecasting/link/6a412af80da2335e0b5a0777/download?_tp=eyJjb250ZXh0Ijp7ImZpcnN0UGFnZSI6InB1YmxpY2F0aW9uIiwicGFnZSI6InB1YmxpY2F0aW9uIiwicHJldmlvdXNQYWdlIjoiX2RpcmVjdCJ9fQ&quot;,&quot;text&quot;:&quot;Download Research Paper&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.researchgate.net/publication/406514438_Compositional_RNN_Approach_to_Accurate_Stock_Price_Forecasting/link/6a412af80da2335e0b5a0777/download?_tp=eyJjb250ZXh0Ijp7ImZpcnN0UGFnZSI6InB1YmxpY2F0aW9uIiwicGFnZSI6InB1YmxpY2F0aW9uIiwicHJldmlvdXNQYWdlIjoiX2RpcmVjdCJ9fQ"><span>Download Research Paper</span></a></p><p>The paper proposes a feature-wise compositional recurrent neural network approach for multivariate stock-price forecasting. The five OHLCV inputs are modeled separately with stacked LSTM, GRU, or SRU recurrent layers, regularized with dropout, fused by concatenation, optionally processed by additional recurrent layers and dense layers, and optimized using either Random Search (RS) or Grey Wolf Optimizer (GWO). The study evaluates 54 configurations on daily Hang Seng Index data from Yahoo Finance. The reported best result is an LSTM-GWO configuration, followed by GRU-GWO and SRU-GWO configurations.</p><h2>Implementation Assumptions</h2><ul><li><p>The five input channels are ordered Open, High, Low, Close, Volume.</p></li><li><p>The 20-day and 40-day values are treated as input window lengths, not forecast horizons.</p></li><li><p>Because the paper does not identify the target or horizon, the implementation exposes target<em>column and forecast</em>horizon as configuration choices; the tutorial must label these as implementation decisions.</p></li><li><p>Scaling defaults to training-only fitting over [0, 0.95] to reduce leakage, while documenting that the paper does not specify the fitting partition.</p></li><li><p>The paper's opaque configuration labels such as LSTM-GWO (1-1-0-1) are retained as benchmark labels and are not decoded into architecture parameters.</p></li><li><p>The training loss, gradient optimizer, search spaces, GWO budget, and metric formulas are configurable implementation decisions rather than asserted paper facts.</p></li><li><p>SRU is represented behind a local interface because the paper does not specify an SRU variant or API.</p></li><li><p>Yahoo Finance acquisition is optional for offline reproducibility; local OHLCV files and deterministic synthetic data are supported.</p></li><li><p>No reported benchmark is treated as verified reproduction output.</p></li><li><p>No canonical equations were supplied, so no equation-level LaTeX or equation_id mapping is asserted.</p></li><li><p>Verification is limited to planned static and semantic checks; execution, generated tests, and local verification are disabled by policy.</p></li></ul><h2>Scope, evidence status, and reproduction target</h2><p>What should happen when five market measurements describe the same trading day? The paper&#8217;s central idea is to avoid sending all of them immediately into one undifferentiated recurrent pathway. Instead, Open, High, Low, Close, and Volume each receive their own temporal pathway. The resulting representations are then joined before the final forecast. A useful intuition is to treat the five pathways as specialists: each studies one signal over time, and a prediction head combines their reports.</p><p>Here, <code>OHLCV</code> means the ordered features Open, High, Low, Close, and Volume. The generated model expects tensors in <code>batch &#215; time &#215; feature</code> layout. For a 20-day input window, for example, one batch has shape <code>batch &#215; 20 &#215; 5</code>; for the alternative window, the shape is <code>batch &#215; 40 &#215; 5</code>. The final dimension must retain the canonical OHLCV order.</p><h3>What the paper establishes</h3><p>The supplied paper context supports a feature-wise compositional recurrent method. Each OHLCV stream is processed independently by a stacked LSTM, GRU, or SRU family. Dropout is used for regularization, the five representations are fused through concatenation, and optional recurrent and dense layers produce the stock-price prediction. The paper also compares two outer hyperparameter strategies:</p><ul><li><p><strong>Random Search (RS)</strong> samples candidate configurations stochastically.</p></li><li><p><strong>Grey Wolf Optimizer (GWO)</strong> maintains a population of candidates and updates them according to their relative objective values.</p></li></ul><p>The paper reports 54 evaluated configurations and identifies an LSTM-GWO configuration as its best reported result. It also reports GRU-GWO and SRU-GWO benchmark configurations. These are useful comparison targets, but they do not fully describe how to construct the corresponding models.</p><p>The phrase &#8220;univariate encoding&#8221; in the extracted paper should be read carefully. The implementation target is not one single-variable forecasting model. It is feature-wise encoding: five separate one-channel sequences are modeled and then fused into a multivariate representation.</p><h3>What remains underdetermined</h3><p>An exact reproduction cannot be recovered from the supplied context alone. Several decisions needed by executable code are absent or ambiguous:</p><ul><li><p>The Yahoo Finance symbol, date range, adjustment policy, and missing-row treatment are not uniquely specified.</p></li><li><p>The target column and forecast horizon are not identified. The paper names stock-price prediction, but does not establish that the target is <code>Close</code> or that the horizon is one day.</p></li><li><p>The recurrent widths, exact layer counts, hidden-state extraction rule, dense widths, activations, and SRU implementation are missing.</p></li><li><p>The training loss, gradient optimizer, learning-rate schedule, stopping rule, and checkpoint policy are unspecified.</p></li><li><p>RS bounds, sampling distributions, trial counts, and validation details are absent.</p></li><li><p>GWO population size, iteration budget, bounds, initialization, discrete-parameter encoding, and update specification are absent.</p></li><li><p>Metric formulas and conventions, including percentage-error zero handling and evaluation scale, are not supplied.</p></li></ul><p>The labels <code>LSTM-GWO (1-1-0-1)</code>, <code>GRU-GWO (2-1-1-1)</code>, and <code>SRU-GWO (2-2-0-0)</code> therefore remain opaque benchmark labels. They must not be decoded into presumed layer counts, dropout switches, or dense-layer structures.</p><h3>How the generated package responds</h3><p>The generated package treats the paper as a method specification plus a set of reported targets, not as a complete executable recipe. <code>DataConfig</code> exposes the target, horizon, window length, split, and scaling range. <code>ModelConfig</code> exposes the recurrent family and architecture choices. <code>TrainingConfig</code>, <code>SearchConfig</code>, and <code>EvaluationConfig</code> make training, search, and metric conventions explicit. <code>validate_reproduction_config</code> provides a common boundary for checking these records.</p><p>The model assembly is represented by <code>CompositionalRNN</code> and <code>build_compositional_model</code>. Their intended responsibilities are to create five independent recurrent encoders, concatenate their representations, apply the configured fusion treatment, and produce the configured output dimension. RS and GWO remain separate outer-search interfaces, but both are intended to evaluate candidates through a shared training and validation objective.</p><p>This separation is important. A derived implementation decision&#8212;such as using <code>Close</code> as the target, a one-step horizon, or a particular loss&#8212;can be changed when better evidence becomes available without changing the paper&#8217;s feature-wise architecture or the public configuration interfaces. It also prevents a convenient default from being mistaken for a fact reported by the paper.</p><h3>Worked example: an explicit tutorial configuration</h3><p>For a small paper-oriented setup, choose <code>window_length=20</code>, <code>target_column='Close'</code>, and <code>forecast_horizon=1</code> in <code>DataConfig</code>. These values are tutorial choices: the paper supports 20-day input windows, but it does not specify <code>Close</code> as the target or a one-step forecast horizon. A corresponding <code>ModelConfig</code> can select <code>LSTM</code> and an explicitly documented number of stream layers and units, while a <code>TrainingConfig</code> records the chosen loss, optimizer, batch size, epochs, and seed.</p><p>The resulting input contract is <code>batch &#215; 20 &#215; 5</code>, and the configured output dimension determines the prediction shape. That shape contract describes the scaffold; it does not demonstrate that the configuration reproduces the paper&#8217;s reported LSTM-GWO result.</p><h3>Evidence status</h3><p>This section and the generated package distinguish three kinds of statements:</p><ol><li><p><strong>Paper facts:</strong> feature-wise OHLCV processing, recurrent-family comparison, concatenation fusion, RS and GWO comparison, and the reported benchmark labels and values.</p></li><li><p><strong>Derived explanations:</strong> the &#8220;five specialists&#8221; intuition and the interpretation of separate pathways as feature-wise rather than single-variable modeling.</p></li><li><p><strong>Implementation decisions:</strong> target and horizon defaults, exact architecture settings, preprocessing policies, training choices, search budgets, and metric conventions.</p></li></ol><p>The run policy disabled execution, local static verification, semantic code verification, generated test execution, tutorial verification, and final quality review. Consequently, this package should be described as a configurable reproduction scaffold, not as a verified implementation or a successful reproduction of the published results.</p><h2>Data contract: HSI OHLCV acquisition and provenance</h2><p>How can a forecasting model learn from market data if the source rows, feature order, and cleaning decisions are unclear? The data contract answers that question before any recurrent layer is built. It specifies which five measurements enter the model, preserves their chronological order, and records choices that the paper does not disclose.</p><h3>What the paper specifies&#8212;and what it does not</h3><p>The paper identifies daily Hang Seng Index history obtained through Yahoo Finance/YFinance and names the raw Open, High, Low, Close, and Volume fields as its OHLCV inputs. In the generated package, the canonical order is fixed as <code>Open</code>, <code>High</code>, <code>Low</code>, <code>Close</code>, <code>Volume</code>. A raw table therefore has shape <code>num_days &#215; 5</code>: each row represents one observation and each column represents one input channel.</p><p>The paper does not provide a unique Yahoo Finance ticker, date range, download options, adjusted-price setting, missing-value policy, or duplicate-row policy. These are not minor details. Different tickers or date ranges change the observations, while adjusted and unadjusted prices can produce different price histories. The generated implementation consequently makes these choices explicit and stores them in <code>AcquisitionMetadata</code> rather than presenting a particular choice as recovered paper methodology.</p><p>The reproduction scope also stays deliberately narrow. It retains the five raw OHLCV fields and does not add technical indicators or other engineered predictors. Technical indicators are mentioned by the paper as possible future work, not as part of the evaluated method.</p><h3>Canonical columns and provenance</h3><p><code>CANONICAL_OHLCV_COLUMNS</code> is the package-level contract used by acquisition and model-facing code:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;dd919546-8fd3-48c4-a3e7-16923a169749&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># Canonical feature order required by the paper-oriented data and model APIs.
OHLCVColumns: TypeAlias = tuple[str, str, str, str, str]
CANONICAL_OHLCV_COLUMNS: OHLCVColumns = (
    "Open",
    "High",
    "Low",
    "Close",
    "Volume",
)</code></pre></div><p>This excerpt is copied from <code>src/compositional_rnn_stock/types.py</code>. <code>OHLCVColumns</code> describes a five-name tuple, while <code>CANONICAL_OHLCV_COLUMNS</code> supplies the concrete order. Preserving this order matters because later model code interprets the final input axis as five channels. Reordering columns&#8212;for example, placing Volume before Close&#8212;would change the meaning of the model input without changing its shape.</p><p>The acquisition record captures the choices surrounding those columns:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6de50c19-bc48-47e8-914e-4f51552b0b3d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class AcquisitionMetadata:
    """Provenance and cleaning-policy record for an OHLCV acquisition.

    The paper does not specify the Yahoo Finance symbol, date range, adjustment
    mode, missing-value policy, or duplicate-row policy.  These fields preserve
    the choices made by a local reproduction instead of treating them as paper
    facts.
    """

    source: str
    symbol: str | None
    start: str | None
    end: str | None
    adjusted: bool | None
    frequency: str
    requested_columns: OHLCVColumns
    missing_value_policy: str
    duplicate_policy: str
    ordering_policy: str</code></pre></div><p>This is an implementation record, not an additional claim about the paper. For a network acquisition, <code>source</code>, <code>symbol</code>, <code>start</code>, <code>end</code>, and <code>adjusted</code> identify the request. <code>requested_columns</code> records the five retained fields, and the remaining policy fields describe how the returned table was accepted. For an offline archive, some fields such as <code>symbol</code> may be <code>None</code>, but the artifact still records that the source was a local CSV.</p><h3>Two acquisition paths</h3><p>The generated adapter supports a network path through <code>load_hsi_ohlcv_from_yfinance</code> and an offline path through <code>load_ohlcv_csv</code>. The dispatcher requires exactly one source, preventing an ambiguous call that supplies both a symbol and a local file:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;efe04d8e-fb79-4e72-89df-415e1198f435&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def acquire_hsi_ohlcv_data(
    *,
    symbol: str | None = None,
    local_path: Path | None = None,
    start: str | None = None,
    end: str | None = None,
    adjusted: bool = False,
) -&gt; tuple[pd.DataFrame, AcquisitionMetadata]:
    """Dispatch to an explicit local-file or Yahoo Finance acquisition mode."""

    if (symbol is None) == (local_path is None):
        raise ValueError("provide exactly one of symbol or local_path")
    if local_path is not None:
        return load_ohlcv_csv(local_path, CANONICAL_OHLCV_COLUMNS)
    return load_hsi_ohlcv_from_yfinance(symbol=symbol, start=start, end=end, adjusted=adjusted)</code></pre></div><p>This excerpt is copied from <code>src/compositional_rnn_stock/data/acquisition.py</code>. The function returns a pandas table plus its provenance record. In local mode, it delegates to <code>load_ohlcv_csv</code>; in Yahoo Finance mode, it delegates to <code>load_hsi_ohlcv_from_yfinance</code>. The optional <code>yfinance</code> import is isolated inside the network function, so an offline workflow does not need that package merely to load an archived file.</p><p>For a local archive, a typical call is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;186b79c8-6291-42b0-9b17-2f36d0e7f89a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">frame, metadata = load_ohlcv_csv(Path("data/hsi_ohlcv.csv"), CANONICAL_OHLCV_COLUMNS)</code></pre></div><p>The generated CSV loader looks for a conventional <code>Date</code>, <code>date</code>, <code>Timestamp</code>, or <code>timestamp</code> column and uses it as the index when present. It then selects and orders only the canonical OHLCV fields. An archive should therefore preserve its date column whenever the identity of trading days matters. If no recognized date column exists, the loader retains the file's row index; that is a documented limitation of the local input rather than evidence that dates were absent from the original paper dataset.</p><p>For Yahoo Finance, the caller must supply the symbol and may supply explicit date bounds and an adjustment choice. The paper does not identify these values, so a tutorial command should not silently substitute a supposedly canonical ticker or time range. The command-line entry point instead requires explicit source and target-related settings, while the acquisition function records the selected source parameters.</p><h3>Validation and cleaning</h3><p>Acquisition first validates the structural contract. <code>validate_ohlcv_table</code> requires a nonempty pandas table, all five columns, numeric finite values, an increasing index, and no duplicate index values. It does not infer missing observations or create rows. The relevant responsibility is summarized by its public interface:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8dadc946-a371-4a14-99eb-faf0700293b9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def validate_ohlcv_table(frame: pd.DataFrame) -&gt; None:
    """Validate a chronological table containing exactly the five OHLCV fields.

    Validation deliberately rejects missing values and duplicates rather than
    fabricating observations.  The paper leaves those policies unspecified, so
    callers must clean or otherwise resolve such records before acquisition
    results are accepted.
    """</code></pre></div><p>The implementation's cleaning layer makes the missing and duplicate policies explicit:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;92ba7b8e-9811-4f63-9762-a7b9a36855f8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">cleaned = clean_ohlcv(frame, missing_policy="drop", duplicate_policy="reject")</code></pre></div><p>This call is the focused usage pattern exposed by <code>src/compositional_rnn_stock/data/cleaning.py</code>. The `</p><h2>Scaling, target definition, and sliding windows</h2><p>How does an ordered table of daily market observations become training data without allowing future information to leak backward? The preprocessing pipeline answers this in three stages: scale the five OHLCV channels, construct overlapping time windows, and split the resulting supervised samples chronologically.</p><p>The paper states that the features are normalized to the interval <code>[0, 0.95]</code> and that the model uses 20-day and 40-day temporal windows. It does not supply a canonical scaling equation, identify the target column, or specify the forecast horizon. The generated implementation therefore makes those choices explicit instead of presenting them as recovered paper facts.</p><h3>Scaling five channels independently</h3><p><code>FeaturewiseMinMaxScaler</code> treats Open, High, Low, Close, and Volume as five separate numeric channels. It stores one minimum and maximum for each channel, so the much larger numerical scale of Volume does not determine the scaling of price features. Its final feature axis must have length five, while arbitrary leading dimensions are allowed. Thus it can process both a raw table shaped <code>num_days &#215; 5</code> and model inputs shaped <code>samples &#215; window_length &#215; 5</code>.</p><p>The conventional feature-wise min-max mapping used by the generated scaler is an implementation decision because no canonical equation was supplied. The implementation also chooses to fit statistics on training observations only. This is a leakage-prevention policy: test observations are transformed using training metadata, rather than helping define that metadata. The paper states the target interval but does not state which partition fits the normalization statistics, so this choice must remain visible in provenance.</p><p>A focused part of the generated scaler shows the fitting contract:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;08b437de-8d8b-4cd2-bb3b-2387c8cce6fc&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">scaler = FeaturewiseMinMaxScaler(feature_range=feature_range)
scaler.fit(training_values)
return scaler</code></pre></div><p>Here, <code>training_values</code> is expected to contain only the earlier chronological observations selected by the caller. The scaler records <code>feature_min_</code>, <code>feature_max_</code>, and feature-specific scale metadata. Constant training features are mapped to the lower bound and restored to their fitted constant during inverse transformation, avoiding division by zero. Values outside the fitted extrema are not clipped by <code>transform</code>; the generated code documents this as an intentional choice so extrapolation remains visible.</p><p>The <code>fit</code>, <code>transform</code>, and <code>inverse_transform</code> methods preserve the input shape. That matters later when a prediction is converted back from normalized units to original price units. If evaluation is performed on original prices, the target column must be identified so that the corresponding feature scaling metadata can be applied.</p><h3>Target and horizon are explicit choices</h3><p>The paper describes stock-price prediction and lists OHLCV predictors, but it does not say whether the target is Close, adjusted Close, another price field, or a multivariate output. It also does not specify how far into the future the target lies. The generated windowing API therefore requires both <code>target_index</code> and <code>forecast_horizon</code>.</p><p><code>window_length</code> describes how many past observations enter one model input. For paper-oriented runs, it is treated as either 20 or 40 trading days. <code>forecast_horizon</code> is different: it describes the offset from the end of the input window to the target observation. The 20-day and 40-day values should not be interpreted as prediction horizons merely because they are temporal quantities.</p><p>For a start position <code>s</code>, the generated implementation uses the input rows from <code>s</code> through the end of the selected window, then takes the target at the configured future position. With <code>window_length=3</code> and <code>forecast_horizon=1</code>, the first input contains rows 0, 1, and 2, and its target is row 3. A horizon of 2 would instead select row 4 for that same first input.</p><p>The key implementation fragment is copied from <code>make_sliding_windows</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;965d4edb-2de2-4d83-97a3-c10d2ad50a1f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">sample_count = day_count - window_length - forecast_horizon + 1
if sample_count &lt;= 0:
    raise ValueError(
        "insufficient observations for the requested window and horizon: "
        f"num_days={day_count}, window_length={window_length}, "
        f"forecast_horizon={forecast_horizon}"
    )

inputs = np.empty((sample_count, window_length, _FEATURE_COUNT), dtype=np.float64)
targets = np.empty(sample_count, dtype=np.float64)

target_timestamps = None if timestamp_array is None else np.empty(sample_count, dtype=timestamp_array.dtype)

for start in range(sample_count):
    end = start + window_length
    target_position = end + forecast_horizon - 1
    inputs[start] = array[start:end]
    targets[start] = array[target_position, target_index]
    if target_timestamps is not None:
        target_timestamps[start] = timestamp_array[target_position]</code></pre></div><p>The resulting <code>WindowedDataset</code> records <code>inputs</code> with shape <code>samples &#215; window_length &#215; 5</code> and scalar <code>targets</code> with shape <code>samples</code>. It also records the selected target column, horizon, and optional target timestamps. The five channels remain in canonical Open, High, Low, Close, Volume order; window construction does not reorder or mix them.</p><p>The following focused example is the same alignment pattern used by the generated contract tests. It uses synthetic arrays only and does not represent Hang Seng Index data:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f9803832-6a67-427e-80bb-bdb23840e87a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">dataset = make_sliding_windows(
    values,
    target_index=3,
    window_length=3,
    forecast_horizon=2,
    timestamps=timestamps,
)

assert dataset.inputs.shape == (4, 3, 5)
assert dataset.targets.shape == (4,)
np.testing.assert_array_equal(dataset.inputs[0], values[0:3])
np.testing.assert_array_equal(dataset.targets, values[4:8, 3])</code></pre></div><p>This example uses target index 3, which corresponds to Close in the canonical order. That index is an example configuration choice, not evidence that the paper definitively forecasts Close.</p><h3>Chronological splitting</h3><p>After windows and targets have been aligned, <code>chronological_split</code> divides complete supervised samples without shuffling. Its boundary is the floor of <code>sample_count * train_fraction</code>. The first partition contains earlier target timestamps, and the second contains later ones. The function rejects a split that would leave either partition empty.</p><p>The generated test expresses the central ordering invariant:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4f80eee3-6307-4114-8c31-ac1c5f7d0aa2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">train, test = chronological_split(dataset, train_fraction=0.6)

expected_train_count = int(dataset.sample_count * 0.6)
assert train.sample_count == expected_train_count
assert test.sample_count == dataset.sample_count - expected_train_count
assert train.sample_count &gt; 0
assert test.sample_count &gt; 0
assert train.timestamps is not None
assert test.timestamps is not None
assert int(train.timestamps[-1]) &lt; int(test.timestamps[0])</code></pre></div><p>These are contract tests supplied in the generated repository, but they were not executed under the authoritative run policy. They describe intended behavior rather than reporting a completed verification.</p><p>The paper reports an 80%/20% chronological partition and gives sample counts of 4,678 training samples and 1,170 test samples. Those counts are useful reproduction targets, but they do not uniquely reveal the raw date range, the number of overlapping windows, or whether the 20-day and 40-day datasets were constructed separately.</p><h3>How the prepared-data pipeline orders its work</h3><p><code>prepare_data</code> records the selected target, horizon, window length, split ratio, scaling range, and cleaning policies. Its relevant flow is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;0be6dce0-5afe-4bd7-9006-6ac8f839fe08&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">observation_split = _observation_split_index(
    raw_values.shape[0], data_config.split_ratio
)

scaler = fit_scaler_on_training_observations(
    raw_values=raw_values,
    split_index=observation_split,
    feature_range=data_config.scaling_range,
)
scaled_values = scaler.transform(raw_values)
dataset = make_sliding_windows(
    values=scaled_values,
    target_index=_target_index(data_config.target_column),
    window_length=data_config.window_length,
    forecast_horizon=data_config.forecast_horizon,
    timestamps=timestamps,
)
train_dataset, test_dataset = chronological_split(
    dataset,
    train_fraction=data_config.split_ratio,
)</code></pre></div><p>Notice the distinction between the observation boundary used to fit the scaler and the later sample split. The generated pipeline first determines an observation-level boundary, fits the scaler on the leading observations, transforms the ordered series, constructs windows, and then partitions the completed samples. This is the package's documented leakage-avoidance design, not a uniquely specified procedure from the paper. Because overlapping windows can span an observation boundary, a stronger reproduction should confirm the intended split order from the original experiment details.</p><p>The pipeline also records <code>scaler_fit_policy</code> as <code>training_observations_only</code>, along with the target column, horizon, observation counts, window sample counts, and scaling range. These records make it possible to replace an assumption later without changing the public interfaces.</p><h3>Returning to original units</h3><p>Normalized inputs are useful for model training, but reported price errors may need to be calculated in original units. <code>FeaturewiseMinMaxScaler.inverse_transform</code> restores all five channels, while a target-specific evaluation adapter can select the configured target feature. This distinction is important because the paper does not specify whether its metrics were calculated on normalized values or inverse-transformed prices.</p><p>The generated offline demonstration follows the same policy on synthetic data: it fits the scaler on an earlier chronological portion, transforms the complete ordered sequence, and constructs 20-day windows. The demonstration is useful for checking shape contracts conceptually, but it is not HSI data and cannot establish the paper's reported performance.</p><p>In summary, the preprocessing contract is precise about shapes and ordering but deliberately honest about unresolved semantics. The implementation preserves five feature channels, treats 20 and 40 as input lengths, aligns each window with an explicitly configured future target, and prevents test observations from fitting the scaler. The target column, forecast horizon, split order, and normalization convention remain choices that must be documented for any reproduction run.</p><h2>The feature-wise compositional recurrent model</h2><p>How can five related market signals contribute to one forecast without being mixed too early? The paper&#8217;s answer is feature-wise composition: Open, High, Low, Close, and Volume each travel through an independent recurrent pathway. Their learned summaries are then placed side by side and passed to a prediction head. This is not five separate forecasting models. The final forecast is multivariate because the five representations are fused before prediction.</p><p>The paper supports this overall structure, but it does not uniquely specify the recurrent widths, exact layer counts, representation extraction rule, dense-layer widths, or activation functions. The generated package therefore implements these details through <code>ModelConfig</code>. The code is a configurable reproduction scaffold, not a verified reconstruction of every hidden architecture choice.</p><h3>Tensor flow through five independent streams</h3><p>The public model input, <code>x</code>, has shape <code>batch &#215; time_steps &#215; 5</code>. The final dimension follows the canonical order <code>Open</code>, <code>High</code>, <code>Low</code>, <code>Close</code>, <code>Volume</code>. A 20-day run therefore uses <code>batch &#215; 20 &#215; 5</code>; a 40-day run uses <code>batch &#215; 40 &#215; 5</code>.</p><p><code>FeatureWiseEncoder</code> splits that final dimension into exactly five tensors. Each slice has shape <code>batch &#215; time_steps &#215; 1</code>, so one recurrent stack receives one feature channel rather than the complete OHLCV vector. The five encoders are separate module instances, which means their parameters are not shared.</p><p>The following excerpt shows the stream-level contract. Notice that the recurrent stack is configured with <code>return_sequences=False</code>; this local choice selects the final time-step representation for fusion. The paper does not say whether the final hidden state, a complete sequence, or another aggregation was used.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;fef49559-a0aa-4f44-90f3-c884d74ee619&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class FeatureStreamEncoder(nn.Module):
    """Encode one OHLCV channel with its own recurrent stack.

    The input contract is ``(batch, time_steps, 1)`` and the output contract is
    ``(batch, hidden_size)``.  A separate instance is created for every OHLCV
    channel, so parameters are not shared between feature streams.
    """

    def __init__(
        self,
        recurrent_family: RecurrentFamily | str,
        hidden_size: int,
        layers: int,
        dropout_rate: float = 0.0,
        feature_name: str | None = None,
    ) -&gt; None:
        super().__init__()
        if isinstance(hidden_size, bool) or not isinstance(hidden_size, int) or hidden_size &lt;= 0:
            raise ValueError(f"hidden_size must be a positive integer; received {hidden_size!r}")
        if isinstance(layers, bool) or not isinstance(layers, int) or layers &lt;= 0:
            raise ValueError(f"layers must be a positive integer; received {layers!r}")
        if feature_name is not None and feature_name not in CANONICAL_OHLCV_COLUMNS:
            raise ValueError(
                f"feature_name must be one of {CANONICAL_OHLCV_COLUMNS!r}; received {feature_name!r}"
            )

        self.feature_name = feature_name
        self.input_size = 1
        self.hidden_size = hidden_size
        self.layer_count = layers
        self.recurrent_family = RecurrentFamily.coerce(recurrent_family)
        self.recurrent = build_recurrent_stack(
            family=self.recurrent_family,
            input_size=self.input_size,
            hidden_size=hidden_size,
            layers=layers,
            return_sequences=False,
        )
        self.dropout = DualDropout(stream_rate=dropout_rate, fusion_rate=0.0)</code></pre></div><p>The important invariant is <code>input_size=1</code>: every stream receives one feature channel. After recurrent processing, each stream produces a fixed-width tensor with shape <code>batch &#215; hidden_size</code>. Stream dropout is then applied. Dropout is a training-time regularizer: it masks values while the model is training and is disabled when the module is in evaluation mode.</p><p><code>FeatureWiseEncoder.forward</code> preserves the channel order with a one-element split along the final axis:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;bac1d6c0-8d48-4bc9-877d-85c46b578bd5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        # Each slice retains its one-channel axis: (batch, time_steps, 1).
        feature_slices = torch.split(x, split_size_or_sections=1, dim=-1)
        if len(feature_slices) != 5:
            raise RuntimeError(f"expected five feature slices, received {len(feature_slices)}")

        encoded_streams = tuple(
            encoder(feature_slice)
            for encoder, feature_slice in zip(self.encoders, feature_slices)
        )</code></pre></div><p>This code does not perform early averaging, summation, or mixing. The five encoded outputs remain separate until the fusion module receives them.</p><h3>One interface for LSTM, GRU, and SRU</h3><p>A recurrent layer processes a sequence while maintaining a learned state that carries information across time steps. LSTM and GRU are the two standard recurrent families exposed by the generated wrapper. The paper also evaluates SRU, but the supplied paper context does not identify an SRU variant or Python API.</p><p><code>RecurrentFamily</code> provides a common selector, and <code>build_recurrent_stack</code> returns a <code>RecurrentStack</code> with the same input and output contract for each family. For a stream, the input is <code>batch &#215; time_steps &#215; 1</code>. With final-time-step reduction, the output is <code>batch &#215; hidden_size</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a02e77ee-3d4f-4fea-9aae-d7c0d893a6f6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class RecurrentFamily(str, Enum):
    """Supported recurrent-family selectors.

    ``SRU`` is included because it is one of the families evaluated by the
    paper.  The paper does not identify an SRU variant or Python API, so the
    local adapter below is an explicit implementation choice rather than a
    claim of exact SRU reproduction.
    """

    LSTM = "LSTM"
    GRU = "GRU"
    SRU = "SRU"</code></pre></div><p>For LSTM and GRU, the wrapper constructs batch-first PyTorch stacks. The local <code>SRUAdapter</code> is different: it is a dependency-free compatibility adapter supplied by the generated implementation. Its documentation explicitly says that it is not a canonical SRU implementation. Consequently, selecting <code>SRU</code> makes the interface available for experimentation, but it must not be described as proof of paper-equivalent SRU behavior.</p><p>The wrapper also validates rank, feature width, positive time steps, and floating-point input. Invalid inputs fail explicitly rather than being silently reshaped. This matters because changing <code>batch &#215; time_steps &#215; features</code> to another ordering would change the meaning of the recurrent computation.</p><h3>Dropout and concatenation fusion</h3><p>The paper describes dual dropout: one stage after feature-specific recurrent processing and another stage around feature fusion. The first stage is clear enough to place after each stream encoder. The second stage is not: &#8220;around fusion&#8221; does not establish whether it is before concatenation, after concatenation, or both. <code>DualDropout</code> and <code>ConcatenationFusion</code> expose this as <code>fusion_placement</code>.</p><p>Concatenation means placing the representations end to end along the final dimension. If each of the five streams has width <code>hidden_size</code>, the fused representation has width <code>5 &#215; hidden_size</code>. The operation preserves the batch dimension and does not combine channels by summation.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;1f183738-ffee-494e-9507-89af8544e0bc&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        # The paper's stated fusion operation is concatenation along the
        # representation dimension; stream representations are never summed.
        fused = torch.cat(processed_streams, dim=-1)

        if self.fusion_placement in {"after_concatenation", "both"}:
            fused = self.fusion_dropout.apply_fusion(fused)
        return fused</code></pre></div><p><code>validate_stream_shapes</code> requires exactly five floating-point tensors, matching leading dimensions, a positive representation width, the same dtype, and the same device. A mismatch such as one stream having a different batch size is rejected before concatenation. The default generated placement applies fusion dropout after concatenation, but that default is an implementation decision rather than a recovered paper setting.</p><h3>Optional post-fusion processing and prediction</h3><p>After fusion, <code>PostFusionHead</code> can either send the fused vector directly to dense layers or apply additional recurrent processing first. The paper allows post-fusion recurrent layers but does not specify how a fused vector should be treated as a sequence. The generated head makes one explicit choice: it adds a single time step, giving a tensor shaped <code>batch &#215; 1 &#215; fused_features</code>, and then uses the final representation from that one-step recurrent stack.</p><p>Dense layers then produce the configured target output. The generated implementation uses ReLU between configurable dense layers, followed by a final linear layer. Both the widths and the activation choice are local implementation decisions. The output shape is <code>batch &#215; output_dimension</code>, where <code>output_dimension</code> must agree with the separately configured target definition.</p><h3>Assembling the complete model</h3><p><code>CompositionalRNN</code> connects the stages in order: validate the input, encode five streams, concatenate them, apply the optional post-fusion head, and validate the prediction shape. Its central forward path is deliberately explicit:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e235d89c-941b-461d-888e-cf0edec5a43b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        # Feature-wise recurrent encoding: five tensors of shape
        # (batch, stream_units), preserving the canonical OHLCV order.
        encoded_streams = self.feature_encoder(x)
        if len(encoded_streams) != self.input_feature_count:
            raise RuntimeError(
                f"expected {self.input_feature_count} encoded streams, "
                f"received {len(encoded_streams)}"
            )

        # Paper architecture stage: fuse representations by concatenation.
        fused = self.fusion(encoded_streams)
        expected_fused_shape = (x.shape[0], self.fused_representation_size)
        if tuple(fused.shape) != expected_fused_shape:
            raise RuntimeError(
                "fusion returned an unexpected shape: "
                f"received {tuple(fused.shape)}, expected {expected_fused_shape}"
            )

        # Optional post-fusion recurrence and dense prediction are delegated to
        # the head; its recurrent representation is a configurable choice.
        prediction = self.post_fusion_head(fused)</code></pre></div><p>The model boundary requires exactly five channels and a positive batch and time dimension. It returns <code>batch &#215; output_dimension</code>; an unexpected output shape raises an error. These checks enforce the paper-supported architecture without pretending that unspecified widths or layer counts have been recovered.</p><h3>Focused configuration example</h3><p>The following excerpt is adapted exactly from the generated model-contract test. It demonstrates a small local configuration with one recurrent layer, eight stream units, one dense layer, and one output. Those values are demonstration settings, not the paper&#8217;s reported best architecture. The <code>20</code> in the input shape represents the paper-oriented window choice, while the target semantics remain configured elsewhere.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;471db507-b274-4515-b92f-ad902ffb4612&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def _make_model_config(
    family: RecurrentFamily | str = RecurrentFamily.LSTM,
    *,
    stream_dropout: float = 0.0,
    fusion_dropout: float = 0.0,
) -&gt; ModelConfig:
    """Build a small deterministic configuration for contract tests.

    These tests verify structural invariants of the implementation rather than
    the paper's unresolved architecture details or opaque benchmark labels.
    """
    return ModelConfig(
        recurrent_family=family,
        stream_layers=1,
        stream_units=8,
        stream_dropout=stream_dropout,
        fusion_dropout=fusion_dropout,
        fusion_placement="after_concatenation",
        post_fusion_layers=0,
        post_fusion_units=8,
        dense_layers=(8,),
        output_dimension=1,
    )</code></pre></div><p>A caller would construct the model and prepare an input with shape <code>batch &#215; 20 &#215; 5</code> as follows:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7d4b44fd-231d-44e8-b51e-55542f5ae465&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">model = build_compositional_model(_make_model_config())
inputs = _make_inputs(batch_size=4, time_steps=20)

predictions = model(inputs)</code></pre></div><p>The intended contract is <code>predictions</code> with shape <code>4 &#215; 1</code> for this configuration. This excerpt was not executed under the run policy, so it demonstrates the interface and expected shapes only; it does not establish runtime correctness or paper-level reproduction.</p><p>The opaque labels in the reported results&#8212;<code>LSTM-GWO (1-1-0-1)</code>, <code>GRU-GWO (2-1-1-1)</code>, and <code>SRU-GWO (2-2-0-0)</code>&#8212;are deliberately not translated into <code>ModelConfig</code> fields. The supplied paper context does not define what those four positions mean. Keeping them as benchmark labels prevents an attractive but unsupported architecture claim.</p><p>The generated test plan reinforces structural invariants such as five independent streams, additive fusion width, common family dispatch, and training-only dropout. Those tests were not executed: local static verification, semantic verification, and code execution were disabled or skipped by policy.</p><h2>Training one candidate without hiding missing specifications</h2><p>What does it mean to train one candidate model in this reproduction? The candidate receives batches of chronological windows, produces predictions, computes a selected scalar loss, and updates its trainable parameters with gradients. After each epoch, a separate validation pass records loss with dropout disabled and without changing the parameters. This inner training loop is distinct from Random Search and Grey Wolf Optimizer (GWO), which choose different candidate configurations outside the loop.</p><h3>Paper facts and implementation choices</h3><p>The paper says that recurrent models are trained and compared using validation performance, and it names learning rate, batch size, and training epochs among the optimized hyperparameters. However, the supplied paper context does not specify the training loss, base gradient optimizer, initialization, learning-rate schedule, random seed, batch-shuffling policy, early-stopping rule, or checkpoint-selection rule. It also does not clearly define how the reported 80%/20% partition relates to the validation loss discussed in the method.</p><p>The generated code makes these gaps visible. <code>TrainingConfig</code> supplies the local choices, while <code>ForecastLoss</code> supports explicit <code>mse</code>, <code>mae</code>, and <code>huber</code> objectives. These are available implementation options, not claims about which loss reproduced the paper. Likewise, <code>make_optimizer</code> supports explicit gradient optimizers for a candidate; this inner optimizer should not be confused with RS or GWO, which operate at the outer configuration-search level.</p><p>The generated pipeline uses the prepared test partition as the validation loader for a one-candidate run. Its source documentation explicitly calls this an implementation decision because the paper does not define a separate validation construction. Consequently, this scaffold should not be described as recovering the paper&#8217;s exact validation protocol.</p><h3>Validating predictions and targets</h3><p>A forecasting loss is meaningful only when each prediction is paired with its corresponding target. <code>validate_loss_inputs</code> requires PyTorch tensors with floating-point dtypes, identical shapes, at least one element, and finite values. A mismatch, empty batch, integer tensor, or non-finite value raises an error rather than allowing silent broadcasting or an invalid objective.</p><p>The public loss interface is <code>ForecastLoss</code>. Its output is a scalar tensor, so it can participate in backpropagation. The generated implementation uses mean reduction for each supported local objective. The following excerpt shows the explicit dispatch; it is copied from <code>src/compositional_rnn_stock/losses.py</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c1584044-f9a0-42f6-a91c-fcd7ef52a886&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        if self.name == "mse":
            loss = F.mse_loss(predictions, targets, reduction="mean")
        elif self.name == "mae":
            loss = F.l1_loss(predictions, targets, reduction="mean")
        else:
            loss = F.smooth_l1_loss(
                predictions,
                targets,
                beta=self.huber_delta,
                reduction="mean",
            )</code></pre></div><p>The inputs have the model&#8217;s configured target shape. In the scalar-target pipeline, the model returns <code>batch &#215; 1</code>, while window datasets expose targets as <code>batch</code>. The training loop performs only this unambiguous single-output reshape; other mismatches are rejected. That safeguard matters because accidental broadcasting could produce a finite-looking loss with incorrect semantics.</p><h3>One epoch: update in training mode, evaluate in validation mode</h3><p>During training, <code>train_one_candidate</code> puts the model in training mode, reads a batch shaped <code>batch &#215; time_steps &#215; 5</code>, computes predictions, aligns the targets, backpropagates the scalar loss, and calls the selected optimizer. The five-channel input contract is checked before the model is used. A training loader with no batches or non-finite values is rejected.</p><p>The core update sequence is shown below. This excerpt is copied from <code>src/compositional_rnn_stock/training/loops.py</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;14f54c2f-f599-4ec3-a981-29bf11f25737&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        for batch in train_loader:
            inputs, targets = _prepare_batch(batch, device, "training")
            optimizer.zero_grad(set_to_none=True)
            predictions = model(inputs)
            targets = _align_targets(predictions, targets)
            loss = loss_fn(predictions, targets)
            loss.backward()
            optimizer.step()</code></pre></div><p>After the training batches, <code>evaluate_loss</code> switches the model to evaluation mode and wraps inference in <code>torch.no_grad()</code>. Evaluation mode is important for the compositional model because dropout must be disabled when validation predictions are measured. The function restores the model&#8217;s previous training state afterward, computes a value-weighted mean over the validation targets, and rejects an empty or non-finite result.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;9b2709bc-bcf3-48d8-981a-dccbe4af6ca4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    was_training = model.training
    model.eval()
    total_loss = 0.0
    total_values = 0
    try:
        with torch.no_grad():
            for batch in loader:
                inputs, targets = _prepare_batch(batch, device, "validation")
                predictions = model(inputs)
                targets = _align_targets(predictions, targets)
                loss = loss_fn(predictions, targets)
                value_count = int(targets.numel())
                total_loss += float(loss.detach().cpu().item()) * value_count
                total_values += value_count
    finally:
        model.train(was_training)</code></pre></div><p>This separation is the main training invariant: validation loss is observed, not optimized directly. The validation loader is never passed to <code>backward()</code> or <code>optimizer.step()</code> by this workflow. The paper&#8217;s exact split and model-selection policy remain unresolved, so the generated code retains the final state after the configured epochs rather than silently inventing a best-checkpoint rule.</p><h3>Optimizer and reproducibility settings</h3><p><code>make_optimizer</code> constructs the inner gradient optimizer from the configured name and learning rate. The generated implementation accepts <code>adam</code>, <code>adamw</code>, and <code>sgd</code> in this function. It validates that the learning rate is positive and finite, that the model has trainable parameters, and that the optimizer name is supported. The paper does not identify which of these, if any, was used.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;54497fcf-4b0a-44d6-8edc-5340f9ba7e1f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    if normalized == "adam":
        return Adam(parameters, lr=learning_rate)
    if normalized == "adamw":
        return AdamW(parameters, lr=learning_rate)
    return SGD(parameters, lr=learning_rate)</code></pre></div><p><code>set_global_seed</code> seeds supported local Python, NumPy, and PyTorch generators. <code>SeedContext</code> additionally snapshots and restores available random states and can request deterministic PyTorch algorithms. These facilities improve the comparability of local experiments, but the paper does not report its seed or deterministic-backend settings. A seed therefore belongs in provenance, not in a claim that the paper used the same setting.</p><h3>Histories and provenance records</h3><p>Each completed epoch becomes an <code>EpochRecord</code> containing its epoch number, training loss, and validation loss. <code>TrainingHistory</code> preserves insertion order and rejects non-increasing epoch numbers. Its <code>best_epoch</code> helper can identify the earliest record with the smallest available validation loss, but the training loop does not automatically restore that epoch&#8217;s weights. Selecting and restoring a best checkpoint would be an additional explicit implementation decision.</p><p>The outer record is <code>TrainingRun</code>. It links the trained <code>CompositionalRNN</code>, its <code>TrainingHistory</code>, the <code>ModelConfig</code>, the <code>TrainingConfig</code>, and a prepared-data identifier. That identifier records acquisition and preprocessing metadata, window length, forecast horizon, target column, and train/test sample counts. This prevents a local loss curve from becoming detached from choices such as <code>Close</code> versus another target, a 20-day versus 40-day window, or a particular scaling policy.</p><h3>Worked example: defining a local candidate run</h3><p>The following configuration fragment is copied from <code>scripts/train_model.py</code>. It selects <code>Close</code> and a one-step horizon, but those are tutorial-level implementation choices because the paper does not identify the target column or forecast horizon.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;cbf4178e-5fe8-4b68-bcbe-5db024175ad1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    data_config = DataConfig(
        symbol=args.symbol,
        local_path=args.local_path,
        start=args.start,
        end=args.end,
        target_column=args.target_column,
        forecast_horizon=args.forecast_horizon,
        window_length=args.window_length,
        split_ratio=args.split_ratio,
        scaling_range=(0.0, 0.95),
    )</code></pre></div><p>A focused call path inside <code>train_one_candidate</code> then creates the selected loss and optimizer and trains for the configured number of epochs. This excerpt is copied from <code>src/compositional_rnn_stock/training/loops.py</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;66c861ec-4f74-4dad-b373-33530bb939a3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    loss_fn = _loss_for_config(config)
    optimizer = make_optimizer(model, config)
    device = _model_device(model)
    history = TrainingHistory()</code></pre></div><p>At the pipeline level, <code>run_training</code> validates the prepared data and configurations, builds chronological loaders, constructs the model, checks that loader inputs have shape <code>batch &#215; time_steps &#215; 5</code>, and delegates to <code>train_one_candidate</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ca0da32f-1468-4955-bceb-9804fdfbc264&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    trained_model, history = train_one_candidate(
        model=model,
        train_loader=train_loader,
        validation_loader=validation_loader,
        config=training_config,
    )</code></pre></div><p>This call defines one local experiment. It does not establish that the loss, optimizer, architecture settings, validation protocol, or final model state match the paper. It also does not execute here: the authoritative run policy disabled code execution and verification.</p><h3>Why this boundary matters</h3><p>The gradient loop answers, &#8220;How do we fit one selected candidate?&#8221; RS and GWO answer a different question, &#8220;Which candidate settings should we try next?&#8221; Keeping those responsibilities separate allows both search methods to use the same loss, loaders, validation objective, and provenance schema. It also makes missing paper specifications inspectable instead of burying them in defaults.</p><p>For this scaffold, the honest output of training is a locally configured model, an epoch-aligned history, and provenance describing the choices. No training result, convergence behavior, or successful reproduction of the paper&#8217;s reported benchmarks is claimed. Verification was skipped under policy, so the code excerpts describe the generated interfaces and intended control flow rather than a checked execution outcome.</p><h2>Random Search and Grey Wolf Optimizer</h2><p>How should a reproduction choose among many possible recurrent-network recipes? Random Search (RS) samples independent candidates, while Grey Wolf Optimizer (GWO) maintains a population of candidates and moves them using the best-ranked candidates as guides. In both cases, a candidate is more than a model family: it can include recurrent-layer count, recurrent units, dropout rates, learning rate, batch size, and training epochs.</p><p>The paper names both RS and GWO and reports that GWO performed better across the LSTM, GRU, and SRU families. However, it does not specify the search bounds, probability distributions, number of trials, wolf population, iteration count, initialization, update settings, invalid-candidate policy, or randomization controls. The generated package therefore treats these as explicit implementation configuration. No displayed setting in this section should be read as the paper's missing experimental setting.</p><h3>One objective shared by both methods</h3><p>The outer optimizer does not train a model directly. Instead, it proposes a typed <code>CandidateConfig</code>. The shared candidate objective builds the corresponding compositional model, creates or receives chronological training and validation loaders, trains the candidate, and records its best validation loss. Validation loss is used here as a local implementation choice because the paper discusses validation performance but does not define the loss or the exact selection rule.</p><p><code>CandidateResult</code> stores the candidate configuration, objective, metrics, provenance, success status, and any failure information. <code>SearchResult</code> stores the selected result and the complete search history. This separation matters: a failed candidate remains visible in the history rather than being silently treated as the best candidate, and the held-out test set is not needed for hyperparameter selection.</p><p>The objective's public protocol is deliberately small. Its candidate has type <code>CandidateConfig</code>, its data bundle must provide training and validation loaders, and its <code>seed</code> is recorded for reproducibility. The generated implementation explicitly records that test data were not used for selection:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f0774190-756b-4bc0-bbe6-b80faa8dd496&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">provenance: dict[str, Any] = {
    "seed": seed,
    "objective": "minimum_validation_loss",
    "selection_basis": "lowest validation loss recorded during training",
    "test_data_used_for_selection": False,
}</code></pre></div><p>This excerpt comes from <code>src/compositional_rnn_stock/search/objectives.py</code>. The objective returns a <code>CandidateResult</code> after <code>train_one_candidate</code> supplies a <code>TrainingHistory</code>; the history must contain a validation loss from which the lowest recorded value can be selected. The code does not claim that this objective is the paper's exact loss or checkpoint policy.</p><p>The same selection helper serves both optimizers. Its <code>minimize</code> argument makes the direction explicit, which is important because validation loss is normally minimized, whereas some metrics are maximized:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7aeb42d0-cde7-4a58-a4c7-327d8dc7ec77&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def select_best_candidate(
    results: Sequence[CandidateResult],
    minimize: bool,
) -&gt; CandidateResult:
    """Select a successful candidate using an explicit objective direction."""
    if not isinstance(minimize, bool):
        raise TypeError(f"minimize must be boolean; received {type(minimize).__name__}")
    if not isinstance(results, Sequence):
        raise TypeError("results must be a sequence of CandidateResult instances")

    validated: list[CandidateResult] = []
    for index, result in enumerate(results):
        if not isinstance(result, CandidateResult):
            raise TypeError(
                f"results[{index}] must be a CandidateResult; "
                f"received {type(result).__name__}"
            )
        if result.success:
            if result.objective is None:
                raise ValueError(
                    f"successful results[{index}] must contain an objective"
                )
            validated.append(result)

    if not validated:
        failure_count = sum(1 for result in results if not result.success)
        raise ValueError(
            "cannot select a best candidate: no successful candidate results "
            f"were available ({failure_count} recorded failures)"
        )

    if minimize:
        return min(validated, key=lambda result: float(result.objective))
    return max(validated, key=lambda result: float(result.objective))</code></pre></div><p>The function accepts a sequence of results, rejects malformed successful records, ignores unsuccessful records for selection, and raises an error when no successful candidate exists. That failure behavior is a derived implementation safeguard, not a reported paper procedure.</p><h3>Typed search spaces</h3><p>A search space must represent different kinds of values. Recurrent-layer counts, units, batch sizes, and epochs are integer-valued. Dropout rates and learning rates are continuous. A model family is categorical when one search spans LSTM, GRU, and SRU. <code>ParameterDomain</code> gives each parameter one of these meanings and validates its bounds or allowed values.</p><p>The following excerpt is copied from <code>src/compositional_rnn_stock/search/space.py</code> and shows the domain constructors exposed by the generated package:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a26d15b1-7fa1-423d-9d21-cababbde7fd4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    @classmethod
    def continuous(cls, lower: float, upper: float) -&gt; "ParameterDomain":
        return cls("continuous", lower=lower, upper=upper)

    @classmethod
    def integer(cls, lower: int, upper: int) -&gt; "ParameterDomain":
        return cls("integer", lower=lower, upper=upper)

    @classmethod
    def categorical(cls, values: Sequence[Any]) -&gt; "ParameterDomain":
        return cls("categorical", values=tuple(values))

    @classmethod
    def discrete(cls, values: Sequence[Any]) -&gt; "ParameterDomain":
        return cls("discrete", values=tuple(values))</code></pre></div><p><code>continuous</code> and <code>integer</code> domains use inclusive lower and upper bounds. <code>categorical</code> and <code>discrete</code> domains hold explicit values. The distinction between categorical and discrete values is useful when documenting intent, although both are represented by positions when a vector-based optimizer needs numeric coordinates.</p><p><code>SearchSpace</code> keeps the named domains in a stable order and exposes encoded bounds. <code>CandidateConfig</code> then groups the decoded values into <code>ModelConfig</code>, <code>TrainingConfig</code>, and any explicitly declared extra parameters. <code>candidate_from_mapping</code> validates a named mapping before constructing that typed record. The paper's list of optimized hyperparameters tells us what should be represented, but not the actual ranges; those ranges must be supplied by the caller.</p><h3>Random Search: independent proposals</h3><p>RS draws one value from every domain for each trial. The generated <code>random_search</code> function accepts a <code>SearchSpace</code>, a callable <code>CandidateObjective</code>, and a search configuration containing a seed, trial count, data bundle, and objective direction. Each trial receives a derived candidate seed, and each result is appended to the history, including failures.</p><p>The core loop is implemented as follows in <code>src/compositional_rnn_stock/search/random_search.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;66b94bf9-159d-4999-a821-cc3720224895&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    history: list[CandidateResult] = []
    for trial_index in range(trial_count):
        candidate = sample_random_candidate(space, rng)
        candidate_seed = seed + trial_index
        try:
            result = objective(candidate, data_bundle, candidate_seed)
            if not isinstance(result, CandidateResult):
                raise TypeError(
                    "objective must return CandidateResult; "
                    f"received {type(result).__name__}"
                )
        except Exception as exc:
            result = CandidateResult(
                configuration=candidate,
                objective=None,
                metrics={},
                provenance={
                    "method": "random_search",
                    "trial_index": trial_index,
                    "seed": candidate_seed,
                    "failure_type": type(exc).__name__,
                },
                success=False,
                error=f"{type(exc).__name__}: {exc}",
            )
        else:
            result.provenance.update(
                {
                    "method": "random_search",
                    "trial_index": trial_index,
                    "sampler_seed": seed,
                    "candidate_seed": candidate_seed,
                }
            )
        history.append(result)</code></pre></div><p>Notice the two seeds recorded in successful results: the sampler seed identifies the sequence of sampled candidates, while the candidate seed identifies the training run. This is a useful provenance design for comparing methods locally. It does not establish that the paper used the same seeding scheme.</p><p>A tutorial experiment might declare three trials for a quick demonstration, but that would be a tutorial budget only. The paper's statement that 54 configurations were evaluated must not be reverse-engineered into a presumed RS trial count or into a presumed division between model families and optimizers.</p><h3>GWO: population-guided proposals</h3><p>GWO works with a matrix of wolf positions. A position is a numeric vector whose coordinates correspond to the ordered parameters in <code>SearchSpace</code>. The best three successful candidates are called alpha, beta, and delta in the generated implementation. Their encoded vectors guide a new population. After each update, positions are clipped to the declared bounds and decoded back into valid typed candidates.</p><p>The generated update function accepts <code>positions</code> with shape <code>population &#215; encoded_dimension</code>. Each leader vector has shape <code>encoded_dimension</code>. The update is intentionally documented as a local GWO-style choice because no canonical GWO equations or settings were supplied in the paper:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;50158be2-0bf6-4374-a9b9-bb3d835d11a4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def update_wolf_positions(
    positions: Any,
    alpha: Any,
    beta: Any,
    delta: Any,
    iteration: int,
    total_iterations: int,
    rng: random.Random,
) -&gt; np.ndarray:
    """Return one local GWO-style update for every wolf.

    ``positions`` has shape ``(population, encoded_dimension)`` and each leader
    has shape ``(encoded_dimension,)``.  The implementation uses the usual
    linearly decreasing exploration coefficient and independent random draws.
    These update rules are implementation decisions, not equations supplied by
    the paper.
    """
    current = np.asarray(positions, dtype=float)
    leader_arrays = [np.asarray(value, dtype=float) for value in (alpha, beta, delta)]
    if current.ndim != 2:
        raise ValueError("positions must have shape (population, encoded_dimension)")
    if current.shape[0] &lt; 1:
        raise ValueError("positions must contain at least one wolf")</code></pre></div><p>The remainder of the function validates leader shapes and iteration bounds, generates independent random coefficients, averages the three leader-guided proposals, and returns an array with the same population-by-dimension shape. The linearly decreasing exploration coefficient, random update details, and alpha/beta/delta ranking are implementation decisions rather than recovered paper settings.</p><h3>Decoding mixed parameters safely</h3><p>A continuous vector cannot directly be used as a batch size or a model-family name. <code>CandidateCodec</code> bridges that gap. It encodes typed values into a stable vector ordering and decodes them by clipping to the domain, rounding integer-like coordinates, selecting categorical positions, and constructing validated <code>ModelConfig</code> and <code>TrainingConfig</code> records.</p><p>The public decoding contract is concise:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b232450f-6957-4cbc-9e4e-9841281d06f6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    def decode(self, vector: Sequence[float] | np.ndarray) -&gt; CandidateConfig:
        array = np.asarray(vector, dtype=float)
        if array.ndim != 1 or array.shape[0] != self.dimension:
            raise ValueError(
                f"encoded candidate must have shape ({self.dimension},); received {array.shape}"
            )
        if not np.all(np.isfinite(array)):
            raise ValueError("encoded candidate must contain only finite values")
        values = {
            name: self.space.domains[name].decode_value(float(array[index]))
            for index, name in enumerate(self.space.parameter_names)
        }
        return self._build_candidate(self.space.validate_candidate(values))</code></pre></div><p>For example, a vector coordinate intended for batch sizes <code>(8, 16, 32)</code> is clipped to the valid index range and rounded to one of those positions. That prevents GWO from producing an invalid fractional batch size. It also means that several nearby continuous positions can decode to the same typed candidate; this is an expected consequence of using a continuous optimizer for mixed parameters.</p><h3>Worked configuration example</h3><p>The following excerpt is a small configuration-only example based on the generated search-contract test. It demonstrates integer, continuous, and discrete domains without asserting that these bounds came from the paper:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2288366c-5ce4-4065-bbf5-e4c42793fa4c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">PARAMETER_DOMAINS = {
    "stream_layers": ParameterDomain.integer(1, 3),
    "units": ParameterDomain.integer(8, 32),
    "dropout": ParameterDomain.continuous(0.0, 0.5),
    "learning_rate": ParameterDomain.continuous(1.0e-4, 1.0e-2),
    "batch_size": ParameterDomain.discrete((8, 16, 32)),
    "epochs": ParameterDomain.integer(1, 4),
}

space = SearchSpace(domains=PARAMETER_DOMAINS)
codec = CandidateCodec(space)</code></pre></div><p>Here, <code>stream_layers</code>, <code>units</code>, <code>batch_size</code>, and <code>epochs</code> are discrete choices, while <code>dropout</code> and <code>learning_rate</code> vary continuously. <code>codec</code> can translate a validated <code>CandidateConfig</code> into a numeric vector for GWO and decode a clipped vector back into a typed candidate. The displayed ranges are deliberately modest tutorial settings; they are not paper-reported bounds.</p><p>In a complete local run, the same <code>space</code> and shared objective can be passed to RS or GWO. The pipeline functions <code>run_random_search</code> and <code>run_gwo_search</code> bind both methods to a chronological training/validation split created from the prepared training partition. Their documented contract excludes <code>PreparedData.test</code> from selection, so the held-out test data are reserved for evaluation after a candidate has been chosen.</p><h3>Practical limits of the search reproduction</h3><p>A fair local comparison requires RS and GWO to use the same feature order, scaling policy, candidate-training routine, validation partition, objective, and failure policy. Only the proposal mechanism should differ. The generated pipeline records the validation fraction as a local choice and shares the candidate objective, but this does not recover the paper's undisclosed validation protocol.</p><p>The paper's GWO superiority claim remains a reported paper result, not a result established by this scaffold. In particular, the generated GWO update, population size, iteration count, bounds, initialization, and discrete decoding cannot be labeled as the paper's implementation. The four-part labels in <code>LSTM-GWO (1-1-0-1)</code>, <code>GRU-GWO (2-1-1-1)</code>, and <code>SRU-GWO (2-2-0-0)</code> also remain opaque; the search code does not decode them.</p><p>The related contract tests are intended to check local invariants such as candidate decoding, seeded RS sampling, GWO shape and bound handling, and the common objective interface. They were not executed in this run. Static verification and semantic code verification were also skipped, so the generated files should be reviewed before relying on them operationally. No search, training run, or reproduction result is claimed here.</p><h2>Held-out evaluation, metrics, and residual analysis</h2><p>How do we know whether a forecast is useful? First, each prediction must be paired with the correct held-out observation. Then both arrays must be evaluated under the same scale and metric conventions. The evaluation layer therefore aligns predictions and targets, optionally restores original price units, computes the seven metrics named by the paper, and summarizes residual behavior.</p><p>This section distinguishes three kinds of statements. The paper names the metrics and discusses error distributions, but it does not provide canonical formulas, denominator rules, residual signs, or plotting conventions. The generated implementation supplies explicit conventional choices for those missing details. The resulting values would be local evaluation outputs&#8212;not verified reproductions of the paper's reported results.</p><h3>From model output to aligned evaluation arrays</h3><p>The generated <code>predict_dataset</code> function runs inference in loader order. It switches the model to evaluation mode so dropout is disabled, uses no-gradient inference, collects scalar predictions and targets batch by batch, and restores the model's previous training-mode flag afterward. Its scalar-target contract accepts arrays shaped <code>samples</code> or <code>samples &#215; 1</code>. It does not reorder or silently truncate samples.</p><p>The key alignment operation is exposed separately as <code>align_predictions_and_targets</code>. It rejects unequal sample counts and preserves the order emitted by the loader. This is why held-out loaders should normally use <code>shuffle=False</code>: a shuffled loader can still produce pairs within a batch, but it no longer represents the original chronological output order.</p><p>The following excerpt is copied from <code>predict_dataset</code>. Notice the two independent checks: the model output and the batch targets are normalized to the scalar-target representation, and their batch counts must agree before either is appended.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e2030c2e-efb3-4221-9e5e-8adae15db365&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">                outputs = model(inputs)
                if not isinstance(outputs, Tensor):
                    raise TypeError(
                        f"model output for batch {batch_index} must be a tensor"
                    )
                predictions, _ = _normalise_scalar_targets(outputs, "predictions")
                batch_targets, _ = _normalise_scalar_targets(targets, "targets")
                if predictions.shape[0] != batch_targets.shape[0]:
                    raise ValueError(
                        f"batch {batch_index} prediction and target counts differ: "
                        f"{predictions.shape[0]} != {batch_targets.shape[0]}"
                    )
                prediction_parts.append(predictions)
                target_parts.append(batch_targets)</code></pre></div><p>The function returns predictions in loader order, while the targets are retained internally for alignment validation. The model input remains the package-wide <code>batch &#215; time &#215; 5</code> contract; this evaluation helper specifically supports the configured scalar-output case. A multivariate target would require an explicit extension rather than an implicit interpretation.</p><h3>Choosing the evaluation scale</h3><p>The paper states that inputs are scaled to <code>[0, 0.95]</code>, but it does not say whether its reported metrics were computed in normalized space or after inverse transformation to price units. The generated pipeline exposes both choices through <code>EvaluationConfig</code>.</p><p>In normalized evaluation, <code>actual_targets</code> and <code>predictions</code> remain in scaled space. RMSE then has normalized units. In original-unit evaluation, <code>inverse_transform_target</code> embeds each scalar target into a five-feature row, applies <code>FeaturewiseMinMaxScaler.inverse_transform</code>, and extracts the configured target column. RMSE can then be interpreted in price units. MAPE, RMSPE, and PBIAS remain percentage-valued under the local conventions, while agreement metrics remain unitless.</p><p>The target column itself is also unresolved by the paper. The generated pipeline requires a configured choice such as <code>Close</code>; it does not assume that the paper's phrase &#8220;stock price&#8221; uniquely means Close or adjusted Close. The forecast horizon is likewise recorded separately from the input <code>window_length</code>.</p><p>The following excerpt is copied from <code>inverse_transform_target</code> and shows why the scaler needs a target index even though the prediction contains only one value per sample.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e01e10f3-791f-46d8-8373-83cc93aad2a3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    embedded = np.zeros((flat_values.shape[0], scaler.n_features), dtype=np.float64)
    embedded[:, target_index] = flat_values
    restored = scaler.inverse_transform(embedded)[:, target_index]

    if len(original_shape) == 1:
        return restored
    return restored.reshape(original_shape)</code></pre></div><p>This operation preserves the prediction shape while using the scaler's five feature-specific statistics. It is an implementation mapping, not a formula recovered from the paper. The scaler must already be fitted, and its fitting policy is recorded as training-only in the generated preprocessing pipeline to reduce leakage.</p><h3>The seven named metrics</h3><p>The paper names <code>R&#178;</code>, <code>RMSE</code>, <code>MAPE</code>, <code>RMSPE</code>, <code>PBIAS</code>, Willmott Index, and <code>NSE</code>. Because no canonical equation records were supplied, the generated <code>metrics.py</code> module documents and implements conventional local definitions. The important practical point is consistency: every candidate must use the same scale, alignment rule, sign convention, and zero-denominator policy.</p><ul><li><p><code>R&#178;</code> summarizes explained variation relative to an actual-value mean baseline. The generated function stores the conventional result on the unit-interval scale by default. A display convention can multiply it by 100, which supports the paper's percentage-style reporting. Thus a paper display of <code>99.2427%</code> is distinct from the underlying unit-interval representation <code>0.992427</code>.</p></li><li><p><code>RMSE</code> measures the typical magnitude of prediction error in the units supplied to the function. It is therefore scale-sensitive: normalized inputs yield normalized-unit RMSE, while inverse-transformed prices yield price-unit RMSE.</p></li><li><p><code>MAPE</code> averages absolute relative errors and returns a percentage. It cannot divide by zero actual values, so <code>EvaluationConfig</code> selects whether such observations raise an error, are ignored, or receive a defined local treatment.</p></li><li><p><code>RMSPE</code> is the root-mean-square version of relative error and uses the same zero-actual policy.</p></li><li><p><code>PBIAS</code> summarizes signed aggregate bias. The generated default residual direction is <code>actual_minus_predicted</code>, so positive PBIAS indicates underprediction under that convention.</p></li><li><p>Willmott Index measures agreement using the generated module's selected standard convention. Its denominator can be undefined for degenerate data, in which case the local implementation returns <code>NaN</code>.</p></li><li><p><code>NSE</code>, or Nash&#8211;Sutcliffe Efficiency, compares squared forecast error with variation around the actual-series mean. Like <code>R&#178;</code>, it is undefined when the actual series has no variation.</p></li></ul><p>The module's report assembler keeps these choices together. The following excerpt is copied from <code>compute_metric_report</code>; it shows that one configuration controls the conventions used by the individual metric functions.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;59a7d27f-00b6-494b-8b1e-48d86a5a7e18&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    _validate_zero_policy(str(zero_policy))
    return {
        "r_squared": r_squared(actual_vector, predicted_vector, str(r2_convention)),
        "rmse": rmse(actual_vector, predicted_vector),
        "mape": mape(actual_vector, predicted_vector, str(zero_policy)),
        "rmspe": rmspe(actual_vector, predicted_vector, str(zero_policy)),
        "pbias": pbias(actual_vector, predicted_vector, str(pbias_convention)),
        "willmott_index": willmott_index(
            actual_vector,
            predicted_vector,
            str(willmott_convention),
        ),
        "nash_sutcliffe_efficiency": nash_sutcliffe_efficiency(actual_vector, predicted_vector),
    }</code></pre></div><p>Before reaching this block, <code>_aligned_arrays</code> converts supported array-like inputs to finite one-dimensional vectors and rejects mismatched shapes. This failure behavior matters: a metric calculated from mispaired samples can look numerically plausible while describing the wrong forecast errors.</p><h3>A small illustrative interface example</h3><p>The following excerpt is adapted directly from the generated metric test's synthetic setup. It demonstrates the intended interface with short local arrays; it is not HSI data, and no numerical result from this example should be compared with the paper's benchmarks. The test uses <code>EvaluationConfig()</code> to select the generated defaults rather than claiming that those defaults were specified by the paper.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;204f3d0c-e283-4ffb-a8fe-d3933470a929&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    actual = np.array([100.0, 105.0, 110.0, 115.0], dtype=float)
    predicted = np.array([99.0, 106.0, 109.0, 116.0], dtype=float)

    # Defaults are deliberately supplied by EvaluationConfig rather than
    # inferred from the paper, whose metric conventions are unspecified.
    report = compute_metric_report(actual, predicted, EvaluationConfig())</code></pre></div><p>A corresponding residual summary can be created from the same aligned arrays. In the generated implementation, <code>compute_residuals(actual, predicted)</code> returns <code>actual - predicted</code> with the same shape, and <code>summarize_residuals</code> flattens the result before calculating its count, mean, population standard deviation, standardized skewness, and excess kurtosis. Fewer than three observations, or zero variance, leave skewness undefined; fewer than four observations, or zero variance, leave excess kurtosis undefined. These estimator and insufficiency rules are implementation decisions because the paper only mentions mean, standard deviation, skewness, and kurtosis without defining their conventions.</p><h3>Residuals and visual analysis</h3><p>Residuals make systematic behavior easier to inspect than a single score. Under the generated <code>actual_minus_predicted</code> convention, positive residuals indicate underprediction and negative residuals indicate overprediction. <code>compare_error_distributions</code> creates a stable table for named residual groups, preserving empty groups with an explicit count and undefined summary values rather than inventing observations.</p><p>The visualization adapters implement the analysis types mentioned by the paper:</p><ul><li><p><code>plot_error_boxplots</code> shows grouped residual spread and central tendency.</p></li><li><p><code>plot_error_violins</code> shows distribution shape.</p></li><li><p><code>plot_metric_comparison</code> gives each metric its own panel so percentage and unitless measures are not placed on one misleading axis.</p></li><li><p><code>plot_pbias_radar</code> displays signed PBIAS values using a locally selected radial convention.</p></li><li><p><code>plot_taylor_summary</code> uses a locally selected correlation-angle and standard-deviation representation.</p></li></ul><p>These functions operate on local predictions and reports. They do not contain the paper's plotted points, styling, axes, or source predictions. Therefore they can support analogous analysis, but they cannot claim exact reproduction of the paper's figures.</p><h3>Evaluate only after model selection</h3><p>The evaluation pipeline enforces the intended experiment order conceptually: train candidates and compare them using the configured validation objective, select a candidate, and only then read the chronological held-out partition for final metrics. <code>evaluate_training_run</code> evaluates a retained trained model. <code>evaluate_search_result</code> requires that a selected trained model be retained; it does not silently retrain from an incomplete search record.</p><p>The resulting <code>EvaluationRun</code> stores predictions, aligned actual targets, residuals, the metric report, and provenance. Provenance includes the evaluation scale, target column, forecast horizon, window length, metric conventions, residual sign, held-out sample count, and preprocessing metadata. This makes a future comparison auditable: a difference in RMSE can be investigated as a possible scale, target, or convention difference rather than treated as unexplained model behavior.</p><p>The paper's seven metrics and its discussion of residual distributions are therefore represented, but not overclaimed. The formulas and conventions are local implementation choices, the target and evaluation scale remain configurable, and no execution or verification occurred in this run. Any future local report must be labeled as a computed result and kept separate from the paper's reported benchmark records.</p><h2>Reported LSTM-GWO, GRU-GWO, and SRU-GWO benchmarks</h2><p>How should you use the paper&#8217;s reported numbers when the original experiment cannot yet be reconstructed exactly? Treat them as reference records, not as expected outputs that automatically validate a new run. The generated package stores the reported values separately from locally computed predictions and metrics, so a later experiment can be compared with the paper without confusing the two.</p><h3>What the paper reports</h3><p>The paper identifies its strongest reported configuration as LSTM-GWO <code>(1-1-0-1)</code>. The four-part label is opaque: the supplied paper context does not define whether its fields represent layer counts, dropout settings, dense layers, or another encoding. The label must therefore remain a benchmark identifier rather than being translated into <code>ModelConfig</code> values.</p><p>The reported LSTM-GWO metrics are:</p><ul><li><p>R&#178;: <code>99.2427%</code></p></li><li><p>RMSE: <code>339.3902</code></p></li><li><p>MAPE: <code>1.1721%</code></p></li><li><p>RMSPE: <code>1.6221%</code></p></li><li><p>PBIAS: <code>0.0523</code></p></li><li><p>Willmott Index: <code>0.9981</code></p></li><li><p>NSE: <code>0.9924</code></p></li></ul><p>The paper also reports the following best labels and associated values:</p><ul><li><p><strong>GRU-GWO `(2-1-1-1)`:</strong> R&#178; <code>99.2322%</code>, RMSE <code>341.7225</code>, MAPE <code>1.1821%</code>, and PBIAS <code>&#8722;0.1357</code>.</p></li><li><p><strong>SRU-GWO `(2-2-0-0)`:</strong> R&#178; <code>99.2009%</code>, RMSE <code>348.6384</code>, and MAPE <code>1.2080%</code>.</p></li></ul><p>The remaining GRU and SRU metric associations are not clear in the supplied extraction. They are intentionally left unavailable rather than inferred from nearby text or reconstructed tables. Likewise, the configuration labels remain undecoded.</p><p>These values also use mixed display conventions. R&#178; is shown by the paper as a percentage, whereas Willmott Index and NSE are shown on a unit-interval scale. A local metrics implementation may store R&#178; internally as <code>0.992427</code> and display it as <code>99.2427%</code>, but that conversion must be recorded explicitly before comparing values.</p><h3>How the generated code preserves the evidence</h3><p><code>PaperBenchmark</code> is an immutable record containing the recurrent family, optimizer, opaque configuration label, metric values, units, and provenance. Its metric schema always contains the seven named metrics, but unavailable values are represented by <code>None</code>. The following excerpt is copied from <code>reported_benchmarks</code> in <code>src/compositional_rnn_stock/evaluation/benchmarks.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;bbd60fe0-c150-4e4f-8747-f80de8076ad0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def reported_benchmarks() -&gt; tuple[PaperBenchmark, ...]:
    """Return the three benchmark records supplied in the paper context.

    Missing GRU and SRU associations remain ``None`` rather than being inferred from
    other reported values. R&#178; is stored in the paper's percentage display convention.
    """
    return (
        PaperBenchmark(
            family="LSTM",
            optimizer="GWO",
            configuration_label="1-1-0-1",
            metrics=_benchmark_metrics(
                r_squared=99.2427,
                rmse=339.3902,
                mape=1.1721,
                rmspe=1.6221,
                pbias=0.0523,
                willmott_index=0.9981,
                nash_sutcliffe_efficiency=0.9924,
            ),
            metric_units=dict(_METRIC_UNITS),
        ),</code></pre></div><p>The function returns records marked as supplied paper targets. It does not train a model, generate predictions, or assert that any local candidate achieved these values. The LSTM record is complete, while the later GRU and SRU records retain <code>None</code> for unavailable associations.</p><p>The generated contract test makes this distinction explicit when it retrieves the records:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;bd70006e-ee4f-47f1-912f-13c2ea31c23b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">benchmarks = reported_benchmarks()

assert [(item.family, item.optimizer, item.configuration_label) for item in benchmarks] == [
    ("LSTM", "GWO", "1-1-0-1"),
    ("GRU", "GWO", "2-1-1-1"),
    ("SRU", "GWO", "2-2-0-0"),
]</code></pre></div><p>This is a planned test file, not evidence of an executed check. Under the authoritative run policy, code execution and verification were disabled.</p><h3>Comparing a future local report</h3><p>The <code>compare_to_report</code> function accepts a local metric mapping and one <code>PaperBenchmark</code>. It produces rows only when both sides have a value. Differences are expressed as local minus reported in the benchmark&#8217;s displayed units. Thus a missing GRU RMSPE value does not become a guessed comparison row, and a local unit-interval R&#178; can be converted to the paper&#8217;s percentage convention when the local report has not declared another unit.</p><p>A comparison is meaningful only when the upstream experiment is also aligned. The provenance should identify the target column, forecast horizon, scaling range and fitting policy, window length, chronological split, recurrent-family configuration, seed, training settings, search settings, evaluation scale, and metric conventions. Without those fields, a numerical difference cannot reveal whether the models or datasets were actually comparable.</p><p>The evaluation pipeline places these comparisons alongside a local <code>EvaluationRun</code>, whose predictions, aligned targets, residuals, metric report, and provenance are kept together. This connects benchmark comparison to the feature-wise compositional model: the local prediction must first come from the five independent OHLCV streams, concatenation fusion, and the selected recurrent and dense head. A benchmark difference cannot repair an architecture or target choice that the paper leaves unspecified.</p><h3>Reported evidence versus local conclusions</h3><p>The paper states that GWO outperformed Random Search across the evaluated LSTM, GRU, and SRU families. That statement is reported evidence from the paper. It is not a result established by this unexecuted scaffold. A local RS-versus-GWO comparison would require the same preprocessing, validation objective, candidate space, and declared search budget for both methods, followed by actual execution and evaluation.</p><p>The safe interpretation is therefore:</p><ol><li><p><code>reported_benchmarks()</code> preserves the supplied reference records.</p></li><li><p><code>compare_to_report</code> provides a labeled comparison mechanism for future local outputs.</p></li><li><p>Missing metrics and opaque configuration labels remain unresolved.</p></li><li><p>No local metric is claimed to match the paper.</p></li></ol><p>The benchmark records are useful precisely because they retain their provenance and limitations. They provide targets for a future, better-specified reproduction rather than proof that the current configurable implementation reproduces the published experiment.</p><h2>Offline synthetic demonstration and artifact provenance</h2><p>How can you learn the repository&#8217;s data and model interfaces without downloading market data? Use the offline demonstration. It creates clearly labeled synthetic OHLCV rows, applies the same five-channel scaling and windowing contracts, builds a configurable compositional model, and inspects the expected tensor shapes. This is a plumbing demonstration only: the synthetic series is not Hang Seng Index data and cannot validate forecasting performance.</p><h3>Why use an offline demonstration?</h3><p>The paper describes daily Hang Seng Index history obtained from Yahoo Finance, but the supplied context does not identify a unique ticker, date range, adjustment policy, or download configuration. Yahoo Finance data can also change as providers revise historical records. An offline example avoids those dependencies while making every local choice visible.</p><p>The generated example uses a deterministic pseudo-random generator when given a selected <code>seed</code>. It creates five columns in the canonical order <code>Open</code>, <code>High</code>, <code>Low</code>, <code>Close</code>, <code>Volume</code>. The values are synthetic prices and volumes; they are not intended to reproduce the empirical distribution of the HSI.</p><p>The example&#8217;s generator is defined in <code>examples/offline_synthetic_demo.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d55bdc58-c420-4eec-8259-188000b3555e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def make_synthetic_ohlcv(num_days: int, seed: int) -&gt; pd.DataFrame:
    """Create deterministic, clearly labeled synthetic OHLCV observations.

    The returned table has shape ``(num_days, 5)`` and canonical column order:
    Open, High, Low, Close, Volume. Values are generated only for this
    demonstration; they are not intended to model the empirical HSI series.
    """
    if isinstance(num_days, bool) or not isinstance(num_days, int) or num_days &lt;= 0:
        raise ValueError("num_days must be a positive integer")
    if isinstance(seed, bool) or not isinstance(seed, int) or seed &lt; 0:
        raise ValueError("seed must be a non-negative integer")

    rng = np.random.default_rng(seed)
    daily_returns = rng.normal(loc=0.0004, scale=0.012, size=num_days)
    close = 18_000.0 * np.exp(np.cumsum(daily_returns))
    open_noise = rng.normal(loc=0.0, scale=35.0, size=num_days)
    open_price = close + open_noise
    intraday_spread = np.abs(rng.normal(loc=55.0, scale=18.0, size=num_days))
    high = np.maximum(open_price, close) + intraday_spread
    low = np.minimum(open_price, close) - intraday_spread
    volume = np.maximum(
        1.0,
        1_000_000.0 + rng.normal(loc=0.0, scale=120_000.0, size=num_days),
    )

    values = np.column_stack((open_price, high, low, close, volume))
    return pd.DataFrame(
        values,
        index=pd.date_range("2020-01-01", periods=num_days, freq="D"),
        columns=list(CANONICAL_OHLCV_COLUMNS),
    )</code></pre></div><p>The function accepts a positive number of rows and a non-negative integer seed. Its output is a table with shape <code>num_days &#215; 5</code>. The validation errors are useful failure cases: a boolean is not accepted as an integer, and invalid row counts or seeds are rejected rather than silently corrected.</p><h3>Scaling an earlier chronological portion</h3><p>The paper states the intended feature range <code>[0, 0.95]</code>, but it does not specify exactly which observations fit the scaling statistics. The demonstration chooses an earlier chronological portion for fitting. This is a leakage-prevention decision made by the implementation, not a recovered paper detail. The remaining observations are transformed with the already fitted scaler.</p><p>The relevant portion of <code>main</code> is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4de1067c-280c-442d-93fb-19aabb527891&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">data_config = DataConfig(
    target_column="Close",
    forecast_horizon=1,
    window_length=20,
    split_ratio=0.8,
    scaling_range=(0.0, 0.95),
)
raw_frame = make_synthetic_ohlcv(num_days=96, seed=7)
raw_values = raw_frame.loc[:, list(CANONICAL_OHLCV_COLUMNS)].to_numpy(
    dtype=np.float64
)

# Fit on an earlier chronological portion to demonstrate a leakage-avoiding
# choice. The paper states the [0, 0.95] range but does not specify fit scope.
scaler_split = int(raw_values.shape[0] * data_config.split_ratio)
scaler = FeaturewiseMinMaxScaler(feature_range=data_config.scaling_range)
scaler.fit(raw_values[:scaler_split])
scaled_values = scaler.transform(raw_values)</code></pre></div><p>Here, <code>target_column='Close'</code>, <code>forecast_horizon=1</code>, and <code>window_length=20</code> are tutorial configuration choices. The paper does not identify the target column or forecast horizon; it only describes 20-day and 40-day input windows. Similarly, the demonstration&#8217;s 96 synthetic rows and seed <code>7</code> are not paper data.</p><p><code>FeaturewiseMinMaxScaler</code> requires a final feature axis of length five. It stores separate scaling metadata for each OHLCV channel, so Volume does not determine the scaling of the price columns. Its <code>inverse_transform</code> method is available when later evaluation should return predictions to original price units. The example does not calculate or present forecast metrics.</p><h3>Building windows and inspecting model shapes</h3><p>After scaling, the example selects the configured target column and constructs overlapping supervised windows:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;be470914-5e1d-4519-95e1-9cf1e0c50378&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">target_index = list(CANONICAL_OHLCV_COLUMNS).index(data_config.target_column)
windowed = make_sliding_windows(
    scaled_values,
    target_index=target_index,
    window_length=data_config.window_length,
    forecast_horizon=data_config.forecast_horizon,
    timestamps=raw_frame.index.to_numpy(),
)

# The model follows the paper's feature-wise structure: five independent
# recurrent streams, concatenation fusion, and a one-value prediction head.
model_config = ModelConfig(
    recurrent_family="LSTM",
    stream_layers=1,
    stream_units=16,
    stream_dropout=0.0,
    fusion_dropout=0.0,
    use_post_fusion_recurrent=False,
    dense_layers=(16,),
    output_dimension=1,
)
model = build_compositional_model(model_config)
model.eval()

batch_size = min(4, windowed.sample_count)
batch_inputs = torch.as_tensor(windowed.inputs[:batch_size], dtype=torch.float32)
with torch.no_grad():
    predictions = model(batch_inputs)</code></pre></div><p><code>make_sliding_windows</code> returns inputs with shape <code>samples &#215; window_length &#215; 5</code> and scalar targets with shape <code>samples</code>. With this tutorial configuration, each input contains 20 time steps and five channels. The target is the selected <code>Close</code> value one step after the input window according to the implementation&#8217;s explicit alignment convention.</p><p><code>build_compositional_model</code> then constructs the feature-wise recurrent model. Each batch input has shape <code>batch &#215; 20 &#215; 5</code>; the model separates the five channels, encodes them independently, concatenates their representations, and returns a prediction with shape <code>batch &#215; output_dimension</code>. The model configuration shown here is an educational choice. It does not decode the paper&#8217;s opaque benchmark label <code>LSTM-GWO (1-1-0-1)</code>, nor does it claim to recover the paper&#8217;s exact layer widths or training recipe.</p><p>The source example prints these contracts:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4fc03318-61c7-4414-b26c-1999d67a53d9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">print("Synthetic demonstration only; no paper result is being reproduced.")
print(f"raw OHLCV shape: {raw_values.shape}")
print(f"scaled OHLCV range: {data_config.scaling_range}")
print(
    "window input shape (samples, time_steps, features): "
    f"{windowed.inputs.shape}"
)
print(f"window target shape: {windowed.targets.shape}")
print(
    "model batch input shape (batch, time_steps, features): "
    f"{tuple(batch_inputs.shape)}"
)
print(f"model prediction shape (batch, output_dimension): {tuple(predictions.shape)}")</code></pre></div><p>These statements document intended dimensions; this run did not execute the example. In particular, no output numbers, predictions, or paper-comparison result should be inferred from the excerpt.</p><h3>Recording provenance and arrays locally</h3><p>A demonstration becomes more useful when its assumptions travel with its outputs. Artifact provenance should identify at least the data source, cleaning policy, scaling range and fit scope, window length, target column, forecast horizon, model family, architecture settings, seed, and evaluation conventions. For real HSI work, the archived raw file and the Yahoo Finance query parameters should also be retained.</p><p>The generated artifact module provides local persistence through <code>save_json_record</code>, <code>save_array</code>, and <code>load_json_record</code>. The README illustrates the JSON-record call as follows:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;097568f9-95a7-4d7b-a563-ace11ba0eeec&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">save_json_record(report, Path('artifacts/report.json'))</code></pre></div><p>A JSON record is appropriate for serializable configuration, acquisition metadata, histories, benchmark labels, and metric reports. Numeric predictions and residuals should be stored with <code>save_array</code>, which preserves their numeric shape and dtype. <code>load_json_record</code> reads a previously saved top-level JSON object. The artifact implementation uses local filesystem operations and rejects implicit overwrites, helping prevent an old experiment from being silently replaced.</p><p>For example, a future evaluation record could contain a local run identifier, <code>window_length</code>, <code>target_column</code>, <code>forecast_horizon</code>, scaler settings, model configuration, seed, metric conventions, and a separate reference to prediction and residual arrays. It should distinguish computed local values from the paper&#8217;s reported benchmark records. The generated interfaces provide the storage operations, while the exact record contents remain an experiment-provenance responsibility.</p><h3>What this demonstration does not establish</h3><p>The synthetic path confirms the intended software contracts conceptually: five ordered channels can be scaled, transformed into windows, and supplied to a feature-wise compositional model. It does not establish that the model forecasts the HSI accurately, that the chosen target and horizon match the paper, or that the reported LSTM-GWO, GRU-GWO, or SRU-GWO values can be reproduced.</p><p>No example execution, artifact creation, test execution, or verification occurred under the run policy. A stronger reproduction would archive the original Yahoo Finance data query and downloaded rows, supply the missing architecture and optimization details, run the configured experiment, and then evaluate predictions under explicitly documented metric conventions. Until then, the offline example is a reproducibility aid&#8212;not evidence of paper-level performance.</p><h2>Verification status, limitations, and responsible reproduction claims</h2><p>How can a carefully organized implementation be useful without overstating what it proves? The answer is to separate <strong>invariants</strong>, <strong>planned checks</strong>, and <strong>verified results</strong>. The generated package records the paper&#8217;s intended data flow, model structure, search interfaces, evaluation conventions, and provenance. However, this run did not execute code or perform verification. Its correct status is therefore <code>verification_skipped</code>: a configurable reproduction scaffold with explicit gaps, not a validated reproduction of the published experiment.</p><h3>What the implementation is designed to preserve</h3><p>Several structural requirements are clear enough to review independently of the paper&#8217;s missing experimental details:</p><ul><li><p>The input feature order is exactly <code>Open</code>, <code>High</code>, <code>Low</code>, <code>Close</code>, <code>Volume</code>.</p></li><li><p>Model inputs use <code>batch &#215; time &#215; 5</code> ordering.</p></li><li><p>Preprocessing constructs chronological windows and keeps training samples before later samples.</p></li><li><p>The scaler targets the stated interval <code>[0, 0.95]</code>. Fitting it on earlier training observations only is a leakage-prevention implementation choice; the paper does not specify its fitting partition.</p></li><li><p>The model creates exactly five independent feature streams.</p></li><li><p>Stream representations are fused by concatenation, rather than by early summation or mixing.</p></li><li><p>Dropout is intended for training mode and should be inactive during evaluation.</p></li><li><p>Predictions and targets must remain sample-aligned before metric calculation.</p></li><li><p>Search candidates must decode to valid typed values for layer counts, units, dropout rates, learning rates, batch sizes, and epochs.</p></li><li><p>RS and GWO should use the same candidate objective, preprocessing, and validation protocol when compared locally.</p></li><li><p>Benchmark labels such as <code>LSTM-GWO (1-1-0-1)</code> remain opaque identifiers and are not translated into architecture settings.</p></li></ul><p>These are implementation contracts and interpretation safeguards. They are not evidence that a particular model reached the paper&#8217;s reported metrics.</p><h3>Planned contract tests are not completed checks</h3><p>The repository contains generated test files intended to make these invariants reviewable later. For data handling, <code>test_scaler_range_and_inverse_round_trip</code>, <code>test_window_target_alignment</code>, and <code>test_chronological_split_has_no_reordering</code> address scaling behavior, window-to-target indexing, and chronological partitioning. The model tests include <code>test_model_accepts_five_feature_input</code>, <code>test_fusion_dimension_is_additive</code>, <code>test_recurrent_family_dispatch_is_consistent</code>, and <code>test_dropout_differs_only_in_training_mode</code>.</p><p>The remaining planned tests cover local metric conventions, candidate encoding, search bounds, and provenance. Examples include <code>test_rmse_zero_for_equal_arrays</code>, <code>test_percentage_metric_zero_policy</code>, <code>test_candidate_codec_validates_discrete_parameters</code>, <code>test_gwo_positions_stay_in_bounds_after_decoding</code>, <code>test_reported_benchmarks_preserve_supplied_values</code>, and <code>test_incomplete_metrics_remain_missing</code>.</p><p>These names describe intended checks, not completed checks. Under the authoritative run policy, test generation and execution were disabled. No test result, passing assertion, successful import, or model output should be inferred from the presence of these files.</p><h3>Static and semantic verification status</h3><p>Static verification and semantic verification address different questions. Static review would inspect syntax, imports, public interfaces, tensor-shape contracts, and configuration propagation without running the experiment. Semantic review would assess whether the implementation&#8217;s behavior matches the intended method, including feature independence, concatenation, dropout mode changes, search-objective consistency, and metric conventions.</p><p>Both forms of verification were skipped here. Local static verification reports the status <code>verification_skipped</code>, and semantic code verification was also skipped. Code execution, test execution, tutorial-section verification, and final quality review were disabled by policy. Consequently, this section must not claim that the generated code is correct, that the tests pass, or that any local result matches the paper.</p><h3>Why exact reproduction remains underdetermined</h3><p>The supplied paper context does not uniquely determine several decisions that materially affect results:</p><ul><li><p>The Yahoo Finance ticker or symbol, date range, download options, adjustment policy, missing-value handling, and duplicate-row handling.</p></li><li><p>The forecast target and horizon. The paper names &#8220;stock price&#8221; and lists OHLCV inputs, but does not identify whether the target is <code>Close</code>, adjusted <code>Close</code>, another price, or a multivariate output.</p></li><li><p>The exact interpretation of the 20-day and 40-day windows beyond treating them as input lengths.</p></li><li><p>The normalization formula, fitting partition, clipping behavior, and inverse-transformation procedure.</p></li><li><p>Recurrent layer counts, widths, activations, hidden-state extraction, dense-layer structure, and exact placement of both dropout stages.</p></li><li><p>The SRU variant and Python implementation or API.</p></li><li><p>The training loss, gradient optimizer, initialization, learning-rate schedule, seed, shuffling policy, stopping rule, and checkpoint-selection rule.</p></li><li><p>RS search domains, distributions, trial budget, and validation protocol.</p></li><li><p>GWO population size, initialization, bounds, iteration count, update schedule, objective, and encoding of discrete parameters.</p></li><li><p>Metric formulas, percentage conventions, zero-denominator handling, bias signs, and the scale used for evaluation.</p></li></ul><p>The paper&#8217;s reported total of 54 configurations does not resolve these omissions. Nor can the opaque four-part labels in the LSTM-GWO, GRU-GWO, and SRU-GWO records be safely decoded from the supplied evidence.</p><h3>A responsible review checklist</h3><p>Before treating a future local run as evidence, review the following items and record each one in the experiment provenance:</p><ol><li><p><strong>Data shape:</strong> confirm that the prepared inputs use <code>batch &#215; time &#215; 5</code> and the canonical OHLCV order.</p></li><li><p><strong>Chronology:</strong> confirm that the split and window-target alignment preserve temporal order and do not use future observations during fitting.</p></li><li><p><strong>Scaling:</strong> record the <code>[0, 0.95]</code> range, the fitting partition, constant-feature policy, and any inverse transformation.</p></li><li><p><strong>Target semantics:</strong> record the target column and forecast horizon instead of assuming that the paper specifies them.</p></li><li><p><strong>Model structure:</strong> record the five independent encoders, recurrent family, representation choice, dropout placement, fusion operation, post-fusion layers, and output dimension.</p></li><li><p><strong>Training:</strong> record the loss, gradient optimizer, learning rate, batch size, epochs, seed, validation protocol, and checkpoint policy.</p></li><li><p><strong>Search:</strong> record the typed domains, objective direction, trial or wolf budget, bounds, discrete decoding, and failure handling.</p></li><li><p><strong>Evaluation:</strong> record whether metrics use normalized or original units, every metric convention, percentage display rules, and zero handling.</p></li><li><p><strong>Benchmarks:</strong> preserve reported values separately from local values, keep missing GRU and SRU associations missing, and leave configuration labels opaque.</p></li><li><p><strong>Verification:</strong> record whether code and tests were actually executed and reviewed; do not substitute planned test names for evidence.</p></li></ol><p>This checklist is a review target, not a statement that the items were checked in this run.</p><h3>What a stronger reproduction would require</h3><p>A stronger claim would require the original data query or an archived copy of the downloaded OHLCV data, the complete architecture and configuration tables, the exact target and forecast horizon, and the full training and validation protocol. It would also require the RS and GWO search spaces and budgets, the selected SRU implementation, canonical metric definitions, and the original prediction or error outputs needed to compare visual analyses.</p><p>After those details were supplied, the implementation could use its provenance records to replace local choices without changing the overall public interfaces. The resulting experiment would still need to be executed, checked, and compared under a documented protocol before it could be described as a reproduction.</p><p>Attention mechanisms, hybrid architectures, and technical indicators are mentioned as future work in the paper. They are outside this reproduction workflow and should not be added to close the evidence gaps.</p><p>The appropriate conclusion at this stage is deliberately modest: the package makes the paper-to-code assumptions visible and organizes the intended pipeline, but neither the generated scaffold nor the supplied benchmark constants establish successful or exact reproduction.</p><p>Use the button or URL below to download the source code.</p>
      <p>
          <a href="/__u/onepagecode.substack.com/p/quant-trading-feature-wise-compositional">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Quant Trading: A BiLSTM-MLP-ReCast Deep Learning Model for Stock Price Forecasting (PyTorch Guide)]]></title><description><![CDATA[Implementing an end-to-end hybrid BiLSTM, MLP, and residual ReCast architecture for high-accuracy stock price prediction.]]></description><link>https://onepagecode.substack.com/p/quant-trading-a-bilstm-mlp-recast</link><guid isPermaLink="false">https://onepagecode.substack.com/p/quant-trading-a-bilstm-mlp-recast</guid><dc:creator><![CDATA[Onepagecode]]></dc:creator><pubDate>Sun, 02 Aug 2026 20:15:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_W6n!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3078fdad-8e80-40a9-be17-640eb18fa447_1254x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Use the button at the end of this article to download the source code.</h2><p>There was no code provided with the paper, the entire implementation is of my own. Use the entire code as a reference.</p><p>Paper URL: https://www.researchgate.net/publication/410381981_A_BiLSTM-MLP-ReCast_Method_for_Stock_Price_Forecasting</p><p>The paper proposes an end-to-end BiLSTM-MLP-ReCast framework for one-step-ahead stock closing-price forecasting. A bidirectional LSTM extracts temporal representations from 20-day OHLCV windows, an MLP applies nonlinear feature transformation, and a ReCast module combines a coarse forecast with a learned residual correction. Training uses joint MSE optimization with Adam. Experiments use normalized Yahoo Finance data for Google, Amazon, and Microsoft over 31 July 2017 to 31 July 2023, with chronological 70:15:15 train-validation-test splits. The reported complete model generally outperforms the listed CNN, LSTM, BiLSTM, MLP-ReCast, BiLSTM-ReCast, and BiLSTM-MLP variants.</p><h2>Implementation Assumptions</h2><ul><li><p>Use Python with PyTorch-style APIs, consistent with the reported environment, but do not require the reported hardware or timing.</p></li><li><p>Use the six listed fields [Open, High, Low, Close, Adjusted Close, Volume] by default while exposing OHLCV as an alternative configuration.</p></li><li><p>Use one scaler per stock fitted only on that stock's training observations.</p></li><li><p>Split raw chronological observations before generating split-local windows to avoid cross-split leakage; document this as an implementation decision because the paper does not specify the boundary convention.</p></li><li><p>Use a 20-step lookback, hidden size 64 per direction, two BiLSTM layers, dropout 0.2, batch size 32, learning rate 1e-5, and 64 epochs as defaults from table_3.</p></li><li><p>Use concatenated terminal forward and backward hidden states as the default BiLSTM representation, producing dimension 2H, while exposing the representation convention in configuration.</p></li><li><p>Require explicit MLP hidden and output widths because the paper does not specify them; provide configuration defaults without presenting them as paper facts.</p></li><li><p>Train the coarse and residual heads jointly using only final-output MSE; do not add a separate residual-supervision loss.</p></li><li><p>Treat Directional Accuracy and MAPE zero handling as configurable conventions because the paper does not define them.</p></li><li><p>Do not implement trading simulation, transaction costs, external indicators, or multi-step forecasting.</p></li><li><p>Do not claim exact reproduction of reported results because data-download behavior, dimensions, seeds, metric conventions, and other details are underspecified.</p></li></ul><h2>Scope, Evidence, and Reproduction Decisions</h2><p>What should this reproduction actually do, and which parts are choices rather than settled facts? The paper describes a one-step-ahead forecasting task: use a recent history of stock observations to predict the next closing price. Its proposed pipeline combines three stages:</p><ol><li><p>A bidirectional LSTM extracts temporal information from a fixed lookback window.</p></li><li><p>An MLP applies a nonlinear feature transformation.</p></li><li><p>A ReCast module adds a learned correction to a coarse forecast.</p></li></ol><p>The correction is part of the differentiable model. It is not a second, independent post-processing procedure.</p><h3>Start with the forecasting object</h3><p>Let <code>X_t</code> denote the historical input sequence ending at forecast origin <code>t</code>. Its sequence length is <code>L</code>, the lookback length, and each step contains <code>D</code> selected features. In the default configuration, <code>L</code> is 20 and <code>D</code> is 6 when the six-field input mode is selected. The paper's first problem-formulation record, <code>eq_01</code>, describes how this length-<code>L</code> sequence is assembled from feature vectors such as <code>x_tau</code>.</p><p>The generated implementation represents this operation through <code>make_windows</code> and <code>make_window_splits</code> in <code>src/bilstm_recast/data/windows.py</code>. For each valid position, those functions preserve a sequence of <code>L</code> rows and associate it with the next chronological closing-price value. The resulting model input has shape <code>[samples, L, D]</code>, while the scalar target has shape <code>[samples, 1]</code>.</p><p>The paper's second problem-formulation record, <code>eq_02</code>, describes the complete learned mapping from <code>X_t</code> to the normalized next-day forecast <code>y_hat</code>. Here, <code>F</code> means the parameterized BiLSTM-MLP-ReCast mapping, and <code>y_hat</code> is one predicted closing value per sample. In code, this mapping is the <code>BiLSTMMLPReCast.forward</code> method, composed and called by <code>run_stock_experiment</code>.</p><p>The equations are referenced here for traceability only. The supplied canonical equation records contain empty LaTeX fields, so no display equations are reproduced or reconstructed from OCR.</p><h3>What is a paper fact?</h3><p>The supplied paper context reports Yahoo Finance observations for Google, Amazon, and Microsoft over 31 July 2017 through 31 July 2023. It describes chronological 70:15:15 training, validation, and test partitions, training-fitted Min-Max normalization, and next-day closing-price prediction.</p><p>The final settings associated with <code>table_3</code> are also paper-provided defaults:</p><ul><li><p>lookback length: 20 trading days;</p></li><li><p>hidden size: 64 per BiLSTM direction;</p></li><li><p>recurrent layers: 2;</p></li><li><p>dropout: 0.2;</p></li><li><p>batch size: 32;</p></li><li><p>Adam learning rate: <code>1e-5</code>;</p></li><li><p>training duration: 64 epochs.</p></li></ul><p>The paper reports complete-model benchmark targets, including Microsoft RMSE <code>0.052044</code> and R2 <code>0.906752</code>, Amazon RMSE <code>0.044958</code> and R2 <code>0.808700</code>, and Google RMSE <code>0.057266</code> and R2 <code>0.788880</code>. These values are reference points from the paper, not guaranteed outputs of this implementation.</p><h3>What is an implementation decision?</h3><p>Several details are not fixed sufficiently to support an exact reproduction claim. The paper lists six input fields&#8212;<code>Open</code>, <code>High</code>, <code>Low</code>, <code>Close</code>, <code>Adjusted Close</code>, and <code>Volume</code>&#8212;but elsewhere labels the proposed input as OHLCV. The generated code keeps that discrepancy visible through <code>FeatureConfig.feature_mode</code>. Its default is <code>six_field</code>, while <code>ohlcv</code> is available as an alternative.</p><p>The MLP hidden and output widths are not specified in the supplied paper context. They therefore remain explicit configuration parameters rather than being presented as reported paper values. The implementation also chooses to split raw observations before constructing split-local windows. This is a leakage-safe interpretation: a training window cannot draw rows from validation or test data, although the paper does not state its exact boundary convention.</p><p>Other unresolved choices include:</p><ul><li><p>external data-download and adjustment settings;</p></li><li><p>ticker names and missing-value handling;</p></li><li><p>initialization, random seeds, and other optimizer details;</p></li><li><p>whether the target is sourced from <code>Close</code> or another close-related field when both are present;</p></li><li><p>the precise definition of Directional Accuracy;</p></li><li><p>the near-zero handling rule for MAPE;</p></li><li><p>the exact architectures and settings for named baselines and ablations.</p></li></ul><p>These are not harmless formatting details. They can change the windows, learned parameters, or reported metrics. The generated project records them as configuration or metadata instead of silently treating one interpretation as established fact.</p><h3>How the local implementation is entered</h3><p>The project intentionally does not implement live Yahoo Finance acquisition. The input boundary is a local CSV, which makes the data source and its preprocessing settings inspectable. The README describes this scope explicitly:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;markdown&quot;,&quot;nodeId&quot;:&quot;bb970e01-b519-4ffd-a281-3000519f9169&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-markdown">Yahoo Finance data can be used by first downloading it externally and supplying a local CSV that satisfies the documented schema.</code></pre></div><p>The command-line entry point requires those local paths. It exposes the feature-mode decision and selected training overrides before delegating to the experiment runner:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;78978b72-ec4f-4c6f-a211-c56c1d889094&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">parser.add_argument(
    "--feature-mode",
    choices=("six_field", "ohlcv"),
    default="six_field",
    help=(
        "Input feature convention: six fields including Adjusted Close, or OHLCV "
        "(default: six_field)."
    ),
)</code></pre></div><p>In <code>scripts/run_reproduction.py</code>, <code>main</code> resolves symbols, builds a validated <code>ReproductionConfig</code>, loads local frames, and calls <code>run_multi_stock_experiment</code>. The intended command shape is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;5a40e3b7-08c8-4ab1-b3a8-bd1c8c4fe8f1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">python scripts/run_reproduction.py \
  --input data/GOOG.csv \
  --symbol GOOG \
  --output artifacts/GOOG</code></pre></div><p>The important point is that this command describes a local reproduction workflow; it does not assert that the original Yahoo Finance download has been duplicated.</p><h3>Follow one stock through the pipeline</h3><p>A single run can be understood as the following sequence:</p><ol><li><p><code>ReproductionConfig.default()</code> supplies the documented model and training defaults.</p></li><li><p><code>validate_config</code> checks feature choices, dimensions, ratios, and metric policies.</p></li><li><p><code>run_stock_experiment</code> validates and date-filters the input frame.</p></li><li><p><code>split_chronologically</code> creates disjoint raw-row partitions.</p></li><li><p>A per-stock scaler is fitted on training rows only and reused for the other partitions.</p></li><li><p><code>make_window_splits</code> creates 20-step, split-local windows and next-step targets.</p></li><li><p>The BiLSTM-MLP-ReCast model produces a final forecast.</p></li><li><p>Joint final-output MSE trains the model, while validation MSE determines model selection.</p></li><li><p>The selected model is evaluated on normalized test targets with the configured metrics.</p></li></ol><p>The orchestration is visible in the generated runner. This excerpt shows the central preparation boundary:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5e14c163-0607-4844-8cd4-154d0baaf60e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">split = split_chronologically(prepared_frame, ratios)
assert_split_integrity(split)

train_values = split.train.loc[:, feature_columns].to_numpy(dtype=np.float64)
scaler = fit_training_scaler(train_values)
lookback = int(config.model.lookback)
window_splits = make_window_splits(
    split.as_dict(),
    scaler,
    feature_columns,
    target_column,
    lookback,
)  # Eq. 1: construct split-local historical windows and next-step targets.</code></pre></div><p>Later in the same function, training, prediction, and evaluation are connected without using test metrics for selection:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;82ef5669-73eb-4f4f-b422-f1b6a2ec4d9b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">history = fit_model(model, train_loader, validation_loader, config.training)  # Eq. 15: final-output MSE training.

predictions, targets = _collect_predictions(model, test_loader, device)  # Eq. 2: one-step forecast mapping.

metrics = evaluate_forecast_metrics(
    predictions,
    targets,
    config.metrics,
    previous_targets=previous_targets,
)  # Eq. 19&#8211;21: normalized MSE, RMSE, and R2 evaluation.</code></pre></div><p>This is an implementation explanation of the paper's pipeline, not evidence that the pipeline has been executed successfully.</p><h3>Worked example: tracing a local frame</h3><p>Suppose a local frame has six selected numeric columns and enough chronological rows for all three partitions. With the default <code>L=20</code>, one window uses rows <code>i</code> through <code>i+19</code> as <code>X_t</code>; row <code>i+20</code> supplies the target closing value. After batching, the model receives <code>[B, 20, 6]</code>.</p><p>The runner first assigns the raw rows to train, validation, and test partitions. It then learns six featurewise scaling ranges from training rows only. The same stored ranges transform validation and test rows. If a validation or test value lies outside a training range, the generated scaler does not silently clip it; it applies the stored transformation and preserves the resulting value for later inspection.</p><p>The model then maps each <code>[20, 6]</code> sequence to one normalized forecast. Training compares that forecast with the normalized target. Finally, the reporting layer calculates normalized metrics and stores the selected feature mode, split policy, representation convention, scaler scope, and metric conventions in the experiment metadata. No numerical result from this example should be labeled a paper result, and no execution is claimed here.</p><p>Synthetic frames are available for a network-free demonstration, but they are demonstration data rather than stock observations from the paper. Likewise, a local CSV may reproduce the expected schema without reproducing the original download settings.</p><h3>Evidence and verification boundary</h3><p>The implementation plan distinguishes two kinds of future review. Static verification would inspect syntax, imports, annotations, file contracts, and traceability from method and equation identifiers to code. Semantic verification would inspect behavior-level invariants such as training-only scaling, tensor shapes, the additive ReCast relation, final-output loss, and validation-only checkpoint selection.</p><p>Both categories were skipped under the authoritative run policy. No code execution, test execution, semantic code verification, tutorial verification, or final quality review occurred. The generated test files are specifications of intended checks, not evidence that those checks passed.</p><p>For this reason, the appropriate claim is that the project provides a reproduction-oriented implementation and makes its assumptions explicit. It does not establish exact reproduction of the paper's reported numbers.</p><h2>Project Structure and Public Interfaces</h2><p>How do you keep a reproduction understandable when it combines data preparation, a recurrent model, residual refinement, training, and several metrics? The generated project treats each responsibility as a separate module. This makes the paper-to-code mapping visible: data code prepares trustworthy windows, model code defines the BiLSTM-MLP-ReCast computation, training code updates parameters, metric code evaluates forecasts, and experiment code coordinates a complete stock run.</p><p>This separation is an implementation decision rather than a claim that the paper prescribes a particular Python package layout. The paper specifies the forecasting pipeline and training objective; the generated project turns those specifications into boundaries that can be inspected and reused.</p><h3>Read the source tree by responsibility</h3><p>The <code>src/bilstm_recast/data/</code> package owns the path from a chronological stock table to supervised windows. Its planned public functions include <code>resolve_feature_columns</code>, <code>split_chronologically</code>, <code>MinMaxFeatureScaler.fit</code>, and <code>make_window_splits</code>. Together, they represent the <code>prepare_stock_windows</code> method: select features explicitly, preserve chronological order, fit normalization on training rows only, and produce input-target arrays.</p><p>The <code>src/bilstm_recast/model/</code> package owns the neural computation. <code>BiLSTMEncoder</code> converts a batch shaped <code>[B,L,D]</code> into a terminal bidirectional representation shaped <code>[B,2H]</code>. <code>FeatureMLP</code> applies the two-layer nonlinear transformation, and <code>ReCastHeads</code> produces the coarse and residual branches. <code>BiLSTMMLPReCast</code> composes those pieces into the complete forward pass described by <code>bilstm_mlp_recast_forward</code>.</p><p>The <code>src/bilstm_recast/training/</code> package owns optimization and validation selection. <code>build_adam_optimizer</code> constructs Adam with the configured learning rate, while <code>train_one_epoch</code>, <code>evaluate_loss</code>, and <code>fit_model</code> separate gradient updates from validation measurement. The loss boundary is <code>compute_forecast_loss</code>, which applies the final-output MSE from <code>compute_forecast_loss</code>; it does not introduce a second residual-supervision objective.</p><p>The <code>src/bilstm_recast/metrics/</code> package computes normalized regression metrics and configurable percentage or directional metrics. The <code>src/bilstm_recast/experiments/</code> package then coordinates these components through functions such as <code>run_stock_experiment</code> and <code>run_multi_stock_experiment</code>. Local loading is intentional: the generated workflow accepts supplied CSV files rather than inventing a live Yahoo Finance API contract, ticker behavior, adjustment mode, or network error policy.</p><p>Finally, <code>scripts/</code> provides command-line entry points, <code>examples/</code> provides a synthetic demonstration, and <code>tests/</code> contains planned contract tests. The test files are specifications of intended behavior; they were not run under the authoritative run policy.</p><h3>Use records to make shapes explicit</h3><p>The low-level shared records live in <code>src/bilstm_recast/types.py</code>. <code>WindowSplit</code> represents prepared supervised data. Its <code>inputs</code> array has shape <code>[samples, lookback, features]</code>, and its <code>targets</code> array is normalized to <code>[samples,1]</code>. This is the data-side form of the model's batch contract: a loader turns those arrays into batches <code>[B,L,D]</code> and <code>[B,1]</code>.</p><p>A focused excerpt shows how the record enforces the target shape without exposing the entire file:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e06c1c58-d354-476f-a0b8-3d1dbb65e6b6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class WindowSplit:
    """A split-local collection of supervised stock windows.

    ``inputs`` has shape ``[samples, lookback, features]`` and ``targets`` has
    shape ``[samples, 1]``.  The targets are normalized when the split was
    produced by the normalized preprocessing pipeline.
    """

    inputs: np.ndarray
    targets: np.ndarray
    name: str = ""</code></pre></div><p>The constructor also checks dimensionality, matching sample counts, numeric types, finite values, and non-empty splits. Those checks are implementation safeguards around the <code>prepare_stock_windows</code> method. They do not replace the leakage policy: the caller must still ensure that splitting and scaler fitting happen in the correct chronological order.</p><p>At the model boundary, <code>ForecastBatch</code> names the three ReCast outputs: <code>final</code>, <code>coarse</code>, and <code>residual</code>. Each must have shape <code>[B,1]</code>, and the record enforces the additive invariant that the final prediction equals the coarse prediction plus the residual correction:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8135305c-ec67-4a86-b547-5ae3d5b80ab9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class ForecastBatch:
    """Outputs from the ReCast branches for one model batch.

    ``final``, ``coarse``, and ``residual`` are tensors with shape ``[B, 1]``.
    The additive invariant ``final == coarse + residual`` is checked at record
    construction time up to floating-point tolerance.
    """

    final: torch.Tensor
    coarse: torch.Tensor
    residual: torch.Tensor</code></pre></div><p>This record connects <code>bilstm_mlp_recast_forward</code> to both <code>compute_forecast_loss</code> and <code>compute_forecast_diagnostics</code>. Training consumes <code>final</code>; diagnostics can inspect all three values. The same shape contract prevents a branch from silently broadcasting across samples or outputs.</p><p>The experiment boundary is represented by <code>ExperimentResult</code>. It groups a stock symbol, the in-memory model, <code>TrainingHistory</code>, <code>MetricResults</code>, and metadata. That metadata is important for this paper because several choices are unresolved: feature mode, split-boundary policy, terminal-state convention, MLP widths, MAPE handling, and Directional Accuracy convention. A numerical result without those decisions would be difficult to interpret.</p><h3>Make failures meaningful at module boundaries</h3><p>The generated <code>src/bilstm_recast/errors.py</code> defines four domain-specific exception types:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;00b39102-dfef-4813-81bf-6156f7e29aaa&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class DataValidationError(Exception):
    """Raised when supplied stock data violates the tabular data contract."""


class ShapeContractError(Exception):
    """Raised when an array or tensor violates a documented shape contract."""


class LeakageError(Exception):
    """Raised when a preprocessing or evaluation operation risks data leakage."""


class ConfigurationError(Exception):
    """Raised when a reproduction configuration is invalid or unsupported."""</code></pre></div><p>These categories clarify where a failure belongs. A missing numeric column is a <code>DataValidationError</code>; an input that is not <code>[B,L,D]</code> is a <code>ShapeContractError</code>; fitting a scaler before the training partition is isolated is a <code>LeakageError</code>; and an invalid feature mode or unsupported convention is a <code>ConfigurationError</code>. The exceptions contain no orchestration logic, so lower-level modules do not need to import the experiment runner.</p><p>That dependency direction is deliberate. Types and errors are foundational. Data modules may depend on them, model modules may depend on configuration and shared records, and experiment orchestration may depend on all of those layers. The reverse direction would make a small record depend on the entire experiment workflow and would make reuse harder.</p><h3>A small interface-level workflow</h3><p>At the public-interface level, a local run is intended to read like this:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;0ab8b666-60d4-4ca8-b85a-50452f33cd76&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from bilstm_recast import ReproductionConfig

config = ReproductionConfig.default()
# validate_config(config)
# result = run_stock_experiment(frame, "GOOG", config)</code></pre></div><p>The package-level <code>ReproductionConfig</code> export and the exact configuration implementation are specified in the project plan, but a generated source excerpt for those files was not supplied in this pipeline context. Therefore, the excerpt above is an interface sketch rather than a claim about verified source text or execution. The generated <code>types.py</code> and <code>errors.py</code> excerpts are the concrete source boundaries available here.</p><p>Conceptually, <code>run_stock_experiment</code> connects the method cards in order: <code>prepare_stock_windows</code> produces split-local windows; <code>bilstm_mlp_recast_forward</code> produces forecasts; <code>compute_forecast_loss</code> supplies the training objective; <code>train_bilstm_mlp_recast</code> selects the best validation state; and <code>evaluate_forecast_metrics</code> reports normalized test metrics. <code>run_multi_stock_experiment</code> repeats that workflow while preserving independent per-stock preparation metadata.</p><h3>What the structure does&#8212;and does not&#8212;verify</h3><p>The generated tests are intended to check contracts such as <code>[samples,L,D]</code> windows, <code>[B,2H]</code> representations, <code>[B,1]</code> forecasts, training-only scaling, final-equals-coarse-plus-residual, and validation-only checkpoint selection. They are not evidence that those checks passed: no tests were executed.</p><p>Static review would inspect syntax, imports, annotations, file boundaries, and traceability from method and equation identifiers to public functions. Semantic review would inspect whether the implementation behavior actually preserves leakage prevention, tensor-shape invariants, joint final-output loss, and validation-based selection. Both static verification and semantic code verification were skipped under the run policy, as were execution and tutorial verification. The project structure therefore documents intended contracts and responsibilities; it does not provide a completed verification result.</p><h2>Configuration and Leakage-Safe Stock Windows</h2><p>How can a forecasting experiment use the past without accidentally allowing future information into training? The answer is to make the data boundary explicit: choose the input columns, preserve chronological order, split raw observations, fit normalization only on training rows, and create each supervised example from a fixed history followed by its next target.</p><p>The paper specifies Yahoo Finance observations from 31 July 2017 through 31 July 2023, a chronological 70:15:15 train-validation-test policy, and a final lookback of 20 trading days. These are paper facts. The generated project adds several decisions that the paper does not fully settle: it accepts local CSV files rather than implementing a live Yahoo Finance download, splits raw rows before constructing windows, defaults to one scaler per stock, and exposes the feature-column ambiguity through configuration.</p><h3>Make the feature choice visible</h3><p>The paper describes six input fields&#8212;Open, High, Low, Close, Adjusted Close, and Volume&#8212;but also labels the proposed input as OHLCV elsewhere. Those descriptions imply different feature counts. The generated configuration does not silently choose between them. Its <code>FeatureConfig</code> exposes <code>feature_mode</code>, with <code>six_field</code> as the default and <code>ohlcv</code> as the alternative. In <code>six_field</code> mode, <code>D</code> is 6; in <code>ohlcv</code> mode, <code>D</code> is 5.</p><p>The target is the configured <code>Close</code> column. This follows the paper's next-day closing-price objective, although the paper does not completely clarify how the target should be distinguished from the presence of Adjusted Close among the inputs. A reproduction should record the selected feature mode and target column in its metadata.</p><p>Here is the relevant configuration excerpt from <code>src/bilstm_recast/config.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e3968d31-4244-4ff5-a877-6abd8b301233&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass
class FeatureConfig:
    """Data-selection and chronological preprocessing configuration.

    The paper describes both six fields including Adjusted Close and the OHLCV
    subset. ``feature_mode`` keeps that ambiguity visible. The default selects
    the six-field description from the data section.
    """

    SIX_FIELD_COLUMNS: ClassVar[Tuple[str, ...]] = (
        "Open",
        "High",
        "Low",
        "Close",
        "Adjusted Close",
        "Volume",
    )
    OHLCV_COLUMNS: ClassVar[Tuple[str, ...]] = (
        "Open",
        "High",
        "Low",
        "Close",
        "Volume",
    )
    feature_mode: str = "six_field"
    feature_columns: Tuple[str, ...] | None = None
    target_column: str = "Close"
    date_column: str = "Date"
    start_date: str = "2017-07-31"
    end_date: str = "2023-07-31"
    split_ratios: Tuple[float, float, float] = (0.70, 0.15, 0.15)</code></pre></div><p>The public method <code>FeatureConfig.resolved_feature_columns</code> returns either the explicitly supplied columns or the columns associated with the selected mode. <code>validate_config</code> then checks that the target is included in those features and that the split ratios sum to 1.0. This cross-check matters because the scaler and window builder expect the target to occupy one position in the same feature matrix.</p><h3>Validate and order the local data</h3><p>The generated loader accepts a local CSV path through <code>load_stock_csv</code> in <code>src/bilstm_recast/data/loaders.py</code>. This is an implementation boundary, not a claim that the paper used CSV files. The paper's source is Yahoo Finance, but the supplied project deliberately avoids inventing a network API, ticker behavior, adjustment mode, or credential requirement. A local file must contain a date column and numeric stock columns; the schema module performs the feature-specific checks afterward.</p><p><code>sort_and_filter_date_range</code> in <code>src/bilstm_recast/data/schema.py</code> applies the configured date interval inclusively and returns a chronologically sorted copy. <code>validate_stock_frame</code> rejects missing dates, duplicate dates, nonnumeric values, and invalid selected data. These failures are preferable to silently filling or reordering data in a way that changes the experiment.</p><p>The intended preparation sequence for one stock is therefore:</p><ol><li><p>Load a local chronological table with <code>load_stock_csv</code>.</p></li><li><p>Resolve <code>six_field</code> or <code>ohlcv</code> with <code>resolve_feature_columns</code>.</p></li><li><p>Filter the inclusive date interval with <code>sort_and_filter_date_range</code>.</p></li><li><p>Validate the selected columns with <code>validate_stock_frame</code>.</p></li><li><p>Split the raw observations chronologically.</p></li><li><p>Fit the scaler only on training rows.</p></li><li><p>Build windows independently inside each split.</p></li></ol><h3>Split raw observations before making windows</h3><p>The paper gives the 70:15:15 chronological policy but does not state whether windows are generated before or after the split. The generated implementation chooses to split raw observations first with <code>split_chronologically</code> in <code>src/bilstm_recast/data/splitting.py</code>. This is a leakage-safe implementation decision.</p><p><code>ChronologicalSplit</code> stores the train, validation, and test frames together with half-open positional ranges. <code>assert_split_integrity</code> checks that the partitions are nonempty, ordered, contiguous, and disjoint. Once split, a training window cannot contain validation rows, and a validation window cannot begin with the tail of the training partition.</p><p>This choice can reduce the number of available windows at the boundaries. That is an intentional trade-off: the implementation gives up cross-boundary examples in exchange for a simple invariant that each split owns all rows used by its windows and targets. The paper does not establish that this is its exact boundary convention, so it should be reported as part of the reproduction configuration.</p><h3>Fit one featurewise scaler on training rows</h3><p>The paper specifies Min-Max normalization to the interval from 0 to 1, with parameters fitted on the training set and reused for validation and test. In the generated code, <code>MinMaxFeatureScaler.fit</code> receives only the training feature array. It stores one minimum and one maximum for each feature. Later calls to <code>transform</code> use those stored values without inspecting the other splits.</p><p>This is the central leakage rule: validation and test values may fall outside the training extrema, but they must not change the normalization ruler. The generated scaler deliberately does not clip such values. Constant training features are handled deterministically by mapping them to zero and restoring them to their training minimum during <code>inverse_transform</code>.</p><p>A focused excerpt from <code>src/bilstm_recast/data/scaling.py</code> shows the training-only API:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;88b45b39-df37-4fdf-8c21-ac24eaa794b8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def fit_training_scaler(train_values: np.ndarray) -&gt; MinMaxFeatureScaler:
    """Fit and return a scaler using only the supplied training values.

    This convenience function intentionally does not accept validation or test
    arrays, making the leakage-prevention boundary explicit at the API level.
    """

    return MinMaxFeatureScaler().fit(train_values)</code></pre></div><p>The scaler preserves the two-dimensional shape <code>[rows, features]</code>. If the selected mode has <code>D</code> features, every transformed array must have shape <code>[rows, D]</code>. <code>make_window_splits</code> reuses the same fitted scaler for all three frames and extracts the scaled target from the configured target-column position.</p><h3>Construct the one-step windows</h3><p>The paper's <code>eq_01</code> describes the historical sequence of length <code>L</code> ending at the forecast origin. No canonical LaTeX was supplied for this equation, so its indexing is explained here rather than reconstructed. For a row index <code>i</code>, the implementation takes feature rows <code>i:i+L</code> as one input and uses the target at row <code>i+L</code> as the next-step label. With the default <code>L=20</code>, the input contains 20 consecutive trading observations and the target is the following observation's normalized close.</p><p>The resulting NumPy arrays have shapes <code>[samples, 20, D]</code> for inputs and <code>[samples, 1]</code> for targets. In the notation used by the paper, one input sequence <code>X_t</code> therefore has shape <code>[L,D]</code>, each <code>x_tau</code> is a feature vector of length <code>D</code>, and <code>y_{t+1}</code> is a scalar normalized closing-price target.</p><p>The core indexing is visible in <code>src/bilstm_recast/data/windows.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4aeb6247-bd7e-44e6-90bc-eaabe5afd0d0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># Eq. 1: each historical sequence ends immediately before its next-step target.
sample_count = feature_array.shape[0] - lookback
window_inputs = np.stack(
    [feature_array[start : start + lookback] for start in range(sample_count)],
    axis=0,
)
window_targets = target_array[lookback:].reshape(sample_count, 1)</code></pre></div><p><code>make_windows</code> validates that the feature and target arrays have the same row count, contain finite numeric values, and include at least <code>lookback + 1</code> rows. <code>validate_window_contract</code> then checks the trailing input shape, target shape, sample alignment, and nonempty output. These checks implement the method card <code>prepare_stock_windows</code> and its <code>eq_01</code> mapping.</p><p>For split-local preparation, <code>make_window_splits</code> performs the same operation independently for <code>train</code>, <code>validation</code>, and <code>test</code>. It requires a fitted <code>MinMaxFeatureScaler</code>, verifies that the scaler's feature count equals the selected column count, transforms each split, and then calls <code>make_windows</code>. The function therefore enforces two separate contracts: normalization ownership belongs to training, while window ownership belongs to each split.</p><h3>Worked example: one small window</h3><p>Suppose a synthetic frame has at least 21 rows and the selected mode has <code>D=6</code>. With <code>L=20</code>, the first valid example uses rows 0 through 19 as its input. Its input shape is <code>[20,6]</code>, and row 20 supplies the target. The next example uses rows 1 through 20 and targets row 21. For <code>R</code> rows in a split, the number of examples is <code>R - 20</code>, so the split arrays have shapes <code>[R-20,20,6]</code> and <code>[R-20,1]</code>.</p><p>The generated test specification expresses the same alignment without claiming that it was executed:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6b9d7704-8a89-4ea7-8c71-7df2e5c5e777&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">inputs, targets = make_windows(feature_values, target_values, lookback)

assert inputs.shape == (rows - lookback, lookback, 2)
assert targets.shape == (rows - lookback, 1)
for sample_index in range(rows - lookback):
    np.testing.assert_array_equal(
        inputs[sample_index], feature_values[sample_index : sample_index + lookback]
    )
    assert targets[sample_index, 0] == target_values[sample_index + lookback]</code></pre></div><p>This excerpt from <code>tests/test_data_contracts.py</code> uses two features and a shorter lookback only to make the indexing easy to inspect. It is not a claim about the paper's final configuration, which uses a 20-day lookback. The test file is a planned specification and was not run under the authoritative run policy.</p><p>Finally, <code>StockWindowDataset</code> in <code>src/bilstm_recast/data/datasets.py</code> copies the prepared arrays into <code>float32</code> tensors. A single item has shape <code>[L,D]</code> and <code>[1]</code>; a <code>DataLoader</code> batch has shape <code>[B,L,D]</code> and <code>[B,1]</code>. Training shuffling, when enabled, changes only the order of already-constructed windows. It does not change their temporal contents or create windows across split boundaries.</p><p>The resulting preprocessing pipeline is therefore precise about what is known and what is chosen: the date interval, chronological split policy, 20-step horizon, and training-only normalization follow the paper; the feature mode, raw-split boundary convention, local-file source contract, and per-stock scaler scope are recorded implementation decisions. No execution, test run, or verification result is claimed here.</p><h2>BiLSTM Temporal Encoder</h2><p>How does a 20-day stock window become a fixed-size feature vector that the MLP can consume? The encoder reads the same sequence in two directions: one recurrent pass follows the original time order, and another follows the reverse order. Their terminal hidden states are then joined into one representation.</p><p>This section separates the paper specification from the implementation interpretation. The paper associates <code>eq_03</code> through <code>eq_07</code> with LSTM gates and forward/backward processing, and <code>eq_08</code> with the bidirectional representation. However, the supplied equation records contain no canonical LaTeX, and the paper does not unambiguously say whether <code>H_t</code> is a final-time-step output, a concatenation of terminal hidden states, or another aggregation. The generated code resolves that ambiguity by selecting the final recurrent layer's terminal forward and backward hidden states.</p><h3>The tensor contract</h3><p>The encoder expects a batch-first tensor with shape <code>[B,L,D]</code>. Here, <code>B</code> is the batch size, <code>L</code> is the lookback length, and <code>D</code> is the number of selected features. Under the default reproduction settings, <code>L</code> is 20. With the default six-field input mode, <code>D</code> is 6; the alternative OHLCV mode has a different feature count.</p><p>The hidden size per direction is <code>H=64</code>, and the number of stacked recurrent layers is <code>N=2</code>. A bidirectional layer has two directions, so PyTorch returns its final hidden state in the shape <code>[N*2,B,H]</code>. With these defaults, that becomes <code>[4,B,64]</code>. The first dimension contains two state entries for each recurrent layer: one forward state and one backward state.</p><p>The implementation maps the paper's <code>eq_16</code> stage&#8212;the BiLSTM transformation from <code>X_t</code> to a temporal representation&#8212;to <code>BiLSTMEncoder.forward</code>. The result has shape <code>[B,2H]</code>, or <code>[B,128]</code> with <code>H=64</code>. This fixed-width vector is the input to the MLP stage described by <code>eq_17</code>.</p><h3>Let the framework own the LSTM gates</h3><p>The paper discusses forget, input, and output gates in <code>eq_03</code>, <code>eq_04</code>, and <code>eq_05</code>, followed by forward and backward recurrent processing in <code>eq_06</code> and <code>eq_07</code>. The generated implementation does not manually code those gate equations. Instead, <code>BiLSTMEncoder</code> uses PyTorch's standard <code>nn.LSTM</code> with <code>bidirectional=True</code>. This is an implementation choice consistent with the method card, but it should not be presented as a literal transcription of every internal LSTM equation: the supplied paper context omits some cell-state details and provides no canonical equation LaTeX.</p><p>The recurrent configuration is established in the constructor:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ccae8251-a8ab-4b79-a951-9a7731e86970&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        self.lstm = nn.LSTM(
            input_size=self.input_dim,
            hidden_size=self.hidden_size,
            num_layers=self.num_layers,
            dropout=dropout,
            batch_first=True,
            bidirectional=True,
        )</code></pre></div><p>Several details matter here. <code>batch_first=True</code> makes the public input layout <code>[B,L,D]</code>. <code>hidden_size</code> is the width <code>H</code> of each direction, not the width after concatenation. <code>num_layers=2</code> creates a two-layer stack. In the framework's LSTM behavior, the configured dropout is applied between stacked recurrent layers; with the paper's reported default, this is dropout <code>0.2</code>. <code>bidirectional=True</code> creates the forward and reverse state pairs required for the representation convention.</p><h3>Validate the sequence before recurrent processing</h3><p><code>BiLSTMEncoder</code> rejects inputs that do not satisfy the configured shape contract. It checks that the input is a floating-point tensor with three dimensions, that its sequence length equals the configured lookback, and that its feature width equals the configured input dimension. Thus, the model boundary catches a window with 19 rows or the wrong number of columns before it reaches the LSTM.</p><p>The public methods are intentionally small. <code>forward</code> returns only the representation, while <code>encode_with_states</code> also exposes the framework hidden-state tensor for inspection:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;780608c0-b260-4542-95e7-4e26508502a8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    def encode_with_states(self, x: torch.Tensor) -&gt; Tuple[torch.Tensor, torch.Tensor]:
        """Return the terminal representation and final hidden state tensor.

        The second return value is PyTorch's ``h_n`` with shape
        ``[num_layers * 2, batch, hidden_size]``.  Cell states are intentionally
        kept internal because the planned encoder API exposes the recurrent
        hidden state used to construct the representation.
        """
        self._validate_input(x)

        # Eq. 16: the bidirectional encoder maps each [B, L, D] window to a temporal representation.
        # Eqs. 03-07: PyTorch's bidirectional LSTM supplies the gate and direction recurrences.
        _, hidden_state = self.lstm(x)

        # Eq. 08: concatenate the terminal forward and backward states of the final layer.
        representation = extract_terminal_bidirectional_state(
            hidden_state,
            layers=self.num_layers,
            hidden_size=self.hidden_size,
        )
        return representation, hidden_state</code></pre></div><p>The unused first result from <code>self.lstm(x)</code> is the sequence of outputs at every time step. The encoder instead uses <code>hidden_state</code>, PyTorch's final hidden-state tensor, because the selected representation convention is based on terminal states. <code>encode_with_states</code> returns both the resulting <code>[B,2H]</code> representation and the original <code>[N*2,B,H]</code> state tensor.</p><h3>Selecting the final layer and both directions</h3><p>When <code>N=2</code>, the code must not accidentally use the first recurrent layer's states. The helper <code>extract_terminal_bidirectional_state</code> computes the start of the final layer's pair, selects the forward and backward entries, and concatenates them along the feature dimension. Its shape checks make the expected framework layout explicit:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;454c85ea-9504-4546-bacb-5688e6968a49&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    final_layer_start = (layers - 1) * _BIDIRECTIONAL_FACTOR
    forward_state = hidden[final_layer_start]
    backward_state = hidden[final_layer_start + 1]

    # Eq. 8: concatenate the terminal forward and backward hidden states.
    representation = torch.cat((forward_state, backward_state), dim=-1)</code></pre></div><p>For <code>N=2</code>, <code>final_layer_start</code> is <code>2</code>. Therefore, entries <code>hidden[2]</code> and <code>hidden[3]</code> are selected from a state tensor shaped <code>[4,B,H]</code>. Each selected tensor has shape <code>[B,H]</code>; concatenating them produces <code>[B,2H]</code>. With <code>H=64</code>, the result is <code>[B,128]</code>.</p><p>The helper also validates the complete state layout before indexing it:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b3944d6d-1db8-4a9a-8972-7524001a09e4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def validate_terminal_state_shape(
    hidden: torch.Tensor,
    layers: int,
    batch_size: int,
    hidden_size: int,
) -&gt; None:
    """Validate the PyTorch bidirectional-LSTM state layout.

    Parameters
    ----------
    hidden:
        Final hidden-state tensor with shape ``[layers * 2, batch, H]``.
        PyTorch stores the forward and reverse states consecutively for each
        recurrent layer.
    layers:
        Number of recurrent layers represented in ``hidden``.
    batch_size:
        Expected batch dimension.
    hidden_size:
        Hidden width for one direction.
    """</code></pre></div><p>This code documents the direction ordering used by the implementation: for each layer, the forward state is followed by the reverse state. That ordering is an implementation-level contract. It is important because concatenating the wrong pair, or selecting states from different layers, would change the input presented to the MLP even if the outer tensor shape still looked correct.</p><h3>Worked shape example</h3><p>Consider one training batch with the paper's default-style settings: batch size <code>B=32</code>, lookback <code>L=20</code>, six input features <code>D=6</code>, two recurrent layers <code>N=2</code>, and hidden size <code>H=64</code> per direction.</p><ol><li><p>The input to <code>BiLSTMEncoder</code> is <code>[32,20,6]</code>.</p></li><li><p>The bidirectional, two-layer LSTM returns <code>h_n</code> with shape <code>[N*2,B,H]</code>, which is <code>[4,32,64]</code>.</p></li><li><p>The helper selects the final layer's forward state <code>[32,64]</code> and backward state <code>[32,64]</code>.</p></li><li><p>Concatenation produces <code>H_t</code> with shape <code>[32,128]</code>.</p></li><li><p>The MLP receives <code>[32,128]</code>, preserving one representation per batch item.</p></li></ol><p>The six-feature count in this example follows the generated default configuration, not an unqualified resolution of the paper's data ambiguity. If OHLCV mode is selected, only <code>D</code> changes at the input boundary; the recurrent output remains <code>[B,2H]</code> because <code>H</code> is configured independently of the feature count.</p><h3>What the contract catches</h3><p>The generated test specification <code>test_terminal_state_has_two_direction_width</code> constructs a state tensor with shape <code>[layers*2,batch,hidden]</code>, applies <code>extract_terminal_bidirectional_state</code>, and checks both the <code>[B,2H]</code> shape and the expected final-layer concatenation. The model contract tests also require complete-model outputs to retain the scalar forecast shape <code>[B,1]</code> and reject an incorrect lookback or feature width.</p><p>These tests are specifications, not completed verification results. Under the run policy, no tests, code execution, static verification, or semantic code verification occurred. The intended invariants are nevertheless clear:</p><ul><li><p>inputs are floating-point tensors shaped <code>[B,L,D]</code>;</p></li><li><p>the framework hidden state is shaped <code>[N*2,B,H]</code>;</p></li><li><p>the selected representation is <code>[B,2H]</code>;</p></li><li><p>the final recurrent layer supplies both directional states; and</p></li><li><p>the terminal-state convention is recorded as an implementation decision for the ambiguous <code>eq_08</code>.</p></li></ul><p>The encoder therefore provides the temporal half of the paper's hybrid design without claiming more mathematical detail than the supplied evidence supports. Its output is now ready for the nonlinear MLP transformation, where the <code>[B,128]</code> representation is mapped to the ReCast heads.</p><h2>MLP Feature Transformation and ReCast Heads</h2><p>How does the model turn a bidirectional temporal summary into a forecast that can correct its own coarse estimate? The answer has two stages. First, an MLP nonlinearly transforms the BiLSTM representation. Then two parallel linear heads produce a coarse forecast and an additive correction. Their sum is the final prediction.</p><p>The paper specifies this structure, but it does not specify the MLP hidden width or output width. The generated implementation therefore makes both widths explicit configuration parameters. They are implementation decisions, not paper-reported hyperparameters.</p><h3>From temporal representation to nonlinear features</h3><p>Let <code>H_t</code> denote the representation emitted by the preceding BiLSTM stage. With hidden size <code>H = 64</code> per direction and the selected terminal-state convention, <code>H_t</code> has width <code>2H = 128</code>. The MLP receives a batch of these vectors with shape <code>[B, 128]</code>, where <code>B</code> is the batch size.</p><p>The paper's <code>eq_09</code> describes the first fully connected transformation followed by an elementwise ReLU activation. In the implementation, the intermediate result is called <code>hidden</code> and corresponds to the planned symbol <code>z_t</code>. ReLU is the only nonlinear activation in this MLP stage. The paper's <code>eq_10</code> describes a second fully connected transformation, whose output is called <code>m_t</code> in the method description.</p><p>The generated <code>FeatureMLP</code> keeps those operations separate and checks the input width before applying them:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c197e140-c819-4315-8cdc-8fa74479cc24&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class FeatureMLP(nn.Module):
    """Apply the paper's two-layer nonlinear feature transformation."""

    def __init__(self, input_dim: int, hidden_dim: int, output_dim: int) -&gt; None:
        super().__init__()
        validate_mlp_dimensions(input_dim, hidden_dim, output_dim)

        self.input_dim = int(input_dim)
        self.hidden_dim = int(hidden_dim)
        self.output_dim = int(output_dim)

        self.first_linear = nn.Linear(self.input_dim, self.hidden_dim)
        self.activation = nn.ReLU()
        self.second_linear = nn.Linear(self.hidden_dim, self.output_dim)

    def forward(self, h: torch.Tensor) -&gt; torch.Tensor:
        """Map a batch of BiLSTM representations to MLP features."""
        if not isinstance(h, torch.Tensor):
            raise ShapeContractError("MLP input must be a torch.Tensor")
        if h.ndim != 2:
            raise ShapeContractError(
                "MLP input must have shape [batch, input_dim]; "
                f"received tensor with shape {tuple(h.shape)}"
            )
        if h.shape[1] != self.input_dim:
            raise ShapeContractError(
                f"MLP input feature width must be {self.input_dim}; "
                f"received {h.shape[1]}"
            )

        hidden = self.activation(self.first_linear(h))
        output = self.second_linear(hidden)
        return output</code></pre></div><p>The important ordering is <code>Linear -&gt; ReLU -&gt; Linear</code>. The second linear layer has no activation after it, so <code>m_t</code> can contain any real-valued feature, including negative values. The output shape is <code>[B, mlp_output_dim]</code>. Since the paper does not report <code>mlp_hidden_dim</code> or <code>mlp_output_dim</code>, a reproduction should record their chosen values in its configuration and metadata rather than presenting them as facts from the paper.</p><p>The implementation validates finite floating-point input and rejects a mismatched feature width. These checks enforce the interface between <code>BiLSTMEncoder</code> and <code>FeatureMLP</code>; they do not change the mathematical method.</p><h3>Two forecasts from the same representation</h3><p>The ReCast module consumes <code>m_t</code>, not the original BiLSTM state. It has two parallel linear heads:</p><ul><li><p>The coarse head produces <code>y_hat_c</code>, a scalar preliminary forecast for each sample. This is the role described by <code>eq_11</code>.</p></li><li><p>The residual head produces <code>r_hat</code>, a scalar learned correction. This is the role described by <code>eq_13</code>.</p></li><li><p>The final forecast <code>y_hat</code> is their elementwise sum, as described by <code>eq_14</code> and <code>eq_18</code>.</p></li></ul><p>Both branch outputs have shape <code>[B, 1]</code>. The shared input and matching output shapes are important: each sample's correction must be added to the coarse forecast from that same sample. The generated <code>ReCastHeads</code> class expresses this directly:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;905bd941-aa01-494a-acc5-28795df236f0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class ReCastHeads(nn.Module):
    """Parallel coarse and residual scalar forecast heads."""

    def __init__(self, input_dim: int) -&gt; None:
        super().__init__()
        if input_dim &lt;= 0:
            raise ShapeContractError("input_dim must be positive")

        self.input_dim = input_dim
        self.coarse_head = nn.Linear(input_dim, 1)
        self.residual_head = nn.Linear(input_dim, 1)

    def forward(self, m: torch.Tensor) -&gt; ForecastBatch:
        _validate_representation(m, self.input_dim)

        coarse = self.coarse_head(m)
        residual = self.residual_head(m)
        final = combine_recast_branches(coarse, residual)
        return ForecastBatch(final=final, coarse=coarse, residual=residual)</code></pre></div><p>The two heads are independent parameterized linear maps, but they receive the same MLP representation. The residual branch is not a separately trained recurrent network and has no additional activation in the default implementation. Most importantly, the generated implementation does not require a precomputed residual target for the forward pass.</p><p>The branch-combination helper makes the additive invariant explicit and rejects incompatible shapes or devices:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c50995f6-0f23-4883-8b94-5efccb9a70c5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def combine_recast_branches(
    coarse: torch.Tensor, residual: torch.Tensor
) -&gt; torch.Tensor:
    """Add aligned coarse and residual forecast branches."""
    _validate_forecast_branch(coarse, "coarse")
    _validate_forecast_branch(residual, "residual")
    if coarse.shape != residual.shape:
        raise ShapeContractError(
            "coarse and residual forecasts must have identical shapes; "
            f"got {tuple(coarse.shape)} and {tuple(residual.shape)}"
        )
    if coarse.device != residual.device:
        raise ShapeContractError("coarse and residual forecasts must share a device")

    return coarse + residual</code></pre></div><p>This addition preserves the computation graph. Gradients from the final forecast remain connected to both heads, the MLP, and the BiLSTM encoder. Consequently, the default joint MSE objective can train the complete architecture end to end. The code does not claim that the coarse branch is independently accurate; its role is to provide a component around which the learned correction can refine the forecast.</p><h3>How the complete forward method connects the pieces</h3><p><code>BiLSTMMLPReCast.forward</code> composes the three stages in the same order as the paper's overall pipeline. It validates an input with shape <code>[B, L, D]</code>, obtains the temporal representation, transforms it with the MLP, and passes that representation to the ReCast heads:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;68d77230-05b9-4388-b3a2-14f7dcbe0032&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def forward(
    self, x: torch.Tensor, return_branches: bool = True
) -&gt; Union[ForecastBatch, torch.Tensor]:
    """Compute a one-step forecast from a normalized input window."""
    if not isinstance(return_branches, bool):
        raise TypeError("return_branches must be a bool")
    self.validate_input(x)

    temporal_representation = self.encoder(x)
    mlp_representation = self.mlp(temporal_representation)
    outputs = self.recast(mlp_representation)
    if return_branches:
        return outputs
    return outputs.final</code></pre></div><p>The complete model therefore maps each normalized lookback window to a normalized scalar next-day forecast. Returning a <code>ForecastBatch</code> is useful during training and analysis because it preserves <code>final</code>, <code>coarse</code>, and <code>residual</code>. Passing <code>return_branches=False</code> exposes only the final <code>[B, 1]</code> prediction for consumers that do not need branch diagnostics.</p><p><code>build_default_model</code> checks that the preprocessing feature count matches <code>config.input_dim</code>. With the six-field default, the path is conceptually <code>[B, 20, 6]</code> to <code>[B, 128]</code> to <code>[B, mlp_output_dim]</code> and finally to three <code>[B, 1]</code> tensors. If OHLCV mode is selected instead, the input width must be configured consistently; it cannot be silently substituted at model construction time.</p><h3>Worked example: tracking widths</h3><p>Suppose <code>B = 32</code>, <code>L = 20</code>, <code>D = 6</code>, and the BiLSTM uses <code>H = 64</code> per direction. The terminal representation has width <code>2H = 128</code>, so its batch shape is <code>[32, 128]</code>.</p><p>An explicit, illustrative MLP configuration might use the width sequence <code>128 -&gt; mlp_hidden_dim -&gt; mlp_output_dim</code>. The two names in the middle are configuration values, not reported paper values. If they were configured as 64 and 32, respectively, the MLP would produce <code>[32, 64]</code> after the first linear-ReLU transformation and <code>[32, 32]</code> after the second linear transformation. Each ReCast head would then map <code>[32, 32]</code> to <code>[32, 1]</code>, and the final result would also be <code>[32, 1]</code>.</p><p>The shape relationship, rather than those illustrative widths, is the essential contract:</p><ul><li><p>BiLSTM representation: <code>[B, 2H]</code>.</p></li><li><p>MLP output: <code>[B, mlp_output_dim]</code>.</p></li><li><p>Coarse forecast: <code>[B, 1]</code>.</p></li><li><p>Estimated correction: <code>[B, 1]</code>.</p></li><li><p>Final forecast: <code>[B, 1]</code> and equal to the branch sum.</p></li></ul><p>The generated test <code>test_final_equals_coarse_plus_residual</code> exercises this last relationship with small tensors. It is a specification of the intended invariant; the supplied run policy states that tests and code verification were not executed.</p><h3>Advanced detail: diagnostic residual versus learned correction</h3><p>The paper's <code>eq_12</code> defines a conceptual coarse residual as the target minus the coarse forecast. The generated <code>compute_coarse_residual</code> utility can calculate this quantity when a target is available. It should not be confused with <code>r_hat</code>, which is the residual head's learned prediction.</p><p><code>build_recast_diagnostics</code> exposes the three model outputs and optionally adds the conceptual coarse residual:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b5aac57c-92d4-4ec0-a771-d8950027c6ea&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">if target is not None:
    # Eq. 12: the conceptual residual is the target minus the coarse forecast.
    diagnostics["coarse_residual"] = compute_coarse_residual(target, coarse)

validate_diagnostic_invariants(diagnostics)
return diagnostics</code></pre></div><p>This diagnostic path is for inspection. The default implementation uses only the final forecast for optimization, so it does not introduce an auxiliary loss for matching <code>r_hat</code> to the conceptual residual. The diagnostic validator also checks the <code>eq_14</code> invariant that the final output equals the coarse output plus the learned correction.</p><p>Because all supplied canonical equation records are empty, this section intentionally references <code>eq_09</code> through <code>eq_18</code> by role without reproducing display formulas. Reconstructing LaTeX from damaged OCR or memory would exceed the evidence supplied for this reproduction.</p><h2>Joint MSE Training and Validation Selection</h2><p>How should the model learn when its prediction has two parts? The practical answer is: judge only the final forecast. The ReCast module produces a coarse estimate and a learned correction, but the generated implementation does not require the correction to match a separately prepared residual target. Instead, both branches contribute to the final prediction, and one mean squared error objective sends gradients through the complete BiLSTM-MLP-ReCast network.</p><p>This section follows the paper's <code>eq_12</code> and <code>eq_15</code> roles and the generated implementations of <code>compute_forecast_loss</code> and <code>fit_model</code>. The supplied equation records contain no canonical LaTeX, so no display equations are reproduced here. The mathematical meanings are explained in prose without reconstructing formulas from OCR.</p><h3>Diagnostic residual versus training target</h3><p>The paper describes a conceptual residual, identified here as <code>e_{t+1}</code>: the true normalized next-day value <code>y_{t+1}</code> minus the coarse forecast <code>y_hat_c</code>. This quantity answers an interpretive question: how far was the coarse branch from the target? In the generated code, <code>compute_coarse_residual</code> exposes that calculation for diagnostics.</p><p>It is important not to confuse this diagnostic with an additional supervised label. The default implementation does not create a second dataset of residual targets, and it does not add a separate residual loss. The learned residual branch produces <code>r_hat</code> from the MLP representation, while the training objective evaluates the final sum of the coarse and residual predictions.</p><p>The implementation makes this distinction explicit:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2127518b-c3e0-40c7-a25f-51c866fa9fc0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">
def compute_coarse_residual(
    target: torch.Tensor,
    coarse: torch.Tensor,
) -&gt; torch.Tensor:
    """Return the conceptual coarse residual ``target - coarse``.

    This diagnostic corresponds to Eq. 12.  It is not an auxiliary training
    target and is deliberately separate from the final-output MSE used for
    optimization.  The returned tensor has shape ``[batch, 1]``.
    """
    validate_loss_shapes(coarse, target)
    target_column = _as_column(target, "target")
    coarse_column = _as_column(coarse, "coarse")

    # Eq. 12: conceptual residual diagnostic, target minus coarse forecast.
    return target_column - coarse_column</code></pre></div><p>Both arguments are scalar-per-sample tensors. The helper accepts either <code>[batch]</code> or <code>[batch, 1]</code>, then canonicalizes them to <code>[batch, 1]</code>. It rejects empty, non-finite, differently shaped, differently typed, or differently located tensors. These checks are implementation safeguards around the paper's scalar-target setting; they are not additional modeling assumptions.</p><h3>The sole optimization objective</h3><p>The paper's <code>eq_15</code> identifies mean squared error over the final forecasts. For a batch, <code>y_i</code> is the normalized observed target, <code>y_hat_i</code> is the normalized final prediction, <code>N</code> is the number of scalar samples being averaged, and <code>Theta</code> denotes all trainable parameters. The loss <code>L(Theta)</code> is therefore a scalar mean of squared final-forecast errors.</p><p>In code, <code>compute_forecast_loss</code> implements that standard mean reduction:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a6268303-9987-4433-9ed0-7251b1c3cc2a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def compute_forecast_loss(
    prediction: torch.Tensor,
    target: torch.Tensor,
) -&gt; torch.Tensor:
    """Compute the joint final-forecast mean squared error.

    Only the final prediction participates in this objective; the coarse and
    residual branches receive gradients through their summed forecast in the
    model's forward pass.  The returned value is a scalar tensor suitable for
    ``backward()``.
    """
    validate_loss_shapes(prediction, target)
    prediction_column = _as_column(prediction, "prediction")
    target_column = _as_column(target, "target")

    # Eq. 15: mean squared error over the final forecast samples.
    return torch.mean((prediction_column - target_column) ** 2)</code></pre></div><p>The inputs have matching <code>[batch, 1]</code> meaning after canonicalization, and the return value has rank zero. Because the prediction is the ReCast sum, gradients flow backward through the residual head, coarse head, MLP, and BiLSTM whenever the model is used in training mode. No branch is optimized in isolation.</p><p>This also explains why the branch outputs should remain available even though only the final output enters the loss. A diagnostic can inspect whether the coarse forecast is systematically high or low and whether the learned correction is substantial, without changing the objective that trains the model.</p><h3>Adam configuration</h3><p>The paper's final configuration in <code>table_3</code> reports a learning rate of <code>1e-5</code>, batch size 32, and 64 epochs. These values are reproduction settings supplied by the paper. The generated optimizer helper uses the learning rate explicitly and leaves other Adam options at framework defaults because the supplied context does not specify weight decay, scheduler settings, or other optimizer modifications.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;1abfa6e2-a48c-4d6a-b33f-b322eda88e41&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def build_adam_optimizer(
    parameters: Iterable[torch.nn.Parameter],
    learning_rate: float,
) -&gt; torch.optim.Optimizer:
    """Construct an Adam optimizer for trainable model parameters."""
    validate_optimizer_config(learning_rate)
    parameter_list = list(parameters)
    if not parameter_list:
        raise ConfigurationError("parameters must contain at least one parameter")

    for index, parameter in enumerate(parameter_list):
        if not isinstance(parameter, torch.nn.Parameter):
            raise ConfigurationError(
                f"parameters[{index}] is not a torch.nn.Parameter instance"
            )
        if not parameter.requires_grad:
            raise ConfigurationError(
                f"parameters[{index}] does not require gradients"
            )

    return torch.optim.Adam(parameter_list, lr=float(learning_rate))</code></pre></div><p>The excerpt shows two responsibilities. First, <code>validate_optimizer_config</code> requires a finite positive learning rate. Second, the helper ensures that the optimizer receives at least one trainable <code>torch.nn.Parameter</code>. The code does not claim that the paper specified initialization, random seeds, gradient clipping, weight decay, or a scheduler; those choices remain outside the supplied evidence.</p><h3>One training batch</h3><p>The training method <code>train_bilstm_mlp_recast</code> maps to the training loop in <code>src/bilstm_recast/training/loops.py</code>. Each loader batch contains inputs with shape <code>[B, L, D]</code> and targets with shape <code>[B, 1]</code>. Here, <code>B</code> is the current batch size, <code>L</code> is the lookback length&#8212;20 in the default configuration&#8212;and <code>D</code> is the selected feature count.</p><p>The essential sequence for one batch is:</p><ol><li><p>Move the input and target tensors to the configured device.</p></li><li><p>Clear accumulated gradients with <code>optimizer.zero_grad</code>.</p></li><li><p>Run the model forward pass and extract its final <code>[B, 1]</code> forecast.</p></li><li><p>Compute the final-output MSE against the normalized target.</p></li><li><p>Backpropagate with <code>loss.backward()</code>.</p></li><li><p>Update every participating parameter with <code>optimizer.step()</code>.</p></li></ol><p>The generated loop expresses the central part of this procedure as follows:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;3147ff1c-a494-4873-a740-702a294cb677&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        optimizer.zero_grad(set_to_none=True)
        # Eq. 15: final-output mean squared error is the sole training objective.
        prediction = _prediction_tensor(model(inputs))
        if prediction.shape[0] != targets.shape[0]:
            raise ShapeContractError(
                "model prediction and target batches must have equal sample counts"
            )
        if prediction.dtype != targets.dtype:
            raise ShapeContractError(
                f"prediction and target dtypes must match; got {prediction.dtype} and {targets.dtype}"
            )

        loss = compute_forecast_loss(prediction, targets)
        if not torch.isfinite(loss).item():
            raise FloatingPointError("training loss became non-finite")
        loss.backward()
        optimizer.step()</code></pre></div><p>The checks preserve the scalar-output contract and reject a non-finite objective before an optimizer update. They do not establish that a run converges; no execution or semantic verification occurred under the authoritative run policy.</p><h3>Epoch averages and unequal final batches</h3><p>The training loop records a sample-weighted epoch loss. After computing a batch mean, it multiplies that mean by the batch's sample count and adds the count to <code>total_samples</code>. The returned epoch value is <code>total_loss / total_samples</code>. This prevents a smaller final batch from receiving the same influence as a full batch merely because it is one batch.</p><p>The same final-output MSE is used for validation, but the model switches to evaluation mode and the pass runs inside <code>torch.no_grad()</code>. Validation therefore measures the current parameters without creating gradient updates. The generated <code>evaluate_loss</code> function still validates input shapes, finite values, prediction dtype, and sample alignment, but it does not call <code>backward</code> or <code>optimizer.step</code>.</p><h3>Selecting the best validation state</h3><p>The paper's training description supports selecting the configuration with the lowest validation MSE. The generated <code>fit_model</code> follows that policy for the configured number of epochs. It records one training loss and one validation loss per epoch, copies the model state whenever validation loss improves, and restores the copied state after the loop.</p><p>The relevant control flow is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8d0a7968-bcc8-4c0a-a096-b079673fa807&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    best_validation_loss = float("inf")
    best_state: dict[str, torch.Tensor] | None = None

    for _ in range(config.epochs):
        train_loss = train_one_epoch(model, train_loader, optimizer, device)
        validation_loss = evaluate_loss(model, validation_loader, device)
        train_losses.append(train_loss)
        validation_losses.append(validation_loss)

        if validation_loss &lt; best_validation_loss:
            best_validation_loss = validation_loss
            # Clone the state so later optimizer updates cannot mutate the checkpoint.
            best_state = copy.deepcopy(model.state_dict())

    if best_state is None:
        raise RuntimeError("no validation checkpoint was produced")
    model.load_state_dict(best_state)</code></pre></div><p><code>best_validation_loss</code> is initialized to infinity so the first finite validation result can become the initial checkpoint. <code>copy.deepcopy</code> matters: without an independent copy, later optimizer updates could alter the object intended to represent the earlier best state. After the final epoch, <code>load_state_dict</code> restores the parameters associated with the lowest observed validation MSE.</p><p>The test set is intentionally absent from this selection loop. It should be evaluated only after the best validation state has been chosen. This separation prevents test observations from influencing training or model selection.</p><p>The lower-level <code>BestCheckpoint</code> and <code>select_best_checkpoint</code> interfaces provide the same policy for explicit state histories. <code>select_best_checkpoint</code> pairs each validation score with a state, rejects empty or mismatched histories, and chooses the smallest finite nonnegative validation loss. <code>save_checkpoint</code> and <code>load_checkpoint</code> serialize that selected state locally; they do not introduce a test-based criterion.</p><h3>Worked trace</h3><p>For one batch, imagine <code>x</code> has shape <code>[32, 20, 6]</code> and <code>y</code> has shape <code>[32, 1]</code>. The model produces a <code>ForecastBatch</code> containing <code>final</code>, <code>coarse</code>, and <code>residual</code>, each with shape <code>[32, 1]</code>. The conceptual computation is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d57760c2-baea-4eca-8070-dd18698ecfce&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">outputs = model(x)
loss = compute_forecast_loss(outputs.final, y)
loss.backward()
optimizer.step()</code></pre></div><p>The important point is that <code>outputs.final</code> is the quantity passed to the loss. The coarse and residual tensors are not discarded by the model, but neither is assigned a separate target by default. Their contribution is learned jointly through the final prediction.</p><h3>What is and is not established</h3><p>The generated tests include specifications such as <code>test_training_uses_final_output_loss</code>, <code>test_validation_does_not_update_parameters</code>, and <code>test_best_checkpoint_uses_lowest_validation_loss</code>. They document intended mathematical and training invariants, but they were not executed. Static verification and semantic code verification were also skipped under the run policy. Consequently, this section explains the planned behavior and the supplied code mapping; it does not claim successful execution, convergence, or reproduction of the paper's reported metrics.</p><p>Likewise, the paper's learning rate, batch size, and epoch count are documented facts, while the absence of a scheduler, weight decay, gradient clipping, seed, and initialization policy is an evidence boundary. Exact convergence and exact agreement with reported results cannot be inferred from the training loop alone, especially because data acquisition and several architecture and preprocessing details remain underspecified.</p><h2>Normalized Metrics and ReCast Diagnostics</h2><p>How do we tell whether the forecast is useful, and how can we inspect what the ReCast branches are doing? The implementation answers these questions separately. Final forecasts are evaluated with regression and percentage metrics, while the coarse and residual branches are exposed for diagnosis. Neither diagnostic output nor an evaluation metric changes the training objective.</p><p>The paper states that evaluation uses normalized targets. Therefore, the generated reporting path compares normalized predictions with normalized targets and does not silently inverse-transform them. This matters: a metric computed after restoring the original price scale would not be directly comparable with the paper's stated normalized-scale protocol.</p><p>The supplied records for <code>eq_19</code>, <code>eq_20</code>, <code>eq_21</code>, <code>eq_12</code>, <code>eq_13</code>, and <code>eq_14</code> contain no canonical equation LaTeX. The descriptions below therefore explain their mathematical roles and code mappings without reconstructing formulas from OCR or memory.</p><h3>MSE, RMSE, and R2</h3><p>Mean squared error, or MSE, averages the squared difference between each normalized target <code>y_i</code> and its normalized prediction <code>y_hat_i</code>. Squaring makes large errors count more heavily and produces a nonnegative scalar. In the generated code, <code>mse_metric</code> accepts either a one-dimensional array with shape <code>[n]</code> or a singleton-column array with shape <code>[n,1]</code>, validates alignment and finiteness, and returns one Python <code>float</code>. This implements the role assigned to <code>eq_19</code>.</p><p>Root mean squared error, or RMSE, is the square root of MSE. It remains on the normalized scale, but its units correspond more directly to the normalized target than squared error does. <code>rmse_metric</code> deliberately calls <code>mse_metric</code> rather than duplicating its validation and reduction logic. This maps to <code>eq_20</code>.</p><p>R2 compares the prediction error with a baseline that always predicts the observed-target mean <code>y_bar</code>. A value near 1 indicates close agreement with the observations; a value of 0 means no improvement over that mean baseline under the standard definition, and a negative value is possible when predictions are worse than the baseline. <code>r2_metric</code> uses the test targets' observed mean, matching the role of <code>eq_21</code>.</p><p>The implementation also makes the otherwise undefined constant-target case explicit. If all targets have zero variance, a perfect prediction returns <code>1.0</code>; an imperfect prediction returns <code>0.0</code>. This is an implementation policy for producing finite output, not an additional paper claim.</p><p>Here is the focused implementation for the three regression metrics:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;eccb1b3b-5978-47b0-bbe2-849f261fcf69&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def mse_metric(targets: np.ndarray, predictions: np.ndarray) -&gt; float:
    """Compute mean squared error on aligned target and prediction arrays.

    This is the standard mean over per-sample squared errors and corresponds
    to the evaluation objective identified as Eq. 19 in the supplied paper
    context. Inputs may be ``[n]`` or singleton-column ``[n, 1]`` arrays.
    """
    target_vector, prediction_vector = _validated_vectors(targets, predictions)
    # Eq. 19: normalized mean squared error.
    squared_errors = np.square(prediction_vector - target_vector)
    return float(np.mean(squared_errors, dtype=np.float64))


def rmse_metric(targets: np.ndarray, predictions: np.ndarray) -&gt; float:
    """Compute root mean squared error from the normalized input arrays."""
    # Eq. 20: root mean squared error is the square root of Eq. 19.
    return float(np.sqrt(mse_metric(targets, predictions)))


def r2_metric(targets: np.ndarray, predictions: np.ndarray) -&gt; float:
    """Compute standard coefficient of determination using the observed mean.

    For a non-constant target, this returns ``1 - SSE/SST`` and therefore may
    be negative when predictions are worse than the observed-mean baseline.
    When all targets are equal, the conventional denominator is zero. This
    implementation uses a finite, explicit policy: a perfect prediction
    returns ``1.0`` and any imperfect prediction returns ``0.0``. The policy
    avoids an undefined NaN while preserving the intuitive perfect-fit case.
    """
    target_vector, prediction_vector = _validated_vectors(targets, predictions)
    # Eq. 21: coefficient of determination relative to the observed target mean.
    observed_mean = float(np.mean(target_vector, dtype=np.float64))
    residual_sum_of_squares = float(
        np.sum(np.square(target_vector - prediction_vector), dtype=np.float64)
    )
    total_sum_of_squares = float(
        np.sum(np.square(target_vector - observed_mean), dtype=np.float64)
    )

    if total_sum_of_squares == 0.0:
        if residual_sum_of_squares == 0.0:
            return _ZERO_VARIANCE_R2_PERFECT
        return _ZERO_VARIANCE_R2_IMPERFECT

    return float(1.0 - residual_sum_of_squares / total_sum_of_squares)</code></pre></div><p>The comments identify the mapping to <code>eq_19</code> through <code>eq_21</code>; they are not reconstructed equation text. The strict input validation also protects the important invariant that predictions and targets remain one-to-one and on the same scale.</p><h3>A small metric example</h3><p>Suppose the normalized targets are <code>[0.0, 1.0]</code> and the predictions are <code>[1.0, 3.0]</code>. The squared errors are <code>1.0</code> and <code>4.0</code>, so MSE is their mean, <code>2.5</code>. RMSE is therefore the square root of <code>2.5</code>, approximately <code>1.5811</code>. For a separate R2 example, targets <code>[1.0, 2.0, 3.0]</code> and predictions <code>[1.0, 2.0, 2.0]</code> give an R2 of <code>0.5</code> under the observed-mean definition. These are hand-worked illustrations of the metric functions, not model results and not paper benchmarks.</p><h3>MAPE: relative error needs a policy</h3><p>Mean absolute percentage error, or MAPE, divides each absolute prediction error by the magnitude of its target and averages the resulting percentages. Normalization makes near-zero targets possible, so an unqualified MAPE implementation could divide by zero or produce an unstable number. The paper reports MAPE but does not specify how such denominators are handled.</p><p>The generated <code>mape_metric</code> therefore requires both a positive <code>epsilon</code> and an explicit <code>zero_policy</code>. The available policies are:</p><ul><li><p><code>epsilon_floor</code>: replace denominators smaller than <code>epsilon</code> with <code>epsilon</code>.</p></li><li><p><code>omit</code>: exclude near-zero target samples, requiring at least one remaining sample.</p></li><li><p><code>raise</code>: reject the calculation if any target is below the threshold.</p></li></ul><p>This is an implementation decision, and the selected policy should be retained in experiment metadata. The result is multiplied by 100, so it is reported as a percentage-like value. It must not be presented as exactly reproducing the paper's MAPE convention, because that convention is absent.</p><p>The policy boundary is explicit in the generated file:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;3886eeab-efbf-4a49-8ae3-c94ba27a0442&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">_SUPPORTED_ZERO_POLICIES: Final[frozenset[str]] = frozenset(
    {"epsilon_floor", "omit", "raise"}
)</code></pre></div><p>For example, the supplied test specification uses targets <code>[0.0, 2.0]</code>, predictions <code>[1.0, 1.0]</code>, <code>epsilon=1.0</code>, and <code>zero_policy="epsilon_floor"</code>. It expects <code>75.0</code> percent under that declared rule. The same input with <code>zero_policy="raise"</code> is expected to raise <code>DataValidationError</code>. These expectations describe the written test contract; tests were not executed under the run policy.</p><h3>Directional Accuracy is convention-dependent</h3><p>Directional Accuracy, or DA, asks whether predicted and actual movement directions agree. However, the paper does not define what constitutes a direction. Possible interpretations include comparing each forecasted change with the previous observed close, comparing consecutive actual and predicted sequences, or comparing the signs of absolute levels. These choices can produce different scores.</p><p>The generated <code>directional_accuracy</code> function refuses to hide that ambiguity. It supports two named conventions:</p><ul><li><p><code>change_vs_previous</code>: compare <code>actual - previous_actual</code> with <code>predicted - previous_actual</code>.</p></li><li><p><code>sign_of_values</code>: compare the signs of the supplied actual and predicted levels directly. This is an explicit alternative, not a claim about the paper's definition.</p></li></ul><p>The first convention requires a <code>previous_actual</code> array aligned with the target and prediction arrays. If it is missing, validation raises <code>DataValidationError</code> rather than silently guessing a prior value. Scores are returned as percentage-like values in <code>[0,100]</code>; zero signs match only other zero signs.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;454a1bb5-7490-44ba-b60b-45038f78f83e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def directional_accuracy(
    actual: np.ndarray,
    predicted: np.ndarray,
    previous_actual: np.ndarray | None,
    convention: str,
) -&gt; float:
    """Compute Directional Accuracy as a percentage in the range ``[0, 100]``.

    Under ``change_vs_previous``, actual direction is ``actual - previous`` and
    forecast direction is ``predicted - previous``.  Under ``sign_of_values``,
    the signs of the two supplied levels are compared directly.  In both
    conventions, zero signs are handled deterministically: two zero signs
    match, while a zero sign and a nonzero sign do not.

    The source paper does not define this metric's convention, so callers must
    pass one explicitly and should persist that choice in experiment metadata.
    """
    validate_directional_inputs(actual, predicted, previous_actual, convention)

    actual_vector = _as_finite_vector(actual, "actual")
    predicted_vector = _as_finite_vector(predicted, "predicted")

    if convention == "change_vs_previous":
        # Direction is measured relative to the preceding observed target.
        previous_vector = _as_finite_vector(previous_actual, "previous_actual")
        actual_direction = np.sign(actual_vector - previous_vector)
        predicted_direction = np.sign(predicted_vector - previous_vector)
    else:
        actual_direction = np.sign(actual_vector)
        predicted_direction = np.sign(predicted_vector)

    matches = actual_direction == predicted_direction
    return float(np.mean(matches, dtype=float) * 100.0)</code></pre></div><p>The aggregate <code>evaluate_forecast_metrics</code> function reports DA as undefined when the configured convention is <code>change_vs_previous</code> but no preceding values are supplied. Thus, a missing DA value means that the required convention inputs were not available; it is not silently treated as zero.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e4a0579a-dbc2-4b35-a0f4-ae7d1b535921&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">if previous_targets is None and metric_config.directional_convention == "change_vs_previous":
    directional: float | None = None
else:
    directional = _metric_value(
        directional_accuracy(
            targets,
            predictions,
            previous_targets,
            metric_config.directional_convention,
        ),
        "directional_accuracy",
    )</code></pre></div><h3>Aggregate reporting stays on the normalized scale</h3><p><code>evaluate_forecast_metrics</code> combines MSE, RMSE, R2, MAPE, and DA into the repository's <code>MetricResults</code> record. It records the scale and the metric conventions, including the MAPE epsilon, MAPE zero policy, directional convention, and whether DA was defined. <code>format_metric_report</code> then labels the output as normalized, while <code>compare_with_paper_targets</code> calculates differences from supplied benchmark values without claiming reproduction.</p><p>This separation is important. The paper's reported values&#8212;for example, Microsoft RMSE <code>0.052044</code> and R2 <code>0.906752</code>&#8212;are comparison targets from the supplied results, not expected outputs guaranteed by the generated implementation. Data-download settings, architecture widths, initialization, seeds, split boundaries, and metric conventions remain incomplete.</p><h3>ReCast diagnostics are not another evaluation metric</h3><p>The ReCast method has three useful branch-level quantities: the coarse forecast, the learned residual correction, and their final sum. The generated <code>build_recast_diagnostics</code> function exposes these tensors with shape <code>[batch,1]</code>. If a target is supplied, it also computes the conceptual coarse residual: target minus coarse forecast. This corresponds to the diagnostic role of <code>eq_12</code>.</p><p>The residual branch output itself corresponds to <code>eq_13</code>, while the final branch sum corresponds to <code>eq_14</code>. The diagnostics do not introduce a second loss. The default training path still optimizes only the final forecast MSE, as described in the previous section.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8dad3f01-3cb5-47a9-8c4d-8b52520f813b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">if target is not None:
    # Eq. 12: the conceptual residual is the target minus the coarse forecast.
    diagnostics["coarse_residual"] = compute_coarse_residual(target, coarse)

validate_diagnostic_invariants(diagnostics)
return diagnostics</code></pre></div><p><code>validate_diagnostic_invariants</code> checks that all branch tensors are floating-point <code>[batch,1]</code> values with matching shapes, devices, and dtypes. It also checks the central ReCast invariant exactly: the final tensor must equal the coarse tensor plus the residual tensor. The conceptual coarse residual is intentionally different from the learned correction: it is calculated from the target and coarse forecast for inspection, whereas the residual head produces a prediction from the MLP representation.</p><p>This distinction prevents a common interpretation error. A diagnostic residual can reveal how far the coarse branch was from the target, but it does not mean the implementation trained the residual head against a separately supervised residual target.</p><h3>What this evaluation layer establishes&#8212;and what it does not</h3><p>The metric layer establishes a reproducible software contract for aligned normalized arrays, explicit edge-case handling, and branch diagnostics. It does not resolve the paper's missing MAPE or DA definitions. It also does not inverse-transform values, simulate trades, or evaluate transaction costs; those activities are outside the reproduction scope.</p><p>The generated metric tests specify hand-computable relationships such as RMSE being the square root of MSE, R2 using the observed mean, and the coarse residual having the target-minus-coarse sign. The configuration tests also require a directional convention and preserve the paper targets as labeled records. These tests were generated as specifications but were not run. Static verification, semantic code verification, and code execution were all skipped under the authoritative run policy.</p><h2>Experiment Runner, Reports, and Benchmark Targets</h2><p>How do the separate preprocessing, model, training, and metric modules become one reproducible stock experiment? The experiment runner acts as the conductor. It accepts one stock table, applies the configured decisions in order, trains the complete model, evaluates the held-out test windows, and stores enough metadata to explain what happened. A paper benchmark is then used as a labeled reference point&#8212;not as proof that a local run reproduced the paper.</p><h3>From one stock table to one result record</h3><p>The paper describes separate experiments for Google, Amazon, and Microsoft. The generated implementation treats each stock independently. This means each stock receives its own chronological split and its own featurewise scaler fitted only to that stock's training rows. Sharing a scaler across stocks would be an additional design choice not specified by the paper.</p><p>The public entry point is <code>run_stock_experiment</code> in <code>src/bilstm_recast/experiments/runner.py</code>. Its input is a pandas <code>DataFrame</code>, a non-empty stock <code>symbol</code>, and a validated <code>ReproductionConfig</code>. Its output is an <code>ExperimentResult</code> containing the trained model, training history, normalized test metrics, and metadata. The runner rejects the wrong input types, invalid symbols, malformed split ratios, incompatible feature dimensions, empty evaluation loaders, and prediction tensors that do not have shape <code>[batch, 1]</code>.</p><p>The sequence of operations is deliberately explicit:</p><ol><li><p>Apply the configured date range and validate the selected columns.</p></li><li><p>Split raw rows chronologically into train, validation, and test partitions.</p></li><li><p>Fit a scaler on training feature rows only.</p></li><li><p>Transform each partition and construct split-local lookback windows.</p></li><li><p>Create loaders and build the model using the selected feature count.</p></li><li><p>Train with final-output MSE and select using validation MSE.</p></li><li><p>Collect test predictions in chronological loader order.</p></li><li><p>Compute normalized metrics and attach the preprocessing and ambiguity decisions to the result.</p></li></ol><p>The following excerpt shows the central preparation and training handoff. The comments identify the paper mappings, while the split-before-window behavior remains an implementation decision made to enforce a clear leakage boundary.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8a3a55e4-e948-4dd3-8594-31651e434668&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">split = split_chronologically(prepared_frame, ratios)
assert_split_integrity(split)

train_values = split.train.loc[:, feature_columns].to_numpy(dtype=np.float64)
scaler = fit_training_scaler(train_values)
lookback = int(config.model.lookback)
window_splits = make_window_splits(
    split.as_dict(),
    scaler,
    feature_columns,
    target_column,
    lookback,
)  # Eq. 1: construct split-local historical windows and next-step targets.

train_loader = make_data_loader(
    window_splits["train"], int(config.training.batch_size), shuffle=True
)
validation_loader = make_data_loader(
    window_splits["validation"], int(config.training.batch_size), shuffle=False
)
test_loader = make_data_loader(
    window_splits["test"], int(config.training.batch_size), shuffle=False
)

if int(config.model.input_dim) != len(feature_columns):
    raise ConfigurationError(
        "Model input_dim does not match the selected feature mode: "
        f"input_dim={config.model.input_dim}, features={len(feature_columns)}."
    )
model = build_default_model(len(feature_columns), config.model)
device = _training_device(config)
model.to(device)
history = fit_model(model, train_loader, validation_loader, config.training)  # Eq. 15: final-output MSE training.</code></pre></div><p>Here, the input windows follow the <code>eq_01</code> method contract: each sample is a sequence with shape <code>[L, D]</code>, batched as <code>[B, L, D]</code>, and paired with a next-step scalar target. The model implements the learned mapping described by <code>eq_02</code>; its final output is obtained only after the BiLSTM, MLP, and ReCast stages described earlier. The <code>fit_model</code> call uses the <code>eq_15</code> training objective and returns a <code>TrainingHistory</code> record rather than a paper result claim.</p><p>The test split is deliberately absent from the training call. It is used only after model selection:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;3433311b-9dda-43e8-a83c-891ff751ff77&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">predictions, targets = _collect_predictions(model, test_loader, device)  # Eq. 2: one-step forecast mapping.

metrics = evaluate_forecast_metrics(
    predictions,
    targets,
    config.metrics,
    previous_targets=previous_targets,
)  # Eq. 19&#8211;21: normalized MSE, RMSE, and R2 evaluation.</code></pre></div><p><code>_collect_predictions</code> switches the model to evaluation mode, disables gradient tracking, preserves batch order, and concatenates prediction and target arrays. It fails explicitly if the loader has no batches or if the model returns a non-scalar output. <code>evaluate_forecast_metrics</code> then maps the aligned normalized arrays to the regression metrics associated with <code>eq_19</code>, <code>eq_20</code>, and <code>eq_21</code>. The runner does not inverse-transform these values, matching the paper's stated normalized evaluation policy.</p><h3>Metadata is part of the result</h3><p>A numerical metric without its preprocessing decisions is difficult to interpret. The runner therefore records choices such as the feature mode, selected columns, target column, split policy, scaler scope, terminal-state convention, residual-loss policy, and metric scale. It also records raw row counts, window sample counts, scaler extrema, and the validation-selection rule.</p><p>This distinction is important. The paper supplies the 20-day window, chronological split ratios, and training-only normalization policy. The generated project additionally records that it uses split-local windows, one scaler per stock, concatenated terminal bidirectional states, and a single final-output MSE. Those latter details make the implementation reproducible as a software decision, but they should not be presented as unambiguous paper facts.</p><p>For multiple stocks, <code>run_multi_stock_experiment</code> simply calls the single-stock runner once per mapping entry and returns a dictionary keyed by symbol. Its independent calls ensure that one stock's scaler or result metadata is not silently reused for another stock.</p><h3>Local command-line workflow</h3><p>The generated command-line entry point is <code>scripts/run_reproduction.py</code>. It requires local CSV paths and intentionally performs no network acquisition. This is an implementation boundary: the paper names Yahoo Finance as the data source, but the supplied code does not invent a downloader API, ticker adjustment behavior, or credentials workflow.</p><p>A local invocation has this conceptual form:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;2021aa66-afc6-4c04-8147-3492c4e80b36&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">python scripts/run_reproduction.py \
  --input data/GOOG.csv \
  --symbol GOOG \
  --output artifacts/GOOG</code></pre></div><p>The parser also exposes the six-field versus OHLCV choice and optional overrides for epochs, batch size, learning rate, and device. Before delegation, <code>main</code> validates the configuration, loads the local frames, prints the selected decisions, runs the experiments, and saves each result. The default values remain the reported <code>table_3</code> settings unless explicitly overridden.</p><p>The corresponding public call sequence can be read as follows:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;06561bc9-113d-47f3-a2ba-32fb484f18cc&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from pathlib import Path

from bilstm_recast.data.loaders import load_stock_csv
from bilstm_recast.experiments.runner import run_stock_experiment
from bilstm_recast.io.results import save_experiment_result

frame = load_stock_csv(Path("data/GOOG.csv"))
result = run_stock_experiment(frame, "GOOG", config)
save_experiment_result(result, Path("artifacts/GOOG"))</code></pre></div><p>This excerpt is a workflow illustration, not a report that the command was run. <code>load_stock_csv</code> returns a local pandas frame and fails at the input boundary when the path or required schema is invalid. <code>run_stock_experiment</code> performs the full preparation, training, and evaluation sequence. <code>save_experiment_result</code> writes local artifacts and rejects unsupported or non-finite values rather than silently stringifying them.</p><p>The serialization module writes an <code>experiment_result.json</code>, metadata, history files, metric files, and optional diagnostics. The following focused excerpt shows the explicit local-directory behavior:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8b2c9064-905b-4da4-9acc-d986878602d0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def save_experiment_result(result: ExperimentResult, directory: Path) -&gt; None:
    """Write an experiment result and its reportable components locally.

    The primary ``experiment_result.json`` preserves every dataclass field,
    including configuration and ambiguity metadata. Additional files make
    histories, metrics, and diagnostics convenient to inspect. Model modules
    or optimizer objects are not accepted as serializable fields; callers
    should use the checkpoint utilities for model parameters.
    """
    if not isinstance(directory, Path):
        directory = Path(directory)
    if not isinstance(result, ExperimentResult):
        raise DataValidationError("result must be an ExperimentResult instance")

    try:
        directory.mkdir(parents=True, exist_ok=True)
    except OSError as exc:
        raise DataValidationError(f"could not create result directory {directory}: {exc}") from exc</code></pre></div><p>The saved metadata should be read alongside the metrics. In particular, it identifies whether the run used <code>six_field</code> or <code>ohlcv</code>, which target column was selected, how MAPE near-zero values were handled, and which Directional Accuracy convention was requested. Without those fields, two numerically different runs could be mistaken for equivalent reproductions.</p><h3>Reported targets versus calculated metrics</h3><p>The generated <code>src/bilstm_recast/experiments/targets.py</code> module stores paper values in immutable <code>PaperTarget</code> records. These values are comparison references only. They are not used for training, checkpoint selection, or automatic pass/fail decisions.</p><p>The supplied context gives complete-model MSE, RMSE, and R2 values for Google in <code>table_4</code>, Amazon in <code>table_5</code>, and Microsoft in <code>table_6</code>. It also supplies full-model MAPE values for those stocks in <code>table_8</code>. The target registry includes only numerical values explicitly present in the supplied evidence; it does not invent missing baseline metrics or Directional Accuracy records.</p><p>A target lookup combines records from different source tables when the stock and variant match:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;0a16e0e9-f8c6-4dbe-9186-188e8ccdbee3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def lookup_target(symbol: str, variant: str) -&gt; dict[str, float] | None:
    """Return reported metrics for a symbol/model pair, or ``None`` if absent.

    Records from multiple source tables are combined, which allows the full
    model's regression metrics from tables 4--6 and its separately reported
    MAPE from table 8 to be returned together. The returned dictionary is a new
    mutable copy and can therefore be safely edited by a caller.
    """
    canonical_symbol = _canonical_symbol(symbol)
    canonical_variant = _canonical_variant(variant)
    combined: dict[str, float] = {}

    for target in paper_targets():
        if (
            target.symbol == canonical_symbol
            and _canonical_variant(target.variant) == canonical_variant
        ):
            combined.update(target.metrics)

    return combined or None</code></pre></div><p>For example, a conceptual comparison for the Microsoft complete model would look up <code>MSFT</code> and <code>BiLSTM-MLP-ReCast</code>, then compare the resulting dictionary with the calculated <code>MetricResults</code>. The reporting function <code>compare_with_paper_targets</code> is intended to return numerical differences; it does not turn those differences into a reproduction claim. A close value can still arise from a different data adjustment setting, a different MLP width, a different random initialization, or another unresolved choice.</p><h3>Worked example: one local stock, end to end</h3><p>Consider a local <code>GOOG.csv</code> file that satisfies the selected schema. The workflow is:</p><ol><li><p><code>load_stock_csv</code> reads the table without contacting Yahoo Finance.</p></li><li><p><code>run_stock_experiment(frame, "GOOG", config)</code> filters the configured date interval and chooses the explicit feature mode.</p></li><li><p><code>split_chronologically</code> creates ordered train, validation, and test row partitions.</p></li><li><p><code>fit_training_scaler</code> learns per-feature extrema from the training rows only.</p></li><li><p><code>make_window_splits</code> produces <code>[samples, 20, D]</code> inputs and <code>[samples, 1]</code> normalized next-close targets within each partition.</p></li><li><p><code>build_default_model</code> constructs the configured BiLSTM-MLP-ReCast network.</p></li><li><p><code>fit_model</code> trains for the configured number of epochs and retains the lowest-validation-MSE state according to the generated training policy.</p></li><li><p>The selected model produces normalized test predictions, which are passed to <code>evaluate_forecast_metrics</code>.</p></li><li><p><code>save_experiment_result</code> writes metrics, history, and decisions under <code>artifacts/GOOG</code>.</p></li><li><p><code>lookup_target("GOOG", "BiLSTM-MLP-ReCast")</code> retrieves the supplied paper targets for comparison, if desired.</p></li></ol><p>This example explains the data flow but does not provide calculated numbers: no execution occurred under the run policy. Likewise, a synthetic demonstration would exercise the same interfaces while remaining demonstration-only rather than evidence about Yahoo Finance or the paper's reported tables.</p><h3>What the benchmark comparison can and cannot say</h3><p>A comparison can say that a calculated local metric differs from, or is numerically near, a supplied reported value under the documented configuration. It cannot by itself establish exact reproduction. The supplied context leaves data-download and adjustment behavior, feature selection, target-column interpretation, split boundaries, MLP widths, initialization, random seeds, MAPE handling, Directional Accuracy, and baseline definitions incomplete.</p><p>The same caution applies to <code>table_7</code> and <code>table_8</code>. The paper names ablation and comparative variants, but the available records do not fully specify every architecture or provide every numerical value. The generated target registry therefore preserves known values and leaves unavailable comparisons explicit rather than fabricating them.</p><p>The planned static review would inspect imports, annotations, public interfaces, and equation-to-code traceability. Semantic review would inspect behavior such as training-only scaling, test-free selection, output alignment, and the final-forecast loss. Both verification categories were skipped under the authoritative run policy, as were execution and test runs. Consequently, this runner and its reports should be treated as a documented reproduction implementation, not as a verified reproduction of the paper's benchmark results.</p><h2>Baselines, Ablations, and Local Demonstration</h2><p>How should a reproduction handle a baseline that the paper names but does not describe well enough to rebuild? The safest answer is to keep the baseline visible, record why it is unavailable, and refuse to substitute an invented architecture. Otherwise, a comparison may look faithful while actually measuring a different model.</p><p>The paper reports comparisons involving CNN, LSTM, BiLSTM, MLP-ReCast, BiLSTM-ReCast, BiLSTM-MLP, and the complete BiLSTM-MLP-ReCast model. The supplied context gives a sufficiently detailed contract for the complete model: a bidirectional LSTM, a two-layer MLP, two parallel linear ReCast heads, and a final forecast trained with MSE. It does not provide all layer configurations, state-reduction rules, widths, or output heads needed to reproduce the other variants exactly.</p><p>This distinction concerns both <code>train_bilstm_mlp_recast</code> and <code>evaluate_forecast_metrics</code>. The training method can optimize the complete model, and the evaluation method can measure its predictions, but neither method can make an under-specified baseline precise. The generated implementation therefore treats the complete model as executable and the remaining named variants as explicit unavailable entries.</p><h3>Make support status part of the API</h3><p><code>VariantSpec</code> in <code>src/bilstm_recast/experiments/variants.py</code> is metadata for one architecture. It records the variant name, description, whether it uses the BiLSTM, MLP, and ReCast components, and whether it is supported. An unsupported entry must also carry an <code>unsupported_reason</code>. This turns an evidence limitation into a visible configuration result rather than a hidden default.</p><p>The supported registry contains only the complete architecture:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;0f0d9bc5-a792-4702-a88e-67866c62cba8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def supported_variant_specs() -&gt; dict[str, VariantSpec]:
    """Return variants whose architecture is sufficiently specified.

    The complete BiLSTM-MLP-ReCast model is the only variant with enough detail
    in the supplied context for a faithful implementation.  CNN, LSTM, BiLSTM,
    MLP-ReCast, BiLSTM-ReCast, and BiLSTM-MLP are reported by the paper, but
    their exact layer configurations and/or ablation construction are missing.
    """

    complete = VariantSpec(
        name=_COMPLETE_VARIANT_NAME,
        description=(
            "Two-layer bidirectional LSTM, two-layer MLP, and parallel coarse "
            "and residual linear ReCast heads."
        ),
        uses_bilstm=True,
        uses_mlp=True,
        uses_recast=True,
    )
    return {complete.name: complete}</code></pre></div><p>The name <code>bilstm_mlp_recast</code> identifies the executable complete model. The implementation's default model still follows the paper's one-step mapping represented by <code>eq_02</code>, and its training loop still applies the final-output objective represented by <code>eq_15</code>. No new equation is reproduced here because the supplied canonical LaTeX record for both identifiers is empty.</p><p><code>build_variant_model</code> checks the feature dimension and configuration before constructing <code>BiLSTMMLPReCast</code>. If a caller supplies an unsupported specification, it raises <code>UnsupportedVariantError</code> rather than silently mapping that request to the complete model. This is important for ablation fairness: using the full model under a different label would invalidate the comparison.</p><h3>Keep paper-named ablations without fabricating them</h3><p><code>build_ablation_plan</code> in <code>src/bilstm_recast/experiments/ablation.py</code> creates an <code>AblationPlan</code>. The plan contains the complete model and the paper-named alternatives. Each unavailable alternative has a concrete reason, such as missing CNN kernel and pooling details or an unspecified non-ReCast forecast head.</p><p>The plan separates executable and non-executable entries through <code>supported()</code> and <code>unsupported()</code>. <code>run_supported_ablations</code> executes only entries with an explicit contract. In the generated implementation, that means the complete BiLSTM-MLP-ReCast model is the only executable entry. <code>record_ablation_note</code> provides a place to preserve an unresolved design decision instead of losing it in an experiment log.</p><p>This is an implementation policy derived from the evidence boundary, not a claim that the paper's ablation experiments are unimportant. Tables 7 and 8 remain relevant comparison references, but the supplied context does not specify enough architecture detail to reproduce every row safely. A future implementation could add a baseline once its exact input representation, layers, widths, state handling, and output mapping were established.</p><h3>Treat benchmark records as references, not expected outputs</h3><p><code>PaperTarget</code> in <code>src/bilstm_recast/experiments/targets.py</code> stores a symbol, model variant, metric mapping, and source table identifier. <code>paper_targets()</code> includes only numerical values supplied in the context. For example, the complete-model regression targets include the reported values for Google, Amazon, and Microsoft from <code>table_4</code>, <code>table_5</code>, and <code>table_6</code>; the supplied context also provides complete-model MAPE values from <code>table_8</code>.</p><p><code>lookup_target</code> merges available records for a symbol and variant and returns a caller-owned dictionary. That makes it possible for reporting code to compare calculated normalized metrics with the reported records while keeping their provenance visible. The values are targets for comparison, not guaranteed outputs of this implementation. Differences do not by themselves identify which unspecified preprocessing or training choice caused them.</p><h3>Use synthetic data only to exercise the pipeline</h3><p>When Yahoo Finance source files are unavailable, the generated project provides local synthetic data. <code>make_synthetic_stock_frame</code> creates a deterministic business-day table with <code>Date</code>, <code>Open</code>, <code>High</code>, <code>Low</code>, <code>Close</code>, <code>Adjusted Close</code>, and <code>Volume</code>. It enforces positive numeric values and basic OHLC relationships. <code>make_synthetic_stock_frames</code> creates one independent frame per symbol label using stable symbol-derived seed offsets.</p><p>These frames are not observations for Google, Amazon, or Microsoft. They are demonstration inputs only. In particular, a metric calculated from them is not evidence for the paper's reported results, and it should not be compared as though it came from the paper's date interval.</p><p>The module's documentation makes that boundary explicit:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;05f9ec19-33a8-48e5-baa3-43d229052acc&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">"""Deterministic synthetic stock-like data for local demonstrations.

The generated data is not Yahoo Finance data and must not be used as a claim of
paper-result reproduction.  It exists only to exercise the local preprocessing
and modeling pipeline when source data is unavailable.
"""</code></pre></div><h3>Walk through the local demonstration</h3><p><code>examples/minimal_local_demo.py</code> provides the planned end-to-end walkthrough. It creates a deterministic synthetic frame, obtains <code>ReproductionConfig.default()</code>, adjusts the small demonstration split, validates the configuration, and sends the frame to <code>run_stock_experiment</code>.</p><p>The split adjustment is deliberate. With only 120 rows, a 70:15:15 split would leave fewer than 20 rows in each 15% partition. Under the generated split-local window policy, validation and test partitions therefore need a larger share to contain a complete 20-step window. The production reproduction default remains the paper's 70:15:15 policy; the example uses 65:17.5:17.5 solely for its small demonstration dataset.</p><p>The central setup and call are:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2e36bb3a-0729-4b50-b896-06b824e1254c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">frame = make_synthetic_stock_frame(num_rows=num_rows, seed=0)

config = ReproductionConfig.default()
demo_feature_config = replace(
    config.feature,
    # This is a demonstration-only split adjustment necessitated by the
    # small default dataset and split-local 20-step windows.
    split_ratios=(0.65, 0.175, 0.175),
)
config = replace(config, feature=demo_feature_config)
validate_config(config)

# Eq. 01: the runner constructs ordered lookback windows and next-step targets.
# Eq. 02: the runner trains and evaluates the complete forecasting mapping.
# Eq. 15: training uses the final forecast's joint mean squared error.
# Eq. 19, Eq. 20, Eq. 21: normalized MSE, RMSE, and R2 are reported by the runner.
result = run_stock_experiment(frame, symbol="SYNTHETIC", config=config)</code></pre></div><p>The call follows the same conceptual path as a real local stock file: schema and date handling, chronological splitting, training-only scaling, split-local 20-step windows, model construction, joint training, validation-based selection, and normalized test reporting. The returned <code>ExperimentResult</code> is an experiment record, not a paper benchmark. This walkthrough describes the code path; it does not claim that the demonstration was executed under the current run policy.</p><p>The example also calls <code>torch.manual_seed(0)</code> as part of its local setup. That is a demonstration choice in the generated file, not a seed reported by the paper and not a guarantee of identical results across all PyTorch backends. More generally, the paper leaves initialization, seeds, data-download behavior, and several preprocessing details unresolved.</p><h3>What a fair extension would require</h3><p>To implement a named baseline responsibly, add its specification only after resolving the missing contracts. At minimum, record the input shape, temporal or convolutional layers, hidden or channel widths, sequence reduction, output head, dropout behavior, and training settings. Then add a distinct <code>VariantSpec</code>, model builder, and tests for its output shape and objective. Do not reuse the complete model merely to fill a table row.</p><p>The same rule applies to ablations. A paper label such as <code>BiLSTM-MLP</code> indicates which modules are present, but it does not by itself determine the exact output layer or dimensions. The generated plan therefore preserves the label and reason for unavailability while allowing only the fully specified architecture to run.</p><p>The practical result is a narrower but more honest reproduction scope: the complete BiLSTM-MLP-ReCast model can be traced through <code>eq_02</code> and <code>eq_15</code>, evaluated with the configured metric method, and demonstrated on clearly labeled local synthetic data. The named comparison and ablation targets remain useful references, but unsupported architectures are not presented as reproduced experiments.</p><h2>Invariant Checklist, Verification Status, and Reproduction Limits</h2><p>How can you decide whether a forecasting reproduction is trustworthy before interpreting its metrics? Start with the plumbing rather than the benchmark table. Confirm that future rows did not enter training, tensor shapes remain consistent, the final forecast is actually the sum of the two ReCast branches, and the test set is used only after model selection.</p><p>This section is a review checklist, not a verification report. Under the run policy, no code, tests, static checks, semantic checks, tutorial checks, or final review were executed.</p><h3>The mathematical traceability boundary</h3><p>The supplied equation records for <code>eq_01</code>, <code>eq_08</code>, <code>eq_14</code>, <code>eq_15</code>, <code>eq_19</code>, <code>eq_20</code>, and <code>eq_21</code> contain empty canonical LaTeX fields. They are therefore referenced here by role only; no display equations are reproduced or reconstructed from OCR.</p><p>The relevant meanings are still clear enough to guide review. <code>eq_01</code> concerns constructing an <code>L</code>-step historical window and its next-step target. <code>eq_08</code> concerns joining the forward and backward BiLSTM terminal states. <code>eq_14</code> concerns adding the coarse and residual forecasts. <code>eq_15</code> is the joint mean squared error objective. <code>eq_19</code>, <code>eq_20</code>, and <code>eq_21</code> represent MSE, RMSE, and the standard observed-mean coefficient of determination, respectively.</p><p>These roles map to <code>make_windows</code> in <code>src/bilstm_recast/data/windows.py</code>, <code>extract_terminal_bidirectional_state</code> in <code>src/bilstm_recast/model/representation.py</code>, <code>combine_recast_branches</code> in <code>src/bilstm_recast/model/recast.py</code>, <code>compute_forecast_loss</code> in <code>src/bilstm_recast/losses.py</code>, and the metric functions in <code>src/bilstm_recast/metrics/regression.py</code>.</p><h3>Shape and additive invariants</h3><p>The model's primary input contract is <code>[B,L,D]</code>: <code>B</code> is batch size, <code>L</code> is the lookback length, and <code>D</code> is the selected feature count. With the default configuration, <code>L = 20</code>. The terminal hidden-state tensor supplied by a two-layer bidirectional LSTM has shape <code>[N*2,B,H]</code>, which becomes <code>[4,B,64]</code> when <code>N = 2</code> and <code>H = 64</code>.</p><p>The selected final recurrent layer contributes one <code>H</code>-wide state for each direction. Concatenating them produces <code>[B,2H]</code>, or <code>[B,128]</code> under the default hidden size. The MLP then produces a configured feature width, and each ReCast head must return <code>[B,1]</code>. The target supplied to the loss must have a compatible scalar shape.</p><p>A central invariant is that the final forecast equals the coarse forecast plus the residual correction. The intended contract test is <code>test_final_equals_coarse_plus_residual</code> in <code>tests/test_model_contracts.py</code>; it checks the operation represented by <code>eq_14</code>. This is an algebraic interface requirement, not a claim that the generated test was run.</p><p>The terminal-state convention is itself an implementation decision. The paper's <code>eq_08</code> description does not unambiguously distinguish terminal hidden states from final-time-step outputs or another aggregation. The generated implementation chooses final-layer terminal states and records that choice in configuration and result metadata. A future review should ensure that the layer and direction indexing actually preserves this documented convention.</p><h3>Data and leakage checklist</h3><p>For each stock, a future enabled review should check the following sequence:</p><ul><li><p>The input table is chronological, filtered to the stated paper interval when applicable, and validated for required numeric columns and missing values.</p></li><li><p>The selected feature mode is recorded explicitly. The default is the six-field set containing <code>Open</code>, <code>High</code>, <code>Low</code>, <code>Close</code>, <code>Adjusted Close</code>, and <code>Volume</code>; <code>ohlcv</code> is an available five-field alternative. This resolves an ambiguity in the paper without pretending that the paper resolved it.</p></li><li><p>Raw observations are partitioned into disjoint chronological train, validation, and test portions using the configured 70:15:15 policy.</p></li><li><p>The featurewise Min-Max scaler is fitted only on training rows. Its parameters are reused unchanged for validation and test rows.</p></li><li><p>Windows are constructed locally within each split. Each <code>[L,D]</code> input is followed by the next chronological closing-price target, yielding arrays shaped <code>[samples,L,D]</code> and <code>[samples,1]</code>.</p></li><li><p>No training input or target contains validation or test observations under the selected split-boundary policy.</p></li></ul><p>The intended checks are specified in <code>tests/test_scaler_fits_training_rows_only</code>, <code>test_chronological_splits_are_disjoint</code>, and <code>test_windows_have_next_step_targets</code> in <code>tests/test_data_contracts.py</code>. They describe leakage and indexing invariants, but they do not constitute completed verification.</p><p>Splitting before window construction is a generated implementation decision. The paper specifies chronological splitting but does not state whether windows may cross split boundaries. The split-local policy sacrifices boundary-crossing samples in favor of an explicit no-leakage rule.</p><h3>Training and selection checklist</h3><p>The default training policy comes from the paper's reported final settings: batch size 32, Adam learning rate <code>1e-5</code>, and 64 epochs, alongside the model settings described above. The training loop should use the final forecast, not the conceptual coarse residual, as its optimization signal.</p><p><code>compute_coarse_residual</code> represents the diagnostic quantity <code>target - coarse forecast</code>, associated with <code>eq_12</code>. It is not a second target in the default reproduction. <code>compute_forecast_loss</code> applies one mean MSE to the final forecast, corresponding to <code>eq_15</code>, so gradients can reach the encoder, MLP, coarse head, and residual head jointly.</p><p>A future review should confirm that:</p><ul><li><p>Training batches run in model training mode and receive gradient updates.</p></li><li><p>Validation runs use evaluation mode and do not update parameters.</p></li><li><p>The best state is selected using the lowest validation MSE.</p></li><li><p>Test predictions are generated only after the selection decision.</p></li><li><p>No unspecified scheduler, weight decay, gradient clipping, early stopping, seed, or initialization policy is silently presented as a paper requirement.</p></li></ul><p>The intended semantic cases are named <code>test_training_uses_final_output_loss</code>, <code>test_validation_does_not_update_parameters</code>, and <code>test_best_checkpoint_uses_lowest_validation_loss</code> in <code>tests/test_training_selection.py</code>. Again, these are planned checks recorded in the generated project, not successful test results.</p><h3>Metric checklist</h3><p>The reporting path keeps predictions and targets on the normalized scale by default, matching the paper's stated evaluation policy. The regression helpers in <code>src/bilstm_recast/metrics/regression.py</code> are intended to preserve the following relationships:</p><ul><li><p>MSE, associated with <code>eq_19</code>, is nonnegative and measures the mean squared prediction error.</p></li><li><p>RMSE, associated with <code>eq_20</code>, is the square root of MSE.</p></li><li><p>R2, associated with <code>eq_21</code>, uses the mean of the observed targets as its baseline and requires explicit handling when target variance is zero.</p></li></ul><p>The tests <code>test_mse_is_mean_squared_error</code> and <code>test_rmse_is_sqrt_of_mse</code> in <code>tests/test_loss_metrics.py</code> specify hand-computable relationships. They were not executed.</p><p>MAPE and Directional Accuracy require additional policy choices. Normalized targets can approach zero, so <code>mape_metric</code> must use a declared near-zero rule such as an epsilon floor, omission, or rejection. The paper reports MAPE but does not define this behavior. Similarly, <code>directional_accuracy</code> requires a stated convention: for example, comparing predicted and actual changes relative to a previous value. The paper does not specify whether direction compares consecutive changes, forecast versus prior close, or another quantity.</p><p>The configuration test <code>test_directional_accuracy_requires_convention</code> is intended to prevent the implementation from silently inventing a missing prior value or convention. A future report should preserve the selected MAPE and Directional Accuracy policies alongside the metric values.</p><p>Branch diagnostics are separate from these metrics. <code>build_recast_diagnostics</code> can expose the coarse forecast, learned correction, final forecast, and target-minus-coarse diagnostic. It must not introduce an auxiliary training objective merely because those quantities are being inspected.</p><h3>A compact future-review example</h3><p>Consider one synthetic experiment created by <code>make_synthetic_stock_frame</code> and passed through the planned local runner. A future enabled verification run could apply this checklist without treating synthetic output as paper evidence:</p><ol><li><p>Confirm that the selected feature mode and model dimensions are present in the configuration metadata.</p></li><li><p>Confirm disjoint chronological partitions and training-only scaler parameters.</p></li><li><p>Confirm that each window has shape <code>[20,D]</code> and that its target is the immediately following closing value.</p></li><li><p>Pass a batch shaped <code>[B,20,D]</code> through the model and inspect <code>[B,1]</code> coarse, residual, and final outputs.</p></li><li><p>Confirm the exact additive relation between final, coarse, and residual outputs.</p></li><li><p>Confirm that the training loss consumes the final output and that validation selection does not inspect test metrics.</p></li><li><p>Confirm normalized MSE, RMSE, and R2 relationships, then record explicit MAPE and Directional Accuracy conventions.</p></li><li><p>Store the resulting metadata and distinguish calculated values from paper targets.</p></li></ol><p>This is a review plan, not a claim that the synthetic demonstration or any of these checks has run.</p><h3>Static versus semantic verification</h3><p>Static verification asks whether files can be inspected for structural consistency: syntax, imports, annotations, public interfaces, dependency direction, configuration defaults, and equation-to-code traceability. The local static-verification record is <code>verification_skipped</code>; no static checks were performed under the run policy.</p><p>Semantic verification asks a different question: whether behavior follows the method-level invariants. Examples include training-only scaling, correct next-step indexing, <code>[B,2H]</code> terminal representations, <code>[B,1]</code> branch outputs, final additive refinement, final-output MSE, and validation-only checkpoint selection. The semantic verifier was also skipped, and no repair path was requested.</p><p>Consequently, the generated test files should be read as intended specifications. There is no claim here that imports, tests, numerical relationships, or end-to-end execution succeeded.</p><h3>What prevents an exact paper reproduction?</h3><p>Several unresolved details can change the numerical result even when the high-level architecture is implemented faithfully:</p><ul><li><p>Yahoo Finance ticker names, download dates, adjustment behavior, and missing-observation handling are not fully specified.</p></li><li><p>The paper alternates between six listed fields and OHLCV, and does not completely distinguish the target's <code>Close</code> source from the presence of <code>Adjusted Close</code>.</p></li><li><p>The exact rule for windows at split boundaries is not stated.</p></li><li><p>MLP hidden and output widths are missing.</p></li><li><p>Initialization, random seeds, data-loader shuffling, and other training details are not fixed.</p></li><li><p>MAPE zero handling and Directional Accuracy definitions are absent.</p></li><li><p>The exact configurations of the named baseline and ablation variants are incomplete.</p></li><li><p>Canonical equation LaTeX was not supplied, and several extracted equations are damaged or ambiguous.</p></li></ul><p>The scope also deliberately excludes external indicators, multi-step forecasting, trading simulation, transaction costs, slippage, and portfolio evaluation. The generated project is a statistical one-step forecasting reproduction, not a trading system.</p><p>Finally, the reported Microsoft complete-model values&#8212;RMSE <code>0.052044</code> and R2 <code>0.906752</code>&#8212;along with the Google and Amazon values in <code>table_4</code> through <code>table_6</code>, are comparison targets from the paper. They are not guaranteed outputs, acceptance thresholds, or evidence that a local run reproduced the study. A meaningful comparison requires documenting the data source, feature mode, split policy, architecture widths, metric conventions, and all other choices that affect the result.</p><p>Use the button or URL below to download the source code.</p>
      <p>
          <a href="/__u/onepagecode.substack.com/p/quant-trading-a-bilstm-mlp-recast">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Quant Trading: AI Feature Engineering & Alpha Generation with LLMs (LightGBM Guide)]]></title><description><![CDATA[Implementing an end-to-end retrieval-augmented LLM feature discovery and long-short equity engine in Python.]]></description><link>https://onepagecode.substack.com/p/quant-trading-ai-feature-engineering</link><guid isPermaLink="false">https://onepagecode.substack.com/p/quant-trading-ai-feature-engineering</guid><dc:creator><![CDATA[Onepagecode]]></dc:creator><pubDate>Sat, 01 Aug 2026 20:04:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_W6n!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3078fdad-8e80-40a9-be17-640eb18fa447_1254x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Use the button at the end of this article to download the source code!</h2><p>Today we are learning about this paper: https://arxiv.org/abs/2602.00196</p><p>The paper evaluates retrieval-augmented large-language-model feature discovery for cross-sectional U.S. equity selection. GPT-4.1 generates executable, schema-constrained Python transformations from analyst-estimate, options, and price-volume data. Features are screened for implementability and point-in-time validity, then used individually or with vendor features in LightGBM models. Predictions are cross-sectionally standardized, winsorized, converted into dollar-neutral long-short portfolios, and evaluated primarily using forward-shifted returns from t+1 to t+2 over January 2019 through December 2024. Two generation routes are compared: fixed structured prompting and DSPy programmatic prompting optimized with MIPRO. The paper reports that accurate retrieval context, AI/traditional signal complementarity, and implementation choices materially affect performance.</p><h2>Implementation Assumptions</h2><ul><li><p>Use pandas, NumPy, SciPy, LightGBM, statsmodels, and optional cvxpy with CLARABEL.</p></li><li><p>Treat all supplied equation records as semantic identifiers only because their canonical LaTeX fields are empty.</p></li><li><p>Use signal<em>date, execution</em>date, and return_date explicitly throughout the pipeline.</p></li><li><p>Use January 2015&#8211;December 2018 only for discovery, validation, and hyperparameter selection, and January 2019&#8211;December 2024 for headline out-of-sample reporting.</p></li><li><p>Represent missing vendor data, unavailable external services, unspecified thresholds, and OCR-damaged optimization details through explicit configuration and audit metadata rather than invented defaults.</p></li><li><p>The implementation is a reproduction scaffold and cannot claim numerical agreement without the unavailable raw data, prompts, generated feature corpus, schemas, and hyperparameters.</p></li></ul><h2>Scope, Evidence Limits, and Reproduction Decisions</h2><p>What does it mean to reproduce this paper when the original vendor files, prompts, generated features, and exact model settings are unavailable? The practical answer is to reproduce the experiment's <em>architecture and decision points</em> without claiming that its reported numerical results have been recreated.</p><p>The paper studies retrieval-augmented language-model feature discovery for cross-sectional U.S. equity selection. Its workflow combines point-in-time daily data, generated structured-data features, LightGBM prediction, cross-sectional score processing, dollar-neutral portfolio construction, and delayed-return evaluation. The generated Python package mirrors those stages and keeps their unresolved choices visible.</p><h3>Three kinds of evidence</h3><p>It helps to separate three categories throughout this tutorial:</p><ul><li><p><strong>Paper facts.</strong> The paper uses a discovery period from January 2015 through December 2018 and reserves January 2019 through December 2024 for headline out-of-sample reporting. It compares conventional and AI-generated features, uses LightGBM, and evaluates long-short portfolios under timing and cost variations.</p></li><li><p><strong>Implementation decisions.</strong> The package exposes choices such as the target interval, implementation lag, smoothing window, rebalance frequency, cost model, market-impact coefficient, and optimization interpretation through configuration objects.</p></li><li><p><strong>Unavailable inputs.</strong> The supplied evidence does not include the complete EDI, TrueBeats, or SpiderRock schemas; the raw vendor data; the full generated-feature corpus; exact GPT-4.1 prompts; the retriever and knowledge base; DSPy/MIPRO settings; LightGBM hyperparameters; or factor data.</p></li></ul><p>Consequently, the implementation is best understood as an auditable reproduction scaffold. It can document how a faithful run should be organized, but it cannot establish numerical agreement with the paper.</p><h3>Why timing is configuration, not a hidden default</h3><p>The paper uses both ordinary <code>t</code>-to-<code>t+1</code> notation and a headline implementable evaluation based on returns from <code>t+1</code> to <code>t+2</code>. The package therefore names the signal and realized-return dates explicitly: <code>signal_date</code> is when features and predictions are formed, <code>execution_date</code> is when the configured implementation delay places the trade, and <code>return_date</code> identifies the later realized-return observation.</p><p><code>DateSplit</code> records the chronological boundary between discovery and out-of-sample data. <code>TimingConfig</code> records the target convention and implementation lag. These objects are immutable dataclasses, so a configuration cannot be silently altered after it has been created.</p><p>The following excerpt is copied from <code>src/stock_selection/config.py</code>. Notice that the two target names are explicit and that <code>implementation_lag</code> is validated separately from the target label.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f915f395-e1dd-4b1f-804c-ff1ce356f8cd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class TimingConfig:
    """Define the target interval and implementation lag from a signal date.

    ``t_to_t_plus_1`` is the standard one-period comparison.  The
    ``t_plus_1_to_t_plus_2`` target represents the paper's headline
    implementation-lag evaluation.  The lag is retained separately so that
    execution-lag sensitivity experiments remain explicit.
    """

    target: str = "t_to_t_plus_1"
    implementation_lag: int = 0
    signal_date_column: str = "signal_date"
    execution_date_column: str = "execution_date"
    return_date_column: str = "return_date"
    price_column: str = "adjusted_close"
    delisting_return_column: Optional[str] = "delisting_return"

    def __post_init__(self) -&gt; None:
        allowed = {"t_to_t_plus_1", "t_plus_1_to_t_plus_2"}
        if self.target not in allowed:
            raise ValueError(f"target must be one of {sorted(allowed)}")
        if isinstance(self.implementation_lag, bool) or not isinstance(
            self.implementation_lag, Integral
        ):
            raise TypeError("implementation_lag must be a nonnegative integer")
        if self.implementation_lag &lt; 0:
            raise ValueError("implementation_lag must be nonnegative")</code></pre></div><p>This is an implementation contract, not a claim that the supplied paper resolves every timing detail. A run should label lag zero as the standard or non-implementable comparison when appropriate, and label positive-lag configurations separately.</p><h3>Costs, smoothing, and optimization remain visible</h3><p>The same principle applies beyond timing. The paper contains two reported market-impact coefficient conventions: <code>k=0.3</code> in one transaction-cost description and <code>k=0.2</code> in another. The generated <code>CostConfig</code> does not merge these values. Instead, a caller can create two separately named configurations and preserve the selected coefficient in the audit record.</p><p>Likewise, smoothing and rebalancing are not interchangeable. A 5-day or 21-day trailing prediction average is a different experiment from unsmoothed predictions, and monthly rebalancing is different from daily rebalancing. The configuration must retain both the smoothing window and rebalance frequency so that cost and performance summaries cannot accidentally combine incompatible variants.</p><p>The optional optimization module requires even more caution. Several displayed optimization expressions and constraints are OCR-damaged in the supplied evidence. <code>OptimizationConfig</code> therefore requires an explicit <code>objective_interpretation</code> before optimization can be enabled. This is a deliberate refusal to present an inferred formulation as the paper's recovered equation.</p><p><code>UniverseConfig</code> records another boundary between paper intent and local data availability. It can request the paper's maximum of 2,500 eligible names per date, security-type exclusions, and optional liquidity filters, but the actual behavior depends on whether the local source files contain the required columns.</p><p><code>ModelConfig</code> similarly holds provider-neutral LightGBM settings. The paper identifies LightGBM and chronological estimation, but does not supply all objective, tree, learning-rate, regularization, missing-value, or retraining details. Those values must be supplied by the user rather than invented by this tutorial.</p><h3>Worked example: two timing configurations</h3><p>The following example creates one conventional target configuration and one headline delayed-target configuration. The code uses public symbols from <code>src/stock_selection/config.py</code>; it is an illustration of configuration construction, not an executed result.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;251f0714-2203-415b-ac78-61bbeac97254&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from stock_selection.config import (
    CostConfig,
    ExperimentConfig,
    TimingConfig,
    validate_config,
)

standard = ExperimentConfig(
    configuration_name="standard_t_to_t_plus_1",
    timing=TimingConfig(
        target="t_to_t_plus_1",
        implementation_lag=0,
    ),
    costs=CostConfig(model="none"),
)

headline = ExperimentConfig(
    configuration_name="headline_t_plus_1_to_t_plus_2",
    timing=TimingConfig(
        target="t_plus_1_to_t_plus_2",
        implementation_lag=1,
    ),
    costs=CostConfig(
        model="position_level",
        impact_coefficient=0.3,
        aum=100_000_000.0,
    ),
)

validate_config(standard)
validate_config(headline)</code></pre></div><p>The first object describes the ordinary one-period comparison. The second expresses the delayed target and records a position-level cost convention with an explicit AUM and <code>k=0.3</code>. To represent the paper's other reported convention, create another configuration with <code>impact_coefficient=0.2</code> and a distinct <code>configuration_name</code>; do not overwrite or silently reinterpret the first one.</p><p>Each configuration should retain at least its date boundaries, target name, implementation lag, cost model, impact formulation and coefficient, smoothing window, rebalance frequency, universe policy, model parameters, and optimization interpretation. These values are the experiment's provenance, not merely convenience arguments.</p><h3>Keys, audit records, and synthetic data</h3><p>The package uses pandas DataFrames for security-date panels. <code>Panel</code> and <code>PredictionTable</code> are type aliases that document the expected key columns without forcing a particular DataFrame index. Source panels retain <code>CWIQ Code</code> and a date column; prediction tables conventionally retain <code>CWIQ Code</code> and <code>signal_date</code>.</p><p><code>AuditRecord</code> stores route, source-context, validation, and configuration metadata. It is intended for generated features and other pipeline artifacts. <code>MetricResult</code> stores named values together with metadata such as date coverage and convention choices. These containers do not supply missing paper information; they provide a place to record what the caller actually selected.</p><p>When original data are unavailable, <code>SyntheticDataConfig</code> and <code>make_synthetic_sources</code> create deterministic EDI-like, TrueBeats-like, SpiderRock-like, universe, and factor-like tables. Their values are artificial. They are useful for demonstrating key alignment and pipeline plumbing, but they are not substitutes for the paper's vendor data and cannot reproduce its empirical results.</p><p>The local command-line entry point in <code>scripts/run_reproduction.py</code> exposes <code>build_argument_parser</code> and <code>main</code>. It accepts either <code>--synthetic</code> or user-supplied local source paths, prepares local data through the pipeline, and emits metadata. Its documented behavior intentionally does not imply external API access, completed feature generation, fitted models, or numerical verification.</p><h3>Verification boundary</h3><p>The authoritative run policy disabled local static verification, test generation, code execution, semantic code verification, tutorial verification, and final quality review. The supplied verification records therefore report that those stages were skipped. This section makes no claim that the generated code executed successfully, passed tests, or reproduced any reported Sharpe ratio or other metric.</p><p>The intended review strategy is still clear: inspect syntax and dependencies, trace keys and shapes, check chronological masks and target shifts, verify portfolio invariants, preserve cost and optimization labels, and retain audit metadata. Those are review criteria, not completed checks in this run.</p><p><strong>Advanced detail.</strong> A full numerical reproduction would additionally require the original vendor data and release timestamps, complete schemas, the exact retrieval corpus and retriever, GPT-4.1 prompts and generated feature code, DSPy/MIPRO compilation settings, LightGBM parameters, factor files, and resolved interpretations of the cost and optimization inconsistencies. Without them, the responsible claim is an evidence-grounded implementation architecture rather than a verified empirical replication.</p><h2>Paper-to-Code Architecture and Dependency Graph</h2><p>How should a multi-stage quantitative experiment be organized so that one stage cannot quietly use information from a later stage? The generated implementation treats the reproduction as a sequence of guarded transformations:</p><ol><li><p>prepare a point-in-time security-date panel;</p></li><li><p>discover and validate features during the discovery period;</p></li><li><p>freeze the accepted feature definitions;</p></li><li><p>standardize inputs and fit chronological models;</p></li><li><p>transform predictions into portfolio scores and weights;</p></li><li><p>join those weights to later realized returns;</p></li><li><p>evaluate performance and, optionally, costs, inference, optimization, and robustness.</p></li></ol><p>This ordering is an implementation interpretation of the paper's workflow, not evidence that the complete experiment has run. The supplied code provides interfaces and orchestration, while the original vendor data, complete schemas, generated feature corpus, external generation services, and exact model settings are unavailable.</p><h3>The package boundary</h3><p>The generated package separates the workflow into modules with different responsibilities. The separation is useful because each boundary has a different correctness question:</p><ul><li><p><strong>Data modules</strong> ask whether rows, identifiers, dates, releases, universe membership, and return targets are valid.</p></li><li><p><strong>Feature modules</strong> ask whether a generated transformation is aligned, schema-constrained, point-in-time, and sufficiently usable.</p></li><li><p><strong>Model modules</strong> ask whether training observations precede prediction observations and whether feature columns have a stable order.</p></li><li><p><strong>Portfolio modules</strong> ask whether predictions are converted into the intended long-short exposure without using future returns.</p></li><li><p><strong>Evaluation modules</strong> ask whether metrics use the selected execution horizon and the appropriate cross-sectional or time-series grouping.</p></li><li><p><strong>Cost and smoothing modules</strong> ask whether trading activity and friction assumptions are kept separate from gross returns.</p></li><li><p><strong>Optimization and robustness modules</strong> are optional analyses that reuse the core artifacts rather than silently changing the primary experiment.</p></li></ul><p>The central orchestrator is <code>ReproductionPipeline</code> in <code>src/stock_selection/pipeline.py</code>. Its public methods correspond to the major temporal boundaries:</p><ul><li><p><code>ReproductionPipeline.prepare_data()</code> aligns local sources, applies universe filters, and constructs target columns.</p></li><li><p><code>ReproductionPipeline.run_discovery()</code> operates on the discovery partition and freezes the resulting <code>FeatureRegistry</code>.</p></li><li><p><code>ReproductionPipeline.run_oos()</code> materializes frozen features, fits requested model variants, and evaluates out-of-sample predictions.</p></li></ul><p>A focused excerpt shows the intended order without reproducing the whole file:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;897f9e7b-bfef-4c65-b7df-385fda9270f2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    def prepare_data(self, sources: Mapping[str, pd.DataFrame]) -&gt; pd.DataFrame:
        """Align local sources, filter the investable universe, and build targets.

        The alignment is keyed by ``CWIQ Code`` and ``date``.  No source is
        forward-filled here.  Target columns are added only after filtering,
        and their signal, execution, and return dates remain explicit.
        """
        panel = align_sources(sources, self.config)
        universe = self._universe()
        panel = filter_security_types(panel, universe)
        panel = select_largest_universe(panel, universe)
        panel = apply_liquidity_filters(panel, universe)
        panel = build_return_targets(panel, self._timing())</code></pre></div><p>The important detail is not merely the function names. <code>align_sources()</code> runs before universe selection and target construction; <code>build_return_targets()</code> runs after filtering; and the resulting rows retain explicit timing columns. This prevents the implementation from burying execution timing inside a later portfolio function.</p><p>The next boundary is discovery. The pipeline requires a caller-supplied feature generator because the paper does not specify a complete provider implementation:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;bb8cb97d-e814-41d4-9e4f-c37d433bcc82&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    def run_discovery(self, panel: pd.DataFrame) -&gt; FeatureRegistry:
        """Discover, validate, and freeze generated features on discovery data only."""
        if not isinstance(panel, pd.DataFrame):
            raise TypeError("panel must be a pandas DataFrame")
        generator = self._get("feature_generator", "generator", default=None)
        if generator is None or not hasattr(generator, "generate"):
            raise ValueError(
                "ExperimentConfig must provide a local feature_generator; "
                "the supplied paper context does not define a default provider"
            )</code></pre></div><p>This is a deliberate guard. It does not substitute a fabricated GPT-4.1 client, prompt, retriever, or DSPy/MIPRO configuration. A structured-prompting adapter and a DSPy/MIPRO adapter can both satisfy the generator interface, but their route, retrieval context, and generation metadata remain distinguishable.</p><p>After discovery, the registry is frozen before out-of-sample materialization. The freeze is a temporal control: it prevents feature definitions selected using 2015&#8211;2018 information from being changed while processing 2019&#8211;2024 observations. The pipeline records the registry and its manifest in <code>artifacts</code>, alongside configuration and stage metadata.</p><h3>Stable data contracts between stages</h3><p>The shared aliases <code>Panel</code> and <code>PredictionTable</code> are defined in <code>src/stock_selection/types.py</code>. They are pandas <code>DataFrame</code> contracts rather than special container classes:</p><ul><li><p>A <code>Panel</code> represents security-date observations and retains <code>CWIQ Code</code> and a date column.</p></li><li><p>A <code>PredictionTable</code> conventionally contains <code>CWIQ Code</code>, <code>signal_date</code>, and a prediction column.</p></li><li><p>A <code>FeatureRegistry</code> stores accepted feature definitions and their validation history.</p></li><li><p>A <code>MetricResult</code> stores named values together with metadata such as date coverage or convention choices.</p></li><li><p>An <code>AuditRecord</code> stores route, source context, validation decisions, and configuration information for a generated feature or pipeline artifact.</p></li></ul><p>These objects preserve identity across transformations. <code>CWIQ Code</code> identifies the security, while the date fields distinguish when information was observed, when a position may be executed, and when a return is realized. A numerical column without those keys is not sufficient for this workflow: it cannot be safely joined to later returns or compared across model variants.</p><p>The type containers are intentionally lightweight. For example, the generated <code>AuditRecord</code> does not prescribe a particular retrieval engine. Its <code>source_context</code> is opaque because the paper does not supply canonical document chunks, embedding settings, or retriever behavior. That design records provenance without pretending that an unspecified implementation has been recovered.</p><h3>Artifact flow through the orchestrator</h3><p>The following excerpt is copied from <code>ReproductionPipeline.run_oos()</code> and shows the downstream order:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2c032170-fe6b-4cc9-a5bd-d4c162783f6f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        parts = make_date_split(panel, self._date_split())
        assert_no_split_overlap(parts)
        oos = parts.get("oos") or parts.get("out_of_sample") or parts.get("oos_data")
        if oos is None:
            raise KeyError("make_date_split did not return an out-of-sample partition")
        enriched = materialize_feature_sets(oos, registry)
        baseline, ai = self._feature_columns(enriched, registry)
        requested = self._get("model_configurations", "feature_configurations", default=("baseline", "ai", "combined"))
        results: dict[str, Any] = {}
        for configuration in tuple(requested):
            results[str(configuration)] = self._run_variant(enriched, str(configuration), baseline, ai)
        self.artifacts["oos_results"] = results
        self._record_stage("run_oos", rows=int(len(oos)), configurations=list(results))
        return results</code></pre></div><p>Several safeguards are visible here. The split is checked before out-of-sample processing, the registry must already be frozen, and features are materialized only after selecting the OOS partition. The default requested configurations distinguish <code>baseline</code>, <code>ai</code>, and <code>combined</code>; this is feature-set selection, not yet ensemble averaging. A separate function, <code>average_standardized_predictions()</code>, is intended for aligned prediction variants such as structured and DSPy outputs.</p><p>Each variant produces intermediate artifacts rather than returning only one final number. The planned result structure contains predictions, weights, portfolio returns, metrics, and model parameters. This makes it possible to inspect where a discrepancy arises: input selection, model chronology, score conversion, return alignment, or metric convention.</p><h3>Dependencies are optional by design</h3><p>The package declaration in <code>pyproject.toml</code> keeps the core tabular dependencies separate from stages that require additional libraries:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;toml&quot;,&quot;nodeId&quot;:&quot;23249ecd-4055-41d5-8023-e892d46b5923&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-toml">[project.optional-dependencies]
model = [
    "lightgbm&gt;=4.0",
]
statistics = [
    "scipy&gt;=1.10",
    "statsmodels&gt;=0.14",
]
optimization = [
    "cvxpy&gt;=1.4",
    "clarabel&gt;=0.7",
]</code></pre></div><p>NumPy and pandas are required for the panel and numerical operations. LightGBM belongs to the model extra. SciPy supports rank-based statistics, while statsmodels supports HAC-oriented inference wrappers. cvxpy and CLARABEL are isolated because portfolio optimization is optional and the paper's optimization displays are OCR-damaged. Separating these dependencies also makes the architecture honest: installing the package does not imply that vendor data, a language-model provider, a retrieval index, or a solver configuration has been supplied.</p><h3>Paper facts, derived structure, and implementation decisions</h3><p>The paper fact is that the method combines generated features, conventional features, LightGBM models, portfolio formation, and several robustness analyses. The derived architectural explanation is that these components should communicate through explicit, keyed artifacts and chronological boundaries. The implementation decisions include the following:</p><ul><li><p><code>ReproductionPipeline</code> retains intermediate artifacts in an <code>artifacts</code> dictionary rather than hiding them in local variables.</p></li><li><p><code>ExperimentConfig</code> stores timing, cost, smoothing, model, and optimization choices so they can be audited.</p></li><li><p>Structured prompting and DSPy/MIPRO are separate generation routes but share validation, registry, and downstream evaluation interfaces.</p></li><li><p>Cost analysis, optimization, and robustness are optional branches; they should consume frozen core outputs rather than retune the primary OOS experiment.</p></li><li><p>The generated package records unresolved choices instead of silently resolving them, including standard versus delayed targets, daily versus monthly rebalancing, and the paper's separate <code>k=0.2</code> and <code>k=0.3</code> impact conventions.</p></li></ul><p>The architecture therefore supports a faithful local reproduction plan, but it does not establish that the full paper experiment ran. Under the authoritative run policy, code execution, local static verification, semantic code verification, and numerical verification were not performed. The generated files should be read as an auditable design and set of implementation interfaces, not as evidence of successful execution or reproduced empirical results.</p><h2>Data Contracts, Point-in-Time Alignment, Universe Construction, and Target Timing</h2><p>How can a model use information available on a signal date without accidentally seeing the future? The central rule is simple: a row identified by <code>CWIQ Code</code> and <code>date</code> may contain only information known at that date. A later return is a target for training or evaluation, not an input feature.</p><p>The paper's data workflow aligns EDI price-volume data with TrueBeats and SpiderRock observations, filters an investable U.S. equity universe, and separates discovery dates from out-of-sample dates. The generated implementation turns those requirements into explicit pandas boundaries. This is an implementation architecture, not a numerically verified reconstruction: the local loaders require user-supplied files and do not contact vendors or external services.</p><h3>1. Establish a stable source contract</h3><p>The first boundary is the source schema. <code>DatasetSchema</code> records the required fields, optional fields, and the source-specific names of the identifier and date columns. Because the supplied paper context does not contain complete vendor schemas, the caller must provide those declarations rather than relying on undocumented aliases.</p><p><code>validate_schema</code> checks that required fields exist, that <code>CWIQ Code</code> values are present and non-empty, and that date values are parseable. It deliberately does not reject duplicate security-date rows: the source may define multiple observations on a date, and the later alignment stage must decide whether that ambiguity is acceptable.</p><p>The following excerpt is copied from <code>src/stock_selection/data/schema.py</code>. Notice that the common key columns are added to the caller's required columns before validation.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;aff28651-916b-46af-a67a-7b72774860e4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">required = tuple(dict.fromkeys((*schema.required_columns, *schema.key_columns)))
missing = _missing_columns(frame, required)
if missing:
    raise ValueError(
        f"{schema.name!r} is missing required columns: {', '.join(missing)}"
    )

_validate_identifier_series(frame[schema.identifier_column], schema.identifier_column)

dates = pd.to_datetime(frame[schema.date_column], errors="coerce", utc=True)
invalid_dates = dates.isna()
if invalid_dates.any():
    count = int(invalid_dates.sum())
    raise ValueError(
        f"Date column {schema.date_column!r} contains {count} unparseable value(s)"
    )</code></pre></div><p><code>normalize_keys</code> then maps identifiers to stripped pandas strings and dates to timezone-naive, midnight-normalized timestamps. It preserves row count, missingness, and column values; it does not sort, deduplicate, forward-fill, or impute observations. That preservation matters because missing vendor data must remain visible to later point-in-time decisions.</p><h3>2. Load local data only</h3><p><code>load_table</code> supports local CSV, compressed CSV, Parquet, and Feather files. It validates the supplied schema, canonicalizes custom key names to <code>CWIQ Code</code> and <code>date</code>, and applies deterministic sorting. <code>load_source_bundle</code> requires the path and schema mappings to contain exactly the same source names, preventing a file from being silently interpreted under the wrong schema.</p><p>This is an implementation decision made because the required EDI, TrueBeats, SpiderRock, and knowledge-base services are not supplied. The loader does not download data or infer vendor fields. A complete reproduction therefore begins with locally materialized files and user-declared schemas.</p><h3>3. Align sources without forward-filling</h3><p><code>align_sources</code> normalizes every source and rejects duplicate <code>(CWIQ Code, date)</code> keys. EDI is used as the anchor when it is present; otherwise, the first supplied source becomes the anchor. Other sources are left-joined on the two canonical keys. Each joined source receives an <code>available_&lt;source&gt;</code> flag, so a missing vendor row can be distinguished from a vendor row whose payload happens to contain missing values.</p><p>The relevant behavior is visible in this excerpt from <code>src/stock_selection/data/alignment.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b48dd5db-7d38-4231-9395-0a85a2b18697&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">for source_name, source_frame in normalized_sources.items():
    if source_name == anchor_name:
        continue
    prepared = _rename_collisions(source_frame, existing_columns, source_name)
    availability_column = f"available_{_source_suffix(source_name)}"
    if availability_column in existing_columns or availability_column in prepared.columns:
        raise ValueError(
            f"availability column {availability_column!r} conflicts with source data"
        )
    prepared[availability_column] = True
    panel = panel.merge(
        prepared,
        how="left",
        on=list(_KEY_COLUMNS),
        sort=False,
        validate="one_to_one",
    )
    panel[availability_column] = panel[availability_column].fillna(False).astype(bool)</code></pre></div><p>The <code>validate="one_to_one"</code> merge is an important guard. It requires each source to have at most one row for each security-date key after the earlier duplicate check. The implementation does not fill a missing TrueBeats or SpiderRock value with a later observation. If the source supplies release timestamps, <code>validate_point_in_time</code> compares each release date with the signal <code>date</code> and rejects a row released afterward. Missing release metadata is preserved rather than treated as proof of availability.</p><h3>4. Filter the investable universe</h3><p>After alignment, the pipeline can apply eligibility and liquidity policies. <code>filter_security_types</code> excludes configured security types and can require an explicit U.S. common-equity eligibility column. Missing security types or eligibility values are not silently classified as eligible; the configured missing-value policy determines whether they are excluded, retained, or rejected.</p><p><code>select_largest_universe</code> applies the market-cap limit independently for each date. This is a per-date ranking, not one ranking computed over the entire sample. The default reproduction target is at most 2,500 names per date when sufficient valid observations exist. Ties are resolved deterministically by <code>CWIQ Code</code>, and missing market capitalization follows an explicit policy.</p><p>The core selection operation from <code>src/stock_selection/data/universe.py</code> is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;242c9d80-b4b6-4549-aaf1-3ce420fe73c1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">working = working.sort_values(
    [
        _DATE_COLUMN,
        "__universe_missing_cap",
        "__universe_market_cap",
        _SECURITY_ID_COLUMN,
    ],
    ascending=[True, True, False, True],
    kind="mergesort",
)
selected = working.groupby(_DATE_COLUMN, sort=False, group_keys=False).head(
    config.max_names
)
return selected.drop(
    columns=["__universe_market_cap", "__universe_missing_cap"]
).copy()</code></pre></div><p><code>apply_liquidity_filters</code> is optional and requires explicit ADV or spread columns whenever corresponding thresholds are configured. This keeps liquidity screening auditable: the pipeline cannot claim that a security passed a threshold when the source column was absent.</p><h3>5. Make signal and target dates explicit</h3><p>The paper uses one-period return notation in equation record <code>eq_3_4_return</code>, and portfolio aggregation is represented by equation record <code>eq_3_4_portfolio_return</code>. Their canonical LaTeX fields are empty in the supplied evidence, so no formula is reconstructed here. Semantically, the first record identifies a forward return constructed from price observations, while the second maps weights formed at one date to realized returns over a later interval.</p><p>The implementation makes that timing visible through four columns:</p><ul><li><p><code>signal_date</code>: when the feature and prediction are formed;</p></li><li><p><code>execution_date</code>: the beginning of the configured realized interval;</p></li><li><p><code>return_date</code>: the end of that interval;</p></li><li><p><code>realized_return</code>: the target consumed by portfolio evaluation.</p></li></ul><p><code>TimingConfig</code> distinguishes the standard <code>t_to_t_plus_1</code> target from the paper's headline implementable <code>t_plus_1_to_t_plus_2</code> target. An <code>implementation_lag</code> shifts both ends of the interval. This distinction prevents the common mistake of labeling a return from <code>t</code> to <code>t+1</code> as though it were the delayed implementable target.</p><p><code>construct_log_return</code> aligns a forward return to each source row by shifting within each security's chronological history. It does not synthesize missing calendar dates. <code>build_return_targets</code> also retains a <code>standard_return</code>, while <code>target_return</code> follows the selected timing configuration. Where a terminal delisting return is supplied, the implementation treats it as an additional terminal adjustment; the exact vendor convention remains a data-policy decision because the paper's source fields are not supplied.</p><p>A focused excerpt from <code>src/stock_selection/data/returns.py</code> shows how the configured offsets become date columns:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;3907357e-cbdd-4ac4-a9ef-f1874f4e976b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">source_dates = pd.to_datetime(result[date_column])
start_offset = (0 if timing.target == "t_to_t_plus_1" else 1) + timing.implementation_lag
end_offset = (1 if timing.target == "t_to_t_plus_1" else 2) + timing.implementation_lag

result[timing.signal_date_column] = source_dates
result[timing.execution_date_column] = _shift_by_security(
    source_dates, frame, start_offset, date_column
)
result[timing.return_date_column] = _shift_by_security(
    source_dates, frame, end_offset, date_column
)
result["target_start_date"] = result[timing.execution_date_column]
result["target_end_date"] = result[timing.return_date_column]
result["target_start_offset"] = start_offset
result["target_end_offset"] = end_offset</code></pre></div><p>For the delayed target with zero additional lag, a signal on date <code>t</code> points to an <code>execution_date</code> at <code>t+1</code> and a <code>return_date</code> at <code>t+2</code>. With a positive implementation lag, both offsets move later. Rows without enough future observations receive missing target dates and returns rather than fabricated values.</p><h3>Worked example: a delayed target on a small panel</h3><p>The generated package includes <code>SyntheticDataConfig</code> and <code>make_synthetic_sources</code> in <code>src/stock_selection/data/synthetic.py</code> for deterministic demonstrations. Those values are teaching fixtures, not EDI, TrueBeats, or SpiderRock data and cannot reproduce the paper's empirical results.</p><p>A local demonstration can conceptually follow this sequence:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;dfe83bbb-6106-4604-b121-e84c1d86bff2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">sources = make_synthetic_sources(synthetic_config)
panel = align_sources(sources, experiment_config)
panel = filter_security_types(panel, universe_config)
panel = select_largest_universe(panel, universe_config)
timed = build_return_targets(panel, timing_config)</code></pre></div><p>Suppose the panel contains security <code>A</code> on three successive observations. Under <code>target="t_plus_1_to_t_plus_2"</code> and <code>implementation_lag=0</code>, the first row retains its original <code>signal_date</code>; its <code>execution_date</code> is the second observation and its <code>return_date</code> is the third. The first row can therefore receive a delayed realized return, while the final row has no later interval and remains missing as a target. The security-date keys remain attached throughout.</p><p>The supplied test file makes this invariant concrete. The following excerpt is copied from <code>tests/test_data_and_timing.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6fc08d44-a74f-4e37-8754-2c7a876aad00&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">first_observation = timed.loc[
    (timed["CWIQ Code"] == "A") &amp; timed["date"].eq(dates[0])
].iloc[0]
assert first_observation["signal_date"] == dates[0]
assert first_observation["execution_date"] == dates[1]
assert first_observation["return_date"] == dates[2]
assert first_observation["execution_date"] &gt; first_observation["signal_date"]
assert first_observation["return_date"] &gt; first_observation["execution_date"]</code></pre></div><p>This test is a planned deterministic check over synthetic rows, not evidence that the test was executed. Under the authoritative run policy, code execution, local static verification, and semantic code verification were skipped.</p><h3>6. Split discovery from out-of-sample dates</h3><p><code>make_date_split</code> uses inclusive boundaries from <code>DateSplit</code>. The paper's discovery period is January 2015 through December 2018, and the headline out-of-sample period is January 2019 through December 2024. The implementation returns separate <code>discovery</code> and <code>oos</code> frames and excludes dates outside those configured periods. <code>assert_no_split_overlap</code> checks date-level disjointness, which is stronger than merely checking DataFrame indexes.</p><p>This separation is a correctness constraint, not just an organizational preference. Feature discovery, validation, and hyperparameter selection must use discovery data only. The accepted feature definitions and model choices must be fixed before headline out-of-sample evaluation. Split boundaries are inclusive within each period but must remain chronologically disjoint across periods.</p><h3>Implementation decisions and remaining limits</h3><p>The paper facts are the use of point-in-time data, eligibility and market-cap filtering, chronological discovery/OOS separation, and a delayed <code>t+1</code>-to-<code>t+2</code> evaluation convention. The generated implementation decisions are to reject duplicate source keys, retain availability flags, use observation-based shifts within each security, and make missing-data policies explicit.</p><p>Several details cannot be recovered from the supplied evidence: complete vendor schemas, corporate-action and delisting field definitions, release-date precedence when multiple records exist, and the exact treatment of missing market capitalization. These should remain configuration and audit metadata rather than hidden assumptions. Most importantly, the data layer must never use a future release as a feature, and a future return must remain a target only.</p><h2>Generated-Feature Contracts, Retrieval, Structured Prompting, and DSPy/MIPRO</h2><p>How can a language-model-generated transformation become a trustworthy model feature rather than an untracked piece of code? The key is to make the generator replaceable while making the feature contract strict. Regardless of whether a candidate comes from structured prompting or DSPy/MIPRO, it must declare its inputs, return exactly one row-aligned pandas <code>Series</code>, and retain provenance about its route and retrieval context.</p><p>This section implements the <code>llm_feature_generation_and_validation</code> method. The paper fact is that GPT-4.1 generates executable, schema-constrained transformations using retrieved vendor documentation, with separate structured-prompting and DSPy/MIPRO routes. The implementation decision is to place both routes behind provider-neutral interfaces. The supplied evidence does not specify the exact prompts, retriever, embeddings, document chunks, retrieval count, DSPy signature, demonstrations, MIPRO settings, or scoring weights. Those details therefore remain caller-supplied rather than being invented.</p><h3>The one-Series feature contract</h3><p>A generated feature is evaluated on a security-date panel: rows represent observations for securities identified by <code>CWIQ Code</code> and date. The feature function receives the panel and must return one scalar value for every input row. &#8220;One Series&#8221; means a one-dimensional pandas object, not a new table, a collection of columns, or a separately indexed result.</p><p>The contract is expressed by <code>FeatureFunction</code>, while <code>GeneratedFeature</code> adds the metadata needed for auditing. The following excerpt is copied from <code>src/stock_selection/features/contracts.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;9a95a55f-1ba1-4ced-b785-1b9b28d9cf0c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@runtime_checkable
class FeatureFunction(Protocol):
    """Callable contract for one generated, row-aligned feature.

    Implementations receive a security-date ``DataFrame`` and must return exactly
    one numeric ``Series``.  The returned series must preserve the input row
    ordering and index; it must not add, remove, or reorder observations.
    """

    def __call__(self, frame: pd.DataFrame) -&gt; pd.Series:
        """Compute one feature value for every input row."""
        ...</code></pre></div><p>The <code>GeneratedFeature</code> record stores the callable, its <code>name</code>, declared <code>required_columns</code>, economic <code>rationale</code>, generation <code>route</code>, <code>context_version</code>, and optional audit metadata. Its call boundary checks that the input is a pandas <code>DataFrame</code>, rejects missing declared columns, executes the callable, and passes the result to <code>align_feature_series</code>.</p><p><code>align_feature_series</code> checks both length and index equality. Equal lengths are not sufficient: a reordered <code>Series</code> could assign one security's value to another security-date row. The function therefore preserves a strict key and row-order invariant. A validation failure should stop the candidate rather than silently reindexing it.</p><p><code>FeatureResult</code> pairs a materialized <code>Series</code> with a validation report. Its <code>accepted</code> property reads the recorded status and defaults to false when no acceptance decision has been recorded. This is important because the paper does not supply universal thresholds for coverage, similarity, diversity, IC, or Sharpe; those decisions must remain explicit in later validation configuration.</p><h3>Worked example: wrapping a local transformation</h3><p>The paper discusses recurring patterns such as cross-sectional ranking, volatility normalization, momentum adjustment, and interactions. The generated <code>src/stock_selection/features/templates.py</code> provides local educational templates for those patterns. They are not recovered GPT-4.1 outputs and do not constitute the paper's complete feature corpus.</p><p>For example, <code>cross_sectional_rank</code> validates a numeric column and date column, ranks observations within each date, and returns a <code>Series</code> with the original index and row order. A feature wrapper can declare that dependency and attach provenance:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;18da6bf1-3c47-4774-af7d-49876513d938&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">feature = GeneratedFeature(
    name="eps_estimate_cross_sectional_rank",
    function=lambda frame: cross_sectional_rank(
        frame, "eps_estimate", "date"
    ),
    required_columns=("eps_estimate", "date"),
    rationale="Educational date-local ranking template.",
    route="local_template",
    context_version="tutorial-v1",
)</code></pre></div><p>This snippet is an implementation example built from the generated public symbols; it does not claim that this exact feature was generated by the paper's model. The declared column list gives a schema validator a boundary to inspect, while the rationale and route allow later reports to distinguish a local template from an externally generated candidate.</p><p>The other templates follow the same one-Series rule. <code>volatility_normalize</code> sorts within a recognized security identifier and date, uses a bounded trailing window, and returns missing values for incomplete or zero-dispersion windows. <code>momentum_adjust</code> compares the current value with a bounded prior observation within each security. <code>interaction_feature</code> multiplies two aligned numeric columns. Their exact normalization and momentum conventions are implementation choices because the supplied paper context reports patterns but not a complete feature definition.</p><h3>Retrieval is a separate, local abstraction</h3><p>Retrieval-augmented generation requires documentation context describing permitted columns and vendor fields. In a full reproduction, that context would come from the relevant knowledge base. The generated <code>KnowledgeBase</code> in <code>src/stock_selection/generation/retrieval.py</code> deliberately implements only a local, explicit policy called <code>exact_record</code>. It loads text files, searches query terms, ranks matching records deterministically, and performs no network access, embedding lookup, or inferred chunking.</p><p>That distinction matters. <code>KnowledgeBase.retrieve</code> is useful for demonstrating the interface, but it is not evidence that the paper used exact-record matching. The paper's retrieval implementation is unspecified, so this class should be treated as a replaceable adapter.</p><p>The retrieval boundary also exposes the broken-versus-corrected documentation comparison without hard-coding an undocumented corruption mechanism:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a3f95aab-90ca-4447-a7ea-63a461b422fd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def make_corrupted_variant(documents: Sequence[str], policy: str) -&gt; list[str]:
    """Apply an explicitly named documentation ablation policy.

    The paper does not specify how its corrupted knowledge base was produced.
    Consequently, only ``identity`` is implemented here, which preserves the
    supplied records and is useful for a control route.  Any other policy must
    be implemented by the caller rather than being silently approximated.
    """
    if not isinstance(policy, str) or not policy.strip():
        raise ValueError("policy must be a non-empty string")
    records = list(documents)
    if any(not isinstance(document, str) for document in records):
        raise TypeError("All documentation records must be strings")
    if policy == "identity":
        return records
    raise RetrievalConfigurationError(
        f"Corruption policy {policy!r} is unspecified by the supplied paper context; "
        "provide an explicit local implementation"
    )</code></pre></div><p>The <code>identity</code> policy is a control, not a reconstruction of the paper's broken knowledge base. To reproduce that ablation, a user must supply and label an explicit local transformation. The resulting route and context version should be retained in audit metadata so that corrected and corrupted documentation are never confused.</p><h3>Shared generation context and batches</h3><p><code>GenerationContext</code> packages the schema descriptor, retrieved text, route, context version, and discovery metadata. The schema is intentionally flexible because the supplied evidence does not define a canonical vendor-schema object. The caller may provide a <code>DatasetSchema</code> or another explicit descriptor, but the generator must treat it as the permitted field boundary.</p><p><code>GenerationBatch</code> contains a tuple of <code>GeneratedFeature</code> candidates, audit records, route, context version, and batch metadata. <code>validate_generation_batch</code> checks that candidates do not silently claim a different route or retrieval context. It does not decide whether a feature is sufficiently sparse, diverse, or predictive; those are later feature-validation stages.</p><p>The provider-neutral interface is small:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;370e4e88-a812-4981-92d3-9d0c22982f3a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@runtime_checkable
class FeatureGenerator(Protocol):
    """Provider-neutral interface for structured or programmatic generation.

    Implementations may use a local deterministic provider or an explicitly
    configured external adapter in another module.  The paper's exact GPT-4.1
    API, prompts, retrieval configuration, and DSPy/MIPRO settings are not
    supplied here, so this interface does not prescribe any of them.
    """

    def generate(self, context: GenerationContext) -&gt; GenerationBatch:
        """Generate candidate feature definitions and their audit records."""

        ...</code></pre></div><p>The output is intentionally metadata-rich rather than an untracked source-code string. A provider can differ, but the downstream validator receives the same candidate representation.</p><h3>Structured prompting route</h3><p>The structured route uses <code>StructuredPromptConfig</code> to record an explicitly chosen local request format: allowed operations, whether negative lags are forbidden, a maximum rolling window, and the one-Series output contract. <code>build_structured_request</code> serializes the schema, retrieved context, constraints, and discovery metadata. It does not retrieve documents or call GPT-4.1.</p><p><code>LocalProvider</code> is the injected provider boundary. <code>generate_structured_features</code> builds the request, asks the provider for a <code>GenerationBatch</code>, and applies <code>validate_generation_batch</code>. This makes the route suitable for a deterministic local adapter or a separately configured external service without claiming that the paper's exact prompt has been recovered.</p><p>A useful property of this design is that structured prompting cannot bypass provenance checks merely because it is the simpler route. Candidates still carry the route and context version and must later pass the common feature validator.</p><h3>DSPy/MIPRO route</h3><p>The DSPy route is separate because its programmatic prompting and optimization process differs conceptually from a fixed structured template. <code>DSPyConfig</code> records the signature name, route name, demonstration identifiers, optimizer name, optimizer settings, scoring settings, context version, and a required <code>program_factory</code>.</p><p>The <code>program_factory</code> is an explicit implementation decision. Without it, <code>compile_dspy_program</code> raises <code>DSPyConfigurationError</code> rather than pretending to know the missing DSPy signature or MIPRO setup. Before compilation, <code>_validate_discovery_feedback</code> checks that feedback does not extend beyond the latest discovery observation. This protects the paper's chronological boundary: generation feedback and feature selection belong to January 2015 through December 2018, not the headline out-of-sample period.</p><p>After compilation, <code>generate_dspy_features</code> accepts any object with <code>generate(context)</code>. The adapter may return a <code>GenerationBatch</code> directly or a sequence of <code>GeneratedFeature</code> objects, which <code>_coerce_batch</code> normalizes into the common batch format. Thus structured and DSPy candidates converge at the same provenance contract before feature-level validation.</p><h3>Putting both routes together</h3><p>A complete local experiment would follow this conceptual sequence:</p><ol><li><p>Define a user-supplied schema and retrieve local documentation.</p></li><li><p>Build a <code>GenerationContext</code> with a route and context version.</p></li><li><p>Generate a structured or DSPy <code>GenerationBatch</code>.</p></li><li><p>Validate every candidate's declared columns, compilation behavior, temporal restrictions, rolling windows, coverage, similarity, and diversity.</p></li><li><p>Register accepted candidates and freeze the registry before out-of-sample materialization.</p></li></ol><p>For the worked example, the <code>cross_sectional_rank</code> wrapper can be attached to a <code>GenerationBatch</code> alongside provider-produced candidates. A structured provider receives the serialized request from <code>build_structured_request</code>; a DSPy adapter receives the same kind of <code>GenerationContext</code> through its compiled program. Both outputs must contain <code>GeneratedFeature</code> instances and both are checked by <code>validate_generation_batch</code> before entering the shared downstream path.</p><p>This shared boundary is the main reproduction decision: the generator is interchangeable, but the feature's input declaration, aligned output, route, context version, rationale, and validation record are not. It also makes audit review possible. A later model result can be traced back to the exact accepted feature definitions and documentation context used during discovery.</p><h3>Advanced detail: what remains unresolved</h3><p>The paper's retrieval-quality result establishes that documentation quality matters, but it does not provide enough information to reconstruct the corrupted-document transformation or the corrected retrieval index. Likewise, the DSPy/MIPRO comparison does not supply the complete signature, demonstrations, optimizer schedule, or scoring weights. These are not minor formatting details: they can change which candidates are generated and accepted.</p><p>Accordingly, the generated interfaces expose those values as configuration and metadata. They should not be silently filled with plausible defaults and should not be described as the paper's exact implementation. No code execution or verification is claimed for these excerpts; the authoritative run policy skipped local static verification and semantic code verification.</p><h2>Feature Validation, Screening, and Registry Freezing</h2><p>How does a generated feature earn the right to enter model training? Treat validation as a gate rather than a single test. A candidate must reference permitted columns, respect the signal-time information boundary, use bounded rolling windows, return an aligned value for every input row, and meet any explicitly configured coverage and similarity rules.</p><p>The paper fact is that generated features are screened for schema validity, point-in-time correctness, sparsity, similarity, and diversity before they are used downstream. The generated implementation turns those requirements into a common validation layer and an auditable <code>FeatureRegistry</code>. The registry is populated during discovery and frozen before out-of-sample materialization. This is an implementation architecture; the supplied evidence does not provide the paper's acceptance thresholds or complete generated-feature corpus.</p><h3>Static inspection is only one gate</h3><p><code>validate_feature_schema</code> checks the declared <code>required_columns</code> of a <code>GeneratedFeature</code> against a user-supplied <code>DatasetSchema</code>. It does not infer missing vendor documentation. <code>compile_feature</code> then obtains inspectable Python source, parses it with <code>ast</code>, rejects imports and selected unsafe runtime constructs, and invokes Python's compiler. It does not execute the candidate.</p><p>That distinction matters. Syntax and restricted-construct checks can reject obvious problems, but they do not prove that arbitrary generated code is safe, economically meaningful, or free of every possible look-ahead path. The generated file explicitly describes compilation as static and says that it is not a substitute for a process-level sandbox.</p><p>A focused excerpt shows the boundary:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;17e85de4-fad8-4646-907c-98faaa7d5ad8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def compile_feature(feature: GeneratedFeature) -&gt; None:
    """Statically compile a candidate while rejecting obvious unsafe constructs.

    Compilation validates Python syntax and bytecode compilation only; it is not a
    substitute for a process-level sandbox and does not execute the callable.
    """
    if not isinstance(feature, GeneratedFeature):
        raise TypeError("feature must be a GeneratedFeature")
    tree = _source_tree(feature)
    for node in ast.walk(tree):
        if isinstance(node, (ast.Import, ast.ImportFrom)):
            raise ValueError(
                f"feature {feature.name!r} contains imports; imports are not permitted"
            )
        if isinstance(node, ast.Call):
            if isinstance(node.func, ast.Name) and node.func.id in _FORBIDDEN_CALLS:
                raise ValueError(
                    f"feature {feature.name!r} uses forbidden call {node.func.id!r}"
                )</code></pre></div><p>The important implementation choices are visible in the excerpt: imports are rejected, a fixed set of dangerous calls is rejected, and the function is not run by <code>compile_feature</code>. Runtime alignment is handled separately by the generated-feature contract and by <code>FeatureRegistry.materialize</code>.</p><h3>Temporal checks require explicit metadata</h3><p>The paper requires non-negative lags and bounded rolling windows, but it does not define a universal static analyzer capable of recovering those properties from every possible Python transformation. The generated implementation therefore reads feature metadata such as <code>lags</code>, <code>max_lag</code>, <code>release_columns</code>, <code>rolling_windows</code>, and <code>uses_future_data</code>.</p><p><code>check_point_in_time_lags</code> rejects negative or invalid lag declarations and can compare release-date columns with the frame's <code>date</code> column. <code>check_rolling_windows</code> rejects expanding or otherwise unbounded windows, missing windows, non-integral windows, and non-positive windows. These checks are meaningful only when the generator or feature author records the relevant metadata accurately, so the metadata itself belongs in the audit trail.</p><p>This is an implementation decision, not a recovered detail of the paper's code. A production deployment would also need a restricted execution environment and a stronger review of generated operations. The generated implementation does not claim that its AST inspection alone establishes complete safety.</p><h3>Coverage, similarity, and diversity</h3><p>After the static and temporal gates pass, <code>discover_features</code> materializes the candidate and calls <code>measure_sparsity</code>. In this package, sparsity is the fraction of values that are missing, nonnumeric, or nonfinite. The function returns <code>1.0</code> for an empty Series, making an empty candidate maximally unusable rather than silently acceptable.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e302a885-1f52-4237-ad4f-c6e8058d0f27&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def measure_sparsity(values: pd.Series) -&gt; float:
    """Return the fraction of values that are missing, nonnumeric, or nonfinite."""
    if not isinstance(values, pd.Series):
        raise TypeError("values must be a pandas Series")
    if len(values) == 0:
        return 1.0
    numeric = pd.to_numeric(values, errors="coerce")
    usable = numeric.notna() &amp; np.isfinite(numeric.to_numpy(dtype=float, na_value=np.nan))
    return float(1.0 - usable.mean())</code></pre></div><p>The sparsity threshold is not hard-coded by the paper context. <code>discover_features</code> looks for configuration names such as <code>max_feature_sparsity</code>, <code>feature_sparsity_threshold</code>, or <code>sparsity_threshold</code>; if none is supplied, it records the measured sparsity without applying an inferred cutoff. This prevents an undocumented numerical choice from being mistaken for a paper specification.</p><p>Similarity suppression is also configurable. The generated implementation uses absolute pairwise Pearson correlation on finite, pairwise-overlapping values. That metric is explicitly an implementation choice because the paper says that similarity suppression is used but does not identify the metric or threshold. Candidates with too few comparable observations are retained rather than automatically treated as duplicates.</p><p>The paper also mentions diversity. The supplied generated files do not define a separate diversity score or acceptance formula. In this reproduction architecture, retaining non-similar candidates is the available nonredundancy mechanism; a distinct diversity procedure would need an additional, explicitly configured policy rather than an invented default.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;56b47cc0-a916-49e8-bec2-c29b9f500eef&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def suppress_similar_features(
    features: Sequence[FeatureResult], threshold: float
) -&gt; list[FeatureResult]:
    """Retain accepted features while suppressing absolute Pearson correlation above threshold.

    Pearson correlation on pairwise finite observations is used as an explicit
    implementation choice because the paper specifies similarity suppression but
    does not supply its metric. Features with no comparable finite observations are
    retained rather than silently treated as duplicates.
    """
    if not isinstance(threshold, (int, float, np.integer, np.floating)):
        raise TypeError("threshold must be numeric")
    if not math.isfinite(float(threshold)) or not 0.0 &lt;= float(threshold) &lt;= 1.0:
        raise ValueError("threshold must be finite and lie in [0, 1]")

    retained: list[FeatureResult] = []
    for candidate in features:
        if not isinstance(candidate, FeatureResult):
            raise TypeError("features must contain FeatureResult instances")
        if not candidate.accepted:
            continue</code></pre></div><p>The rest of the function compares each accepted candidate with already retained candidates and suppresses it when the absolute correlation reaches the configured threshold. Because the threshold is supplied by the caller, two runs can deliberately preserve different screening policies while still recording which policy was used.</p><h3>The discovery pipeline and its audit trail</h3><p><code>discover_features</code> is the orchestration point. It obtains a <code>GenerationBatch</code> from the provider-neutral generator, validates the batch against its <code>GenerationContext</code>, checks every candidate, and registers both accepted and rejected candidates. Rejected candidates are not discarded from the audit history: their validation reports retain rejection reasons.</p><p>The registry is therefore more than a list of column names. Each entry retains the <code>GeneratedFeature</code>, its validation report, required columns, rationale, generation route, retrieval-context version, feature metadata, and other audit fields. This makes it possible to distinguish a feature rejected for a missing schema column from one rejected for look-ahead metadata or excessive configured sparsity.</p><p>A simplified view of the discovery boundary is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;092edf49-d3e3-4673-b161-b1bf776dd585&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">registry = FeatureRegistry()
accepted_results: list[FeatureResult] = []
reports: dict[str, ValidationReport] = {}

for feature in batch.candidates:
    report, values = _validate_candidate(frame, feature, schema, config)
    reports[feature.name] = report
    if bool(report.get("accepted", False)) and values is not None:
        accepted_results.append(_new_result(feature, values, report))
    registry.register(feature, report)</code></pre></div><p>After optional similarity suppression, the generated pipeline calls <code>registry.freeze()</code>. This ordering is essential: discovery feedback and screening belong to the January 2015&#8211;December 2018 discovery period, while the January 2019&#8211;December 2024 out-of-sample phase must use the resulting definitions without adding or replacing features.</p><h3>Freezing prevents out-of-sample mutation</h3><p><code>FeatureRegistry.register</code> rejects duplicate names and rejects every mutation after freezing. <code>FeatureRegistry.materialize</code> executes only entries whose reports have <code>accepted=True</code>, preserves the input key columns, and checks that each accepted callable returns a Series with the same length and index as the input frame.</p><p>The registry's mutation boundary is concise:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;55180ef5-9527-4c2b-987f-8da00facf854&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def freeze(self) -&gt; None:
    """Freeze accepted feature definitions before out-of-sample evaluation."""
    self._frozen = True


def materialize(self, frame: pd.DataFrame) -&gt; pd.DataFrame:
    """Execute accepted features and return keys plus aligned feature columns.

    Each callable is invoked through :class:`GeneratedFeature`, which
    enforces declared-column availability and one-Series row alignment.
    The returned rows retain the input order and the conventional
    ``CWIQ Code``/``date`` keys when those keys are present.
    """
    if not isinstance(frame, pd.DataFrame):
        raise TypeError("frame must be a pandas DataFrame")

    key_columns = self._key_columns(frame)
    output = frame.loc[:, key_columns].copy()
    for name, entry in self._entries.items():
        if not bool(entry.validation_report.get("accepted", False)):
            continue
        try:
            values = entry.feature(frame)
        except Exception as exc:
            raise RuntimeError(
                f"failed to materialize accepted feature {name!r}"
            ) from exc</code></pre></div><p><code>materialize_feature_sets</code> adds a further guard: it refuses to materialize an unfrozen registry and checks that materialization has not changed the row count or key values. The result is a key-preserving frame containing the accepted AI feature columns, ready for later model-feature selection.</p><h3>Worked example: one accepted candidate and one rejected candidate</h3><p>Suppose a discovery frame contains <code>CWIQ Code</code>, <code>date</code>, and <code>signal</code>. The supplied test file defines <code>_identity_feature</code>, which returns <code>frame["signal"]</code> without changing the input index. Wrapped as a <code>GeneratedFeature</code> with <code>required_columns=("signal",)</code>, it satisfies the one-Series alignment contract when called on the test frame.</p><p>A second candidate can use the same kind of callable but declare <code>metadata={"lags": [-1]}</code>. The function <code>check_point_in_time_lags</code> rejects that metadata with a negative-lag error. The candidate's report should remain registered with <code>accepted=False</code> and a rejection reason, while the valid candidate can remain eligible for registration as accepted.</p><p>The conceptual sequence is:</p><ol><li><p>Build or load the discovery-period frame and a user-supplied <code>DatasetSchema</code>.</p></li><li><p>Generate candidates through the selected route.</p></li><li><p>Run <code>validate_feature_schema</code>, <code>compile_feature</code>, <code>check_point_in_time_lags</code>, and <code>check_rolling_windows</code>.</p></li><li><p>Materialize candidates that pass those gates and measure sparsity.</p></li><li><p>Apply configured similarity suppression, if requested.</p></li><li><p>Register every candidate and its report.</p></li><li><p>Freeze the registry.</p></li><li><p>Call <code>materialize_feature_sets</code> on a later panel only after the registry is frozen.</p></li></ol><p>The supplied <code>tests/test_feature_contracts.py</code> expresses the intended semantic checks for alignment, negative-lag rejection, and frozen-registry immutability. Those tests were not executed under the authoritative run policy, so they are examples of planned checks rather than reported verification results.</p><p><strong>Advanced detail &#8212; discovery feedback and reproducibility.</strong> The paper's DSPy/MIPRO route can use discovery-period predictive or portfolio feedback, but such feedback must not include headline out-of-sample observations. The registry should therefore retain the discovery date range, generation route, retrieval-context version, validation thresholds, and model or evaluation configuration alongside each feature definition. Without those records, a later run might appear to use the same feature names while silently changing the code, context, or screening policy.</p><p>The central invariant is simple: no candidate that was not validated and frozen during discovery should affect out-of-sample feature materialization. The implementation makes that invariant inspectable, while leaving unspecified thresholds, similarity definitions, and production sandbox details explicit rather than pretending that the paper supplied them. Verification remains limited here: local static verification and semantic code verification were skipped, and no code execution or numerical validation is claimed.</p><h2>Daily Standardization, Feature Configurations, and Walk-Forward LightGBM</h2><p>How should the model compare securities on a given day without letting a later day influence that comparison? The implementation uses two chronological safeguards. First, each input feature is standardized within its own date cross-section. Second, the LightGBM model is fitted only on permitted historical observations before predicting the next eligible cross-section.</p><p>This section connects the paper's feature-set and prediction framework to <code>standardize_features_daily</code>, <code>select_model_features</code>, <code>select_hyperparameters_discovery</code>, and <code>fit_predict_walk_forward</code>. The paper identifies these operations through <code>eq_3_4_feature_sets</code>, <code>eq_3_4_prediction</code>, and <code>eq_3_4_estimation</code>. Their supplied canonical LaTeX fields are empty, so those identifiers are used as semantic anchors only; no equation display is reconstructed here.</p><h3>Standardize within each date</h3><p>Suppose <code>X</code> is a feature matrix with shape <code>(n_samples, n_features)</code>. Each row represents one security-date observation, identified elsewhere by <code>CWIQ Code</code> and <code>date</code>. For a selected feature column, <code>standardize_features_daily</code> groups rows by <code>date</code>, computes the mean and population dispersion inside each group, and replaces each valid value with its date-local standardized value.</p><p>The important implementation consequence is that the statistics for 2020-01-02 are calculated only from securities observed on 2020-01-02. They are not calculated from the entire sample, and they are not learned from future dates. The function returns a copy with the same rows, index, and columns as the input. Missing feature values remain missing.</p><p>The paper requires daily cross-sectional standardization, but does not specify the dispersion denominator or the behavior of a constant cross-section. The generated code therefore makes both choices visible. It uses population dispersion (<code>ddof=0</code>) and accepts <code>"zero"</code>, <code>"nan"</code>, or <code>"raise"</code> as the <code>zero_dispersion_policy</code>. These are implementation decisions, not recovered paper details.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;da1f234b-e259-4dc2-b6fc-f2461704ba06&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># Grouping is date-local and preserves the original row order.
date_groups = result.groupby(date_column, sort=False, dropna=False)

for column in columns:
    values = result[column].astype(float)
    means = date_groups[column].transform("mean")
    # Eq. 3.4 standardization: date-local centering and scaling of each feature.
    dispersions = date_groups[column].transform(
        lambda group: group.std(ddof=0)
    )
    valid_count = date_groups[column].transform("count")
    undefined_dispersion = valid_count.gt(0) &amp; (
        dispersions.isna() | dispersions.eq(0.0)
    )</code></pre></div><p>The code validates that selected columns are numeric and contain no infinite values. A constant date-feature group has no usable dispersion. Under the default <code>"zero"</code> policy, its valid observations become neutral zero values; an all-missing group remains missing. This avoids silently creating infinities or artificial cross-sectional direction.</p><p>The supplied test file <code>tests/test_standardization_and_portfolio.py</code> illustrates the intended invariant with deterministic data. It checks that a date with values 1, 2, and 3 has mean zero and population standard deviation one after transformation, while a later constant cross-section receives the configured zero fallback:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5cf7895a-d581-4b30-959f-89fbaa190a0b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">standardized = standardize_features_daily(
    frame,
    feature_columns=["feature"],
    date_column="date",
    zero_dispersion_policy="zero",
)

first_date = standardized.loc[
    standardized["date"].eq(pd.Timestamp("2019-01-02")), "feature"
]
# Eq. 3.4 standardization: statistics are computed within each date group.
assert first_date.mean() == pytest.approx(0.0)
assert first_date.std(ddof=0) == pytest.approx(1.0)</code></pre></div><p>This test is supplied as a planned or local semantic check, not as evidence that the test suite was executed. The run policy disabled execution and verification.</p><h3>Choose the feature configuration explicitly</h3><p>The paper distinguishes conventional baseline features, AI-generated features, and their combination. In the implementation, <code>select_model_features</code> receives a DataFrame, a <code>configuration</code> string, and two ordered column lists: <code>baseline_columns</code> and <code>ai_columns</code>. It returns a matrix with shape <code>(n_rows, n_selected_features)</code> and preserves the input row index.</p><p>The semantic operation identified by <code>eq_3_4_feature_sets</code> is column selection or concatenation. For <code>"baseline"</code>, only baseline columns are selected. For <code>"ai"</code>, only AI columns are selected. For <code>"combined"</code>, the selected columns are <code>baseline + ai</code>, in that order. Missing, duplicated, or nonnumeric columns are rejected rather than silently coerced.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7069b46e-bcbe-4570-b939-eb62a98f2b58&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">if normalized == "baseline":
    selected = baseline
elif normalized == "ai":
    selected = ai
else:
    # Eq. 3.4 feature sets: combined features are concatenated, not averaged.
    selected = baseline + ai

if len(set(selected)) != len(selected):
    raise ValueError(
        "baseline and AI feature sets contain duplicate column names; "
        "combined model columns must be unambiguous"
    )
return frame.loc[:, selected].copy()</code></pre></div><p>Concatenation is different from ensemble averaging. <code>average_standardized_predictions</code> accepts multiple prediction tables, standardizes each table by signal date, inner-aligns them on <code>CWIQ Code</code> and date, and averages the standardized prediction columns. It does not concatenate model input columns. This distinction matters: combined features give one model more input variables, whereas an ensemble averages outputs from separately produced prediction tables.</p><p>The function uses an inner alignment intentionally. A security-date key absent from one component is not treated as a zero prediction. The output contains the common keys and a <code>prediction</code> column representing the average of the component standardized predictions. This behavior is an implementation choice that makes coverage explicit.</p><h3>Fit only on the permitted past</h3><p>For the model stage, <code>X</code> has shape <code>(n_samples, n_features)</code>, <code>y</code> is a length-<code>n_samples</code> forward-return target, and <code>dates</code> is a length-<code>n_samples</code> series of normalized signal dates. The paper's <code>eq_3_4_prediction</code> identifies the nonlinear LightGBM mapping from a feature vector to a return prediction. In code, this is the estimator's <code>predict</code> call after fitting. <code>eq_3_4_estimation</code> corresponds to the estimator's <code>fit</code> call on the allowed historical training window.</p><p>The generated <code>WalkForwardSpec</code> makes the chronology configurable. Its <code>window</code> can be <code>"expanding"</code> or <code>"rolling"</code>; <code>max_train_dates</code> limits a rolling window; <code>min_train_dates</code> prevents fitting on an undersized history; <code>training_lag</code> removes the most recent completed signal dates from training; and <code>retrain_frequency</code> controls how often the model is refitted. The paper uses time-aware estimation but does not fully resolve every rolling-versus-expanding or retraining detail, so these settings remain explicit.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;aec8b6bc-ffa6-4661-9adc-8d4ea8ff71b7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class WalkForwardSpec:
    """Configuration for chronological LightGBM estimation.

    Dates are treated as ordered signal dates.  A model trained for a date is
    never allowed to use observations on that date or later.  ``training_lag``
    removes additional completed signal dates from the end of the training
    window when an implementation lag is required.
    """

    window: str = "expanding"
    max_train_dates: int | None = None
    min_train_dates: int = 20
    training_lag: int = 0
    retrain_frequency: int = 1</code></pre></div><p>For a current date position, <code>_training_positions</code> ends the training range before that date, then applies <code>training_lag</code>. In an expanding window, the start remains the beginning of the available history. In a rolling window, the start is moved forward when <code>max_train_dates</code> is supplied. Rows whose targets or features are unavailable are excluded from fitting; current-date rows with missing features are excluded from prediction.</p><p>The model loader imports LightGBM only when fitting is requested. If the optional dependency is absent, <code>_load_lightgbm</code> raises an informative <code>ImportError</code>; the generated code does not substitute another model. The estimator defaults are exposed through the <code>params</code> mapping. Although the generated scaffold supplies a deterministic baseline parameter mapping when no discovery candidates are provided, the paper's exact tree count, learning rate, depth, regularization, early stopping, and missing-value choices are not specified by the supplied evidence.</p><h3>Discovery-only hyperparameter selection</h3><p><code>select_hyperparameters_discovery</code> is deliberately separate from <code>fit_predict_walk_forward</code>. It receives discovery-period <code>X</code>, <code>y</code>, and <code>dates</code>, plus a <code>WalkForwardSpec</code>. Candidate parameter mappings must be supplied explicitly. If there is more than one candidate, the function evaluates chronological validation dates using training dates before each validation date and selects the candidate with the lowest mean squared error. If no candidates are supplied, it returns the documented deterministic baseline configuration and performs no hidden tuning.</p><p>This implements the paper's separation between discovery and headline out-of-sample reporting: January 2015 through December 2018 may be used for feature discovery, validation, and hyperparameter selection, while January 2019 through December 2024 must not tune the reported model. The function is not a claim that the paper's exact nested-fold design has been recovered; the supplied context does not provide its complete fold definitions.</p><h3>Produce a <code>PredictionTable</code></h3><p><code>fit_predict_walk_forward</code> loops through sorted signal dates. For each date, it builds a permitted historical mask, fits or reuses a LightGBM model according to <code>retrain_frequency</code>, and predicts the current cross-section. The returned table has one row per eligible security-date prediction and includes:</p><ul><li><p><code>CWIQ Code</code>, the security identifier;</p></li><li><p><code>signal_date</code>, the date at which the prediction is formed;</p></li><li><p><code>prediction</code>, the raw model output; and</p></li><li><p><code>row_position</code>, an audit-oriented reference to the input row.</p></li></ul><p>The function rejects duplicate security-date outputs and non-finite predictions. It does not convert predictions into weights; that belongs to the subsequent score and portfolio stage. Keeping raw predictions separate preserves the distinction between model output, daily cross-sectional score processing, and later realized-return alignment.</p><p>A small worked configuration might select baseline and AI columns as follows:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7197a080-f3a5-4d0d-b525-836f6a912a2c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">X = select_model_features(
    panel,
    configuration="combined",
    baseline_columns=["value_signal", "price_signal"],
    ai_columns=["ai_rank", "ai_interaction"],
)
Xz = standardize_features_daily(
    X.assign(date=panel["date"]),
    feature_columns=list(X.columns),
    date_column="date",
    zero_dispersion_policy="zero",
)

spec = WalkForwardSpec(
    window="expanding",
    min_train_dates=20,
    training_lag=1,
    retrain_frequency=1,
)</code></pre></div><p>Here <code>X</code> has four selected features, while <code>Xz</code> retains the same number of rows and selected feature columns after date-local standardization. The <code>training_lag=1</code> setting is an explicit implementation choice for keeping the most recent signal date out of the training window; the target construction must separately define whether the realized interval is standard or the headline delayed interval. The resulting <code>PredictionTable</code> is then ready for date-wise prediction z-scoring and portfolio construction.</p><h3>Advanced detail: ordering and pooled observations</h3><p>The paper states that model inputs are standardized daily, while some generated features may themselves be cross-sectional ranks. The supplied evidence does not fully specify whether ranking occurs before or after the general feature-standardization pass. The pipeline should therefore record that ordering in its configuration rather than treating it as self-evident.</p><p>Likewise, LightGBM is fit on pooled security-date rows, but the paper does not state whether dates or securities receive explicit weighting. The generated implementation uses the rows that pass the chronological masks and leaves weighting to the configured estimator. This is a reproduction decision, not a recovered paper parameter.</p><p>The central invariant remains unchanged: model fitting for a signal date may use only permitted historical rows, and the prediction table must retain the security and signal-date keys needed by the later portfolio stage.</p><h2>Prediction Scores, Winsorization, Dollar-Neutral Weights, and Realized Returns</h2><p>How does a model prediction become a tradable long-short portfolio? The model produces a relative signal for each security on a <code>signal_date</code>; it does not directly specify how many dollars to buy or sell. The portfolio layer therefore standardizes predictions within each date, limits extreme scores, normalizes positive and negative signals separately, and joins the resulting weights to a later realized-return interval.</p><p>This section implements the <code>prediction_score_and_portfolio_weights</code> method. The paper fact is that predictions are cross-sectionally standardized, winsorized to the interval <code>[-3, 3]</code>, converted into dollar-neutral long-short weights, and evaluated against forward returns. The generated code keeps the signal date separate from the later return date so that the timing choice&#8212;standard <code>t</code> to <code>t+1</code> or the headline implementable <code>t+1</code> to <code>t+2</code> interval&#8212;cannot be hidden inside the weight calculation.</p><h3>From raw predictions to daily scores</h3><p>Let a raw prediction be the model's scalar output for one security on one signal date. The implementation calls the resulting standardized value <code>alpha</code>. For each date, <code>cross_sectional_prediction_zscore</code> computes the mean and dispersion using only that day's securities. Thus, a prediction is interpreted relative to its contemporaneous cross-section rather than relative to the entire historical panel.</p><p>The paper identifies this operation as <code>eq_3_4_standardization</code>. Its canonical LaTeX is not supplied, so no formula is reconstructed here. Semantically, <code>\hat r_{i,t+1}</code> denotes the raw prediction for security <code>i</code> formed at date <code>t</code>; <code>\bar r_t</code> is the date-<code>t</code> cross-sectional mean; <code>\sigma_t</code> is the corresponding dispersion; and <code>\alpha_{i,t}</code> is the standardized scalar score. The security index <code>i</code> ranges over the names available on that date, so the transformation is a date-grouped operation rather than a single global normalization. In code, the mapping is <code>cross_sectional_prediction_zscore(predictions, prediction_column, date_column)</code> in <code>src/stock_selection/portfolio/scores.py</code>.</p><p>The generated implementation makes two choices explicit. Missing predictions remain missing, while a valid cross-section with zero or undefined dispersion receives score zero. It also uses pandas' sample standard deviation (<code>ddof=1</code>) for each date. The paper's extracted equation does not specify the denominator convention or the zero-dispersion fallback, so these are implementation decisions that should remain documented in the experiment configuration.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;0c927c20-e648-4dd2-8a2e-f2cefb654d2d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    groups = predictions[date_column]
    means = raw.groupby(groups, sort=False, dropna=False).transform("mean")
    # pandas' sample standard deviation is used for each cross-section.  For a
    # singleton or constant cross-section it is undefined or zero, respectively.
    dispersion = raw.groupby(groups, sort=False, dropna=False).transform(
        lambda values: values.std(ddof=1)
    )

    alpha = pd.Series(np.nan, index=predictions.index, dtype=float, name="alpha")
    valid_predictions = raw.notna() &amp; means.notna()
    nonzero_dispersion = dispersion.notna() &amp; (dispersion &gt; 0.0)
    regular = valid_predictions &amp; nonzero_dispersion

    # Eq. 3.4: cross-sectional prediction standardization.
    alpha.loc[regular] = (
        (raw.loc[regular] - means.loc[regular]) / dispersion.loc[regular]
    ).to_numpy()</code></pre></div><p>The important invariant is alignment: the returned <code>Series</code> has the same index and row order as <code>predictions</code>. A value in <code>alpha</code> still belongs to the same <code>CWIQ Code</code> and signal date as the raw prediction that produced it.</p><h3>Winsorization limits extreme signals</h3><p>A standardized score can still be unusually large when a security is far from the daily cross-sectional center. The paper's <code>eq_3_4_winsorization</code> operation clips the standardized score to <code>[-3, 3]</code>. Here, <code>\alpha_{i,t}</code> is the unbounded standardized score and <code>\tilde\alpha_{i,t}</code> is the clipped score. Both are scalar values associated with one security-date observation. The code mapping is <code>winsorize_scores(alpha, lower=-3.0, upper=3.0)</code>.</p><p>The function preserves missing values, validates finite bounds, and applies elementwise clipping. It does not replace the score with a rank or drop the observation. The default bounds are visible in the generated file:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;89f6d103-f041-4fbf-9172-74272918311b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">DEFAULT_WINSOR_LOWER: Final[float] = -3.0
DEFAULT_WINSOR_UPPER: Final[float] = 3.0</code></pre></div><p>A focused test records the intended invariant without depending on paper data:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5d8a3a64-852e-41a0-b719-6e0b37547621&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    # Eq. 3.4 winsorization: standardized scores are clipped to [-3, 3].
    assert clipped.iloc[:-1].min() == pytest.approx(-3.0)
    assert clipped.iloc[:-1].max() == pytest.approx(3.0)
    assert np.all(clipped.dropna().between(-3.0, 3.0))
    assert pd.isna(clipped.iloc[-1])</code></pre></div><p>This test is supplied code, but it was not executed under the authoritative run policy. It expresses the intended bound rather than reporting a completed verification result.</p><h3>Separate long and short normalization</h3><p>The next question is how a clipped score becomes a portfolio weight. The paper identifies this transformation as <code>eq_3_4_weights</code>. For security <code>i</code> on date <code>t</code>, <code>w_{i,t}</code> is the signed portfolio weight. The positive part <code>\tilde\alpha^+_{i,t}</code> contributes to the long side, while <code>\tilde\alpha^-_{i,t}</code> denotes the nonnegative magnitude of a negative score. The index <code>j</code> represents another security in the same date cross-section, because each side is normalized against the aggregate score magnitude of all eligible names that day.</p><p>The canonical LaTeX for this equation is unavailable, so the following is a semantic description rather than a recovered display formula. Positive scores are divided by the total positive score and assigned a positive sign. Negative magnitudes are divided by the total negative magnitude and assigned a negative sign. With the default <code>side_exposure=1.0</code>, the long side has gross exposure one and the short side has gross exposure one whenever that side contains a nonzero score. The resulting total absolute exposure is therefore two when both sides are populated, while net exposure is zero.</p><p><code>src/stock_selection/portfolio/weights.py</code> implements this policy with separate totals:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4b466508-82ed-4da0-b3d0-517832b6e85e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    # Eq. 3.4: score-proportional long and short portfolio weights.
    long_totals = positive.groupby(result[date_column], sort=False, dropna=False).transform("sum")
    short_totals = negative_magnitude.groupby(result[date_column], sort=False, dropna=False).transform("sum")

    long_weights = pd.Series(0.0, index=result.index, dtype=float)
    short_weights = pd.Series(0.0, index=result.index, dtype=float)
    has_long = long_totals &gt; 0.0
    has_short = short_totals &gt; 0.0
    long_weights.loc[has_long] = exposure * positive.loc[has_long] / long_totals.loc[has_long]
    short_weights.loc[has_short] = -exposure * negative_magnitude.loc[has_short] / short_totals.loc[has_short]

    result[WEIGHT_COLUMN] = long_weights + short_weights</code></pre></div><p>The implementation decision for an empty side is to assign zero weight rather than inventing a fallback. Missing scores also receive zero weight. This is practical and auditable, but it is not a separately specified paper convention; an alternative reproduction should expose a different policy explicitly.</p><h4>Worked example: one five-security cross-section</h4><p>Suppose the clipped scores on one signal date are as follows:</p><ul><li><p>Security <code>A</code>: <code>2</code></p></li><li><p>Security <code>B</code>: <code>1</code></p></li><li><p>Security <code>C</code>: <code>0</code></p></li><li><p>Security <code>D</code>: <code>-1</code></p></li><li><p>Security <code>E</code>: <code>-3</code></p></li></ul><p>The positive score total is <code>3</code>, so <code>A</code> receives a long weight of <code>2/3</code> and <code>B</code> receives <code>1/3</code>. The negative magnitudes total <code>4</code>, so <code>D</code> receives <code>-1/4</code> and <code>E</code> receives <code>-3/4</code>. <code>C</code> has zero weight. The long weights sum to <code>1</code>, the short weights sum to <code>-1</code>, and the net portfolio weight is zero. This illustrates why the paper's descriptions of &#8220;unit gross leverage&#8221; and total gross exposure must be read carefully: the generated default uses one unit on each side, not one unit in total absolute exposure.</p><p>The same operation is available for a single cross-section through <code>construct_naive_weights(alpha)</code>. That function is the implementation mapping for <code>eq_appendix_B_naive_weights</code>, whose supplied equation also has no canonical LaTeX. It is useful when comparing the direct score-proportional portfolio with the optional optimization module.</p><h3>Keep ensemble averaging separate from feature concatenation</h3><p>The file <code>src/stock_selection/model/features.py</code> contains <code>average_standardized_predictions</code>, which handles an ensemble of prediction tables. This is different from selecting a combined feature matrix with <code>select_model_features</code>.</p><p>Feature concatenation occurs before model fitting: baseline columns are placed before AI columns in the combined matrix. Ensemble averaging occurs after separate model predictions have been standardized by signal date. The function inner-aligns the prediction tables on <code>CWIQ Code</code> and date, then averages the standardized values. The intersection policy is deliberate: a security absent from one model variant is not silently treated as a zero prediction.</p><p>The core alignment and averaging step is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ada38244-ad02-4060-bedb-459a9c56cdf2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">        selected["_standardized"] = _standardize_by_date(
            selected[value_column], selected[date_column]
        )
        prepared.append(
            selected.loc[:, [identifier, date_column, "_standardized"]]
        )

    assert common_keys is not None
    identifier, date_column = common_keys
    combined = prepared[0].rename(columns={"_standardized": "_prediction_0"})
    for position, table in enumerate(prepared[1:], start=1):
        combined = combined.merge(
            table.rename(columns={"_standardized": f"_prediction_{position}"}),
            on=[identifier, date_column],
            how="inner",
            sort=False,
            validate="one_to_one",
        )</code></pre></div><p>The output <code>prediction</code> column is therefore an average of component standardized predictions. It can then pass through the same score, clipping, and weight stages as a single model prediction.</p><h3>Join signal-date weights to realized returns</h3><p>Portfolio return aggregation is the implementation mapping for <code>eq_3_4_portfolio_return</code>. The paper's symbols are <code>R^{port}_{t+1}</code> for the portfolio return, <code>w_{i,t}</code> for the weight formed at the signal date, and <code>r_{i,t+1}</code> for a realized security return. In the headline delayed evaluation, the realized interval is instead configured as <code>t+1</code> through <code>t+2</code>; the important invariant is that the weight is formed before the return interval it evaluates.</p><p><code>aggregate_portfolio_return(weights, returns)</code> joins on <code>CWIQ Code</code> and either <code>signal_date</code> or <code>date</code>, depending on the shared columns. The return table must contain one return value per security-date key. A missing realized return contributes no weighted observation, and a security absent from the return table contributes no realized-return term. Those are explicit join behaviors, not evidence that the paper specified a particular missing-data policy.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e64f0d89-5053-4333-b39c-d8f7b64327df&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    merged = left.merge(right, on=key_columns, how="left", validate="one_to_one")
    merged["_contribution"] = merged[WEIGHT_COLUMN].astype(float) * merged[return_column].astype(float)
    merged.loc[merged[return_column].isna(), "_contribution"] = np.nan

    # Eq. 3.4: portfolio return is the weighted aggregation of realized returns.
    portfolio_returns = merged.groupby(date_column, sort=True)["_contribution"].sum(min_count=1)
    portfolio_returns.name = "portfolio_return"
    return portfolio_returns.astype(float)</code></pre></div><p>For the worked example, suppose the later realized returns are <code>0.03</code> for <code>A</code>, <code>0.01</code> for <code>B</code>, <code>0.02</code> for <code>D</code>, and <code>-0.01</code> for <code>E</code>. Applying the weights above gives a portfolio return of approximately <code>0.0258</code> for that signal date. This is a derived arithmetic illustration, not a paper result. In a real run, the <code>returns</code> table would be built by the configured target-construction stage, and its key would identify the return interval corresponding to the selected timing policy.</p><h3>What to preserve when securities disappear</h3><p>Daily universes change. A security may have a missing quote, fail a liquidity filter, delist, or lack a return observation over the evaluation interval. The generated functions reject duplicate security-date keys and retain explicit key columns, which makes these cases visible. They do not silently forward-fill an unavailable return or renormalize weights after a missing realized observation.</p><p>That policy is conservative for auditability, but the paper context does not fully specify whether missing securities should trigger weight renormalization, a zero contribution, or exclusion of the entire date. A numerical reproduction must choose and record one policy in its evaluation configuration. It must also keep the return horizon explicit: standard <code>t</code> to <code>t+1</code> and implementable <code>t+1</code> to <code>t+2</code> should not be combined under one unlabeled output.</p><h3>Practical checklist</h3><p>Before passing portfolio returns to metrics, inspect these invariants:</p><ul><li><p>scores are grouped by <code>signal_date</code>, not standardized across future dates;</p></li><li><p>winsorized scores remain within <code>[-3, 3]</code>;</p></li><li><p>positive weights are nonnegative and negative weights are nonpositive;</p></li><li><p>each populated side has the configured exposure, normally <code>1.0</code>;</p></li><li><p>both populated sides sum to zero net exposure;</p></li><li><p>weights retain <code>CWIQ Code</code> and signal-date keys;</p></li><li><p>realized returns are joined for the configured later interval; and</p></li><li><p>missing predictions, missing returns, and disappearing securities follow recorded policies.</p></li></ul><p>The supplied tests in <code>tests/test_standardization_and_portfolio.py</code> express these intended checks through <code>test_winsorization_bounds_are_three</code>, <code>test_weights_have_zero_net_exposure_and_unit_sides</code>, and <code>test_portfolio_return_uses_aligned_future_returns</code>. The run policy disabled test execution, local static verification, semantic code verification, and numerical verification, so these tests should be read as planned or supplied checks rather than completed evidence.</p><h2>Metrics, Execution-Lag Analysis, HAC Inference, and Factor Attribution</h2><p>How do we tell whether a cross-sectional signal is useful, and how do we ensure that the answer uses the same trading delay as the portfolio? The evaluation layer answers two related questions. First, does the ranking of securities agree with subsequent returns? Second, does the resulting portfolio produce consistent returns through time? The generated implementation keeps these questions separate: <code>compute_daily_spearman_ic</code> measures cross-sectional ranking quality, while <code>annualized_sharpe</code> and the portfolio-return functions measure the time series of portfolio outcomes.</p><p>The paper facts are that information coefficients use daily cross-sectional Spearman correlations, Sharpe is annualized with 252 trading periods, and execution-lag experiments include lag 0, 1, 2, 3, 5, and 10. The implementation decisions are equally important: lag alignment is explicit, zero-dispersion statistics become undefined rather than being assigned an invented value, and hit-rate and drawdown conventions are selected through function arguments. The supplied verification stages were skipped, so the following describes the generated code and its intended behavior, not an executed result.</p><h3>Information coefficient: evaluate the ranking within each date</h3><p>The symbol <code>IC_t</code> denotes the information coefficient on date <code>t</code>. In this implementation it is a Spearman rank correlation computed separately for each date between a score and the realized return assigned to that date. Spearman correlation compares ranks rather than raw magnitudes, which makes it appropriate for asking whether higher-scored securities tend to have higher subsequent returns.</p><p>This is the semantic implementation of <code>eq_3_4_ic</code>. Its inputs are two pandas DataFrames keyed by <code>CWIQ Code</code> and either <code>signal_date</code> or <code>date</code>. The score table should contain a numeric column such as <code>score</code>, <code>alpha</code>, or <code>prediction</code>; the return table should contain a numeric column such as <code>realized_return</code> or <code>forward_return</code>. The output is a date-indexed one-dimensional Series named <code>daily_ic</code>. It is not a pooled correlation across all securities and dates.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e529250e-c16a-48f6-bd68-a78ea23f23e1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def compute_daily_spearman_ic(scores: pd.DataFrame, returns: pd.DataFrame) -&gt; pd.Series:
    """Compute date-wise cross-sectional Spearman information coefficients.

    The score and realized-return tables must identify observations by security and
    date.  Rows with missing values are omitted within their date, and dates with
    fewer than two usable securities or constant ranks produce ``NaN``.  The
    return interval is determined by the caller's timing configuration; this
    function does not shift or infer a forward horizon.
    """
    scores = _require_frame(scores, "scores")
    returns = _require_frame(returns, "returns")
    if SECURITY_KEY not in scores.columns or SECURITY_KEY not in returns.columns:
        raise KeyError(f"Both inputs must contain {SECURITY_KEY!r}")
    date_column = _choose_date_column(scores, returns)
    score_column = _choose_numeric_column(
        scores,
        ("score", "alpha", "prediction", "clipped", "standardized_prediction"),
        {SECURITY_KEY, date_column},
        "scores",
    )
    return_column = _choose_numeric_column(
        returns,
        ("realized_return", "forward_return", "return", "log_return"),
        {SECURITY_KEY, date_column},
        "returns",
    )</code></pre></div><p>Notice the timing boundary in the docstring: this function does not shift returns. The caller must first construct the appropriate target, such as the paper's delayed <code>t+1</code> to <code>t+2</code> interval, and then pass that aligned target. Duplicate security-date keys raise an error because silently aggregating duplicates could change the cross-section. Dates with fewer than two usable observations, or with constant ranks, produce <code>NaN</code> rather than a misleading correlation.</p><h3>Portfolio returns and annualized Sharpe</h3><p>The paper's portfolio-return quantity, identified here by <code>eq_3_4_portfolio_return</code>, is the date-level aggregation of security returns using weights formed at the signal date. <code>evaluate_portfolio_returns</code> delegates that aggregation to the shared portfolio utility. Its input weights must therefore be aligned with realized security returns for the intended future interval. The function does not infer whether those returns represent standard or delayed execution.</p><p>The symbol <code>SR</code> denotes annualized Sharpe. The generated <code>annualized_sharpe</code> function uses the paper's 252-period convention: it divides the mean daily portfolio return by its sample standard deviation and applies the annualization factor. The exact canonical LaTeX for <code>eq_3_4_sharpe</code> was not supplied, so this paragraph and the function name provide the mapping without reconstructing a display formula.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c6823664-e81a-42c2-a762-1a2fdebc7973&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def annualized_sharpe(
    portfolio_returns: pd.Series,
    periods_per_year: int = 252,
) -&gt; float:
    """Return the annualized daily-return Sharpe ratio.

    This implements the semantic mapping of ``eq_3_4_sharpe`` using the paper's
    252-period convention.  Missing observations are excluded.  A series with
    fewer than two observations or zero sample dispersion has undefined Sharpe
    and is represented by ``numpy.nan`` rather than by an invented value.
    """
    returns = _validate_numeric_series(portfolio_returns, "portfolio_returns")[
        :
    ].dropna()
    if isinstance(periods_per_year, bool) or not isinstance(periods_per_year, int):
        raise TypeError("periods_per_year must be an integer")
    if periods_per_year &lt;= 0:
        raise ValueError("periods_per_year must be positive")
    if len(returns) &lt; 2:
        return float("nan")
    dispersion = float(returns.std(ddof=1))
    if not np.isfinite(dispersion) or dispersion == 0.0:
        return float("nan")
    return float(np.sqrt(periods_per_year) * returns.mean() / dispersion)</code></pre></div><p>The important shape is one finite or missing return per evaluation date. Missing observations are omitted. A single observation, or a sample with zero standard deviation, has no defined sample Sharpe in this implementation and returns <code>numpy.nan</code>. This is an explicit implementation policy because the paper does not specify a zero-volatility fallback.</p><p><code>compute_hit_rate</code> and <code>compute_max_drawdown</code> expose additional conventions rather than hiding them. The available hit-rate choices are <code>positive_days</code>, <code>nonnegative_days</code>, and <code>absolute_days</code>. Drawdown can use a compounded wealth path or a cumulative arithmetic path. The paper does not fully specify these conventions, so a metric table should record the selected value. In particular, these functions should not be treated as evidence that one convention is uniquely required by the paper.</p><h3>Worked example: a small date-grouped evaluation</h3><p>Consider a synthetic panel with several securities on each of two signal dates. Each row contains <code>CWIQ Code</code>, <code>signal_date</code>, a model <code>score</code>, and a realized return already aligned to the chosen execution interval. The intended sequence is:</p><ol><li><p>Pass the score and return tables to <code>compute_daily_spearman_ic</code>.</p></li><li><p>Inspect one IC value per signal date; the calculation is cross-sectional within each date.</p></li><li><p>Use signal-date weights and the same realized-return interval with <code>evaluate_portfolio_returns</code>.</p></li><li><p>Pass the resulting date-indexed portfolio-return Series to <code>annualized_sharpe</code>, <code>compute_hit_rate</code>, and <code>compute_max_drawdown</code>.</p></li></ol><p>For example, a score that ranks securities correctly on both dates may produce positive daily IC values even if the portfolio return series is volatile. Conversely, a portfolio can have a positive Sharpe while its daily cross-sectional IC is inconsistent. These are different diagnostics, not interchangeable summaries. No numerical output is asserted here because the synthetic example was not executed under the run policy.</p><h3>Execution-lag sensitivity</h3><p>An execution lag specifies how many later observations separate the signal from the return used for evaluation. Lag 0 is retained for comparison but is labeled non-implementable in the generated robustness output. Positive lags represent delayed evaluation scenarios. This distinction matters because the paper's headline implementable evaluation uses a delayed return interval, while other notation describes a standard one-period return.</p><p><code>run_execution_lag_sensitivity</code> receives fixed prediction and return tables plus an explicit sequence of lags. It returns a DataFrame with <code>execution_lag</code>, predictive metrics, observation counts, and a Boolean <code>lag_0_non_implementable</code> flag. The implementation uses positional shifting within each security because the supplied context does not define a complete calendar-day or holiday convention. That is an implementation choice and should remain visible in audit metadata; it is not a recovered paper rule.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4f12867c-987d-461b-b584-58fbba367749&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def run_execution_lag_sensitivity(
    predictions: pd.DataFrame,
    returns: pd.DataFrame,
    lags: Sequence[int],
) -&gt; pd.DataFrame:
    """Evaluate fixed predictions over explicitly supplied execution lags."""
    if isinstance(lags, (str, bytes)):
        raise TypeError("lags must be a sequence of integers")
    lag_values = list(lags)
    if not lag_values:
        raise ValueError("lags must contain at least one lag")
    if len(set(lag_values)) != len(lag_values):
        raise ValueError("lags must not contain duplicates")
    rows: list[dict[str, Any]] = []
    for lag in lag_values:
        metrics = _lagged_evaluation(predictions, returns, lag)
        rows.append({"execution_lag": lag, **metrics, "lag_0_non_implementable": lag == 0})
    return pd.DataFrame(rows)</code></pre></div><p>The helper validates duplicate lags and preserves the requested order. It does not retune the model for each lag. This is important for an alpha-decay analysis: the question is how a fixed signal behaves as execution is delayed, not which lag produces the best newly optimized model.</p><h3>HAC inference: uncertainty around a time series</h3><p>Daily returns and daily IC observations can be serially correlated. Newey-West heteroskedasticity-and-autocorrelation consistent, or HAC, inference adjusts the uncertainty of an estimated mean or regression coefficient for that dependence. It changes the standard error and test statistic; it does not change the underlying portfolio-return or IC series.</p><p><code>HACConfig</code> records <code>max_lags</code>, confidence level, bootstrap replicates, expected block length, and random seed. <code>newey_west_summary</code> returns a <code>MetricResult</code> containing an estimate, HAC standard error, t statistic, and observation count. The paper specifies five Newey-West lags for the factor-attribution regression, but it does not uniformly specify the bandwidth for Sharpe or IC inference. Therefore, a five-lag configuration is appropriate for that specified factor regression, while callers must explicitly choose and label the bandwidth for other summaries.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2dbf0e35-1b84-45c4-8b8f-0ba26a28339e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class HACConfig:
    """Configuration for HAC and stationary-bootstrap summaries.

    ``max_lags`` is intentionally explicit because the paper specifies five
    lags for factor attribution but does not specify a bandwidth for Sharpe or
    IC inference.  Callers should therefore record their chosen value rather
    than relying on an inferred paper default.
    """

    max_lags: int = 5
    confidence_level: float = 0.95
    bootstrap_replicates: int = 1_000
    block_length: float = 5.0
    random_seed: int | None = 0</code></pre></div><p>The generated <code>stationary_bootstrap_interval</code> is optional. It resamples finite observations in their supplied chronological order using the configured expected block length and returns a percentile interval for the sample mean. The paper does not specify the bootstrap block-length rule, so <code>block_length</code> is configuration metadata rather than a paper-recovered constant. Missing or non-finite observations are omitted by the helper before inference.</p><h3>Optional factor attribution</h3><p>Factor attribution asks whether strategy returns can be explained by exposures to supplied risk factors. The paper describes Fama-French five factors plus momentum and uses five HAC lags for this regression. The generated <code>align_factors</code> function accepts local factor data and a nonnegative positional lag, performs an inner date join, and does not interpolate or forward-fill missing factor dates. <code>run_factor_attribution</code> then includes every non-return column in the regression and reports alpha, factor loadings, standard errors, t statistics, R-squared, and alignment metadata.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;15b329d5-1e1c-4b89-b39c-66d34ba65445&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def run_factor_attribution(
    aligned: pd.DataFrame,
    hac_lags: int = 5,
) -&gt; MetricResult:
    """Run optional factor attribution with Newey-West errors.

    The supplied factor frame is expected to contain the five Fama-French
    factors plus momentum when those columns are available.  No factor names are
    invented here: every non-return column is included in the regression.  The
    default HAC bandwidth is five lags, matching the paper's factor-attribution
    specification.  If a short sample cannot support five lags, the effective
    bandwidth is reduced and recorded explicitly.

    The exact external factor source, holiday treatment, and forward-shifted
    date convention are not supplied by the paper and must be resolved by the
    caller before this function is used for numerical reproduction.
    """</code></pre></div><p>The factor frame must be supplied by the caller; it is not included in the generated package. The exact factor source, holiday calendar, and forward-shifted date convention are likewise unresolved. If the sample is too short to use all requested lags, the generated function reduces the effective bandwidth and records both requested and used values. That behavior is transparent implementation handling, not a claim about the paper's intended small-sample rule.</p><h3>Aggregating experiments without hiding coverage</h3><p><code>aggregate_experiment_metrics</code> converts <code>MetricResult</code> objects into an auditable table while preserving metric values and metadata. It does not reweight or average experiment results. This is appropriate because the paper's grand composites involve multiple experiments and datasets, while the supplied evidence does not fully specify coverage weighting or missing-date treatment.</p><p><code>average_experiment_scores</code> serves a different purpose: it aligns prediction variants on the intersection of <code>CWIQ Code</code> and date keys, standardizes each variant within date, and computes an arithmetic mean. It does not forward-fill or outer-join missing predictions. The output records input coverage and the number of contributing variants, so an ensemble score is not mistaken for a directly observed score on every possible key.</p><h3>Practical evaluation checklist</h3><p>Before interpreting a metric table, confirm the following:</p><ul><li><p>The score and realized return share security-date keys, with the return interval determined before metric computation.</p></li><li><p>IC is grouped by date and is not calculated on the pooled panel.</p></li><li><p>Portfolio returns use weights formed at the signal date and returns from the configured later interval.</p></li><li><p>Sharpe uses the documented 252-period annualization and records its zero-dispersion policy.</p></li><li><p>Lag 0 is labeled non-implementable, while positive lags are explicitly identified.</p></li><li><p>Hit-rate, drawdown, compounding, HAC bandwidth, bootstrap block length, and factor-date alignment are recorded as conventions.</p></li><li><p>Factor attribution is run only when caller-supplied factor data are available.</p></li><li><p>No paper-level numerical result is claimed from these functions without the unavailable original inputs and a completed execution and verification process.</p></li></ul><p>The generated evaluation modules provide the numerical boundaries needed by the larger pipeline, but they do not by themselves establish empirical reproduction. Under the authoritative run policy, no code execution, test execution, local static verification, semantic code verification, or numerical comparison occurred.</p><h2>Smoothing, Rebalancing, Turnover, Liquidity Filters, and Transaction Costs</h2><p>How can a strategy reduce trading without accidentally using future information? The paper treats this as a separate implementation question: smooth the prediction signal before forming scores, rebalance either daily or monthly, measure every change in positions, and subtract explicitly defined trading costs from gross returns. These choices are alternatives to compare, not defaults that can be silently combined.</p><p>This section implements the <code>transaction_costs_turnover_and_smoothing</code> method. The paper describes trailing smoothing windows of none, 5, 10, and 21 days, liquidity filters based on average dollar volume and bid-ask spreads, and position-level spread and market-impact costs. The generated implementation preserves those choices as separate functions and records the impact coefficient explicitly.</p><h3>Smoothing comes before score formation</h3><p>The paper's ordering is important: smoothing is applied to raw predictions before cross-sectional standardization and portfolio formation. For a security identified by <code>CWIQ Code</code>, <code>smooth_predictions</code> sorts observations by security and date, then computes a trailing moving average. A window of 5 therefore uses the current observation and preceding observations for that security; it must not use a later prediction merely because that later row is present in the table.</p><p>This is an implementation decision embodied by <code>src/stock_selection/portfolio/smoothing.py</code>. The function preserves the original row order and all non-prediction columns, which makes it possible to join the smoothed result back to the original security-date keys.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4784c0c7-4c98-4dba-84cd-f2030565e5b7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">smoothed = (
    result.groupby(_SECURITY_COLUMN, sort=False, dropna=False)["prediction"]
    .rolling(window=int(window), min_periods=1)
    .mean()
    .reset_index(level=0, drop=True)
)
result["prediction"] = smoothed.to_numpy(dtype=float)</code></pre></div><p>The <code>window</code> argument is either <code>None</code> or a positive integer. <code>None</code> leaves predictions unchanged; the robustness plan compares it with 5, 10, and 21. Missing predictions remain subject to the function's input validation and downstream missing-data policy. The generated function does not claim to reproduce the paper's complete vendor-data handling, because those source details are not supplied.</p><p>The following excerpt from <code>scripts/run_robustness.py</code> shows how smoothing and rebalancing are deliberately crossed as labeled variants. This is useful because a 21-day daily-rebalance experiment is not the same experiment as a 21-day monthly-rebalance experiment.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;38ed4dac-9075-4fa4-8383-492d1e01a8ec&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def _prediction_variants(predictions: pd.DataFrame) -&gt; dict[str, pd.DataFrame]:
    variants: dict[str, pd.DataFrame] = {}
    for window in (None, 5, 10, 21):
        label = "none" if window is None else str(window)
        smoothed = smooth_predictions(predictions, window)
        for frequency in ("daily", "monthly"):
            scheduled = apply_rebalance_schedule(smoothed, frequency)
            variants[f"smoothing_{label}__rebalance_{frequency}"] = scheduled
    return variants</code></pre></div><h3>Daily and monthly rebalancing are different experiments</h3><p><code>apply_rebalance_schedule</code> accepts <code>daily</code> or <code>monthly</code>. Daily rebalancing returns a validated copy. Monthly rebalancing uses the first observed date in each calendar month as the formation date, then carries the most recent previously formed weight forward for each security. A security with no prior formation remains missing rather than being populated from a future observation.</p><p>That last behavior is a leakage safeguard. If a name first appears halfway through a month, the implementation cannot use its later monthly formation to fill an earlier date. When turnover is calculated, callers may interpret missing positions as zero, but that interpretation must be explicit.</p><p>The paper reports daily and monthly variants with different smoothing and cost conventions. Consequently, a reported monthly result must not be described as evidence about the daily-rebalanced strategy. The generated command-line tool writes each variant under a distinct name and includes metadata indicating that smoothing and rebalancing remain separate.</p><h3>Turnover includes entries and exits</h3><p>The paper's turnover identifier is <code>eq_3_10_turnover</code>. Its supplied equation record has no canonical LaTeX, so the implementation follows the stated meaning: on each date, sum the absolute changes in weights across the union of current and previous positions. This union is essential. A security that disappears contributes its previous weight as an exit, and a newly appearing security contributes its current weight as an entry.</p><p><code>compute_turnover</code> in <code>src/stock_selection/costs/turnover.py</code> accepts long-form DataFrames containing <code>CWIQ Code</code>, one supported date column, and one supported weight column. It validates duplicate keys, aligns the current and previous tables, fills absent positions with zero, and groups absolute changes by date.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;50091471-9440-43d9-ba30-2cebb0f10d27&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">merged["weight_current"] = merged["weight_current"].fillna(0.0)
merged["weight_previous"] = merged["weight_previous"].fillna(0.0)
turnover = (
    merged.assign(
        absolute_change=(
            merged["weight_current"] - merged["weight_previous"]
        ).abs()
    )
    .groupby("date", sort=True)["absolute_change"]
    .sum()
    .rename("turnover")
)</code></pre></div><p>A focused test in <code>tests/test_metrics_and_costs.py</code> demonstrates the intended entry-and-exit accounting:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5fa0faa1-52ad-47f8-b385-fcab1f776508&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># On 2019-01-02, A changes by 0.25 and B enters at 0.50.
# On 2019-01-03, A exits from -0.40 and B enters at 0.40.
assert turnover.loc[pd.Timestamp("2019-01-02")] == pytest.approx(0.75)
assert turnover.loc[pd.Timestamp("2019-01-03")] == pytest.approx(0.80)</code></pre></div><p>This is a planned or supplied semantic check, not a reported execution result. Under the run policy, tests and code execution were not performed.</p><h3>Spread costs and liquidity inputs</h3><p>The paper's <code>eq_3_10_spread</code> record describes a one-way half bid-ask spread expressed relative to midpoint price. <code>compute_spread_cost</code> in <code>src/stock_selection/costs/spread.py</code> requires aligned numeric <code>bid</code>, <code>ask</code>, and <code>midpoint</code> Series. It rejects negative quotes, reversed quotes, nonpositive midpoints, and mismatched indexes. Incomplete quote rows remain missing rather than being silently assigned zero cost.</p><p>The generated implementation maps the semantic description to the following focused operation:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c2ace42e-d2d6-44a7-9de2-2b73fc80e3c5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">result.loc[complete] = (
    (ask_values.loc[complete] - bid_values.loc[complete])
    / (2.0 * midpoint_values.loc[complete])
)</code></pre></div><p>Liquidity filtering happens upstream of cost estimation when enabled. The paper describes excluding securities with median daily dollar volume below $1 million or bid-ask spread above 50 basis points. The implementation plan requires those thresholds to be supplied through <code>UniverseConfig</code>; it does not assume that every local input contains <code>ADV</code>, spread, or quote fields. If the necessary columns are absent, the caller must choose an explicit policy rather than pretending that the filter was applied.</p><p><code>AUM</code> also matters for converting portfolio weights into dollar position sizes. The paper reports a $100 million robustness case, but the exact mapping from turnover to traded notional and the treatment of entry, exit, and rebalance costs are not fully specified. Therefore, an implementation should retain <code>AUM</code>, position-size convention, and trade-timing convention in its audit metadata.</p><h3>Position-level market impact: preserve the ambiguity</h3><p>The paper's <code>eq_3_10_impact</code> record describes impact as depending on volatility and position size relative to average dollar volume. The accompanying prose describes a square-root position-to-ADV form, but the supplied canonical LaTeX is unavailable. More importantly, the paper gives two impact coefficients: Section 3.10 states <code>k = 0.3</code>, while the robustness discussion uses <code>k = 0.2</code>.</p><p>The generated <code>compute_market_impact</code> function therefore requires both <code>coefficient</code> and <code>formulation</code>. It currently accepts the explicitly named formulation <code>sqrt_position_over_adv</code> and rejects an unrecognized formulation instead of silently selecting one.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;fe4d9647-67d1-4cb8-bad2-4589be665140&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">impact = coefficient_value * volatility_values * np.sqrt(position_values / adv_values)
result = pd.Series(impact, index=volatility_values.index, name="market_impact")</code></pre></div><p>This code is an implementation mapping for <code>eq_3_10_impact</code>, not a recovery of the missing equation text. <code>volatility</code> is nonnegative, <code>position_dollars</code> is a nonnegative absolute position size, and <code>adv</code> must be strictly positive. The output is one nonnegative one-way impact cost per aligned security-date row.</p><p>The robustness script keeps the two reported configurations distinct:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d5e645b7-7087-47f2-aba2-a1112ee41fa1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">for coefficient in (0.2, 0.3):
    impact = compute_market_impact(
        merged["volatility"],
        merged[weight_column].abs(),
        merged["adv"],
        coefficient,
        "sqrt_position_over_adv",
    )</code></pre></div><p>The <code>eq_3_10_total_cost</code> identifier refers to combining one-way spread and impact into a round-trip cost. <code>compute_round_trip_cost</code> doubles the sum of the two validated one-way components. The resulting configuration must state whether costs are being applied per traded position, per rebalance, or through another explicitly selected convention.</p><h3>Worked example: compare smoothing and cost configurations</h3><p>Suppose a prediction table contains several dates for securities <code>A</code> and <code>B</code>. To compare an unsmoothed signal with a 5-day signal, call <code>smooth_predictions(predictions, None)</code> and <code>smooth_predictions(predictions, 5)</code>. Feed each output separately into the score-standardization and portfolio-weight stages described earlier. Do not smooth already standardized scores if the experiment is intended to follow the paper's stated ordering.</p><p>For the cost side, use the same aligned <code>volatility</code>, absolute <code>position_dollars</code>, and <code>adv</code> inputs twice: once with <code>coefficient=0.2</code> and once with <code>coefficient=0.3</code>. Compute the spread once from the quote data, then call <code>compute_round_trip_cost</code> for each impact result. The two outputs are separate sensitivity artifacts; neither is the uniquely correct paper value given the internal inconsistency.</p><p>A simple turnover case makes the key accounting visible. If a previous portfolio has security <code>A</code> at weight <code>0.25</code>, while the current portfolio changes <code>A</code> to <code>0.50</code> and adds <code>B</code> at <code>-0.50</code>, the date's turnover includes both the <code>0.25</code> change in <code>A</code> and the <code>0.50</code> entry in <code>B</code>. If <code>A</code> later disappears, its prior weight is still included as an exit through the union alignment.</p><h3>What this implementation does&#8212;and does not&#8212;establish</h3><p>The generated functions provide auditable mechanics for trailing smoothing, rebalance scheduling, union-based turnover, quote validation, semantic impact calculation, and round-trip costs. They do not establish the paper's reported net Sharpe ratios or cost statistics. Full reproduction would still require the original position data, quote and liquidity histories, AUM and trade conventions, complete daily portfolio records, and a resolved choice between the paper's incompatible impact configurations.</p><p>The equation identifiers <code>eq_3_10_spread</code>, <code>eq_3_10_impact</code>, <code>eq_3_10_total_cost</code>, and <code>eq_3_10_turnover</code> are used here as semantic anchors because their canonical LaTeX fields are empty. No formula has been reconstructed or displayed. Local static verification and semantic code verification were skipped by policy, and no code execution or numerical verification is claimed.</p><h2>Optional Convex Optimization, Risk Models, Sector Neutrality, and Concentration</h2><p>When should portfolio weights be chosen by an optimizer rather than assigned directly from prediction scores? The intuition is to let the optimizer balance three competing goals: exposure to the model's signal, the cost of changing positions, and portfolio risk. This is an optional extension of the paper's naive score-proportional portfolio, not a replacement for it.</p><p>The paper describes a convex portfolio-optimization experiment with a standardized prediction vector <code>alpha</code>, portfolio weights <code>w</code>, previous weights <code>w_prev</code>, position costs <code>c</code>, a covariance or risk model <code>Sigma</code>, and nonnegative penalty parameters <code>lambda_tc</code> and <code>lambda_risk</code>. It also discusses long/short exposure, maximum positions, optional sector neutrality, and concentration diagnostics. However, the supplied optimization displays are OCR-damaged. Their equation records&#8212;<code>eq_5_4_optimization</code>, <code>eq_appendix_B_constraints</code>, <code>eq_appendix_B_diag_risk</code>, <code>eq_appendix_B_factor_risk</code>, <code>eq_appendix_B_sector_covariance</code>, <code>eq_appendix_B_sector_neutrality</code>, <code>eq_appendix_B_effective_n</code>, <code>eq_appendix_B_sector_tilts</code>, and <code>eq_appendix_B_naive_weights</code>&#8212;contain no canonical LaTeX.</p><p>Therefore, this section distinguishes three things:</p><ul><li><p><strong>Paper fact:</strong> optimization trades prediction exposure against turnover costs and risk, with neutrality and position constraints discussed as possible controls.</p></li><li><p><strong>Implementation decision:</strong> the generated code adopts a documented long-minus-short interpretation and an <code>alpha_minus_turnover_minus_risk</code> objective interpretation.</p></li><li><p><strong>Evidence limit:</strong> this implementation does not claim to have recovered the paper's exact damaged objective or constraints, and it was not executed under the run policy.</p></li></ul><h3>The optimization inputs and outputs</h3><p><code>solve_portfolio_optimization</code> accepts five inputs. <code>alpha</code> is a length-<code>n</code> standardized prediction vector, where <code>n</code> is the number of securities in the current cross-section. <code>previous_weights</code> and <code>costs</code> are also length-<code>n</code> vectors. <code>risk_inputs</code> is a mapping containing the inputs for either a diagonal or sector-factor risk model. <code>config</code> must expose an optimization configuration with explicit interpretations, penalty values, exposure limits, and solver selection.</p><p>The result is an <code>OptimizationResult</code> dataclass. Its <code>weights</code>, <code>w_plus</code>, and <code>w_minus</code> arrays all have shape <code>(n,)</code>. The first is the net portfolio, while the latter two are its nonnegative long and short decompositions. The result also stores solver <code>status</code>, objective components, diagnostics, and the complete selected configuration. This metadata matters because two apparently similar optimized portfolios may use different risk models, side exposures, position bounds, or interpretations of the damaged notation.</p><p>The following excerpt shows the public result structure. It is copied from <code>src/stock_selection/optimization/solver.py</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;273264a0-fa03-4cc9-84d4-0b6bba773b7d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class OptimizationResult:
    """Auditable result from the optional convex portfolio optimizer.

    Vectors have shape ``(n_securities,)`` and use the input security order.
    The optimization is intentionally explicit about its interpretation of the
    paper's OCR-damaged objective and constraints.
    """

    weights: np.ndarray
    w_plus: np.ndarray
    w_minus: np.ndarray
    status: str
    objective_value: float
    prediction_exposure: float
    transaction_cost_penalty: float
    risk_penalty: float
    diagnostics: dict[str, Any]
    configuration: dict[str, Any]</code></pre></div><p>The implementation requires the configuration fields <code>objective_interpretation</code> and <code>constraint_interpretation</code>. It accepts the exact labels <code>alpha_minus_turnover_minus_risk</code> and <code>separate_unit_sides_net_bound</code>; otherwise it raises an error instead of silently selecting a formula. This is deliberate: an explicit failure is more auditable than an undocumented repair to OCR-damaged mathematics.</p><h3>Long and short decomposition</h3><p>The generated constraint builder represents each net weight as the difference between two nonnegative vectors. <code>w_plus</code> contains long-side amounts and <code>w_minus</code> contains short-side amounts. The net vector is <code>w = w_plus - w_minus</code>, with separate side-exposure constraints. In the default interpretation, each populated side has <code>side_exposure=1.0</code>. Consequently, the total absolute gross exposure can be two even though each side has one unit of exposure; this should not be confused with a statement that total gross exposure is one.</p><p><code>max_position</code> bounds each net security position. The appendix method card gives <code>0.02</code> as an implementation default, but the generated solver treats the value as configuration rather than embedding it as an unchangeable paper constant. The constraint function also validates vector lengths and nonnegative finite scalar settings.</p><p>Here is the focused constraint construction from <code>src/stock_selection/optimization/constraints.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a22d431d-112b-4fa5-9358-078560d345a5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    # Eq. appendix_B_constraints: nonnegative long/short decomposition.
    constraints: list[Any] = [w_plus &gt;= 0, w_minus &gt;= 0]
    # Eq. appendix_B_constraints: net weights equal long exposure minus short exposure.
    constraints.append(w == w_plus - w_minus)
    # Eq. appendix_B_constraints: separate long and short exposure constraints.
    constraints.extend(
        [
            cp.sum(w_plus) == side_exposure_value,
            cp.sum(w_minus) == side_exposure_value,
        ]
    )
    # Eq. appendix_B_constraints: maximum absolute position bound per security.
    constraints.extend(
        [
            w &lt;= max_position_value,
            w &gt;= -max_position_value,
        ]
    )</code></pre></div><p>The <code>problem</code> argument of <code>add_long_short_constraints</code> is retained for compatibility with the planned solver interface, but constraints are returned as a list rather than attached to an already-created CVXPY problem. <code>add_sector_neutrality_constraints</code> then copies that list and appends one zero-net-exposure constraint for every distinct sector label.</p><h3>Prediction, turnover, and risk penalties</h3><p>The implementation's selected interpretation of <code>eq_5_4_optimization</code> maximizes prediction exposure while subtracting a turnover-cost penalty and a risk penalty. Here <code>alpha @ w</code> is the scalar prediction exposure. The auxiliary vector <code>turnover</code> has shape <code>(n,)</code> and is constrained to be at least both <code>w - w_prev</code> and its negative, which represents the absolute change in each position without using a nonconvex absolute-value expression directly. The cost vector <code>c</code> supplies a nonnegative per-position cost, and <code>lambda_tc</code> controls its penalty.</p><p>Risk is controlled by <code>lambda_risk</code> and one of two explicit models:</p><ul><li><p><strong>Diagonal risk:</strong> <code>volatility</code> is a length-<code>n</code> vector. The implementation penalizes squared volatility-scaled positions. This is the semantic mapping for <code>eq_appendix_B_diag_risk</code>.</p></li><li><p><strong>Sector-factor risk:</strong> <code>loadings</code> has shape <code>(n, k)</code>, <code>covariance</code> has shape <code>(k, k)</code>, and <code>idiosyncratic_variance</code> has shape <code>(n,)</code>. The loadings map security weights to <code>k</code> factor exposures, while the covariance prices factor risk and the idiosyncratic vector prices security-specific risk. This maps to <code>eq_appendix_B_factor_risk</code>.</p></li></ul><p>The risk functions accept NumPy arrays and, when CVXPY is installed, CVXPY-compatible expressions. They validate dimensions, finite values, nonnegative variances, covariance symmetry, and positive semidefiniteness. These checks protect the intended convex formulation, but they do not establish that the formulation is identical to the paper's damaged display.</p><p>The sector covariance helper makes the stated sector-volatility and correlation assumptions explicit. For <code>n_sectors</code> sectors, <code>sector_volatility</code> is a nonnegative scalar and <code>correlation</code> must lie in <code>[-1, 1]</code>; the resulting matrix has the corresponding variance on its diagonal and equicorrelated covariance off the diagonal. The paper context states a sector volatility of <code>0.02</code> and pairwise correlation of <code>0.30</code>, but the function accepts them as parameters so alternative documented configurations remain possible.</p><h3>Optional sector neutrality</h3><p>Sector neutrality is not automatically imposed. If <code>sector_neutral</code> is enabled, <code>risk_inputs</code> must contain one hashable sector label for every security. <code>add_sector_neutrality_constraints</code> groups positions by label and requires the sum of net weights within each group to be zero. Missing labels, mismatched lengths, and unhashable labels fail explicitly.</p><p>This is a portfolio constraint, not a feature transformation. It changes the feasible set of <code>w</code>; it does not alter <code>alpha</code>, the model predictions, or the earlier score-standardization step. The comparison should therefore use identical prediction, smoothing, rebalance, cost, and evaluation settings when contrasting a naive portfolio with an optimized or sector-neutral portfolio.</p><p>The supplied test file expresses the intended invariant using a small four-security example:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;eada6e7b-c77f-4a18-9c42-daa6f864dcbd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    w.value = np.array([0.15, -0.15, 0.25, -0.25])
    sector_violations = [np.asarray(constraint.violation, dtype=float) for constraint in constraints[7:]]

    assert len(sector_violations) == 2
    assert all(np.allclose(violation, 0.0) for violation in sector_violations)</code></pre></div><p>This excerpt comes from <code>tests/test_optimization_invariants.py</code>. It is a planned semantic check, not a reported execution result. Under the authoritative run policy, test execution and semantic verification were disabled.</p><h3>Risk construction and solver-status handling</h3><p>The solver creates CVXPY variables for <code>w</code>, <code>w_plus</code>, <code>w_minus</code>, and the auxiliary turnover vector. It selects the risk expression from the configured <code>risk_model</code>, adds the long/short and optional sector constraints, and asks for the explicitly selected <code>CLARABEL</code> solver. After the solve, it accepts only <code>optimal</code> or <code>optimal_inaccurate</code> statuses and rejects missing or non-finite returned vectors.</p><p>That status guard is important. A numerical optimizer's return value is not automatically a valid portfolio: the problem may be infeasible, unbounded, or unable to produce values. The generated implementation records the status before exposing weights to later evaluation. It also records prediction exposure, the weighted transaction-cost penalty, the weighted risk penalty, side exposures, net exposure, maximum absolute position, and sector tilts.</p><p>The command-line robustness script delegates optimization to this configured solver rather than constructing hidden defaults:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a786f99a-9ee5-4d3b-9f6c-9b895bfce65d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    # Eq. 5.4 and Appendix B are intentionally delegated to the explicitly
    # configured solver; OCR-damaged interpretations must not be inferred here.
    result = solve_portfolio_optimization(
        payload["alpha"],
        payload["previous_weights"],
        payload["costs"],
        payload["risk_inputs"],
        payload["config"],
    )</code></pre></div><p>This excerpt is copied from <code>scripts/run_robustness.py</code>. The input pickle must supply <code>alpha</code>, <code>previous_weights</code>, <code>costs</code>, <code>risk_inputs</code>, and <code>config</code>; the script does not contact an external service or claim that CLARABEL was run.</p><h3>Naive weights and concentration diagnostics</h3><p>The optimizer should be compared with the paper's naive prediction-proportional portfolio, represented by <code>eq_appendix_B_naive_weights</code>. The naive constructor uses separate positive and negative score totals, matching the long/short logic described earlier in this tutorial. The optimized portfolio introduces additional constraints and penalties; it is not simply another score normalization.</p><p>The diagnostics module reports three useful quantities:</p><ul><li><p><code>effective_n</code> measures breadth. Because the appendix display is OCR-damaged but its prose defines effective breadth as inverse Herfindahl concentration, <code>compute_effective_n</code> uses the documented interpretation <code>1 / sum of squared weights</code>. It does not renormalize the input weights, so the caller's exposure convention is preserved.</p></li><li><p><code>maximum_abs_position</code> is the largest absolute security weight.</p></li><li><p><code>sector_tilts</code> is the sum of absolute net sector exposures, corresponding to <code>eq_appendix_B_sector_tilts</code>.</p></li></ul><p>The inverse-HHI choice is an implementation interpretation, not recovered canonical equation text. It should be recorded in metadata, as <code>summarize_concentration</code> does:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b24f28b7-c57a-477d-9ce6-858666548388&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    return MetricResult(
        name="portfolio_concentration",
        values={
            "effective_n": effective_n,
            "maximum_abs_position": max_position,
            "sector_tilts": sector_tilts,
        },
        metadata={
            "effective_n_convention": "inverse_hhi_sum_squared_weights",
            "sector_tilts_convention": "sum_absolute_net_sector_exposures",
            "weight_column": _WEIGHT_COLUMN
            if _WEIGHT_COLUMN in weights.columns
            else "sole_numeric_column",
            "sector_column": _DEFAULT_SECTOR_COLUMN,
            "weights_renormalized": False,
        },
    )</code></pre></div><p>This focused excerpt is copied from <code>src/stock_selection/optimization/diagnostics.py</code>. <code>compute_effective_n</code>, <code>compute_sector_tilts</code>, and <code>summarize_concentration</code> validate numeric weights and aligned sector labels before returning scalar diagnostics in a <code>MetricResult</code>.</p><h3>Worked example: a four-security portfolio</h3><p>Consider four securities with a hypothetical net-weight vector <code>[0.15, -0.15, 0.25, -0.25]</code>. Assign the first two to <code>technology</code> and the last two to <code>utilities</code>. A compatible long/short decomposition could use positive long amounts <code>[0.15, 0, 0.25, 0]</code> and short amounts <code>[0, 0.15, 0, 0.25]</code>. The net vector is the long vector minus the short vector. Each sector sums to zero, so the optional sector-neutrality constraints are satisfied for this example.</p><p>The example is only a shape and invariant demonstration. It does not specify <code>alpha</code>, <code>lambda_tc</code>, <code>lambda_risk</code>, a covariance matrix, previous positions, or a solver result. Consequently, it cannot establish which portfolio an optimizer would select. A real run would need all of those inputs, an explicit optimization configuration, and a usable solver status.</p><p>For the same vector, <code>compute_effective_n</code> applies the documented inverse-HHI convention without rescaling the signed weights. <code>compute_sector_tilts</code> groups the weights by sector, sums each group's net exposure, takes absolute values, and adds them. In this balanced example, the sector-tilt diagnostic is zero. These diagnostics describe concentration and neutrality; they do not measure predictive performance.</p><h3>Implementation boundary</h3><p>The optional optimization module is useful for testing portfolio-design hypotheses, but it is the least directly recoverable part of the supplied paper evidence. The generated files make the unresolved choices visible through <code>objective_interpretation</code>, <code>constraint_interpretation</code>, <code>risk_model</code>, penalty values, exposure settings, position bounds, sector-neutrality flags, and solver selection. They also preserve the distinction between naive and optimized portfolios.</p><p>No optimization equation is displayed here because every supplied optimization equation record has an empty canonical LaTeX field. Reconstructing a formula from the semantic summary would violate the evidence contract. The package therefore provides a documented implementation interpretation, not proof of exact mathematical recovery.</p><p>The supplied local static-verification record says verification was skipped, and semantic code verification was also skipped. No code execution, solver execution, test execution, numerical verification, or reproduction of the paper's reported results is claimed.</p><h2>Size, Volatility, Lag, Feature Importance, and Aggregate Robustness Workflows</h2><p>How can you tell whether a portfolio result is broadly supported rather than driven by one company-size group, one market regime, or one timing choice? Robustness analysis reuses the frozen signal and portfolio pipeline under controlled partitions and delays. It changes the question being measured, not the model that was selected.</p><p>The paper describes robustness checks by market capitalization, volatility regime, execution lag, feature importance, and portfolio concentration. The implementation therefore keeps partition labels, lag labels, feature-importance types, and coverage metadata in the outputs. These analyses are descriptive: they must not feed back into feature selection or out-of-sample hyperparameter tuning.</p><h3>What is fixed, and what varies?</h3><p>The core experiment should already have produced predictions, weights, realized returns, and fitted models. A robustness run reuses those artifacts while varying one declared dimension:</p><ul><li><p><strong>Size:</strong> assign each date's eligible securities to small, mid, and large market-cap groups.</p></li><li><p><strong>Volatility:</strong> assign observations to configured volatility regimes. The generated implementation uses a global quantile policy over the supplied frame, so the caller must restrict the frame or provide a different point-in-time-safe policy when future information would otherwise enter the breakpoints.</p></li><li><p><strong>Execution lag:</strong> compare explicit lags such as 0, 1, 3, 5, and 10. Lag 0 is retained as a reference but labeled non-implementable.</p></li><li><p><strong>Feature importance:</strong> summarize importance values from already fitted LightGBM models. This helps interpret a model; it is not evidence of incremental predictive validity.</p></li><li><p><strong>Concentration:</strong> report effective breadth, maximum absolute position, and sector tilts alongside performance metrics.</p></li></ul><p>The method card <code>portfolio_concentration_and_robustness</code> describes this reuse principle. The supplied equation records <code>eq_3_4_sharpe</code>, <code>eq_appendix_B_effective_n</code>, and <code>eq_appendix_B_sector_tilts</code> are semantic anchors only: their canonical LaTeX fields are empty, so no formula is reconstructed here. In code, <code>annualized_sharpe</code>, <code>compute_effective_n</code>, and <code>compute_sector_tilts</code> are the corresponding mappings.</p><h3>Daily market-cap partitions</h3><p><code>assign_market_cap_terciles</code> preserves the input row alignment and computes size groups separately for each date. Within a date, it stably sorts valid market-cap values and divides the ordered observations into three rank bands. Stable sorting makes tied values deterministic without using information from another date.</p><p>The following excerpt is copied from <code>src/stock_selection/robustness/partitions.py</code>. Notice that the date groups are obtained before sorting and that missing or invalid market-cap values are not assigned a label.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;82886e10-15c7-401e-b83a-3540005372d8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    result = np.full(len(frame), np.nan, dtype=object)
    labels = np.array(["small", "mid", "large"], dtype=object)

    for positions in _position_groups(frame, date_column):
        valid_positions = positions[finite[positions]]
        if valid_positions.size == 0:
            continue

        # Stable sorting makes tied market caps deterministic without using
        # information outside the current date cross-section.
        ordered = valid_positions[
            np.argsort(market_cap[valid_positions], kind="mergesort")
        ]
        n = ordered.size
        ranks = np.arange(n, dtype=float)
        band_numbers = np.floor(3.0 * ranks / n).astype(int)
        band_numbers = np.minimum(band_numbers, 2)
        result[ordered] = labels[band_numbers]

    return pd.Series(result, index=frame.index, name="market_cap_tercile")</code></pre></div><p>This is a date-local breakpoint policy, which is important for point-in-time analysis. It does not guarantee that every date has observations in all three groups: a date with fewer than three valid securities can legitimately have an empty band. The returned <code>Series</code> has the same index as the input <code>DataFrame</code>, allowing the labels to be attached without changing row identity.</p><h3>Volatility regimes are an explicit policy</h3><p>The paper mentions volatility-regime robustness, but the supplied evidence does not uniquely specify the volatility variable, quantile frequency, or breakpoint timing. <code>assign_volatility_regimes</code> therefore accepts the volatility column and quantile probabilities as arguments. In the generated implementation, breakpoints are calculated globally over the supplied frame. That is an implementation decision, not a recovered paper rule.</p><p>The function documents the limitation directly:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ba283c23-8b1b-47fc-8df5-6cf0a09f670e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    This function has no date argument, so its breakpoint policy is explicitly
    global over the provided frame rather than daily.  Callers requiring a
    point-in-time or date-local policy must invoke it on a time-restricted
    frame or implement the corresponding grouped workflow.  Quantile values
    are probabilities in ``[0, 1]``; regime labels are ``regime_0`` through
    ``regime_n`` in ascending volatility order.  Missing values remain missing.</code></pre></div><p>For a genuine out-of-sample robustness run, compute or supply regime breakpoints using only the allowed information. For example, a caller could apply this function separately to a historically restricted frame, or implement a grouped expanding-breakpoint procedure. Do not treat the global default as automatically safe for a full 2015&#8211;2024 panel.</p><h3>Partitioned evaluation without retuning</h3><p><code>run_partitioned_evaluation</code> accepts a prepared data frame, a same-index partition <code>Series</code>, and an evaluator callback. For each nonmissing label, it passes a copy of the corresponding subset to the callback and records the label and row count with the returned <code>MetricResult</code> values. The evaluator may accept either the subset and its label or only the subset.</p><p>The key invariant is that this function does not fit a new model or select new parameters. It evaluates each partition using the caller's existing core configuration. The output retains the partition label, so a result such as a Sharpe value is interpretable as &#8220;the value for this named segment,&#8221; rather than an unlabeled aggregate.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2e40c414-9da6-4737-8004-f7df5a75f8af&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def run_partitioned_evaluation(
    data: pd.DataFrame,
    partition: pd.Series,
    evaluator: Callable[..., MetricResult],
) -&gt; pd.DataFrame:
    """Evaluate fixed core settings separately for each supplied partition.

    ``partition`` must align to ``data`` by index.  The evaluator receives a
    copy of each non-missing partition and may accept either the subset alone
    or the subset plus its partition label.  Robustness partitions are
    descriptive only: this function never fits or retunes a model.
    """</code></pre></div><p>The function requires the partition to have the same length and index as <code>data</code>. That shape check matters: a positionally shifted label could silently assign a security to the wrong size or volatility group. Empty or missing partitions are skipped rather than converted into an implicit category.</p><h4>Worked example: size and lag labels</h4><p>A local synthetic demonstration can create a small panel with <code>CWIQ Code</code>, <code>date</code>, and <code>market_cap</code>, then pass the panel's market-cap column and date column to <code>assign_market_cap_terciles</code>. The resulting <code>market_cap_tercile</code> series can be attached to the panel and supplied to <code>run_partitioned_evaluation</code> together with an evaluator that uses the already-created predictions and returns.</p><p>The same frozen predictions and realized returns can then be sent to <code>run_execution_lag_sensitivity</code> with the explicit lag list <code>[0, 1, 3, 5, 10]</code>. The expected artifact is a table labeled by <code>execution_lag</code>; it should also retain the boolean <code>lag_0_non_implementable</code> marker. No numerical output is asserted here, because the code was not executed and synthetic data are not paper observations.</p><h3>Execution-lag sensitivity</h3><p><code>run_execution_lag_sensitivity</code> is designed for alpha-decay comparisons. It validates that lags are nonnegative, unique integers and delegates each lag to the internal lagged evaluator. The generated helper aligns returns by positional observation within each security. This is an implementation choice because the supplied paper context does not formalize whether an execution lag means calendar days, trading observations, or another date convention.</p><p>The output should therefore be read with its configuration metadata. A lag of 0 is useful as a timing reference, but the paper's implementable headline setting is a positive delay. Keep standard <code>t</code>-to-<code>t+1</code> evaluation, delayed <code>t+1</code>-to-<code>t+2</code> evaluation, and the broader lag-sensitivity table as distinct experiment variants.</p><h3>Feature importance and concentration diagnostics</h3><p>Feature importance summarizes what fitted models used internally. <code>summarize_feature_importance</code> supports models exposing LightGBM booster importance and records both <code>split</code> and <code>gain</code> importance types when available. The resulting table contains <code>feature</code>, <code>importance_type</code>, <code>mean_importance</code>, and <code>model_count</code>. Aggregating across models is useful only when feature names and model roles are comparable; it does not replace out-of-sample evaluation.</p><p>Concentration diagnostics answer a different question: how broadly is the portfolio spread? <code>compute_effective_n</code> uses the inverse-Herfindahl interpretation stated in the paper's prose for <code>eq_appendix_B_effective_n</code>. Because the displayed equation is OCR-damaged, this is explicitly a documented interpretation rather than recovered equation text. The function does not renormalize weights, so the caller's exposure convention&#8212;such as one unit on each long and short side&#8212;is preserved.</p><p><code>compute_sector_tilts</code> groups weights by sector, sums net exposure within each group, takes the absolute value of each sector total, and adds those absolute exposures. The function requires one nonmissing sector label per position. <code>summarize_concentration</code> combines effective breadth, maximum absolute position, and sector tilts in a <code>MetricResult</code> with convention metadata.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;bff315a2-9e54-44c1-8030-a058deb2e980&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def compute_effective_n(weights: pd.Series) -&gt; float:
    """Compute inverse-HHI portfolio breadth.

    The paper's appendix display is OCR-damaged.  The prose explicitly describes
    effective breadth as the inverse of the Herfindahl concentration measure, so
    this function uses ``1 / sum(weights**2)`` as a documented interpretation,
    rather than recovered equation LaTeX.
    """
    values = _as_finite_weights(weights)
    # Eq. appendix_B_effective_n: documented inverse-HHI interpretation.
    concentration = float(np.square(values.to_numpy(dtype=float)).sum())
    if concentration == 0.0:
        return 0.0
    return 1.0 / concentration</code></pre></div><p>The zero-weight result is an implementation convention. It avoids an undefined diagnostic but should be retained in metadata if such a portfolio occurs. Compare concentration metrics under the same return, smoothing, rebalance, and cost settings; otherwise a difference in breadth can be confused with a difference in evaluation protocol.</p><h3>Aggregating experiment variants</h3><p>The paper reports composites across multiple experiments. <code>aggregate_experiment_metrics</code> converts <code>MetricResult</code> objects into one row per experiment and preserves metric metadata under <code>metadata_</code> columns. It does not automatically reweight or average results, because coverage and missing-date treatment are not fully specified.</p><p>For prediction variants, <code>average_experiment_scores</code> uses the intersection of security-date keys, standardizes each input table within date, and averages the aligned standardized values. It deliberately does not forward-fill or apply implicit coverage weighting. This is an implementation policy that should be recorded when comparing structured, DSPy, baseline, combined, or ensemble variants.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;bb908bba-a2d6-46d9-bb7a-29beb20fb8c4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">    merged = prepared[0].rename(columns={"__standardized__": "standardized_prediction_0"})
    for index, table in enumerate(prepared[1:], start=1):
        merged = merged.merge(
            table.rename(columns={"__standardized__": f"standardized_prediction_{index}"}),
            on=[SECURITY_KEY, "__aggregate_date__"],
            how="inner",
            validate="one_to_one",
        )

    score_columns = [f"standardized_prediction_{index}" for index in range(len(prepared))]
    merged = merged.dropna(subset=score_columns).sort_values(
        ["__aggregate_date__", SECURITY_KEY], kind="mergesort"
    )
    merged["experiment_count"] = len(score_columns)
    merged["average_prediction"] = merged[score_columns].mean(axis=1)</code></pre></div><p>The <code>experiment_count</code> column records how many variants contributed to each output row. The function's attributes also retain input coverage and the missing-data policy. This makes the aggregation auditable instead of presenting a composite as though every experiment covered identical securities and dates.</p><h3>Running the local robustness command</h3><p><code>scripts/run_robustness.py</code> provides a local command-line entry point for smoothing, lag, turnover, cost, and optional optimization artifacts. It keeps configurations such as daily versus monthly rebalancing and impact coefficients <code>k=0.2</code> versus <code>k=0.3</code> separately labeled. It reads local tables or a local optimization input pickle and does not contact external services.</p><p>A practical run policy is to create one output directory per configuration matrix and record:</p><ul><li><p>the source artifact paths and date coverage;</p></li><li><p>the partition rule and breakpoint data used for size or volatility groups;</p></li><li><p>the execution-lag convention;</p></li><li><p>smoothing and rebalance frequency;</p></li><li><p>cost formulation and impact coefficient;</p></li><li><p>model and feature-registry identifiers;</p></li><li><p>feature-importance type; and</p></li><li><p>concentration conventions for effective breadth and sector tilts.</p></li></ul><p>These records allow a reader to distinguish a genuine robustness comparison from a collection of results produced under silently different assumptions.</p><h3>Limits and verification boundary</h3><p>The robustness architecture is supported by the generated interfaces, but the supplied verification stages were skipped by policy. No code execution, test execution, static verification, semantic code verification, or numerical comparison was performed. Accordingly, this section describes intended responsibilities and explicit implementation decisions, not verified outputs.</p><p>The most important safeguards are procedural: compute size groups using date-local information, make volatility breakpoints point-in-time safe, preserve partition and lag labels, reuse frozen models rather than retuning them, and report concentration beside performance. Robustness analysis should explain how a result changes&#8212;not quietly search for a more favorable result.</p><h2>Synthetic Tutorial Path, Audit Metadata, and Verification Boundary</h2><p>What can you learn from a local run when the original vendor data and language-model services are unavailable? You can trace the mechanics of the reproduction safely: create a keyed panel, apply timing and universe policies, preserve configuration metadata, and identify the artifacts that later feature generation and modeling would consume. You cannot use that demonstration to claim that the paper's features, Sharpe ratios, or other empirical results were reproduced.</p><p>This final section connects the local synthetic-data path to the evidence boundary. The paper facts are the chronological workflow and its timing requirements. The implementation decisions are the deterministic synthetic fixtures, local-only command-line interface, artifact recording, and explicit failure when required external components are absent. No execution or numerical verification was performed under the authoritative run policy.</p><h3>A local data path, not a paper-data substitute</h3><p><code>SyntheticDataConfig</code> and <code>make_synthetic_sources</code> in <code>src/stock_selection/data/synthetic.py</code> create artificial EDI-like, TrueBeats-like, SpiderRock-like, universe, and factor-like tables. Each source uses <code>CWIQ Code</code> and <code>date</code> keys, stable ordering, and deterministic values from a configured seed. The generated file explicitly describes these values as tutorial demonstrations rather than vendor data.</p><p>A focused excerpt shows the intended boundary:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;92993229-cbfd-49ef-b12a-bc1ddc2f16e4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class SyntheticDataConfig:
    """Configuration for deterministic tutorial-only panel data.

    The generated observations are artificial and are not vendor data from the
    paper.  All source tables use ``CWIQ Code`` and ``date`` as their key
    columns and are sorted by date followed by security identifier.
    """

    start_date: DateLike = "2015-01-01"
    end_date: DateLike = "2019-12-31"
    n_securities: int = 50
    n_edi_features: int = 2
    n_truebeats_features: int = 3
    n_spiderrock_features: int = 3
    seed: int = 7</code></pre></div><p>The dates, feature counts, and seed are demonstration controls. They do not reproduce the paper's vendor coverage or generated-feature corpus. <code>make_synthetic_sources(config)</code> returns a mapping containing <code>edi</code>, <code>truebeats</code>, <code>spiderrock</code>, <code>universe</code>, and <code>factors</code>; the factors are also synthetic and should not be used as the paper's factor data.</p><p>The synthetic source builder preserves useful structural properties. It creates unique security-date keys, generates adjusted prices and volume-like fields, supplies market capitalization and security-type metadata, and sorts the source rows. Those properties make it suitable for illustrating key alignment and timing checks. They do not establish that the corresponding production data policies&#8212;corporate-action treatment, delisting handling, or vendor release timing&#8212;match the paper.</p><h3>Worked example: invoking the local preparation path</h3><p>The generated command-line entry point is <code>scripts/run_reproduction.py</code>. Its <code>--synthetic</code> mode constructs the local fixtures, creates an explicit experiment configuration, and calls <code>ReproductionPipeline.prepare_data</code>. A minimal invocation is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;14575cfc-c32e-4e1c-92f7-50608a733971&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">python scripts/run_reproduction.py --synthetic</code></pre></div><p>This command is an implementation example, not an execution result. Under the run policy, it was not run. The script's description makes the stopping point explicit:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e096c3bc-1179-4829-b8b3-27c3e74c2c9a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">parser = argparse.ArgumentParser(
    description=(
        "Prepare local or deterministic synthetic stock-selection data. "
        "This entry point performs no network access and does not claim "
        "numerical reproduction of the paper's reported results."
    )
)</code></pre></div><p>The preparation flow is deliberately narrower than the full paper workflow. Its <code>main</code> function constructs configuration, loads either synthetic or user-supplied local sources, prepares the panel, and writes metadata:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b8568414-1015-4917-878f-fcd64518fefb&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">config = _make_experiment_config(args)
sources = _load_sources(args)
pipeline = ReproductionPipeline(config)
panel = pipeline.prepare_data(sources)
payload = {
    "status": "prepared_local_data",
    "note": (
        "This command prepares data through the local pipeline only. "
        "It does not call external APIs, run feature generation, fit a model, "
        "or claim numerical verification."
    ),
    "source_names": sorted(sources),
    "rows": int(len(panel)),
    "securities": int(panel["CWIQ Code"].nunique()),
    "dates": int(panel["date"].nunique()),
    "configuration": pipeline.artifacts.get("configuration", {}),
    "artifacts": pipeline.artifacts,
}</code></pre></div><p>Readers should notice what is absent: this path does not silently create a GPT-4.1 provider, retrieve vendor documentation, discover features, fit LightGBM, or report portfolio metrics. Those stages require caller-supplied components. In particular, <code>ReproductionPipeline.run_discovery</code> requires a feature generator and raises an error when one is not configured, because the supplied evidence does not define a default provider.</p><h3>What the pipeline records</h3><p><code>ReproductionPipeline</code> retains intermediate artifacts rather than returning only a final number. Its configuration snapshot records the timing, universe, and other experiment choices. Each stage is appended to the <code>stages</code> collection with metadata such as row counts, security counts, target choice, implementation lag, and accepted feature names.</p><p>For a completed feature-discovery stage, the important audit chain would be:</p><ol><li><p><strong>Source identity and schema.</strong> Record the local file or dataset version, required columns, identifier column, date column, and source availability policy.</p></li><li><p><strong>Timing configuration.</strong> Preserve <code>signal_date</code>, <code>execution_date</code>, <code>return_date</code>, target horizon, and implementation lag. Standard <code>t</code> to <code>t+1</code> evaluation and the headline delayed <code>t+1</code> to <code>t+2</code> evaluation must remain separate configurations.</p></li><li><p><strong>Feature provenance.</strong> For every <code>GeneratedFeature</code>, retain the route, rationale, required columns, retrieval-context version, generated-code identity, and validation report.</p></li><li><p><strong>Registry state.</strong> Record which candidates were rejected, which were accepted, and when the <code>FeatureRegistry</code> was frozen. Out-of-sample materialization must use the frozen definitions.</p></li><li><p><strong>Model configuration.</strong> Preserve feature configuration&#8212;baseline, AI-only, combined, or ensemble&#8212;along with standardization policy, walk-forward specification, and LightGBM parameters selected during discovery.</p></li><li><p><strong>Evaluation configuration.</strong> Record score clipping, side exposure, return alignment, smoothing window, rebalance frequency, cost model, impact coefficient, inference settings, and any optimization interpretation.</p></li></ol><p><code>AuditRecord</code> and <code>MetricResult</code> are the shared typed containers planned for this purpose. They are metadata contracts, not evidence that a model or metric was successfully computed in this run.</p><h3>Replacing synthetic sources with local files</h3><p>The same command-line entry point accepts repeated <code>NAME=PATH</code> arguments instead of <code>--synthetic</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;02fa2225-05ba-41e9-8d96-30ef880695ec&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">python scripts/run_reproduction.py \
  --source edi=data/edi.parquet \
  --source truebeats=data/truebeats.parquet \
  --source spiderrock=data/spiderrock.parquet \
  --source universe=data/universe.parquet</code></pre></div><p>These paths are illustrative inputs to the local loader interface; they are not supplied by the paper context. <code>load_source_bundle</code> reads local tables and validates the canonical key columns through user-supplied <code>DatasetSchema</code> objects. The loader does not contact a vendor or infer missing schema definitions.</p><p>Before replacing the synthetic inputs, confirm that the local files define the fields required by the selected pipeline stages. At minimum, the preparation path needs consistent <code>CWIQ Code</code> and <code>date</code> keys. Target construction additionally needs a configured price field and any applicable delisting-return information. Universe selection may require security type, U.S. common-equity eligibility, market capitalization, ADV, spread, and related metadata. Feature generation requires the permitted vendor columns and their point-in-time release information.</p><h3>The verification boundary</h3><p>The repository contains tests such as <code>test_forward_target_is_after_signal_date</code>, <code>test_feature_returns_one_aligned_series</code>, <code>test_weights_have_zero_net_exposure_and_unit_sides</code>, and <code>test_cost_models_preserve_coefficient_separation</code>. They express intended semantic checks for timing, alignment, portfolio invariants, and the separate <code>k=0.2</code> and <code>k=0.3</code> cost configurations. The optimization tests similarly preserve explicit interpretations for long/short decomposition, sector neutrality, and inverse-HHI effective breadth.</p><p>Those tests must not be described as executed results. The authoritative run policy disabled code execution, test generation, local static verification, tutorial-section verification, semantic code verification, and final quality review. The supplied verification records therefore say that local static verification was skipped and semantic code verification was skipped. No claim is made that the generated files import, run, pass tests, or reproduce numerical outputs.</p><p>A useful distinction is:</p><ul><li><p><strong>Planned static review</strong> would inspect syntax, imports, dependency ordering, key columns, shapes, and configuration branches without running the experiment.</p></li><li><p><strong>Planned semantic review</strong> would inspect whether functions preserve timing, alignment, neutrality, clipping, cost-label, and optimization invariants.</p></li><li><p><strong>Execution verification</strong> would run the package and tests on available inputs.</p></li><li><p><strong>Numerical reproduction</strong> would compare outputs against the paper's reported results using the original or equivalent data, prompts, generated features, and resolved parameters.</p></li></ul><p>Only the first two are described as intended review strategies here, and even those review stages were skipped in this run. The latter two did not occur.</p><h3>Final replacement checklist</h3><p>Before treating a local-data run as a reproduction attempt, check the following and retain the evidence in the audit manifest:</p><ul><li><p><strong>Keys and timing:</strong> every panel, feature, prediction, weight, and realized-return table preserves <code>CWIQ Code</code> and date keys; release dates do not exceed <code>signal_date</code>; execution and return dates are later than the signal date for implementable variants.</p></li><li><p><strong>Splits:</strong> January 2015 through December 2018 is used for discovery, validation, and hyperparameter selection; January 2019 through December 2024 is reserved for headline out-of-sample reporting.</p></li><li><p><strong>Feature freezing:</strong> generated candidates carry route, retrieval-context version, rationale, required columns, and validation status; the registry is frozen before OOS materialization.</p></li><li><p><strong>Model chronology:</strong> each walk-forward fit uses only permitted historical rows, preserves feature order, and does not tune parameters on OOS observations.</p></li><li><p><strong>Portfolio invariants:</strong> scores are date-wise standardized and clipped to <code>[-3, 3]</code>; populated long and short sides use the configured exposure; net exposure is zero when both sides are available.</p></li><li><p><strong>Evaluation labels:</strong> standard, delayed, and non-implementable lag-0 results are labeled separately; IC is date-wise; Sharpe uses the declared annualization and dispersion conventions.</p></li><li><p><strong>Trading-friction labels:</strong> smoothing and monthly rebalancing are distinct variants; turnover includes entries and exits; static and position-level costs are separate; impact coefficient <code>k=0.2</code> is not merged with <code>k=0.3</code>.</p></li><li><p><strong>Optimization status:</strong> solver status, objective interpretation, constraints, risk model, and OCR-related decisions are retained before optimized weights are used.</p></li><li><p><strong>Evidence status:</strong> synthetic values, local provider adapters, and supplied tests are clearly marked as demonstrations or planned checks rather than paper results.</p></li></ul><h3>What full numerical reproduction would still require</h3><p>A numerical reproduction would need evidence not included in the supplied context: the vendor data and complete schemas; point-in-time release records; the exact GPT-4.1 prompts and generated feature corpus; retrieval documents, indexing, and the corrupted-document transformation; the DSPy signature, demonstrations, MIPRO settings, and feedback scores; exact LightGBM parameters and retraining rules; factor data and alignment conventions; and resolved transaction-cost and optimization interpretations.</p><p>Until those inputs and decisions are supplied, the generated package should be described accurately as an auditable reproduction architecture with a synthetic local path. It teaches how the pieces connect and where leakage or ambiguity must be controlled, but it does not establish the paper's empirical findings.</p><p>Use the button or URL below to download the source code.</p>
      <p>
          <a href="/__u/onepagecode.substack.com/p/quant-trading-ai-feature-engineering">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Quant Research: Do Vision-Language Models (VLMs) Truly Read Candlestick Charts?]]></title><description><![CDATA[Benchmarking multi-modal VLMs against XGBoost baselines for 30-day stock return prediction from OHLCV charts.]]></description><link>https://onepagecode.substack.com/p/quant-research-do-vision-language</link><guid isPermaLink="false">https://onepagecode.substack.com/p/quant-research-do-vision-language</guid><dc:creator><![CDATA[Onepagecode]]></dc:creator><pubDate>Fri, 24 Jul 2026 20:09:09 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1439434768192-c60615c1b3c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ4NTE1NTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Use the URL at the end of this article to download entire source code</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1439434768192-c60615c1b3c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ4NTE1NTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1439434768192-c60615c1b3c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ4NTE1NTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1439434768192-c60615c1b3c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ4NTE1NTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1439434768192-c60615c1b3c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ4NTE1NTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1439434768192-c60615c1b3c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ4NTE1NTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1439434768192-c60615c1b3c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ4NTE1NTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="3888" height="2592" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1439434768192-c60615c1b3c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ4NTE1NTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2592,&quot;width&quot;:3888,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;bull grayscale photo&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="bull grayscale photo" title="bull grayscale photo" srcset="https://images.unsplash.com/photo-1439434768192-c60615c1b3c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ4NTE1NTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1439434768192-c60615c1b3c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ4NTE1NTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1439434768192-c60615c1b3c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ4NTE1NTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1439434768192-c60615c1b3c8?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw2fHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ4NTE1NTh8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@eiskonen">Hans Eiskonen</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><p>The paper introduces a benchmark for evaluating whether vision-language models can use daily and weekly candlestick charts to predict 30-day stock returns. It constructs aligned visual and numerical datasets from HS300 and S&amp;P 500 OHLCV data, evaluates commercial VLMs with confusion-matrix diagnostics and IC/Rank IC time-series metrics, and compares them with an XGBoost numerical baseline. The reported results indicate that VLMs are generally weak in ordinary market conditions and show stronger sensitivity to short-term outcomes than to the explicitly requested 30-day horizon.</p><h2>Implementation Assumptions</h2><ul><li><p>Python 3.11 is used with type annotations and dataclasses.</p></li><li><p>The implementation is local and provider-agnostic; no commercial VLM API client, credentials, network call, or undocumented request schema is implemented.</p></li><li><p>Input OHLCV data are supplied by local CSV or DataFrame adapters; provider-specific acquisition is outside the reproduction core.</p></li><li><p>The paper's canonical Eq. (1) is implemented for forward returns; OCR-damaged equations eq<em>2 through eq</em>7 are represented by named metric functions without claiming recovered LaTeX.</p></li><li><p>Unresolved choices such as directional threshold, zero-return treatment, bias normalization, IC grouping, significance testing, weekly boundary, chart dimensions, XGBoost features, and hyperparameters are represented by explicit configuration objects.</p></li><li><p>The default chart selection excludes records on or after the cutoff date and keeps at most 50 candles, as stated in the paper.</p></li><li><p>The primary label horizon is 30 subsequent trading-day positions; horizon 5 is available for sensitivity analysis.</p></li><li><p>The package supports synthetic local demonstration data because the original constituent lists and complete raw datasets are not supplied.</p></li><li><p>No code execution, test execution, semantic verification, or result verification is claimed under the run policy.</p></li></ul><h2>Scope, evidence boundaries, and repository map</h2><p>What would it take to ask a vision-language model whether it can read a stock chart without accidentally giving it future information? The benchmark's central idea is to present two views of the same historical situation: a detailed daily candlestick chart and a compressed weekly candlestick chart. Both views use the same stock and cutoff date, and the model produces a continuous estimate for a future return.</p><p>A <strong>candlestick chart</strong> summarizes price movement over an interval. Its open, high, low, and close describe the price range and direction, while volume records traded activity. Together, these fields are commonly called <strong>OHLCV</strong>: open, high, low, close, and volume. In this reproduction scaffold, the chart also contains moving-average overlays and a volume panel.</p><p>The <strong>cutoff date</strong> identifies the sample's present point in time. It determines which observations may be used to construct the input and from which point the future target is measured. <strong>Future leakage</strong> means allowing an observation at or after that cutoff into an input that is supposed to represent the past. The most important invariant in the repository is therefore that chart history and numerical-baseline history contain only observations strictly before the cutoff, even though the target itself is computed from a later closing price.</p><h3>What comes from the paper</h3><p>The supplied paper context directly supports several parts of the benchmark design:</p><ul><li><p>Samples pair daily and weekly candlestick charts for one stock and cutoff date.</p></li><li><p>The underlying records are OHLCV data from HS300 and S&amp;P 500 stock universes.</p></li><li><p>Charts include MA5, MA20, and MA90 overlays and volume information.</p></li><li><p>The primary target is a 30-trading-day forward return.</p></li><li><p>The VLM prompt asks about candle bodies and wicks, moving averages, volume, inflection points, and consistency between daily and weekly views.</p></li><li><p>Evaluation includes directional confusion statistics, prediction-bias analysis, IC, Rank IC, ICIR, significance summaries, market-regime analysis, extreme-stock analysis, and alternate-horizon analysis.</p></li></ul><p>Here, <strong>VLM</strong> means vision-language model: a model that accepts visual inputs together with text instructions and returns a textual response. The benchmark is visual-first, but it is not text-free. The model receives a prompt describing the task and chart concepts, including moving-average color semantics. Those instructions are part of the input contract and should not be described as independently supervised visual labels.</p><p>The paper reports dataset totals and experimental tables, including a reported total of 193,524 chart samples. This scaffold treats those numbers as paper claims and reference metadata, not as independently verified outputs. The exact constituent snapshots, the 32 cutoff dates, and the precise test-window conventions are not supplied here.</p><h3>What this repository decides locally</h3><p>A full reproduction needs more detail than the paper context provides. The package therefore exposes missing choices through configuration rather than silently presenting them as recovered facts. Important examples include:</p><ul><li><p>weekly period boundaries and incomplete-week handling;</p></li><li><p>chart dimensions, fonts, axis formatting, and image-compression settings;</p></li><li><p>the feature vector, training split, hyperparameters, and selection procedure for the XGBoost numerical baseline;</p></li><li><p>the threshold that converts continuous predictions and actual returns into rise or fall classes;</p></li><li><p>treatment of zero returns;</p></li><li><p>the normalization used for prediction bias;</p></li><li><p>the grouping, missing-value policy, standard-deviation convention, and significance test for IC and Rank IC.</p></li></ul><p>This distinction matters for run-policy readers. A configurable default is an implementation decision, not evidence that the paper used that value. The scaffold is designed to make these decisions visible in manifests and evaluation metadata so that a later reproduction can replace them when the missing source details become available.</p><h3>How the package is organized</h3><p>The shared data contracts live in <code>src/vlm_candlestick/types.py</code>. <code>OHLCVRecord</code> represents one scalar dated OHLCV observation. <code>SampleKey</code> identifies a market, stock, and cutoff date. <code>ForwardReturnLabel</code> stores the aligned current close, future close, horizon, and scalar return. <code>ChartMetadata</code> identifies a daily or weekly chart artifact. <code>PredictionRecord</code> preserves both raw model text and an optional parsed score, so an invalid response is not silently converted into a number. <code>MetricResult</code> stores metric values, counts, warnings, and policy metadata.</p><p>Configuration is centralized in <code>src/vlm_candlestick/config.py</code>. <code>ChartConfig</code> covers the bar limit, moving-average windows, colors, and rendering choices. <code>LabelConfig</code> describes the trading-day horizon. <code>DirectionConfig</code> makes the unresolved classification threshold and zero policy explicit. <code>WeeklyResampleConfig</code> records the weekly aggregation convention. <code>InferenceConfig</code> captures bounded concurrency and persistence batching. <code>EvaluationConfig</code> records correlation and significance policies, while <code>BaselineConfig</code> keeps the incomplete numerical-baseline specification visible.</p><p>The package helper <code>build_sample_key()</code> in <code>src/vlm_candlestick/__init__.py</code> creates a stable labelled identifier from the market, stock, and cutoff date. Derived chart or prediction records can add frequency or model context. Conceptually, one synthetic example moves through the following identity-preserving chain:</p><ol><li><p>A <code>SampleKey</code> identifies a stock and cutoff.</p></li><li><p><code>forward_return_label_construction()</code> attaches the 30-trading-day target.</p></li><li><p><code>build_paired_charts()</code> creates daily and weekly chart metadata with the same identity.</p></li><li><p><code>parse_vlm_score()</code> converts a replayed response into an optional validated prediction while retaining the raw text.</p></li><li><p>Evaluation functions consume aligned predictions and labels and return a <code>MetricResult</code> with the selected policy recorded.</p></li></ol><p>That chain is an implementation explanation, not a claim that the paper specifies these exact Python classes or function boundaries.</p><h3>The main processing layers</h3><p>The repository follows the benchmark's data dependencies in a roughly chronological order:</p><ul><li><p><code>data/</code> validates local OHLCV frames, performs trading-position lookup, and aggregates daily records into weekly records.</p></li><li><p><code>labels/</code> constructs forward-return targets using the cutoff and a subsequent trading-day position.</p></li><li><p><code>charts/</code> computes moving averages, renders daily and weekly charts, and verifies that the pair shares its identity.</p></li><li><p><code>dataset/</code> stores aligned sample manifests containing the two chart paths and one shared target.</p></li><li><p><code>transport/</code> converts images to RGB, optionally composites a white background, resizes proportionally, and prepares Base64 payloads.</p></li><li><p><code>prompt/</code> builds the paper-derived instructions and parses the tagged numerical response.</p></li><li><p><code>inference/</code> defines a provider-neutral interface, local replay behavior, duplicate protection, five-worker bounded processing, ten-entry batching, and resume support.</p></li><li><p><code>baseline/</code> prepares pre-cutoff numerical windows and an explicit feature schema for an optional XGBoost wrapper.</p></li><li><p><code>evaluation/</code> implements directional, bias, correlation, regime, extreme-stock, and horizon analyses.</p></li><li><p><code>pipeline/</code> provides local entry points for dataset construction, replay or injected inference, and evaluation.</p></li></ul><p>The provider-neutral boundary is deliberate. The paper names commercial VLM configurations, but the supplied material does not provide their authentication procedures, request schemas, model parameters, retry policies, or rate-limit behavior. The generated package therefore does not invent a live API client. A replay provider or injected callable can demonstrate the interfaces without implying that a commercial request was sent.</p><h3>What the local demonstration means</h3><p>The tutorial's synthetic data path is useful for showing how identities and invariants move through the package, but it is not a substitute for the paper's market data. A synthetic stock and cutoff can demonstrate a valid <code>SampleKey</code>, paired chart paths, a parsed score such as <code>0.100</code>, and an evaluation record. It cannot establish the paper's sample counts, model behavior, or reported metrics.</p><p>The same caution applies to verification. The planned repository includes mathematical and static-contract test files, but the authoritative run policy disabled code execution, test generation, local static verification, semantic code verification, tutorial-section verification, and final quality review. Consequently, this section makes no claim that the generated code was executed, that tests passed, or that reported results were reproduced. The code and workflow are presented as a scaffold whose assumptions and boundaries can be reviewed and refined in an authorized later run.</p><h2>OHLCV inputs, trading-day alignment, and the 30-day target</h2><p>How do we define &#8220;30 days later&#8221; without accidentally counting weekends, holidays, or missing observations? In this benchmark, the answer is positional: start at the cutoff date and move forward by 30 observed trading rows. If that future row does not exist, the sample has no valid primary target.</p><p>The paper constructs its input from <strong>OHLCV</strong> data: open, high, low, close, and volume. The scaffold accepts this data from local CSV files or pandas DataFrames rather than acquiring it from TuShare or Yahoo Finance. That boundary is intentional: the supplied material names those providers but does not provide credentials, exact acquisition procedures, constituent snapshots, or missing-data rules.</p><h3>Validate the raw observations first</h3><p>The boundary function <code>validate_ohlcv_frame()</code> in <code>src/vlm_candlestick/data/validation.py</code> checks that a frame contains the six required columns: <code>date</code>, <code>open</code>, <code>high</code>, <code>low</code>, <code>close</code>, and <code>volume</code>. It normalizes dates, converts numeric fields to floating-point values, rejects non-finite values, enforces positive prices and nonnegative volume, checks the OHLC relationships, rejects duplicate dates, and sorts the result chronologically.</p><p>That validation protects later positional indexing. A date sequence containing duplicates or out-of-order rows cannot reliably answer the question &#8220;which observation is 30 trading days later?&#8221; The separate <code>sort_and_deduplicate_ohlcv()</code> helper uses an explicit local policy when deduplication is required: it retains the last input row for each date and then validates the result. This is an implementation decision, not a missing rule recovered from the paper.</p><p>A focused loader excerpt shows how local files are normalized without adding provider-specific behavior:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;be2a6d10-829b-4a4d-9d11-f3f1234e724b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from vlm_candlestick.data.validation import (
    REQUIRED_COLUMNS,
    sort_and_deduplicate_ohlcv,
    validate_ohlcv_frame,
)


def load_ohlcv_csv(path: Path) -&gt; OHLCVTable:
    """Load and normalize one local OHLCV CSV file.

    The CSV must contain ``date``, ``open``, ``high``, ``low``, ``close``, and
    ``volume`` columns. Extra columns are ignored so that the returned frame
    has a stable schema. Duplicate dates are resolved by retaining the last
    row in file order, then the resulting records are validated and sorted
    chronologically.
    """
    file_path = Path(path)
    if not file_path.exists():
        raise FileNotFoundError(f"OHLCV CSV does not exist: {file_path}")

    raw = pd.read_csv(file_path)
    missing = [column for column in REQUIRED_COLUMNS if column not in raw.columns]
    if missing:
        raise ValueError(
            f"OHLCV CSV {file_path} is missing required columns: {missing}"
        )

    selected = raw.loc[:, list(REQUIRED_COLUMNS)]
    normalized = sort_and_deduplicate_ohlcv(selected)
    return validate_ohlcv_frame(normalized)</code></pre></div><p>The complete generated loader also handles empty files and parser errors. The important contract here is the returned chronological table with one row per retained date. Suspensions, interpolation, corporate-action adjustment, and provider-specific repair remain outside the supplied specification.</p><h3>Use trading positions, not calendar arithmetic</h3><p>A trading-day index is an ordered sequence of dates that actually occur in the stock's input data. The function <code>future_position_date()</code> receives that sequence, a cutoff date, and a positive integer <code>horizon</code>. It finds the cutoff's zero-based position and returns the date at <code>cutoff_position + horizon</code>. It returns <code>None</code> when the cutoff is absent or when there are not enough subsequent observations.</p><p>The relevant implementation keeps the distinction explicit:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;9713204f-f0f0-4dce-a951-e06fc78cb0d4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def future_position_date(
    date_index: Sequence[date],
    cutoff: date,
    horizon: int,
) -&gt; date | None:
    """Return the date exactly ``horizon`` observations after ``cutoff``.

    The returned date is strictly later than the cutoff when it exists.  ``None``
    is returned when the cutoff is absent or when insufficient subsequent trading
    observations are available.  The horizon is a count of rows, not calendar
    days, matching the paper's trading-day alignment requirement.
    """
    if isinstance(horizon, bool) or not isinstance(horizon, int) or horizon &lt;= 0:
        raise ValueError("horizon must be a positive integer")

    normalized_index = _normalized_date_index(date_index)
    normalized_cutoff = _as_date(cutoff)
    try:
        cutoff_position = normalized_index.index(normalized_cutoff)
    except ValueError:
        return None

    future_position = cutoff_position + horizon
    if future_position &gt;= len(normalized_index):
        return None
    return normalized_index[future_position]</code></pre></div><p>For example, suppose the available dates are Tuesday, Wednesday, Thursday, Friday, Monday, and the next Tuesday. If the cutoff is Wednesday and the horizon is <code>2</code>, the future observation is Friday: Thursday is one subsequent row and Friday is two. A calendar calculation of two dates would not express this invariant and could land on a weekend or holiday.</p><p>The companion <code>trading_position()</code> function is stricter: it raises an error when the target date is not present, because exact positional lookup is its purpose. In contrast, <code>future_position_date()</code> treats an absent cutoff or insufficient history as an unavailable label and returns <code>None</code>. This difference allows dataset construction to omit incomplete samples without fabricating targets.</p><h3>Keep chart history strictly before the cutoff</h3><p>The visual input and numerical history obey a stronger temporal rule than the label lookup. <code>eligible_history()</code> validates the frame and retains only rows whose date is strictly earlier than the cutoff:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ea33f453-6a57-423b-9b8a-53802bdb3d9f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def eligible_history(frame: pd.DataFrame, cutoff: date) -&gt; pd.DataFrame:
    """Return validated OHLCV rows strictly before ``cutoff``.

    The result is a new chronologically ordered frame.  The strict predicate is
    intentional: the cutoff row and every later row are excluded, preventing
    chart or feature construction from observing the target-date record or
    future data.
    """
    normalized = validate_ohlcv_frame(frame)
    cutoff_date = _as_date(cutoff)
    cutoff_timestamp = pd.Timestamp(cutoff_date)
    result = normalized.loc[normalized["date"] &lt; cutoff_timestamp].copy()
    return result.reset_index(drop=True)</code></pre></div><p>This creates an important distinction. The cutoff close is needed to calculate the future label, but the chart history is interpreted as containing only records before the cutoff. Therefore, the current-date close may participate in target construction while its candle is excluded from the visual history. That convention follows the supplied implementation plan; the paper does not fully clarify whether the current-date candle appears elsewhere in the chart.</p><h3>Construct the canonical forward-return target</h3><p>The paper's primary target is a scalar 30-trading-day forward return. In the equation below, <code>r_{30}</code> is the target, <code>P_t</code> is the positive closing price on the cutoff date, and <code>P_{t+30}</code> is the positive closing price 30 subsequent trading-day positions later. The expression maps directly to <code>construct_forward_return()</code> and its named method-card entry point, <code>forward_return_label_construction()</code>.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;r_{30} = \\frac{P_{t+30} - P_t}{P_t}&quot;,&quot;id&quot;:&quot;2F60F38368&quot;}" data-component-name="LatexBlockToDOM"></div><p>This is a scalar real-valued return, not a vector or a class label. The current and future prices are positive scalars, so the division is defined under the OHLCV validation rules. A positive result indicates that the later close exceeds the cutoff close; a negative result indicates a decrease. The paper's VLM output is separately constrained to the interval <code>[-0.5, 1.0]</code>, but the supplied target definition does not state an equivalent clipping rule for observed returns.</p><p>The generated label function first validates the horizon and frame, locates the exact cutoff row, obtains the future date by positional lookup, and then applies the equation:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2c86fb77-b6de-40c6-af6d-580afdbd95e5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def construct_forward_return(
    frame: OHLCVTable,
    cutoff: date,
    horizon: int = 30,
) -&gt; ForwardReturnLabel | None:
    """Construct one positional forward-return label.

    The cutoff row is the current observation, and ``horizon`` counts subsequent
    rows in the validated chronological trading-date index.  A missing cutoff or
    insufficient future history returns ``None`` rather than fabricating a label.
    """
    _validate_horizon(horizon)
    normalized = validate_ohlcv_frame(frame)
    cutoff_date = _normalize_date(cutoff)

    dates = [timestamp.date() for timestamp in normalized["date"]]
    try:
        cutoff_position = dates.index(cutoff_date)
    except ValueError:
        return None

    future_date = future_position_date(dates, cutoff_date, horizon)
    if future_date is None:
        return None

    future_position = cutoff_position + horizon
    if future_position &gt;= len(normalized):
        return None

    current_close = float(normalized.iloc[cutoff_position]["close"])
    future_close = float(normalized.iloc[future_position]["close"])
    if current_close &lt;= 0.0 or future_close &lt;= 0.0:
        raise ValueError("current and future closing prices must be positive")

    # Eq. 1: r_{30} = \frac{P_{t+30} - P_t}{P_t}
    return_value = (future_close - current_close) / current_close

    return ForwardReturnLabel(
        cutoff_date=cutoff_date,
        future_date=future_date,
        current_close=current_close,
        future_close=future_close,
        horizon=horizon,
        return_value=float(return_value),
    )</code></pre></div><p><code>ForwardReturnLabel</code> records the cutoff date, future date, current close, future close, horizon, and scalar return. The identity of the stock is carried by the surrounding sample key rather than by this label object. Later, the visual sample and numerical baseline must attach this same label to the same stock-and-cutoff identity.</p><h3>A hand-checkable example</h3><p>Consider this shortened close-price sequence:</p><p>| Position | Date | Close | |---:|---|---:| | 0 | 2024-01-02 | 100.0 | | 1 | 2024-01-03 | 101.0 | | 2 | 2024-01-04 | 110.0 | | 3 | 2024-01-05 | 111.0 | | 4 | 2024-01-08 | 120.0 | | 5 | 2024-01-09 | 121.0 |</p><p>With cutoff <code>2024-01-03</code> and horizon <code>2</code>, the function selects position 3, <code>2024-01-05</code>, not <code>2024-01-05</code> because it is two calendar dates later, but because it is the second subsequent observed row. The current close is <code>101.0</code> and the future close is <code>111.0</code>, so the returned scalar is computed from those two values.</p><p>The generated tests express the same reasoning without relying on a long real dataset:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;907309ea-9402-4fed-acb1-25241dccec14&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def test_forward_return_uses_subsequent_trading_position() -&gt; None:
    """Verify that horizon indexing counts observations, not calendar days."""
    frame = _small_ohlcv_frame()
    cutoff = date(2024, 1, 3)
    date_index = [timestamp.date() for timestamp in frame["date"]]

    future = future_position_date(date_index, cutoff, horizon=2)

    # 2024-01-04 is position +1 and 2024-01-05 is position +2.
    assert future == date(2024, 1, 5)

    label = forward_return_label_construction(frame, cutoff, horizon=2)

    assert label is not None
    assert label.cutoff_date == cutoff
    assert label.future_date == date(2024, 1, 5)</code></pre></div><p>The test uses horizon <code>2</code> only to keep the fixture small. It checks the same parameterized procedure used with the paper's primary horizon of <code>30</code>.</p><h3>Optional five-day sensitivity labels</h3><p>The paper additionally examines sensitivity to five-day outcomes. The scaffold does not invent a second canonical equation for that analysis. Instead, <code>forward_return_label_construction(..., horizon=5)</code> uses the same positional lookup and return calculation with <code>H = 5</code>, where <code>H</code> is the positive integer trading-day horizon. The primary call remains <code>horizon=30</code>.</p><p>This parameterization is useful because it prevents two implementations from drifting apart: both horizons use the same cutoff close, the same date-index convention, and the same missing-future policy. Horizon comparisons must still preserve sample identity and report missingness, because a stock may have enough history for a five-day label but not for a thirty-day label.</p><h3>What is guaranteed, and what remains open</h3><p>The temporal guarantees supplied by this scaffold are concrete:</p><ul><li><p>OHLCV rows are validated and ordered before positional lookup.</p></li><li><p>A future label uses exactly the requested number of subsequent observations.</p></li><li><p>Chart and numerical history can use only rows strictly before the cutoff.</p></li><li><p>Missing cutoff or future observations produce no fabricated label.</p></li><li><p>The visual and numerical versions of a stock-and-cutoff sample must share the same target.</p></li></ul><p>Several data decisions remain unresolved by the paper: treatment of suspended stocks, duplicate or missing observations, corporate actions, changing index membership, and the exact cutoff-date list. The generated code records explicit validation and omission behavior, but those choices should not be presented as exact recovery of the original data pipeline. The supplied static and semantic verification stages were skipped under the run policy, and no code or tests are claimed to have been executed.</p><h2>Weekly aggregation, moving averages, and chart construction</h2><p>How can one historical OHLCV series become two useful visual views without changing the stock or accidentally including future data? The benchmark uses a close-up daily view and a compressed weekly view. The daily chart preserves individual trading sessions; the weekly chart combines those sessions into larger candles that can make longer trends easier to see. The two outputs must still share the same stock and cutoff identity.</p><p>The paper specifies the data semantics: weekly <code>open</code> is the first daily open in the period, <code>high</code> is the maximum high, <code>low</code> is the minimum low, <code>close</code> is the last close, and <code>volume</code> is the sum. It also specifies MA5, MA20, and MA90 overlays, a volume panel, and no more than 50 displayed candles. The chart history is interpreted here as strictly pre-cutoff: a row dated on the cutoff or later cannot appear in either input image.</p><p>Some rendering details are not supplied. The week boundary, weekly label date, incomplete-week treatment, dimensions, fonts, and exact plotting style therefore remain implementation decisions in <code>WeeklyResampleConfig</code> and <code>ChartConfig</code>. Recording those choices is important: changing the weekly boundary can change which daily observations belong to a candle and can therefore change the model input.</p><h3>Aggregate daily observations into weekly OHLCV</h3><p><code>aggregate_weekly_ohlcv()</code> in <code>src/vlm_candlestick/data/weekly.py</code> validates the input, applies a configured pandas resampling rule, performs the five field-specific aggregations, and returns a chronologically ordered frame with <code>date</code>, <code>open</code>, <code>high</code>, <code>low</code>, <code>close</code>, and <code>volume</code>. The function is deliberately policy-driven because the supplied paper does not resolve every weekly convention.</p><p>The central aggregation block is short enough to inspect directly:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a6ef43a4-d35b-4eec-ad48-1c94c426a489&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">aggregated = (
    indexed.resample(
        rule,
        label=label,
        closed=policy.closed,
    )
    .agg(
        open=("open", "first"),
        high=("high", "max"),
        low=("low", "min"),
        close=("close", "last"),
        volume=("volume", "sum"),
    )
    .dropna(subset=["open", "high", "low", "close"])
    .reset_index()
    .rename(columns={"date": "date"})
)</code></pre></div><p>The input is a date-indexed OHLCV table with one row per daily observation. The result has one row per retained weekly period. Dropping periods without a usable price prevents an empty resampling bin from becoming a chart candle. The implementation also validates positive prices, nonnegative volume, and basic candle relationships at the input boundary.</p><p><code>filter_weekly_before_cutoff()</code> applies the no-future rule after aggregation. <code>frequency_frame()</code> uses that function for daily data and first aggregates, then filters, for weekly data. This ordering matters: the chart receives only eligible period labels, while the weekly construction policy remains explicit rather than being hidden in a calendar shortcut.</p><p>A small worked example makes the field ownership concrete. Suppose one configured weekly period contains daily opens <code>10, 11, 12, 13, 14</code>, highs <code>12, 15, 14, 16, 17</code>, lows <code>9, 10, 11, 12, 13</code>, closes <code>11, 12, 13, 14, 15</code>, and volumes <code>100, 200, 300, 400, 500</code>. The resulting weekly candle has open <code>10</code>, high <code>17</code>, low <code>9</code>, close <code>15</code>, and volume <code>1,500</code>. This is the behavior targeted by <code>test_weekly_aggregation_fields()</code> in <code>tests/test_weekly_and_charts.py</code>; that test is a planned/generated test artifact, not an executed check in this run.</p><h3>Compute moving averages before truncating the display</h3><p>A moving average is a rolling mean of closing prices. For a frequency-specific series, MA5 uses the most recent five observations, MA20 uses 20, and MA90 uses 90. These are observation windows, so the daily and weekly versions have different time spans even though they use the same window names.</p><p><code>compute_moving_averages()</code> returns columns named <code>MA5</code>, <code>MA20</code>, and <code>MA90</code> when those windows are requested. Its <code>min_periods=window</code> setting leaves the leading values missing until a complete historical window exists. <code>attach_chart_indicators()</code> copies the OHLCV frame and adds the indicator columns without changing the original price or volume values.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7695ab08-7a2c-4978-a2d6-fd0081c1986d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">result = {
    f"MA{window}": close.astype("float64").rolling(
        window=window,
        min_periods=window,
    ).mean()
    for window in normalized_windows
}
return DataFrame(result, index=close.index)</code></pre></div><p>The important sequencing decision is to compute indicators on the complete eligible frequency history before selecting the final bars for display. If the chart shows only the latest 50 candles, an MA90 value can still use earlier eligible observations. Computing the average after truncation would unnecessarily discard that history and change the displayed indicator.</p><h3>Select and render at most 50 eligible bars</h3><p><code>select_last_bars()</code> sorts the prepared frame, removes every row whose date is greater than or equal to the cutoff, and keeps the most recent <code>max_bars</code> rows. It rejects a bar limit above 50, matching the paper's stated display cap. The function returns a compact, chronologically ordered frame for rendering.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a59632aa-6824-416a-a5e7-15ea21da88c9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">result = frame.copy()
result["date"] = pd.to_datetime(result["date"]).dt.normalize()
result = result.sort_values("date", kind="stable")
# Chart tensor flow: eligible OHLCV -&gt; most recent &lt;= 50 bars -&gt; image.
result = result.loc[result["date"] &lt; _normalise_cutoff(cutoff)]
if result.empty:
    raise ValueError("no chart candles occur strictly before the cutoff")
return result.tail(max_bars).reset_index(drop=True)</code></pre></div><p>The strict comparison is the key invariant. A chart may use historical rows before the cutoff, but it must not display the cutoff row or any later row under this implementation's interpretation. A missing eligible history is treated as an error rather than producing an empty image.</p><p><code>render_candlestick_png()</code> draws the selected price candles, moving-average lines, and volume panel into a local PNG. The specified color mapping is MA5 black, MA20 blue, and MA90 purple. Up candles are green and down candles are red in the generated renderer. Figure dimensions, DPI, typography, and other presentation details are configurable local decisions because they are absent from the supplied paper context.</p><h3>Build the aligned daily/weekly pair</h3><p>The pair builder applies the same <code>SampleKey</code> to both artifacts. It prepares daily and weekly frames from the same raw stock history, attaches frequency-specific indicators, selects the recent eligible bars, and writes deterministic paths containing the market, stock, cutoff date, and frequency.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;08fd0633-5458-494c-9c5b-d832aebc266b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">daily_metadata = _render_one_frequency(
    frame=frame,
    key=key,
    frequency="daily",
    output_path=daily_path,
    config=config,
    weekly_policy=weekly_policy,
)
weekly_metadata = _render_one_frequency(
    frame=frame,
    key=key,
    frequency="weekly",
    output_path=weekly_path,
    config=config,
    weekly_policy=weekly_policy,
)

if daily_metadata.key != weekly_metadata.key:
    raise RuntimeError("daily and weekly chart identities do not match")
if daily_metadata.frequency != "daily" or weekly_metadata.frequency != "weekly":
    raise RuntimeError("paired chart frequencies are not correctly assigned")
return daily_metadata, weekly_metadata</code></pre></div><p><code>build_paired_charts()</code> therefore produces exactly one daily and one weekly <code>ChartMetadata</code> record for the requested stock and cutoff. <code>index_chart_pair()</code> then serializes the two paths and identity fields for a manifest. This identity check prevents an especially damaging alignment error: pairing a daily chart from one cutoff with a weekly chart from another.</p><p>In practice, a local workflow can construct a validated daily frame, call <code>aggregate_weekly_ohlcv()</code> with an explicit weekly policy, attach <code>(5, 20, 90)</code> indicators to each frequency, and call <code>build_paired_charts()</code>. The resulting paths might follow the generated layout <code>output_dir / market / stock / cutoff / daily.png</code> and <code>weekly.png</code>. Those paths are implementation choices, but retaining all identity components makes collisions easier to detect.</p><p>The corresponding planned tests check only supported invariants: weekly field ownership, strict pre-cutoff selection, the 50-bar limit, moving-average names and alignment, and shared daily/weekly identity. They do not assert unspecified fonts, dimensions, or plotting aesthetics. The generated chart and test files were not executed or statically verified under the run policy, so this section describes their intended responsibilities rather than claiming a verified rendering result.</p><p>Finally, the chart colors and overlays should not be mistaken for separately supervised labels. They are visual content and prompt-referenced semantics supplied to the VLM. The benchmark's visual input is therefore a paired chart plus textual instructions, not a text-free image-only task in the strictest sense. The next pipeline stage can use these PNG paths to create aligned manifests and provider-neutral image requests without changing the stock, cutoff, or target identity.</p><h2>Aligned manifests and the no-leakage data contract</h2><p>How can you be sure that a model prediction is attached to the correct stock, cutoff date, charts, and future return? Treat each benchmark example as a manifest-backed receipt. The receipt records the market, stock, and cutoff identity; points to exactly one daily chart and one weekly chart; and stores the single forward-return label shared by the visual and numerical versions of the sample.</p><p>This alignment is more than bookkeeping. If the daily chart belongs to one cutoff while the weekly chart or target belongs to another, the experiment no longer measures the intended task. The repository therefore carries one <code>SampleKey</code> through chart generation, label construction, manifest serialization, inference, and evaluation.</p><h3>The sample identity</h3><p>The core identity has three fields: <code>market</code>, <code>stock</code>, and <code>cutoff_date</code>. The market distinguishes universes such as HS300 and S&amp;P 500; the stock identifies the constituent; and the cutoff date identifies the historical point at which the forecast is made. Frequency and model information can be added for derived artifacts, but they do not replace the underlying sample identity.</p><p>The public helper <code>build_sample_key()</code> in <code>src/vlm_candlestick/__init__.py</code> creates a labelled, deterministic string. Its labels make leading-zero stock identifiers unambiguous.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;69f03793-cece-4d02-b2a7-d872b9888a1a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">build_sample_key("HS300", "000001", "2024-01-01")
# 'market=HS300|stock=000001|cutoff=2024-01-01'</code></pre></div><p>That excerpt is the usage shown in the generated module's documentation. The function accepts a <code>date</code>, <code>datetime</code>, or ISO-8601 string, normalizes the cutoff to an ISO date, and rejects empty components or delimiter characters. Optional <code>frequency</code> and <code>model</code> fields distinguish, for example, a daily chart artifact from a model prediction while retaining the same market, stock, and cutoff.</p><h3>One target, two charts</h3><p>The paper's primary target is the 30-trading-day forward return. Its canonical definition is the same target used by the manifest's <code>ForwardReturnLabel</code>; later stages should consume that stored scalar rather than independently recomputing it.</p><p>The following equation defines the purpose of the label. It is the supplied canonical equation for the benchmark's primary horizon and is mapped in the implementation to <code>forward_return_label_construction()</code> and, through the dataset builder, to the label field of <code>BenchmarkSample</code>.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;r_{30} = \\frac{P_{t+30} - P_t}{P_t}&quot;,&quot;id&quot;:&quot;F919C120E1&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here, <code>r_{30}</code> is a scalar real-valued return, <code>P_t</code> is the positive closing price at the cutoff, and <code>P_{t+30}</code> is the positive closing price 30 subsequent trading-day positions later. The formula is not a vector operation: each stock and cutoff produces one scalar target. The chart inputs remain historical views, while the future close is used only to construct the label.</p><p>In particular, the implementation's chart convention excludes rows on or after the cutoff, whereas the target still needs the cutoff close and a later close. This apparent distinction is important: excluding the cutoff candle from the visual history does not mean that the cutoff close is unavailable to label construction. The label constructor is also parameterized, so a five-trading-day sensitivity label can follow the same positional pattern without being treated as a second supplied canonical equation.</p><h3><code>BenchmarkSample</code> enforces the contract</h3><p><code>src/vlm_candlestick/dataset/records.py</code> defines <code>BenchmarkSample</code> as a frozen record containing four components: a <code>SampleKey</code>, daily <code>ChartMetadata</code>, weekly <code>ChartMetadata</code>, and a <code>ForwardReturnLabel</code>. Its validation performs the checks that matter before inference:</p><ul><li><p>the daily artifact must declare frequency <code>daily</code>;</p></li><li><p>the weekly artifact must declare frequency <code>weekly</code>;</p></li><li><p>both chart identities must equal the sample key;</p></li><li><p>the label cutoff must equal the key cutoff;</p></li><li><p>the stored return must agree with its current and future closing prices.</p></li></ul><p>The central identity check is visible in the generated class:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;958b80a8-fdbc-4fe0-8ca1-2e4074084926&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">if self.daily.frequency != "daily":
    raise ValueError("daily chart metadata must have frequency 'daily'")
if self.weekly.frequency != "weekly":
    raise ValueError("weekly chart metadata must have frequency 'weekly'")
if self.daily.key != self.key or self.weekly.key != self.key:
    raise ValueError("both chart identities must match the sample key")
if self.label.cutoff_date != self.key.cutoff_date:
    raise ValueError("label cutoff_date must match the sample key")</code></pre></div><p>This is a useful failure boundary. A mismatch is rejected before a request is assembled, rather than becoming an unexplained model error later. The <code>target</code> property exposes the stored scalar return as a <code>float</code>, while <code>daily_path</code> and <code>weekly_path</code> expose the two local chart paths.</p><p>Use <code>make_benchmark_sample()</code> rather than manually bypassing these checks:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e54c12fe-9c58-44d9-a29e-0427c3d17ef9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">return BenchmarkSample(key=key, daily=daily, weekly=weekly, label=label)</code></pre></div><p>The function is deliberately small because the dataclass owns the invariants. It constructs a validated record; it does not generate charts or recalculate the target. That separation prevents downstream code from silently changing the definition of a sample.</p><h3>Building a manifest from local inputs</h3><p><code>build_benchmark_manifest()</code> in <code>src/vlm_candlestick/dataset/build.py</code> joins the three required ingredients: validated stock frames, eligible cutoff dates, and an output directory for chart artifacts. For each stock and cutoff, it first requests the forward-return label. If the required future trading observation is unavailable, the sample is omitted under the configured missing-data policy, or the builder raises an error when that policy is <code>raise</code>.</p><p>The key portion of the generated builder is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f98e8350-fad5-42b3-bed0-8f222fb259ea&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">label = forward_return_label_construction(
    frame=frame,
    cutoff=cutoff,
    horizon=effective_label_config.horizon,
)
if label is None:
    if effective_label_config.missing_data_policy == "raise":
        raise ValueError(
            f"no complete {effective_label_config.horizon}-trading-day "
            f"label for stock {stock!r} at cutoff {cutoff.isoformat()}"
        )
    continue

key = SampleKey(market=market, stock=stock, cutoff_date=cutoff)</code></pre></div><p>After the label and key exist, the builder calls <code>build_paired_charts()</code> with the same stock, market, and cutoff. That function returns one daily and one weekly <code>ChartMetadata</code> object. <code>make_benchmark_sample()</code> then combines those objects with the label, so the shared identity is checked at the point where the complete record is formed.</p><p>The builder also keeps a set of serialized identities and raises on duplicates. This matters when cutoff lists overlap or when a caller accidentally supplies the same stock more than once. The exact constituent snapshots, 32 cutoff dates, and index-membership changes from the paper are not supplied, so this local builder does not invent or assert them. It processes the frames and cutoffs that the caller provides.</p><h3>JSON Lines as a reproducibility receipt</h3><p><code>write_sample_manifest()</code> serializes one complete sample per line as UTF-8 JSON Lines. Each line contains the key, daily and weekly frequency/path records, and the label's cutoff date, future date, current close, future close, horizon, and return value. JSON Lines is practical here because records can be inspected one at a time, appended or processed incrementally, and loaded without requiring a large nested in-memory object.</p><p>The generated writer uses a temporary file in the destination directory, flushes and synchronizes it, and then replaces the destination:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c9942a6b-9c30-4731-811e-f5f609598274&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">with NamedTemporaryFile(
    mode="w",
    encoding="utf-8",
    dir=path.parent,
    prefix=f".{path.name}.",
    suffix=".tmp",
    delete=False,
) as temporary:
    temporary_name = temporary.name
    if serialized:
        temporary.write("\n".join(serialized))
        temporary.write("\n")
    temporary.flush()
    os.fsync(temporary.fileno())
os.replace(temporary_name, path)</code></pre></div><p>This atomic replacement is an implementation safeguard: an interrupted write should not normally leave the final manifest half-written. It does not solve every possible crash or multi-process coordination problem, but it gives the local workflow a clear persistence boundary.</p><p><code>read_sample_manifest()</code> accepts the generated JSON Lines form and also accepts a JSON array for small hand-written manifests. It reconstructs typed <code>SampleKey</code>, <code>ChartMetadata</code>, and <code>ForwardReturnLabel</code> objects, then passes them through <code>BenchmarkSample</code> validation. Consequently, a malformed or misaligned record is rejected during loading instead of being silently treated as a valid sample.</p><h3>Completeness before inference</h3><p>A manifest record can be structurally valid while a chart file has later been deleted. <code>filter_complete_samples()</code> handles this filesystem-level concern. It retains only records whose daily and weekly paths are regular files and whose stored closing prices are positive. It also rejects duplicate sample identities in the supplied sequence. It does not regenerate missing charts or infer missing labels.</p><p>A practical synthetic walkthrough therefore has this shape:</p><ol><li><p>Generate or load one local stock frame.</p></li><li><p>Select a cutoff with a complete 30-trading-day future horizon.</p></li><li><p>Construct the label using <code>forward_return_label_construction()</code>.</p></li><li><p>Generate daily and weekly chart metadata with the same <code>SampleKey</code>.</p></li><li><p>Call <code>make_benchmark_sample()</code> and inspect <code>sample.target</code>, <code>sample.daily_path</code>, and <code>sample.weekly_path</code>.</p></li><li><p>Write the sample with <code>write_sample_manifest()</code> and restore it with <code>read_sample_manifest()</code>.</p></li><li><p>Run <code>filter_complete_samples()</code> before creating an inference request.</p></li></ol><p>The expected property is not a particular numeric result. It is that the restored record still names the same market, stock, cutoff, chart frequencies, chart paths, horizon, and scalar target. No code execution or reload check was performed under the authoritative run policy; this is the intended workflow and contract, not a claimed verification.</p><p>Finally, the existence of a local manifest must not be confused with reproducing the paper's reported 193,524 chart samples. That count is a paper-reported reference value, while the supplied context omits the exact constituent lists, cutoff dates, and several dataset conventions. A local manifest is reproducible for its explicitly supplied inputs, but exact reproduction of the paper's scale requires those missing specifications.</p><h2>Image preprocessing and provider-neutral VLM requests</h2><p>How do two chart files become a compact, auditable model input? The practical sequence is: prepare each image, encode its binary JPEG bytes as text, build the paper-derived prompt, and place exactly one daily image and one weekly image into a request record. This scaffold stops at that boundary. It does not invent a commercial provider's authentication, network protocol, or request schema.</p><p>A chart can be viewed conceptually as an image with shape <code>(H, W, C)</code>, where <code>H</code> is height, <code>W</code> is width, and <code>C</code> is the number of color channels. The transport code normalizes that representation to RGB, so the final decoded image has three color channels. <strong>Base64</strong> is a text encoding of binary bytes; here, it makes JPEG data suitable for storage in a provider-neutral request dictionary.</p><h3>Normalize the image before encoding</h3><p>The paper specifies RGB conversion, white-background compositing when transparency is present, proportional resizing with LANCZOS, JPEG compression, and Base64 encoding. The generated implementation follows that order. The resize bounding box and JPEG quality are not specified by the paper, so they are explicit local arguments rather than hidden claims about the original experiment.</p><p><code>preprocess_image()</code> returns validated JPEG bytes. Its internal <code>_rgb_image()</code> helper converts transparent images to RGBA, composites them over an opaque white image, and then converts the result to RGB. The focused public function is reproduced below from <code>src/vlm_candlestick/transport/images.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;fc503318-3c84-4a4e-be7f-4e6d8bced6ed&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def preprocess_image(
    path: Path,
    target_size: tuple[int, int] | None = None,
    quality: int = 85,
) -&gt; bytes:
    """Prepare a local chart image as validated JPEG bytes.

    Images are converted to RGB. Any transparency is composited over a white
    background before the optional aspect-ratio-preserving resize. The resize
    target is a maximum bounding box, and the JPEG quality is an explicit local
    implementation choice because the paper does not specify either setting.

    Raises:
        ImageTransportError: if the path cannot be opened or encoded.
        ValueError or TypeError: if transport options are invalid.
    """
    image_path = Path(path)
    _validate_target_size(target_size)
    _validate_quality(quality)

    try:
        with Image.open(image_path) as source:
            source.load()
            rgb = _rgb_image(source)
            resized = _resize_proportionally(rgb, target_size)
            return _encode_jpeg(resized, quality)
    except (OSError, ValueError, RuntimeError) as exc:
        if isinstance(exc, ImageTransportError):
            raise
        raise ImageTransportError(image_path, str(exc)) from exc</code></pre></div><p>Several details matter here. <code>target_size</code> is interpreted as a maximum width and height, not as a forced distortion of the chart. The resize helper computes one scale factor and applies it to both dimensions, preserving the aspect ratio. LANCZOS is used for the resampling filter. <code>_encode_jpeg()</code> then reopens and verifies the encoded artifact before returning it. If the source cannot be opened or encoded, <code>ImageTransportError</code> retains the source path in the exception instead of allowing a missing chart to look like a successful payload.</p><p>For transport, <code>encode_image_base64()</code> performs only the byte-to-text conversion. It does not add a data-URL prefix:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e18b99a0-1a1e-43d4-8e78-64fd5b1d6d6b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def encode_image_base64(jpeg_bytes: bytes) -&gt; str:
    """Encode JPEG bytes as an ASCII Base64 payload.

    This function does not add a data-URL prefix; callers receive only the
    encoded payload suitable for provider-neutral request records.
    """
    if not isinstance(jpeg_bytes, (bytes, bytearray, memoryview)):
        raise TypeError("jpeg_bytes must be bytes-like")
    if not jpeg_bytes:
        raise ValueError("jpeg_bytes must not be empty")
    return base64.b64encode(bytes(jpeg_bytes)).decode("ascii")</code></pre></div><p>The higher-level <code>image_to_base64()</code> function returns both the payload and metadata. Its metadata records the source path, original mode and dimensions, resized dimensions, JPEG quality, byte lengths, and status. On an image-processing failure it returns an empty payload and an error status. This is an implementation choice for auditability: downstream code can preserve the failed sample identifier rather than silently dropping it.</p><p>The planned tests in <code>tests/test_parser_and_transport.py</code> cover the intended invariants without asserting unspecified JPEG settings. They check that a <code>200</code> by <code>100</code> RGB image fitted into a <code>100</code> by <code>100</code> box becomes <code>100</code> by <code>50</code>, and that a transparent source is composited over white. Those tests are present as planned/generated test code, but no test execution was authorized or performed in this run.</p><h3>Build the textual prompt</h3><p>The image bytes are only half of the request. The paper-derived prompt describes the model's role as an unbiased stock trend analyst and asks it to estimate the 30-trading-day forward return. It tells the model what to inspect: candle bodies and wicks, green/up and red/down candle semantics, the MA5/MA20/MA90 overlays, volume, inflection points, and consistency between the daily and weekly views.</p><p>The prompt also imposes a strict output contract: one score in the interval <code>[-0.5, 1.0]</code>, formatted to three decimal places inside <code>&lt;score&gt;NUM&lt;/score&gt;</code>. A score of zero represents an unclear trend, and no explanatory text is requested. <code>build_vlm_prompt()</code> constructs this deterministic text from the stock and cutoff identity; it does not access market data or a model service.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;62796c9a-fd21-4234-957e-ec24a08c9150&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def format_score_instruction() -&gt; str:
    """Return the strict output contract for the VLM score.

    The requested score is a continuous estimate in the paper's stated range.
    A value of zero represents an unclear trend, and the response must contain
    no explanatory text outside the score tag.
    """
    return (
        "Return exactly one score formatted to three decimal places inside "
        "&lt;score&gt;NUM&lt;/score&gt;. NUM must be between -0.5 and 1.0 inclusive. "
        "Use 0.000 when the trend is unclear. Do not return any other text."
    )</code></pre></div><p>These instructions are part of the VLM input contract, not separate supervised labels. For example, asking the model to inspect support-like visual evidence does not create a support/resistance target in the dataset. The request still contains the two charts and one textual instruction string.</p><h3>Assemble exactly two chart payloads</h3><p><code>VLMRequest</code> is a provider-neutral dataclass. It stores one <code>SampleKey</code>, one Base64 payload for the daily chart, one for the weekly chart, and the prompt. Its validation rejects empty or invalid Base64 strings and requires both images to belong to the same request identity. It intentionally excludes the future-return label and numerical baseline features: those values are for evaluation or comparison, not VLM input.</p><p>The assembly function checks the chart frequencies and identity before constructing the request:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;06574c74-2d27-4a12-8852-0d8132aba038&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def build_vlm_request(
    sample: BenchmarkSample,
    daily_b64: str,
    weekly_b64: str,
) -&gt; VLMRequest:
    """Assemble a request for the sample's daily and weekly chart images.

    The prompt is generated from the sample identity and the paper-derived
    prompt contract. The sample label is used only to establish identity and is
    intentionally not copied into the request payload.
    """
    if not isinstance(sample, BenchmarkSample):
        raise TypeError("sample must be a BenchmarkSample instance")
    if sample.daily.frequency != "daily" or sample.weekly.frequency != "weekly":
        raise ValueError("sample must contain one daily and one weekly chart")
    if sample.daily.key != sample.key or sample.weekly.key != sample.key:
        raise ValueError("chart identities must match the sample key")

    prompt = build_vlm_prompt(sample.key.stock, sample.key.cutoff_date)
    return VLMRequest(
        key=sample.key,
        daily_image_base64=daily_b64,
        weekly_image_base64=weekly_b64,
        prompt=prompt,
    )</code></pre></div><p>The important invariant is identity equality: both chart metadata objects must carry the same market, stock, and cutoff identity as the sample. Frequency distinguishes the two payloads, while the shared key prevents a daily chart from one date being paired with a weekly chart from another.</p><p><code>serialize_request()</code> produces a local dictionary with the key fields, a stable sample-key string, an <code>images</code> mapping containing <code>daily</code> and <code>weekly</code>, and the prompt. This structure is useful for inspection, persistence, or writing a future adapter, but it is not claimed to match any commercial API. The paper names commercial VLM configurations, yet the supplied context does not define credentials, authentication, endpoint formats, image fields, rate limits, or model-specific parameters. Consequently, this scaffold implements no network call.</p><h3>A small local example</h3><p>Suppose a previously generated sample has <code>daily.png</code> and <code>weekly.png</code>. A local workflow would preprocess both files, then assemble the request:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;62cdb7e7-7595-40c3-92ff-76b545164552&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">daily_b64, daily_meta = image_to_base64(Path("daily.png"), target_size=(1024, 1024))
weekly_b64, weekly_meta = image_to_base64(Path("weekly.png"), target_size=(1024, 1024))

request = build_vlm_request(sample, daily_b64, weekly_b64)
serialized = serialize_request(request)</code></pre></div><p>The resulting <code>serialized["images"]</code> contains two distinct Base64 strings, while <code>serialized["prompt"]</code> contains the role, horizon, visual analysis instructions, and tagged-score rule. The sample's future label is not present. This separation preserves the benchmark's intended direction: images and prompt go toward inference; the aligned label remains available afterward for evaluation.</p><p>The paper mentions compression-related API cost reduction, but the supplied material does not specify a resize target, JPEG quality, or complete transport configuration. Therefore, this code demonstrates the required processing stages rather than claiming to reproduce that cost figure. Likewise, no output quality or model result is claimed here: image preparation and request assembly were not executed or verified under the run policy.</p><h2>Tagged score parsing and resumable inference orchestration</h2><p>What should happen when a vision-language model (VLM) is asked for one number but returns an explanation, an invalid value, or nothing parseable? A reliable reproduction should not silently guess. It should preserve the raw response, record how parsing ended, and only expose a numeric prediction when that value satisfies the task contract.</p><p>The paper asks the VLM to estimate a 30-trading-day return. In the scaffold, the requested score is therefore associated with the same target definition used by the label pipeline. The canonical target equation is:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;r_{30} = \\frac{P_{t+30} - P_t}{P_t}&quot;,&quot;id&quot;:&quot;4C7A8DD292&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here, <code>r_{30}</code> is a scalar real-valued return, <code>P_t</code> is the positive closing price at the cutoff, and <code>P_{t+30}</code> is the positive closing price 30 subsequent trading-day positions later. The parser does not calculate this target from chart images; it validates the model's reported scalar estimate. The equation's implementation mapping remains <code>construct_forward_return()</code> and <code>forward_return_label_construction()</code> in <code>src/vlm_candlestick/labels/forward_returns.py</code>, while <code>parse_prediction_record()</code> preserves the model-side estimate and its status.</p><p>The paper constrains a valid output to the interval <code>[-0.5, 1.0]</code>, rounded to three decimal places. A positive score represents a bullish prediction, a negative score represents a bearish prediction, and exactly zero represents an unclear trend. The interval is a validation boundary, not a clipping instruction: an output such as <code>1.4</code> is invalid rather than automatically converted to <code>1.0</code>.</p><h3>Parser precedence and failure preservation</h3><p><code>parse_vlm_score()</code> in <code>src/vlm_candlestick/prompt/parser.py</code> applies a deterministic precedence order:</p><ol><li><p>Look for one value in the required <code>&lt;score&gt;NUM&lt;/score&gt;</code> tag.</p></li><li><p>If no tag is present, accept one supported confidence-interval form and use its first endpoint, as described by the supplied parser behavior.</p></li><li><p>Otherwise scan the complete response for one unambiguous numeric value.</p></li><li><p>Reject missing, ambiguous, non-finite, or out-of-range values.</p></li></ol><p>The confidence-interval fallback is parser tolerance rather than the expected output format: the prompt requests no additional text. It exists because the paper's implementation notes describe such a fallback, even though it is in tension with the strict response instruction.</p><p>The central range check is small and deliberately non-coercive:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6ac0396e-9090-4234-8702-c3e932d6f973&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">_MIN_SCORE: Final[float] = -0.5
_MAX_SCORE: Final[float] = 1.0


def validate_score(value: float) -&gt; float:
    """Return a finite score only when it lies in the paper's stated range."""
    try:
        score = float(value)
    except (TypeError, ValueError) as exc:
        raise ValueError("score must be numeric") from exc
    if not math.isfinite(score):
        raise ValueError("score must be finite")
    if not _MIN_SCORE &lt;= score &lt;= _MAX_SCORE:
        raise ValueError("score must lie in [-0.5, 1.0]")
    return score</code></pre></div><p><code>validate_score()</code> accepts one scalar and returns the same value as a float only when it is finite and within the permitted interval. It does not alter the value. <code>format_score()</code> calls this validator before producing the three-decimal representation requested by the prompt.</p><p>The parser records status separately from the numeric value. For example, a tagged response such as <code>&lt;score&gt;0.125&lt;/score&gt;</code> can produce a score of <code>0.125</code> with status <code>parsed_tagged</code>. A response containing multiple plausible numbers receives an ambiguity status, and an out-of-range response receives an <code>out_of_range</code> status. In both failure cases, the original response remains available.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;87343713-6de0-4ceb-8e77-6bb0458ba953&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def parse_prediction_record(
    response: str,
    key: SampleKey,
    model: str,
) -&gt; PredictionRecord:
    """Create a prediction record without discarding an unparseable response."""
    score, status = parse_vlm_score(response)
    # Eq. 1: the prompt's requested score represents the 30-trading-day return.
    return PredictionRecord(
        key=key,
        model=model,
        raw_response=response,
        parsed_score=score,
        status=status,
    )</code></pre></div><p>A <code>PredictionRecord</code> contains the sample identity, model name, raw response, optional scalar score, and parser status. This makes malformed outputs auditable instead of making them disappear during preprocessing. Later evaluation can exclude records whose <code>parsed_score</code> is <code>None</code> while still retaining them for error analysis.</p><p>A small worked example illustrates the distinction:</p><ul><li><p><code>&lt;score&gt;0.125&lt;/score&gt;</code> is a valid tagged score.</p></li><li><p><code>&lt;score&gt;1.400&lt;/score&gt;</code> is rejected as out of range; it is not clipped.</p></li><li><p><code>The likely range is -0.1 to 0.2</code> may use the documented interval fallback and retain the first endpoint.</p></li><li><p><code>The chart is bullish, perhaps 0.1 or 0.2</code> is ambiguous and should not silently select either number.</p></li><li><p><code>No numerical estimate</code> receives a no-number status.</p></li></ul><p>These are parsing behaviors defined by the generated implementation. They are not claims that a particular commercial VLM will produce any of these responses.</p><h3>A provider boundary without an invented API</h3><p>The paper evaluates commercial VLMs, but the supplied material does not define authentication, request parameters, endpoint formats, retries, or rate-limit behavior. The package therefore defines a provider-neutral <code>VLMProvider</code> protocol. Its responsibility is narrow: accept a <code>VLMRequest</code> containing the paired image payloads and prompt, then return raw response text. Parsing remains a separate step.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c2e8ead0-e81e-49a8-8ef5-79b17811364b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@runtime_checkable
class VLMProvider(Protocol):
    """Provider-neutral contract for local or externally adapted inference."""

    def infer(self, request: VLMRequest) -&gt; str:
        """Return raw response text for one request."""
        ...</code></pre></div><p><code>ReplayProvider</code> implements this contract using a local mapping from sample-key strings to recorded response text. It is useful for demonstrations and for developing the persistence path without network access. A missing key raises <code>KeyError</code>, so absent replay data cannot be mistaken for a model prediction.</p><p><code>CallableProvider</code> adapts a local Python callable with the same responsibility. Neither class authenticates, contacts a vendor, parses a score, or fabricates a response. A future provider-specific adapter could implement <code>VLMProvider</code>, but its API behavior would need to come from additional provider documentation rather than from this paper scaffold.</p><h3>Keys, batches, and resume behavior</h3><p>Large inference runs need more than a parser. They need a stable notion of identity. In <code>src/vlm_candlestick/inference/persistence.py</code>, <code>make_prediction_key()</code> combines the model name with the sample key:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;06c34061-77a5-423b-8afd-c0bb170b9430&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def make_prediction_key(record: PredictionRecord) -&gt; str:
    """Return the unique model/sample identity for a prediction record."""
    if not isinstance(record, PredictionRecord):
        raise TypeError("record must be a PredictionRecord")
    return f"{record.model}:{record.key.as_string()}"</code></pre></div><p>The resulting processed key distinguishes, for example, one model's prediction for an HS300 stock and cutoff from another model's prediction for that same sample. The paired daily and weekly images do not receive separate prediction keys because they jointly constitute one benchmark input.</p><p><code>ResultStore</code> persists prediction records in CSV form. Its schema includes market, stock, cutoff date, model, raw response, parsed score, and status. <code>load_processed_keys()</code> reads existing rows and treats both successful and failed records as processed. This is an explicit resume policy: a failed response remains auditable and is not silently replaced on the next run. Retrying failures would require a separate, explicit policy.</p><p>The paper specifies batches of ten entries. <code>append_result_batch()</code> accepts at most ten records, rejects duplicate keys within a batch, skips keys already on disk, and permits a final partial batch. The store uses an in-process re-entrant lock when a shared <code>ResultStore</code> instance performs read-and-append operations. Cross-process locking, crash-atomic replacement, and recovery after a process crash are not claimed because their exact policies are not supplied.</p><h3>Bounded orchestration</h3><p><code>run_inference_jobs()</code> connects the provider, parser, and store. Each <code>InferenceJob</code> contains one <code>VLMRequest</code>, a model name, and the corresponding <code>SampleKey</code>. Before submitting a job, the orchestrator checks persisted keys and acquires a <code>JobLockRegistry</code> reservation. A duplicate job is skipped if it is already persisted or currently reserved in the process.</p><p>The method card specifies at most five concurrent jobs and ten-result persistence batches. The generated orchestration delegates the worker bound and batch size to <code>InferenceConfig</code>, whose default policy is five workers and ten entries. Results are accumulated, flushed whenever ten are available, and flushed again at the end for a partial batch.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8e96b83d-5043-4b57-bb09-0875e20a3989&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def _flush_records(
    pending: list[PredictionRecord],
    store: ResultStore,
    batch_size: int,
) -&gt; None:
    while len(pending) &gt;= batch_size:
        batch = tuple(pending[:batch_size])
        del pending[:batch_size]
        store.append_batch(batch)</code></pre></div><p>The provider call is wrapped so ordinary provider failures become records with status <code>inference_error</code> and preserved exception text. This keeps the job visible to later analysis. The code intentionally does not add retries or rate-limit handling, because neither behavior is specified in the supplied paper context.</p><p>At the command-line boundary, <code>run_manifest_inference()</code> in <code>src/vlm_candlestick/pipeline/run_inference.py</code> reads a local manifest, resolves chart paths, converts the daily and weekly images through the transport helper, builds a <code>VLMRequest</code>, and creates inference jobs. Its default command-line provider mode is <code>replay</code>. Thus the local path demonstrates the complete data flow without claiming live VLM execution.</p><p>A typical local invocation is conceptually:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;f86c5fe3-1d62-4032-b0f5-3ead4622d149&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">python -m vlm_candlestick.pipeline.run_inference \
    --manifest artifacts/manifest.jsonl \
    --output artifacts/predictions.csv \
    --provider replay \
    --replay responses.json</code></pre></div><p>The replay JSON maps sample-key strings to raw response text. The output CSV then preserves both successfully parsed scores and failure records. This arrangement is especially useful when a later evaluation needs to distinguish &#8220;the model predicted zero&#8221; from &#8220;the response could not be parsed&#8221;: the former is a valid scalar expressing uncertainty, while the latter has <code>parsed_score</code> set to <code>None</code> and a failure status.</p><h3>What this scaffold does&#8212;and does not&#8212;guarantee</h3><p>The implementation preserves four important invariants:</p><ul><li><p>A valid score is finite and lies in <code>[-0.5, 1.0]</code>.</p></li><li><p>A raw response is retained whether parsing succeeds or fails.</p></li><li><p>A model/sample identity is not written twice by the same resumable store.</p></li><li><p>Normal completion persists both full ten-entry batches and the final partial batch.</p></li></ul><p>The planned <code>tests/test_inference_persistence.py</code> covers processed-key skipping, duplicate protection, ten-entry and final-batch flushing, failure retention, and the five-worker configuration limit. Under the authoritative run policy, those tests were not executed, and no static or semantic verification was performed. The code excerpts above describe the generated interfaces and intended control flow, not verified runtime behavior.</p><p>Finally, the provider-neutral boundary is a reproduction decision, not a claim that commercial inference is unnecessary. Exact live reproduction would additionally require model release identifiers, credentials, vendor request schemas, retry rules, and rate-limit policies. Those details are outside the supplied evidence, so the local replay path is the honest default for demonstrating parsing and resumable orchestration.</p><h2>Numerical baseline: aligned windows and explicit XGBoost decisions</h2><p>How can a numerical model receive the same forecasting opportunity as the VLM without seeing the chart pixels&#8212;or any future observations? The baseline must use the same stock, cutoff date, and forward-return label as the visual sample. Its inputs are historical daily and weekly OHLCV windows, transformed into a numerical feature vector before an optional XGBoost regressor is fitted.</p><p>This is a fairness and leakage-control requirement, not merely a data-format preference. A chart sample and its numerical counterpart should agree on the sample identity and target. The paper identifies an XGBoost baseline based on OHLCV-derived time-series features, but the supplied paper context does not specify the complete feature vector, window length, train/test split, hyperparameters, random seed, or model-selection procedure. The generated code therefore provides an inspectable local baseline scaffold rather than claiming an exact reconstruction of the paper's baseline.</p><h3>Keep the target aligned with the visual sample</h3><p>The baseline uses the same forward-return definition as the visual task. The paper's canonical target is:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;r_{30} = \\frac{P_{t+30} - P_t}{P_t}&quot;,&quot;id&quot;:&quot;60F0699D55&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here, <code>r_{30}</code> is a scalar real-valued 30-trading-day return, <code>P_t</code> is the positive closing price on the cutoff date, and <code>P_{t+30}</code> is the positive closing price at the 30th subsequent trading-day position. The implementation mapping is <code>construct_forward_return()</code> and <code>forward_return_label_construction()</code> in <code>src/vlm_candlestick/labels/forward_returns.py</code>. The same scalar label must be attached to the visual sample and to the numerical training row.</p><p>The equation does not say to add 30 calendar days. Weekends, holidays, and other absent dates are handled by the trading-position lookup. A sample without the required future observation receives no valid label under the scaffold's explicit missing-label behavior. The five-day sensitivity path uses the same parameterized function with <code>horizon=5</code>; it is an implementation variant for the paper's horizon analysis, not a second supplied canonical equation.</p><h3>Build strictly historical daily and weekly windows</h3><p>The first baseline stage is <code>src/vlm_candlestick/baseline/windows.py</code>. <code>extract_history_window()</code> validates the OHLCV frame, keeps rows whose dates are strictly earlier than the cutoff, and returns the most recent configured rows. <code>build_aligned_baseline_window()</code> performs the same operation for daily data and constructs a weekly history only after removing post-cutoff daily observations.</p><p>That ordering matters. If daily records after the cutoff were aggregated into a weekly candle first, the resulting weekly open, high, low, close, or volume could contain future information. The implementation instead filters first and aggregates second:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;06a72644-5fdf-4e4c-bf81-d1195b3bbd24&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">historical = normalized.loc[
    normalized["date"] &lt; cutoff_timestamp
].reset_index(drop=True)

daily_window = historical.tail(length).reset_index(drop=True)

# Aggregate only pre-cutoff rows so a partial calendar week cannot leak
# post-cutoff OHLCV values into the numerical baseline.
weekly_history = aggregate_weekly_ohlcv(historical, weekly_policy)
weekly_window = weekly_history.loc[
    weekly_history["date"] &lt; cutoff_timestamp
].tail(length).reset_index(drop=True)</code></pre></div><p>The important invariant is that every row in both returned frames is earlier than the cutoff. <code>validate_window_alignment()</code> checks required columns, chronological order, duplicate dates, valid dates, and the strict cutoff predicate. The daily and weekly windows can have fewer rows than requested when the available history is short. The requested length is therefore a configuration decision, not a value recovered from the paper.</p><p><code>WeeklyResampleConfig</code> also records unresolved weekly conventions. The paper describes standard OHLCV aggregation, but the supplied context does not settle the week boundary, label-date convention, holiday handling, or treatment of incomplete weeks. Those choices should remain visible in configuration and metadata rather than being mistaken for paper facts.</p><h3>Turn each window into an inspectable vector</h3><p>The generated <code>FeatureSchema</code> in <code>src/vlm_candlestick/baseline/features.py</code> makes the local feature decision explicit. Its defaults use fields <code>open</code>, <code>high</code>, <code>low</code>, <code>close</code>, and <code>volume</code> for both <code>daily</code> and <code>weekly</code> windows, optionally adding <code>close_return</code>. It also records a window length and a padding value. These are implementation choices because the paper does not provide the complete numerical feature specification.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c03c7271-c75c-462c-acf5-f2e285980fad&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True, slots=True)
class FeatureSchema:
    """Configuration and provenance for the numerical feature vector."""

    window_length: int = 50
    fields: tuple[str, ...] = _DEFAULT_FIELDS
    frequencies: tuple[str, ...] = _DEFAULT_FREQUENCIES
    include_close_returns: bool = True
    pad_value: float = 0.0</code></pre></div><p><code>transform_ohlcv_window()</code> consumes one chronological historical window and returns a one-dimensional pandas <code>Series</code>. Price fields are expressed relative to the latest close in that window, volume uses a <code>log1p</code> transformation scaled by the largest transformed volume, and optional close returns are appended. Short windows are left-padded with the configured value so the vector has stable length. The transformation is deliberately inspectable and does not consume the target, a prediction, or any post-cutoff value.</p><p>For a schema with window length <code>L</code>, <code>F</code> row fields, and <code>R</code> frequencies, the feature vector has <code>R &#215; L &#215; F</code> values, where <code>F</code> includes <code>close_return</code> when enabled. In code, the schema exposes this contract through <code>dimension_per_frequency</code> and <code>dimension</code>. A single sample is represented as a one-dimensional vector; a collection of samples forms a feature matrix <code>X</code> with shape <code>(N, D)</code>, where <code>N</code> is the number of samples and <code>D</code> is the schema-dependent feature dimension.</p><p><code>build_baseline_features()</code> combines the frequency blocks in the declared schema order. It prefixes feature names with the frequency, such as <code>daily_t000_open</code> or <code>weekly_t019_close_return</code>, so the provenance of each value remains visible. <code>feature_schema_manifest()</code> serializes the selected fields, transformations, padding policy, dimensions, and the explicit statement that the feature vector is an implementation decision.</p><h3>Worked example: inspect one aligned feature vector</h3><p>Suppose a local synthetic stock has a cutoff at one of its historical rows. The baseline workflow is conceptually:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ff4b76c2-69cc-4aaf-b681-3fac2ff5c54a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">windows = build_aligned_baseline_window(
    frame=frame,
    cutoff=cutoff,
    length=20,
    weekly_policy=WeeklyResampleConfig(),
)

schema = FeatureSchema(window_length=20)
features = build_baseline_features(windows, schema)
provenance = feature_schema_manifest(schema)</code></pre></div><p>The first call returns two frames, <code>windows["daily"]</code> and <code>windows["weekly"]</code>, both strictly pre-cutoff. The second call converts those frames into one stable feature vector, and the third records how that vector was made. This is useful when comparing experiments: a prediction without its feature-schema metadata is difficult to audit, especially when the source paper leaves the baseline features unspecified.</p><p>The generated test plan <code>tests/test_horizon_and_baseline.py</code> covers this intended contract with <code>test_baseline_windows_exclude_future_rows()</code> and <code>test_feature_schema_is_deterministic()</code>. These are planned tests using deterministic synthetic data; test execution and verification were disabled by the run policy, so no passing result is claimed here.</p><h3>Fit the optional XGBoost wrapper</h3><p>The wrapper in <code>src/vlm_candlestick/baseline/xgboost_model.py</code> separates feature construction from model fitting. <code>XGBoostBaseline.fit()</code> accepts a finite numeric DataFrame <code>X</code> of shape <code>(N, D)</code> and a finite target vector <code>y</code> of shape <code>(N,)</code>. It records the exact feature-column order used during fitting. <code>predict()</code> requires the same columns in the same order and returns a one-dimensional float array with shape <code>(M,)</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;0462502d-9775-45af-be1c-90d50fc46da1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">model = XGBoostBaseline(config)
model.fit(X, target)

predictions = model.predict(X_for_prediction)</code></pre></div><p>The wrapper imports XGBoost only when fitting is requested. If the optional dependency is unavailable, fitting raises an explicit import error rather than silently substituting another model. Invalid matrices, mismatched row counts, duplicate feature names, non-finite values, and feature-column reorderings are rejected. These checks protect the local data contract; they do not establish that the wrapper reproduces the paper's training procedure.</p><p>The convenience functions <code>fit_xgboost_baseline()</code> and <code>predict_xgboost()</code> provide the planned public interface. Configuration values are passed defensively: values supplied through a baseline configuration may be forwarded, while missing values are left to the local XGBoost library defaults. Those defaults are not paper specifications. In particular, the scaffold does not claim the paper's train/test split, hyperparameters, random seed, or model-selection process.</p><h3>What can and cannot be compared</h3><p>The comparison is meaningful only when the visual and numerical rows share market, stock, cutoff, and target identity. Both modalities must exclude observations at or after the cutoff, and both must use the same forward-return construction. Beyond those invariants, the numerical baseline remains underdetermined by the supplied material.</p><p>The paper's reported tables may be used as reference metadata, but they are not independently verified outputs of this scaffold. Likewise, a reported advantage for charts should not be encoded as a universal conclusion: the supplied summary indicates that comparisons can depend on the metric and configuration. A later reproduction would need the missing feature, split, and training specifications before treating an XGBoost result as directly comparable.</p><p>The practical result is an honest baseline boundary: the repository can demonstrate aligned windows, explicit feature provenance, matrix-shape checks, and an optional local XGBoost fit, while clearly labeling the choices that were made because the paper did not provide them.</p><h2>Directional confusion metrics and prediction bias</h2><p>How can a continuous forecast become an auditable statement about direction? A VLM may return a score such as <code>0.125</code> or <code>-0.080</code>, but a confusion matrix needs a discrete answer: rise or fall. That conversion requires a boundary. Because the supplied paper does not state the boundary or the treatment of values exactly at it, the reproduction scaffold makes both choices explicit instead of silently inferring them.</p><p>The paper's task is regression: the model emits a continuous return estimate. Its confusion-matrix analysis is therefore a secondary evaluation view. It converts predictions and actual returns into directional classes, counts four outcomes, and derives accuracy, precision, recall, specificity, and F1. The same directional classes also support prediction-bias analysis, which compares how often a model predicts upward movement with how often upward movement actually occurs.</p><h3>Make the unresolved direction policy visible</h3><p><code>DirectionConfig</code> stores the classification policy. In the generated implementation, <code>threshold=None</code> means that classification is unresolved and must be supplied by the caller. The default <code>zero_policy="exclude"</code> is also a configuration choice, not a paper fact.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8a899cfa-5dd7-439f-ac32-747466f41851&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class DirectionConfig:
    """Continuous-to-direction policy.

    The supplied paper does not state the directional threshold or treatment
    of zero returns. ``threshold=None`` deliberately represents that unresolved
    state and must be replaced before classification.
    """

    threshold: float | None = None
    zero_policy: ZeroPolicy = "exclude"</code></pre></div><p>In this code, the threshold is applied to both predictions and actual returns. Values greater than the threshold receive the upward class <code>1</code>; values below it receive the downward class <code>0</code>. Values exactly equal to the threshold receive class <code>1</code>, class <code>0</code>, or the exclusion sentinel <code>-1</code>, according to <code>zero_policy</code>. The sentinel preserves the original vector shape while allowing later metric functions to omit those observations.</p><p>The <code>classify_direction()</code> function accepts a one-dimensional numeric vector with shape <code>(N,)</code> and returns another vector with shape <code>(N,)</code>. It validates finiteness and requires a resolved threshold before doing any comparison.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a9d70fc4-bfae-40a1-8618-77b0ef5e6147&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">classes = np.full(numeric_values.shape, DOWN_CLASS, dtype=np.int8)
classes[numeric_values &gt; threshold] = UP_CLASS

equal = numeric_values == threshold
if config.zero_policy == "up":
    classes[equal] = UP_CLASS
elif config.zero_policy == "down":
    classes[equal] = DOWN_CLASS
else:
    classes[equal] = EXCLUDED_CLASS</code></pre></div><p>This separation matters operationally. A reproduction can now report, for example, that it used a threshold of <code>0.0</code> and classified equality as down, rather than presenting that choice as if it came from the paper. <code>direction_counts()</code> reports upward, downward, excluded, classified, and total counts so the effect of exclusion remains visible.</p><h3>Count the four confusion outcomes</h3><p>For paired predicted and actual direction vectors, the scaffold uses the paper's conventions:</p><ul><li><p><strong>TP</strong>, or true positive: predicted rise and actual rise.</p></li><li><p><strong>TN</strong>, or true negative: predicted fall and actual fall.</p></li><li><p><strong>FP</strong>, or false positive: predicted rise and actual fall or non-rise.</p></li><li><p><strong>FN</strong>, or false negative: predicted fall and actual rise or non-fall.</p></li></ul><p><code>compute_confusion_counts()</code> removes a pair whenever either side has the exclusion sentinel. It then counts the four remaining combinations. Its <code>ConfusionCounts</code> record enforces the key partition invariant: <code>TP + TN + FP + FN</code> equals the evaluated sample count <code>N</code>. Here, <code>N</code> excludes observations omitted by the configured zero policy, while the original total and exclusion counts remain available in the evaluation metadata.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5138ee63-bcad-4f3b-b3f8-ada6c06ded90&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">tp = int(np.count_nonzero((predicted_valid == UP_CLASS) &amp; (actual_valid == UP_CLASS)))
tn = int(np.count_nonzero((predicted_valid == DOWN_CLASS) &amp; (actual_valid == DOWN_CLASS)))
fp = int(np.count_nonzero((predicted_valid == UP_CLASS) &amp; (actual_valid == DOWN_CLASS)))
fn = int(np.count_nonzero((predicted_valid == DOWN_CLASS) &amp; (actual_valid == UP_CLASS)))
return ConfusionCounts(tp=tp, tn=tn, fp=fp, fn=fn, n=int(valid.sum()))</code></pre></div><p>The paper labels the accuracy, precision, recall, specificity, and F1 displays as equations <code>eq_2</code> through <code>eq_6</code>. Their extracted canonical LaTeX is empty because of OCR damage, so this tutorial does not reproduce those formulas as equations. Instead, <code>compute_confusion_metrics()</code> maps the supplied meanings directly to guarded calculations: <code>eq_2</code> to accuracy, <code>eq_3</code> to precision, <code>eq_4</code> to recall, <code>eq_5</code> to specificity, and <code>eq_6</code> to F1.</p><p>The function does not turn an undefined denominator into an arbitrary zero. If there are no predicted positives, precision is undefined; if there are no actual positives, recall is undefined; and related cases can make F1 undefined. The result stores <code>None</code> for such metrics and names them in <code>metadata["undefined_metrics"]</code> and in warnings.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f703a5da-4d4f-43db-b0ac-a09fdbc98806&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">accuracy = _safe_ratio(counts.tp + counts.tn, counts.n, "accuracy", undefined)
precision = _safe_ratio(counts.tp, counts.tp + counts.fp, "precision", undefined)
recall = _safe_ratio(counts.tp, counts.tp + counts.fn, "recall", undefined)
specificity = _safe_ratio(counts.tn, counts.tn + counts.fp, "specificity", undefined)</code></pre></div><p>The public <code>evaluate_confusion_matrix()</code> function connects the two stages. It validates prediction and target arrays of shape <code>(N,)</code>, classifies both with the same <code>DirectionConfig</code>, computes counts, and adds the threshold, zero policy, class-array shapes, and excluded counts to the result. Thus the metric record contains both the scores and the policy needed to interpret them.</p><h3>Worked example: hand-count the outcomes</h3><p>Suppose we choose <code>threshold=0.0</code> and <code>zero_policy="down"</code>. Consider these predicted and actual returns:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c4b140d0-b2e6-4a1e-a69e-9730f734fbaa&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">predictions = np.array([0.10, 0.02, -0.04, -0.01, 0.08, -0.02])
actuals = np.array([0.06, -0.03, -0.01, 0.04, 0.02, -0.05])
config = DirectionConfig(threshold=0.0, zero_policy="down")

result = evaluate_confusion_matrix(predictions, actuals, config)</code></pre></div><p>The predicted classes are up, up, down, down, up, down. The actual classes are up, down, down, up, up, down. Therefore, <code>TP=2</code> for positions 1 and 5, <code>TN=2</code> for positions 3 and 6, <code>FP=1</code> for position 2, and <code>FN=1</code> for position 4. The four counts sum to six, the number of evaluated observations. The resulting accuracy is the fraction of the six cases that are correct; the other named metrics describe different aspects of positive-case reliability and negative-case recognition.</p><p>This example is not a reported paper result. It only demonstrates how a selected direction policy changes the interpretation of continuous values.</p><h3>Compare directional distributions as bias</h3><p>A confusion matrix evaluates paired correctness. Prediction bias asks a different question: does the model produce too many upward or downward calls overall? The paper's <code>eq_7</code> addresses this distribution comparison and interprets positive bias as optimistic and negative bias as pessimistic. The supplied equation record has no canonical LaTeX and its denominator structure is OCR-damaged, so its exact normalization cannot be recovered from the provided material.</p><p>The generated <code>compute_prediction_bias()</code> function therefore records an explicit local policy: predicted-up proportion minus actual-up proportion, computed on the common classified observations. This is an implementation decision, not a claim that the damaged paper equation has been reconstructed. The result records both <code>normalization_requested</code> and <code>normalization_used</code>, together with raw distribution counts and proportions.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;eadbf0e4-baaa-4037-9734-b76ebfb201d3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">predicted_up_proportion = predicted_counts["up"] / sample_count
actual_up_proportion = actual_counts["up"] / sample_count
# Eq. 7: directional distribution bias under the explicit local policy.
bias = float(predicted_up_proportion - actual_up_proportion)
absolute_bias = abs(bias)
category = bias_category(absolute_bias)</code></pre></div><p>For the example above, the model predicts three upward cases out of six, and the actual series also contains three upward cases. Under this selected policy, the bias is zero and the interpretation is neutral. To see a nonzero case, compare predicted classes <code>[1, 1, 1, 0]</code> with actual classes <code>[1, 0, 0, 0]</code>. The predicted-up proportion is <code>0.75</code>, the actual-up proportion is <code>0.25</code>, and the local bias is <code>0.50</code>, which is positive and therefore optimistic.</p><p>The calibration bands supplied by the paper are applied to absolute bias:</p><ul><li><p>absolute bias up to and including <code>0.05</code>: well-calibrated;</p></li><li><p>above <code>0.05</code> through <code>0.15</code>: slightly biased;</p></li><li><p>above <code>0.15</code> through <code>0.40</code>: moderately biased;</p></li><li><p>above <code>0.40</code>: strongly biased.</p></li></ul><p><code>bias_category()</code> implements those boundaries. It accepts a scalar magnitude and rejects negative, non-finite, or nonnumeric values. The bias function separately handles the no-data case: if no common classified observations remain, it returns undefined metrics and an explicit category rather than assigning a misleading zero.</p><h3>What the tests specify&#8212;and what they do not prove here</h3><p>The planned <code>tests/test_evaluation_metrics.py</code> file captures the intended invariants. For example, it checks the count partition and the named equation mappings without pretending that damaged equation text is available:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4f5d15c6-b5bf-47a0-9587-178bbf4679b0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">assert counts.tp + counts.tn + counts.fp + counts.fn == len(predicted)

result = compute_confusion_metrics(counts)
assert result.metadata["equation_mappings"] == {
    "eq_2": "accuracy",
    "eq_3": "precision",
    "eq_4": "recall",
    "eq_5": "specificity",
    "eq_6": "f1",
}</code></pre></div><p>The same planned test file checks the bias interval boundaries and the explicit <code>eq_7</code> policy. These are implementation contracts for the scaffold, not evidence that the paper's reported confusion-matrix tables or bias values have been reproduced. The class threshold, equality treatment, damaged bias denominator, missing-response handling, and evaluation subsets remain major barriers to exact reproduction.</p><p>Under the authoritative run policy, these tests were not executed. Static verification was disabled, semantic code verification was skipped, and no numerical or result verification was performed. The important practical outcome of this design is therefore not a claimed metric value; it is an auditable path from continuous predictions to directional classes, counts, guarded metrics, and explicitly qualified bias summaries.</p><h2>IC, Rank IC, regimes, extreme stocks, and horizon sensitivity</h2><p>How can we tell whether a model's continuous predictions align with realized returns beyond simply asking whether the direction was correct? A confusion matrix reduces each observation to rise or fall. The analyses in this section retain more information: they measure correlation within groups, inspect market-wide and stock-specific extremes, and compare the requested 30-trading-day outcome with a shorter five-day outcome.</p><p>The paper calls the ordinary grouped Pearson correlation <strong>IC</strong>, or information coefficient. <strong>Rank IC</strong> is the corresponding grouped Spearman rank association: instead of comparing raw values directly, it compares their within-group order. For example, if predictions rank three stocks in the same order as their realized returns, Rank IC can be high even when the numerical scales differ.</p><p>These metrics require an explicit grouping axis. A common choice is one cross-section per cutoff date: predictions for all stocks on a date are compared with those stocks' realized returns, producing one IC value for that date. The supplied paper extraction does not fully resolve whether every IC series uses this date grouping or another time-series grouping. The generated code therefore takes the grouping values as an input and records the selected policy in <code>EvaluationConfig</code>.</p><h3>Compute grouped IC and Rank IC</h3><p><code>validate_aligned_correlation_inputs()</code> in <code>src/vlm_candlestick/evaluation/correlation.py</code> protects the most important data contract: predictions, actual returns, and group labels must all be one-dimensional pandas Series with equal lengths and identical indexes. This prevents a silent row permutation when the three inputs originate from separate tables.</p><p><code>compute_group_ic()</code> then removes missing prediction/return pairs within each group and computes Pearson correlation. <code>compute_group_rank_ic()</code> applies average ranks within each valid group before computing Pearson correlation on those ranks. Groups with too few valid pairs or constant variation are undefined. By default, the implementation omits them; an <code>EvaluationConfig</code> with <code>drop_invalid_groups=False</code> can retain them with a null correlation.</p><p>The central calls are deliberately small:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;cf4b14e9-766a-4936-a688-2de068bb3cfe&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">ic_frame = compute_group_ic(
    predictions, actuals, groups, config=config
)
rank_frame = compute_group_rank_ic(
    predictions, actuals, groups, config=config
)</code></pre></div><p>Each returned frame contains <code>group</code>, <code>correlation</code>, and <code>n</code>. The correlation is a scalar in the interval <code>[-1, 1]</code> when defined, while <code>n</code> records the number of valid paired observations in that group. A null value means that the group did not contain enough usable variation; it does not mean zero predictive skill.</p><h3>Summarize IC, Rank IC, and ratios</h3><p>The paper reports mean, median, ICIR, Rank ICIR, and significance ratios. In the scaffold, <code>summarize_correlation_series()</code> in <code>src/vlm_candlestick/evaluation/summaries.py</code> computes the mean and median of a one-dimensional series of valid group correlations. It also computes a ratio named ICIR&#8212;or Rank ICIR, depending on the caller&#8212;as the mean divided by a configured standard deviation.</p><p>The standard-deviation convention is not supplied by the paper extraction. <code>EvaluationConfig</code> therefore distinguishes <code>sample</code> and <code>population</code> conventions. A series with no valid values, fewer than two values under the sample convention, or zero standard deviation receives an explicit null ratio rather than an infinite value or an invented replacement.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;72f45b68-ba68-4194-9502-eb32c04c21be&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">summary = summarize_correlation_series(ic_values, config)

paired = summarize_ic_and_rank_ic(
    ic_values,
    rank_ic_values,
    config,
)</code></pre></div><p>The result records valid-group counts, undefined metrics, the standard-deviation convention, the grouping policy, and the significance policy. This metadata matters because two implementations can produce different ICIR values from the same IC series if they use different denominators for standard deviation.</p><p><code>compute_significance_ratio()</code> is similarly policy-driven. With <code>significance_method="none"</code>, it returns no significance ratio because the supplied paper context does not identify the statistical test. The optional local <code>positive_fraction</code> policy is descriptive: it reports the fraction of valid correlation values strictly above a configured threshold. It is not presented as the paper's missing significance test.</p><h3>Label market regimes</h3><p>The paper analyzes bull, bear, and sideways market conditions using cross-sectional daily returns. A date is labeled <code>bull</code> when strictly more than 70 percent of classified stocks rise, and <code>bear</code> when strictly more than 70 percent fall. All other dates are <code>sideways</code>.</p><p><code>compute_market_regime_labels()</code> in <code>src/vlm_candlestick/evaluation/regimes.py</code> accepts a date-by-stock return matrix. It applies the configured directional policy, removes missing or excluded observations from the classified denominator, and returns one row per date with the upward and downward fractions, counts, and final regime. A date with no classified observations is rejected because its regime is undefined.</p><p>The strict comparison is visible in the generated implementation:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5f2de958-f387-422f-a109-b0a7da0d0789&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># The paper specifies strict greater-than 70% bull/bear thresholds.
if up_fraction &gt; _MARKET_REGIME_THRESHOLD:
    regime: Regime = "bull"
elif down_fraction &gt; _MARKET_REGIME_THRESHOLD:
    regime = "bear"
else:
    regime = "sideways"</code></pre></div><p>This means exactly 70 percent is not enough for either extreme regime. <code>filter_by_market_regime()</code> can then select prediction records whose dates have a requested label, while <code>market_regime_counts()</code> provides stable counts for <code>bull</code>, <code>bear</code>, <code>sideways</code>, and <code>total</code>.</p><p>The universe used to form the cross-section, treatment of missing stocks, and zero-return policy remain implementation decisions. For example, excluding threshold-equal returns changes the denominator and can change a date's regime. Those choices must remain visible in <code>DirectionConfig</code> and the evaluation report rather than being hidden inside a metric function.</p><h3>Label extreme stocks</h3><p>The extreme-stock analysis turns the same idea around. Instead of asking whether most stocks moved together on a date, it asks whether one stock moved upward or downward on more than 70 percent of its test-period observations.</p><p><code>compute_extreme_stock_labels()</code> in <code>src/vlm_candlestick/evaluation/extremes.py</code> expects dates as rows and stocks as columns. For each stock it reports <code>rising_fraction</code>, <code>falling_fraction</code>, <code>classified_count</code>, <code>observation_count</code>, and a category. A stock is <code>rising</code> or <code>falling</code> only when the relevant fraction is strictly greater than 0.70. Stocks meeting neither condition are <code>other</code>; stocks with no classified observations are <code>insufficient_data</code>.</p><p><code>filter_by_extreme_stock()</code> supports records that identify stocks either through a <code>stock</code> column or through their index. <code>extreme_stock_counts()</code> returns counts for all four categories, including categories with no members. This stable output is useful when assembling comparable reports for several models or subsets.</p><p>The generated test plan uses a hand-checkable ten-observation example: seven upward observations produce exactly 70 percent and therefore remain <code>other</code>, while eight upward observations produce a <code>rising</code> label. The test also checks the analogous falling cases. These are planned tests only; no test execution occurred under the run policy.</p><h3>Compare five-day and 30-day horizons</h3><p>The benchmark's primary label is the 30-trading-day forward return. The paper also studies time sensitivity by comparing predictions with five-trading-day outcomes. This comparison does not replace the primary target. It asks whether a model's predictions are more closely associated with a shorter realized horizon than with the horizon requested in the prompt.</p><p>The canonical target definition is the paper's supplied equation. It maps the cutoff close to the close at a later trading-day position:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;r_{30} = \\frac{P_{t+30} - P_t}{P_t}&quot;,&quot;id&quot;:&quot;6B6B7D835F&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here, <code>r_{30}</code> is a scalar real-valued 30-trading-day return, <code>P_t</code> is the positive closing price at the cutoff, and <code>P_{t+30}</code> is the positive closing price 30 subsequent trading-day positions later. In code, this definition maps to <code>construct_forward_return()</code> and <code>forward_return_label_construction()</code> in <code>src/vlm_candlestick/labels/forward_returns.py</code>. The horizon-sensitivity implementation calls the same constructor with <code>horizon=5</code> for the alternate outcome. The paper does not supply a separate canonical five-day equation.</p><p><code>construct_horizon_returns()</code> in <code>src/vlm_candlestick/evaluation/horizons.py</code> preserves sample identity by returning a one-dimensional Series indexed by <code>SampleKey</code> strings. A missing cutoff or insufficient future history becomes <code>NaN</code>, not a fabricated return. The same cutoff close is used for both horizons, and neither horizon may use a row at or before the cutoff as its future observation.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;574369f1-0508-4006-80af-7fcc26c76468&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">returns_5 = construct_horizon_returns(
    frames, materialized_keys, horizon=5
)
returns_30 = construct_horizon_returns(
    frames, materialized_keys, horizon=30
)
results = _results_frame(
    prediction_series,
    materialized_keys,
    returns_5,
    returns_30,
)
aligned = align_horizon_complete_cases(results)</code></pre></div><p><code>align_horizon_complete_cases()</code> retains incomplete rows and adds <code>complete_5</code>, <code>complete_30</code>, and <code>complete_common</code> flags. This is important because a 30-day target requires more history than a five-day target. <code>evaluate_horizon_sensitivity()</code> reports horizon-specific complete-case summaries and separately reports the stricter common complete-case subset. It groups correlations by <code>cutoff_date</code> in the generated implementation, but that grouping choice remains configurable elsewhere because the paper extraction does not fully settle the grouping axis.</p><p>A compact worked example can use the same synthetic stock and cutoff with one prediction. First construct both labels from the frame and confirm that their <code>current_close</code> fields match. Then place the prediction, five-day return, and 30-day return in one keyed result row. If the frame ends before the 30th subsequent observation, the five-day result may remain available while the 30-day result is missing. The report should show both sample counts rather than comparing the two horizons as if they used identical populations.</p><h3>End-to-end conditional evaluation</h3><p><code>run_evaluation()</code> in <code>src/vlm_candlestick/pipeline/evaluate.py</code> joins local predictions and labels using <code>market</code>, <code>stock</code>, and <code>cutoff_date</code>. It requires an explicit <code>DirectionConfig</code>, computes confusion and bias metrics, builds grouped IC and Rank IC series, and then derives market-regime and extreme-stock labels from the aligned return matrix. The command-line entry point <code>main()</code> requires the direction threshold, zero policy, grouping, standard-deviation convention, and significance method as arguments.</p><p>The intended shape flow is:</p><ul><li><p>prediction and return columns: one-dimensional vectors with length <code>N</code>;</p></li><li><p>grouped IC and Rank IC: one scalar correlation per valid group;</p></li><li><p>date-by-stock return matrix: rows are dates and columns are stocks;</p></li><li><p>regime labels: one record per date;</p></li><li><p>extreme-stock labels: one record per stock;</p></li><li><p>horizon outputs: one five-day and one 30-day scalar return per sample where available.</p></li></ul><p>This organization supports conditional reports without losing provenance. A report can state which grouping, directional threshold, zero treatment, missingness policy, standard-deviation convention, and significance method produced it.</p><h3>Interpretation limits</h3><p>The paper reports stronger sensitivity to five-day outcomes than to the requested 30-day horizon. That is a reported finding from the paper, not a result verified by this scaffold. A local run could produce a different comparison because the constituent universe, cutoff dates, model outputs, missingness rules, and grouping policies may differ.</p><p>Likewise, a high Rank IC does not by itself prove that a model understands candlestick structure. It indicates rank association between predictions and realized returns under the selected grouping and sample definition. Regime and extreme-stock analyses are conditional diagnostics, not independent evidence that a visual model has learned a particular chart concept. The generated code and its planned tests were not executed, statically checked, or semantically verified in this run, so readers should treat the implementation as an auditable reproduction scaffold rather than a verified reproduction of the paper's tables.</p><h2>End-to-end synthetic workflow and reproducibility limitations</h2><p>What does the complete reproduction path look like when no market-data provider or commercial VLM service is available? The repository answers with a small local rehearsal: generate deterministic OHLCV data, align a future-return label, render daily and weekly charts, construct a provider-neutral request, parse a replayed score, and evaluate the resulting prediction. This demonstrates interfaces and invariants; it does not reproduce the paper's real constituents, VLM calls, reported sample counts, or experimental results.</p><h3>Follow one sample from data to metric</h3><p>The end-to-end order is important:</p><ol><li><p><strong>Create or load local data.</strong> <code>generate_synthetic_ohlcv()</code> creates demonstration data, while the local loaders accept CSV files. The paper names TuShare and Yahoo Finance as sources, but acquisition, credentials, and exact constituent snapshots are outside this scaffold.</p></li><li><p><strong>Validate and align observations.</strong> OHLCV rows are normalized, and the cutoff is paired with a future trading-day position rather than a calendar offset.</p></li><li><p><strong>Construct the target.</strong> <code>forward_return_label_construction()</code> creates the scalar label shared by the visual and numerical paths.</p></li><li><p><strong>Build the visual input.</strong> <code>build_paired_charts()</code> creates one daily and one weekly PNG with the same stock, market, and cutoff identity.</p></li><li><p><strong>Persist the sample.</strong> A manifest records the two chart paths and shared label so later stages can join records safely.</p></li><li><p><strong>Prepare inference.</strong> Image transport converts each chart to RGB/JPEG/Base64, and <code>build_vlm_request()</code> places the two payloads alongside the paper-derived prompt.</p></li><li><p><strong>Replay a response locally.</strong> <code>ReplayProvider</code> supplies raw text; this is response injection, not VLM execution.</p></li><li><p><strong>Parse and evaluate.</strong> <code>parse_vlm_score()</code> validates the tagged number, and <code>evaluate_confusion_matrix()</code> applies an explicitly chosen direction policy to local predictions and labels.</p></li></ol><p>The target used in steps 3 and 7 is the paper's canonical forward-return definition:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;r_{30} = \\frac{P_{t+30} - P_t}{P_t}&quot;,&quot;id&quot;:&quot;0DE2D80BEA&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here, <code>r_{30}</code> is a scalar real-valued 30-trading-day return, <code>P_t</code> is the positive closing price at the cutoff, and <code>P_{t+30}</code> is the positive close 30 subsequent trading-day positions later. In code, the mapping is <code>construct_forward_return()</code> and its named wrapper <code>forward_return_label_construction()</code> in <code>src/vlm_candlestick/labels/forward_returns.py</code>. The chart history remains strictly before the cutoff even though the label uses the cutoff close and a later close. A five-day sensitivity label uses the same parameterized implementation pattern; it is not a separate canonical equation supplied by the paper.</p><p>The capstone example is deliberately compact. The following excerpt is copied from <code>examples/local_reproduction.py</code>; notice that it creates synthetic data, chooses local cutoffs, constructs labels, and prepares the paired chart configuration without contacting a provider.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;36d0a7c9-853b-4f1c-b938-b56705b58604&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">frame = generate_synthetic_ohlcv(
    stock=stock,
    start=date(2020, 1, 2),
    periods=180,
    seed=7,
)

# Cutoffs are dates in the local business-day demonstration series.  The
# first four have enough subsequent trading positions for a 30-day label.
cutoffs = make_demo_cutoffs(frame, interval_days=30)[:4]
if len(cutoffs) != 4:
    raise RuntimeError("the synthetic demonstration did not produce four cutoffs")

labels = []
for cutoff in cutoffs:
    label = forward_return_label_construction(frame, cutoff, horizon=30)
    if label is None:
        raise RuntimeError(f"missing synthetic 30-day label for {cutoff.isoformat()}")
    labels.append(label)

chart_config = ChartConfig(max_bars=50, ma_windows=(5, 20, 90))
weekly_policy = WeeklyResampleConfig()
output_dir = Path("artifacts") / "local_demo"</code></pre></div><p>The <code>None</code> check is a meaningful failure boundary: insufficient future history means there is no valid target for that cutoff. The <code>max_bars=50</code> and moving-average windows reflect supplied benchmark behavior, while the weekly policy and output directory remain local configuration choices.</p><p>The next excerpt shows the visual and parser portions of the same example. It is copied exactly from the generated file, including the fixed tagged response. A replayed response demonstrates the parser contract, but it must not be described as a prediction produced by a VLM.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;849a08b1-035c-4d4c-92d3-806a6a38e7ab&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">daily_chart, weekly_chart = build_paired_charts(
    frame=frame,
    stock=stock,
    cutoff=cutoffs[0],
    market=market,
    output_dir=output_dir,
    config=chart_config,
    weekly_policy=weekly_policy,
)
pair_index = index_chart_pair((daily_chart, weekly_chart))

# This is the paper-derived tagged response contract.  Parsing is local;
# it does not imply that a VLM request was made.
parsed_score, parse_status = parse_vlm_score("&lt;score&gt;0.125&lt;/score&gt;")
if parsed_score is None:
    raise RuntimeError(f"the demonstration score could not be parsed: {parse_status}")</code></pre></div><p><code>pair_index</code> is the local artifact index for one daily/weekly pair. The pair must retain one shared identity, while the two paths differ only by frequency. <code>parse_vlm_score()</code> returns both a possible scalar and a status. Invalid, out-of-range, ambiguous, or unparseable text is not silently converted into a plausible number.</p><p>Finally, the example evaluates deterministic local values under an explicit policy:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d5e09bea-13df-43bc-9d7c-f2f6f8c35b0e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># Use deterministic local scores solely to exercise the evaluation API.
# The direction threshold and zero policy are explicit implementation
# decisions because the paper does not specify their exact rule.
predictions = np.asarray([-0.100, 0.050, 0.200, -0.050], dtype=float)
actuals = np.asarray([label.return_value for label in labels], dtype=float)
direction_config = DirectionConfig(threshold=0.0, zero_policy="down")
confusion = evaluate_confusion_matrix(predictions, actuals, direction_config)</code></pre></div><p>The threshold of <code>0.0</code> and the <code>down</code> treatment of zero are implementation decisions, not recovered paper specifications. Consequently, the resulting confusion counts and metrics describe only this synthetic policy and input. They cannot establish model quality or validate any table in the paper.</p><h3>Three local pipeline entry points</h3><p>The package exposes three orchestration boundaries. <code>src/vlm_candlestick/pipeline/build_dataset.py</code> provides <code>run_dataset_build()</code> and its CLI <code>main()</code>: it loads local CSVs, proposes periodic cutoffs, builds labels and paired charts, and writes <code>manifest.jsonl</code>. Its interval and candidate dates should not be confused with the paper's unavailable exact set of 32 cutoff dates.</p><p><code>src/vlm_candlestick/pipeline/run_inference.py</code> provides <code>run_manifest_inference()</code> and a replay-only CLI. It resolves chart paths, performs image preprocessing, creates two-image requests, and delegates bounded processing to <code>run_inference_jobs()</code>. The generated CLI allows at most five workers and supports ten-entry batches, resume behavior, and local replay responses. It does not implement a commercial API schema, authentication, rate-limit handling, or live network calls.</p><p><code>src/vlm_candlestick/pipeline/evaluate.py</code> provides <code>run_evaluation()</code> and the evaluation CLI. It joins prediction and label rows by <code>market</code>, <code>stock</code>, and <code>cutoff_date</code>, then dispatches confusion, bias, grouped correlation, market-regime, and extreme-stock analyses. The command requires explicit direction, grouping, standard-deviation, and significance policies because the paper extraction does not settle those conventions. <code>EvaluationReport</code>, <code>assemble_evaluation_report()</code>, and <code>write_evaluation_report()</code> preserve metric metadata and unresolved-policy provenance.</p><h3>What the planned tests cover</h3><p>The repository contains planned mathematical and contract tests, but their presence is not evidence that they passed. <code>tests/test_labels_and_alignment.py</code> covers positional trading-day lookup, the Eq. (1) calculation, and strict pre-cutoff history. <code>tests/test_weekly_and_charts.py</code> covers weekly field ownership, the 50-bar cap, moving-average names, and daily/weekly identity.</p><p>The transport and parser checks in <code>tests/test_parser_and_transport.py</code> cover RGB conversion, white-background compositing, proportional resizing, tagged scores, range rejection, and preservation of malformed responses. <code>tests/test_inference_persistence.py</code> covers duplicate keys, resume behavior, ten-record batching, partial final batches, and retained failure records using local replay providers.</p><p>The evaluation tests in <code>tests/test_evaluation_metrics.py</code> cover confusion-count partitioning, guarded denominators, named mappings for <code>eq_2</code> through <code>eq_6</code>, the configurable <code>eq_7</code> bias implementation, strict greater-than-70-percent regime and extreme-stock thresholds, Rank IC distinction, and zero-variance summaries. <code>tests/test_horizon_and_baseline.py</code> covers shared cutoff closes for five- and thirty-day labels and no-future numerical windows. Finally, <code>tests/test_static_contracts.py</code> is designed to inspect typed public interfaces, equation-to-function mappings, explicit policy fields, and the absence of a required network client.</p><p>These files express intended invariants, including the fact that vectors have shape <code>(N,)</code>, feature matrices have shape <code>(N, D)</code>, and paired artifacts share identity. They do not turn missing paper specifications into hidden assumptions.</p><h3>Verification status and reproduction boundaries</h3><p>Under the authoritative run policy, code execution, test generation, local static verification, semantic code verification, tutorial-section verification, and final quality review were disabled. The supplied verification records therefore report skipped status. No syntax check, test run, execution result, or semantic confirmation is claimed here. A later authorized pass could perform AST and import checks, execute the planned tests, and review the data-alignment and equation mappings, but that work has not occurred in this run.</p><p>Exact reproduction also remains limited by missing source details. The supplied context does not provide the exact HS300 or S&amp;P 500 constituent snapshots, the complete 32 cutoff dates, changing-membership and missing-record rules, weekly boundaries, incomplete-week conventions, chart dimensions, fonts, rendering parameters, JPEG quality, or resize target. It also omits the complete XGBoost feature vector, window and training procedure, hyperparameters, split, seed, and model-selection rule.</p><p>Further unresolved choices affect evaluation: the continuous-to-direction threshold, zero-return treatment, prediction-bias normalization, IC grouping axis, missing-value handling, standard-deviation convention, significance test, and endpoint inclusivity of the reported test window. These choices are recorded as configuration rather than silently selected as if they came from the paper. Reported dataset totals and table values remain reference claims, not independently verified outputs.</p><p>The paper's discussion of candlestick charts outperforming tabular inputs should likewise be treated cautiously. The supplied summary indicates that the comparison can depend on the metric and that some reported baseline results complicate a universal superiority claim. This scaffold therefore implements comparison interfaces, not a conclusion that visual inputs are better. Its synthetic outputs demonstrate how the pieces connect; they say nothing about whether any commercial VLM truly reads candlesticks.</p><p>Use the button or URL below to download the source code.</p>
      <p>
          <a href="/__u/onepagecode.substack.com/p/quant-research-do-vision-language">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Quant Finance: Option Pricing beyond Black-Scholes via Quantum Mechanics (C++ Code & Research)]]></title><description><![CDATA[Implementing wavefunction-derived probability densities and effective volatility modeling under quantum force potentials in C++17.]]></description><link>https://onepagecode.substack.com/p/quant-finance-option-pricing-beyond</link><guid isPermaLink="false">https://onepagecode.substack.com/p/quant-finance-option-pricing-beyond</guid><dc:creator><![CDATA[Onepagecode]]></dc:creator><pubDate>Tue, 21 Jul 2026 20:05:33 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1616261167032-b16d2df8333b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1OXx8c3RvY2slMjBtYXJrZXR8ZW58MHx8fHwxNzg0NjUzMjg0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Download The source code using the button at the end of this article. </h2><p><strong>The paper did not include any official implementation, so this code was written by me from scratch. While I have tested several parts of it, it may still produce imperfect results. Consider it a demonstration baseline that can be extended and improved further.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1616261167032-b16d2df8333b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1OXx8c3RvY2slMjBtYXJrZXR8ZW58MHx8fHwxNzg0NjUzMjg0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1616261167032-b16d2df8333b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1OXx8c3RvY2slMjBtYXJrZXR8ZW58MHx8fHwxNzg0NjUzMjg0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1616261167032-b16d2df8333b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1OXx8c3RvY2slMjBtYXJrZXR8ZW58MHx8fHwxNzg0NjUzMjg0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1616261167032-b16d2df8333b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1OXx8c3RvY2slMjBtYXJrZXR8ZW58MHx8fHwxNzg0NjUzMjg0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1616261167032-b16d2df8333b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1OXx8c3RvY2slMjBtYXJrZXR8ZW58MHx8fHwxNzg0NjUzMjg0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1616261167032-b16d2df8333b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1OXx8c3RvY2slMjBtYXJrZXR8ZW58MHx8fHwxNzg0NjUzMjg0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="3500" height="2333" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1616261167032-b16d2df8333b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1OXx8c3RvY2slMjBtYXJrZXR8ZW58MHx8fHwxNzg0NjUzMjg0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2333,&quot;width&quot;:3500,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;white paper with green line&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="white paper with green line" title="white paper with green line" srcset="https://images.unsplash.com/photo-1616261167032-b16d2df8333b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1OXx8c3RvY2slMjBtYXJrZXR8ZW58MHx8fHwxNzg0NjUzMjg0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1616261167032-b16d2df8333b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1OXx8c3RvY2slMjBtYXJrZXR8ZW58MHx8fHwxNzg0NjUzMjg0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1616261167032-b16d2df8333b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1OXx8c3RvY2slMjBtYXJrZXR8ZW58MHx8fHwxNzg0NjUzMjg0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1616261167032-b16d2df8333b?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHw1OXx8c3RvY2slMjBtYXJrZXR8ZW58MHx8fHwxNzg0NjUzMjg0fDA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@markusspiske">Markus Spiske</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><p>The paper proposes a quantum mechanics approach to option pricing by mapping the Black-Scholes standard normal density to the ground state of a harmonic oscillator. Additional potentials model market forces, leading to modified probability densities. Option prices are computed by integrating the new density and adjusting the effective volatility. Several force types are analyzed, and the resulting option prices are compared numerically.</p><h2>Implementation Assumptions</h2><ul><li><p>&#8463;&#969; = 1, m = 1, &#969; = 1 for all wavefunction-derived densities (consistent with Appendix Sec. 5.1).</p></li><li><p>Numerical integration domain: [-12, 12] with 20001 points (Simpson's rule) for all unbounded densities.</p></li><li><p>For quantum well, integration domain is [-a, a] with proportionally adjusted resolution.</p></li><li><p>The x&#178; force model includes a compile-time flag USE<em>CORRECTED</em>X2 to optionally replace the paper's erroneous linear-&#946; term with a correct second-order perturbation (derived externally from standard QM).</p></li><li><p>All option pricing uses the Black-Scholes formula structure with &#963;<em>eff and modified N</em>eff(d_eff&#177;).</p></li><li><p>No external dependencies beyond C++17 standard library.</p></li><li><p>Output is CSV files for external plotting; no built-in plotting library.</p></li></ul><h2>Introduction: The Black-Scholes Model and Its Limitations</h2><p>Option pricing is about predicting the future value of an asset. The Black-Scholes model is the classic formula that assumes asset prices follow a random walk with a bell-curve distribution. This section introduces that formula and explains why it works, using simple terms like 'expected payoff' and 'discounting'. We will also look at the C++ implementation that serves as the baseline for the quantum extensions later in the tutorial.</p><h3>The Black-Scholes Formula</h3><p>The Black-Scholes model gives the theoretical price of a European call or put option under the assumption that the underlying asset price follows a geometric Brownian motion with constant volatility. The call price <code>c</code> and put price <code>p</code> are given by the following equations.</p><p>The call option price is</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;c = S_0 N(d_+) - K e^{-rT} N(d_-)&quot;,&quot;id&quot;:&quot;339E6050D7&quot;}" data-component-name="LatexBlockToDOM"></div><p>where <code>S_0</code> is the current asset price, <code>K</code> is the strike price, <code>r</code> is the risk-free interest rate, <code>T</code> is the time to maturity, and <code>N(&#183;)</code> is the cumulative distribution function of the standard normal distribution. The put option price is</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;p = K e^{-rT} N(-d_-) - S_0 N(-d_+)&quot;,&quot;id&quot;:&quot;EBA8E9E5F3&quot;}" data-component-name="LatexBlockToDOM"></div><p>These formulas express the option value as the discounted expected payoff under the risk-neutral measure. The terms <code>N(d_+)</code> and <code>N(d_-)</code> can be interpreted as the risk-adjusted probabilities that the option finishes in-the-money.</p><h3>The Quantiles <code>d_+</code> and <code>d_-</code></h3><p>The arguments <code>d_+</code> and <code>d_-</code> are computed from the model parameters as</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;d_{\\pm} = \\frac{\\ln(S_0/K) + (r \\pm \\sigma^2/2)T}{\\sigma\\sqrt{T}}&quot;,&quot;id&quot;:&quot;2CD321C425&quot;}" data-component-name="LatexBlockToDOM"></div><p>where <code>&#963;</code> is the annualized volatility of the asset. These quantiles arise from the log-normal distribution of the asset price at maturity. The term <code>ln(S_0/K)</code> is the log-moneyness, <code>rT</code> is the drift, and <code>&#963;&#178;T/2</code> is the It&#244; correction. The denominator <code>&#963;&#8730;T</code> standardises the distance to the strike in units of standard deviation.</p><h3>The Standard Normal CDF</h3><p>The cumulative distribution function <code>N(d)</code> is defined as the integral of the standard normal probability density <code>P_BS(x)</code>:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;N(d_{\\pm}) = \\int_{-\\infty}^{d_{\\pm}} P_{\\text{BS}}(x) \\, dx, \\quad P_{\\text{BS}}(x) = \\frac{1}{\\sqrt{2\\pi}} e^{-x^2/2}&quot;,&quot;id&quot;:&quot;DC5888EAB6&quot;}" data-component-name="LatexBlockToDOM"></div><p>This density is the familiar bell curve. In the Black-Scholes world, the log-return is normally distributed, so <code>N(d_+)</code> and <code>N(d_-)</code> are the probabilities that the option expires in-the-money under the risk-neutral measure (adjusted for the drift).</p><h3>Implementation in C++</h3><p>The baseline Black-Scholes pricing is implemented in the header file <code>src/black_scholes_base.h</code>. The code provides functions for the probability density, the cumulative distribution, the computation of <code>d_&#177;</code>, and the option pricing itself.</p><p>First, the standard normal PDF and CDF are implemented using the standard library:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;4ad5ba9f-61b2-4903-80a7-77e9f4dc77df&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">inline double normalPDF(double x) {
    constexpr double inv_sqrt_2pi = 0.3989422804014327; // 1/sqrt(2*pi)
    return inv_sqrt_2pi * std::exp(-0.5 * x * x);
}

inline double normalCDF(double x) {
    return 0.5 * std::erfc(-x * 0.7071067811865476); // 0.7071... = 1/sqrt(2)
}</code></pre></div><p>The CDF uses the complementary error function <code>std::erfc</code> for high precision. The constant <code>0.7071...</code> is <code>1/&#8730;2</code>. This implementation is monotonic and returns values in <code>[0,1]</code>.</p><p>Next, the quantiles <code>d_+</code> and <code>d_-</code> are computed by the function <code>compute_dpm</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;d75ffb0e-7793-4c84-8a87-103482ab041f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">inline std::pair&lt;double, double&gt; compute_dpm(double S0, double K, double r,
                                              double sigma, double T) {
    if (sigma &lt;= 0.0) {
        throw std::invalid_argument("compute_dpm: sigma must be positive");
    }
    if (T &lt;= 0.0) {
        throw std::invalid_argument("compute_dpm: T must be positive");
    }
    if (S0 &lt;= 0.0 || K &lt;= 0.0) {
        throw std::invalid_argument("compute_dpm: S0 and K must be positive");
    }

    double sigma_sqrt_T = sigma * std::sqrt(T);
    double log_moneyness = std::log(S0 / K);
    double drift = r * T;
    double half_var = 0.5 * sigma * sigma * T;

    double d_plus  = (log_moneyness + drift + half_var) / sigma_sqrt_T;
    double d_minus = (log_moneyness + drift - half_var) / sigma_sqrt_T;

    return {d_plus, d_minus};
}</code></pre></div><p>This function validates the inputs (positive volatility, time, and prices) and then directly translates the mathematical formula into code. The <code>half_var</code> term is <code>&#963;&#178;T/2</code>.</p><p>Finally, the option prices are computed by <code>black_scholes_base</code>, which returns a <code>BSResult</code> struct containing the call and put prices, the quantiles, and the CDF values:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;9b2927f1-d58c-4233-9c8e-cd6695d26add&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">struct BSResult {
    double call;      // European call option price
    double put;       // European put option price
    double d_plus;    // d_+ quantile
    double d_minus;   // d_- quantile
    double N_d_plus;  // N(d_+)
    double N_d_minus; // N(d_-)
};

inline BSResult black_scholes_base(double S0, double K, double r, double T,
                                    double sigma) {
    // ... input validation ...
    auto [d_plus, d_minus] = compute_dpm(S0, K, r, sigma, T);
    double N_d_plus  = normalCDF(d_plus);
    double N_d_minus = normalCDF(d_minus);
    double discount = std::exp(-r * T);
    double discounted_strike = K * discount;

    double call = S0 * N_d_plus - discounted_strike * N_d_minus;
    double N_neg_d_plus  = 1.0 - N_d_plus;
    double N_neg_d_minus = 1.0 - N_d_minus;
    double put = discounted_strike * N_neg_d_minus - S0 * N_neg_d_plus;

    return {call, put, d_plus, d_minus, N_d_plus, N_d_minus};
}</code></pre></div><p>The call price is computed exactly as in the equation, and the put price uses the identity <code>N(-d) = 1 - N(d)</code> to avoid redundant CDF evaluations.</p><h3>Worked Example</h3><p>Using the paper's synthetic parameters&#8212;<code>S0 = 20</code>, <code>K = 20</code>, <code>r = 0.1</code> (10%), <code>T = 1</code> year, <code>&#963; = 0.25</code> (25%)&#8212;the baseline call price is approximately 3.022 and the put price is approximately 1.119. The code computes these values with a simple call:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;eb5dc5c2-8168-45e9-945a-106753aa1aac&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">auto [call, put, dp, dm, Np, Nm] = black_scholes_base(20.0, 20.0, 0.1, 1.0, 0.25);
// call &#8776; 3.022, put &#8776; 1.119</code></pre></div><h3>Verification: Put-Call Parity</h3><p>A fundamental consistency check for any option pricing model is put-call parity: <code>c + K e^{-rT} = p + S_0</code>. This relationship holds regardless of the distribution of the underlying, as long as the pricing is arbitrage-free. The implementation includes a static verification function <code>verify_black_scholes_base()</code> that checks this parity (and other properties) to ensure the code is correct. For the sample parameters, the left-hand side and right-hand side agree to within <code>1e-12</code>.</p><h3>Limitations of the Black-Scholes Model</h3><p>The Black-Scholes model is elegant and widely used, but it rests on assumptions that are often violated in real markets: constant volatility, continuous trading, no transaction costs, and normally distributed log-returns. In practice, asset returns exhibit fat tails, skewness, and volatility clustering. The quantum mechanics approach introduced in the paper does not relax the continuous-trading or frictionless-market assumptions; instead, it provides a new way to generate non-Gaussian probability densities by adding "market potentials" to the harmonic oscillator that underlies the standard normal distribution. The following sections will build on this baseline to show how those modifications are implemented in C++.</p><h3>Summary</h3><p>We have reviewed the standard Black-Scholes formulas for European call and put options, the definition of <code>d_&#177;</code>, and the role of the cumulative normal distribution. The C++ implementation in <code>black_scholes_base.h</code> provides a clean, verified baseline that we will later extend with quantum-mechanical probability densities. The next section will establish the analogy between the Black-Scholes density and the ground state of a quantum harmonic oscillator.</p><h2>Quantum Mechanics Analogy: Mapping the Standard Normal Density to a Harmonic Oscillator Ground State</h2><p>Imagine a marble rolling inside a smooth bowl. It tends to stay near the bottom, and if you recorded its position many times, you would find a bell-shaped pattern of positions. The Black-Scholes model assumes that the log-returns of an asset follow a similar bell-shaped (normal) distribution. We can think of the market as having a hidden "potential" that shapes the probability of price moves, just as the bowl determines where the marble is likely to be. This section builds the bridge between the financial model and the quantum mechanical harmonic oscillator.</p><h3>The Harmonic Oscillator Potential</h3><p>In quantum mechanics, a particle of mass <code>m</code> moving in a potential <code>V(x)</code> is described by the stationary Schr&#246;dinger equation. The baseline potential chosen in the paper is the harmonic oscillator:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;V(x) = \\frac{1}{2} m \\omega^2 x^2&quot;,&quot;id&quot;:&quot;E960FBACC0&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here <code>V(x)</code> is the potential energy, <code>m</code> is the mass, <code>omega</code> is the angular frequency, and <code>x</code> is the position. This potential is shaped like a parabola. The corresponding Schr&#246;dinger equation reads</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;-\\frac{\\hbar^2}{2m}\\frac{d^2\\psi}{dx^2} + \\frac{1}{2} m \\omega^2 x^2 \\psi = E \\psi&quot;,&quot;id&quot;:&quot;9E340D62DC&quot;}" data-component-name="LatexBlockToDOM"></div><p>where <code>hbar</code> is the reduced Planck constant, <code>psi(x)</code> is the wavefunction, and <code>E</code> is the energy eigenvalue. The ground state (lowest energy) wavefunction of this system is known exactly:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\psi_g(x) = \\frac{\\alpha}{\\sqrt{\\pi}} e^{-\\alpha^2 x^2/2}, \\quad \\alpha = \\sqrt{\\frac{m\\omega}{\\hbar}}&quot;,&quot;id&quot;:&quot;9662CB50DA&quot;}" data-component-name="LatexBlockToDOM"></div><p>The parameter <code>alpha</code> is the inverse length scale. The square modulus of the wavefunction gives the probability density of finding the particle at position <code>x</code>:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;P_{g,HO}(x) = |\\psi_g(x)|^2 = \\frac{\\alpha^2}{\\pi} e^{-\\alpha^2 x^2}&quot;,&quot;id&quot;:&quot;7EBB81658F&quot;}" data-component-name="LatexBlockToDOM"></div><p>This is a Gaussian (bell curve). Its width is controlled by <code>alpha</code>.</p><h3>Correspondence with the Black-Scholes Density</h3><p>The Black-Scholes model uses the standard normal probability density, which can be written as</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;P_{g,HO}(x) = \\frac{1}{\\sqrt{2\\pi}} e^{-x^2/2} = P_{\\text{BS}}(x) \\quad \\text{with } \\alpha = 1/\\sqrt{2}&quot;,&quot;id&quot;:&quot;7A6312AFCD&quot;}" data-component-name="LatexBlockToDOM"></div><p>Notice that both densities are Gaussians. To make the quantum density match the Black-Scholes density we would like the exponential part to be <code>-x^2/2</code>, which requires <code>alpha^2 = 1/2</code>, i.e., <code>alpha = 1/sqrt(2)</code>. However, with this choice the prefactor becomes <code>(1/2)/pi = 1/(2*pi)</code>, while the standard normal prefactor is <code>1/sqrt(2*pi)</code>. The normalization constants do not agree. The paper's derivation therefore contains an inconsistency in the normalization of the wavefunction.</p><p>For the purpose of option pricing, we side-step this issue and directly adopt the standard normal probability density as the baseline. This is the density that the Black-Scholes model uses and that will be modified later by market forces. In other words, we treat <code>P_BS(x)</code> as our reference ground state probability density, and we do not derive it from a wavefunction that would require a different normalization constant.</p><p>In the code, this baseline density is implemented by the function <code>normalPDF</code>.</p><h3>Code: The Baseline Probability Density</h3><p>The header <code>src/black_scholes_base.h</code> provides the baseline normal PDF:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;38fd775b-ae7d-488b-bc56-f65e8e952085&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">inline double normalPDF(double x) {
    constexpr double inv_sqrt_2pi = 0.3989422804014327; // 1/sqrt(2*pi)
    return inv_sqrt_2pi * std::exp(-0.5 * x * x);
}</code></pre></div><p>This function returns <code>P_BS(x)</code> exactly. It uses a precomputed constant <code>1/sqrt(2*pi)</code> for efficiency.</p><h3>Worked Example</h3><p>Calling <code>normalPDF(0.0)</code> returns approximately <code>0.39894228</code>. This is the peak of the standard normal bell curve, and it corresponds to the probability density at the mean. If we were to construct the wavefunction <code>psi_g</code> with <code>alpha = 1/sqrt(2)</code>, we would obtain a squared modulus with a different peak value; the adopted <code>normalPDF</code> avoids that discrepancy and matches the Black-Scholes model directly.</p><h3>Why This Analogy Matters</h3><p>By mapping the standard normal density to the ground state of a harmonic oscillator, we gain a physical intuition: the parabolic potential keeps the distribution centered and bell-shaped. Adding extra potentials (market forces) will distort this baseline, leading to new probability densities that can capture skewness, fat tails, or hard bounds. The next sections show how to implement those modifications.</p><h3>Advanced Note</h3><p>The harmonic oscillator ground state is a Gaussian because the potential is quadratic. The paper's Eq. (8) provides the wavefunction, but we only ever need the probability density <code>P(x) = |psi(x)|^2</code>. Because of the normalization inconsistency, the code never uses a wavefunction directly; it computes probability densities using the target functional forms, normalizing them numerically when necessary.</p><h2>Market Forces as Potentials: How Hidden Forces Modify the Probability Density</h2><p>Just as a magnet can pull a marble away from the centre of a bowl, hidden &#8220;market forces&#8221; can pull the price distribution away from the standard bell curve. In the quantum analogy, these forces are represented by adding extra terms to the bowl&#8217;s shape&#8212;the potential energy function <code>V(x)</code>. By choosing different potentials, we can generate a rich family of probability densities that go beyond the log&#8209;normal assumption of Black&#8211;Scholes.</p><h3>From Potential to Force</h3><p>In classical physics, a conservative force is the negative gradient of a potential. The paper adopts the same relationship for the market force <code>F(x)</code>:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;F(x) = -\\frac{dV}{dx}&quot;,&quot;id&quot;:&quot;E77AF49AD7&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here <code>F(x)</code> is the market force acting at the dimensionless coordinate <code>x</code> (which corresponds to the log&#8209;return variable in the Black&#8211;Scholes framework), and <code>V(x)</code> is the potential energy. A positive force pushes the distribution to the right; a negative force pushes it to the left. The shape of <code>V(x)</code> determines how the force changes with <code>x</code>.</p><h3>The Baseline Potential</h3><p>The starting point is the harmonic oscillator potential that reproduces the standard normal probability density of Black&#8211;Scholes. Its canonical form is <code>V(x) = &#189; m &#969;&#178; x&#178;</code> (Eq. eq<em>ho</em>potential). Any additional potential <code>V_extra(x)</code> is added to this baseline, so the total potential becomes <code>V_total(x) = &#189; m &#969;&#178; x&#178; + V_extra(x)</code>. The new ground state of the modified Schr&#246;dinger equation then yields a different probability density.</p><h3>Five Types of Market Forces</h3><p>The paper studies five distinct modifications to the potential, each modelling a different kind of market influence. The table below summarises them; the following paragraphs explain each one qualitatively.</p><h4>Constant Force</h4><p>The simplest modification is a linear potential. The canonical form from the paper is:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;V_k(x) = k x&quot;,&quot;id&quot;:&quot;15F81179FA&quot;}" data-component-name="LatexBlockToDOM"></div><p>This corresponds to a constant force <code>F = -k</code>. It tilts the harmonic oscillator bowl, shifting the equilibrium point to a new position. The probability density remains a Gaussian with the same variance, but its peak moves to <code>-x_k</code> (where <code>x_k</code> is proportional to <code>k</code>). In financial terms, a constant force represents a persistent buying or selling pressure that biases the expected return.</p><h4>Linear Force</h4><p>A quadratic addition to the potential models a force proportional to the displacement. The canonical form is:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;V_\\lambda(x) = \\lambda x^2, \\quad \\lambda > 0&quot;,&quot;id&quot;:&quot;5B6E3CF070&quot;}" data-component-name="LatexBlockToDOM"></div><p>This force <code>F = -2&#955; x</code> acts like a spring: it pulls the distribution back toward the origin more strongly than the baseline harmonic oscillator. The result is a Gaussian with a narrower variance. A larger <code>&#955;</code> makes the bowl steeper, reducing the probability of large moves and therefore lowering option prices.</p><h4>x&#178; Force (Cubic Potential)</h4><p>A cubic potential introduces an asymmetry. The canonical form is:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;V_\\beta(x) = \\beta x^3&quot;,&quot;id&quot;:&quot;75E1F660E9&quot;}" data-component-name="LatexBlockToDOM"></div><p>The force <code>F = -3&#946; x&#178;</code> is always negative (for <code>&#946; &gt; 0</code>), pushing the distribution to the left, but the strength grows quadratically with distance. This breaks the left&#8209;right symmetry of the Gaussian, creating a skewed probability density. (The paper&#8217;s perturbative treatment of this case contains a known error, which we will discuss in detail later.)</p><h4>x&#179; Force (Quartic Potential)</h4><p>A quartic potential is symmetric but can change the shape dramatically. The canonical form is:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;V_\\gamma(x) = \\gamma x^4&quot;,&quot;id&quot;:&quot;E31A383E6D&quot;}" data-component-name="LatexBlockToDOM"></div><p>For small <code>&#947;</code>, the force <code>F = -4&#947; x&#179;</code> is weak near the origin and grows rapidly for large <code>|x|</code>. As <code>&#947;</code> increases, the total potential <code>&#189; m &#969;&#178; x&#178; + &#947; x&#8308;</code> can develop a double&#8209;well shape&#8212;two minima separated by a barrier. The ground&#8209;state probability density then becomes bimodal, with peaks on either side of zero. This can model a market that expects a large move in either direction (e.g., ahead of a binary event).</p><h4>Quantum Well (Hard Price Bounds)</h4><p>The infinite square well potential imposes absolute limits on the variable <code>x</code>. The canonical form is:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;V(x) = \\begin{cases} 0 &amp; |x| < a \\\\ \\infty &amp; |x| \\ge a \\end{cases}&quot;,&quot;id&quot;:&quot;F7DF7761E8&quot;}" data-component-name="LatexBlockToDOM"></div><p>Inside the well the particle is free; at the walls the potential is infinite, so the probability density must vanish at <code>x = &#177;a</code> and is exactly zero outside. This models a market with hard price bounds&#8212;for example, a currency peg or a circuit breaker that prevents trading beyond a certain range. The width <code>a</code> controls how much the distribution is squeezed.</p><h3>How the Potentials Modify the Probability Density</h3><p>Each extra potential is added to the harmonic oscillator baseline. The new stationary Schr&#246;dinger equation is solved (exactly or via perturbation theory) to obtain the ground&#8209;state wavefunction <code>&#968;(x)</code>. The probability density is then <code>P(x) = |&#968;(x)|&#178;</code>. Because the potential shapes the wavefunction, different potentials produce different densities:</p><ul><li><p>A <strong>constant force</strong> shifts the Gaussian without changing its width.</p></li><li><p>A <strong>linear force</strong> narrows or broadens the Gaussian.</p></li><li><p>An <strong>x&#178; force</strong> skews the distribution.</p></li><li><p>An <strong>x&#179; force</strong> can split the distribution into two peaks.</p></li><li><p>A <strong>quantum well</strong> truncates the distribution at hard boundaries.</p></li></ul><p>In every case, the new density <code>P(x)</code> replaces the standard normal PDF in the option pricing formula. The effective volatility <code>&#963;_eff</code> is then computed from the quantum standard deviation of <code>x</code> under <code>P(x)</code>, and the cumulative distribution is integrated numerically. The next sections will walk through each model in detail, showing the exact mathematical forms and the corresponding C++ code.</p><p><strong>Worked example (qualitative):</strong> For a constant force <code>F = -k</code>, the potential is <code>V(x) = k x</code>. This tilts the harmonic oscillator bowl, shifting the most likely position to <code>-x_k</code>. The bell curve keeps its shape but moves left or right, changing the probability that the option finishes in&#8209;the&#8209;money.</p><h2>Effective Volatility: Computing &#963;_eff from Quantum Mechanical Expectation Values</h2><p>Volatility measures how much the price jumps around. In the standard Black&#8211;Scholes world, the volatility <code>&#963;</code> is a single number that controls the width of the bell&#8209;shaped distribution of log&#8209;returns. When we introduce a market force, the probability density <code>P(x)</code> changes shape&#8212;it may shift, narrow, or even develop multiple peaks. The paper captures the effect of this new shape on option prices through an <strong>effective volatility</strong>. This section explains what the effective volatility is, how it is computed from the quantum mechanical probability density, and how the C++ code implements the calculation.</p><h3>The Definition of Effective Volatility</h3><p>The paper defines the effective volatility as a rescaling of the original Black&#8211;Scholes volatility by the standard deviation of the quantum probability density. The canonical equation from the paper (Eq. 11) is</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\sigma_{\\text{eff}} = \\sigma \\sigma_{\\text{QM}}, \\quad \\sigma_{\\text{QM}} = \\sqrt{\\langle x^2 \\rangle_{\\text{QM}} - \\langle x \\rangle_{\\text{QM}}^2}&quot;,&quot;id&quot;:&quot;4E3C51D596&quot;}" data-component-name="LatexBlockToDOM"></div><p>In this formula <code>&#963;</code> is the original volatility (for example 0.25), <code>&#963;_QM</code> is the quantum mechanical standard deviation of the dimensionless variable <code>x</code> under the market&#8209;force&#8209;modified probability density <code>P(x)</code>, and <code>&#963;_eff</code> is the effective volatility that will be used in the modified Black&#8211;Scholes formula. The notation <code>&#10216;&#183;&#10217;_QM</code> denotes the expectation value with respect to <code>P(x)</code>: <code>&#10216;x&#10217;_QM</code> is the mean of <code>x</code>, and <code>&#10216;x&#178;&#10217;_QM</code> is the second moment.</p><p>In plain terms <code>&#963;_QM</code> is simply the standard deviation of the new probability curve. If the market force squeezes the distribution, <code>&#963;_QM</code> becomes smaller than 1 and the effective volatility drops. If the force stretches the distribution, <code>&#963;_QM</code> becomes larger than 1 and the effective volatility rises. For the baseline normal density <code>P_BS(x) = (1/&#8730;(2&#960;)) e^{-x&#178;/2}</code> the standard deviation is exactly 1, so <code>&#963;_eff = &#963;</code> and we recover the original Black&#8211;Scholes model.</p><h3>Computing Expectation Values Numerically</h3><p>The paper does not provide analytical formulas for <code>&#10216;x&#10217;</code> and <code>&#10216;x&#178;&#10217;</code> for every force model. Instead the implementation computes them by numerical integration. The algorithm follows the method card <code>effective_volatility_computation</code>:</p><ol><li><p>Choose an integration domain <code>[a, b]</code> that captures virtually all probability mass. For Gaussian&#8209;like densities <code>[-12, 12]</code> is sufficient.</p></li><li><p>Build a uniform grid of <code>N</code> points (default 20001, which is odd and suitable for Simpson&#8217;s rule).</p></li></ol><p>Use the button or URL below to download the source code</p><p>Evaluate <code>P(x)</code>, <code>x&#183;P(x)</code>, and <code>x&#178;&#183;P(x)</code> at each grid point.</p><ol start="3"><li><p>Compute <code>&#10216;x&#10217;</code> and <code>&#10216;x&#178;&#10217;</code> using Simpson integration.</p></li><li><p>Compute <code>&#963;_QM = &#8730;(&#10216;x&#178;&#10217; &#8211; &#10216;x&#10217;&#178;)</code>.</p></li><li><p>Return <code>&#963;_eff = &#963; &#183; &#963;_QM</code>.</p></li></ol><p>The choice of domain and number of points is a practical decision. The interval <code>[-12, 12]</code> covers more than 99.9% of the probability mass for any Gaussian&#8209;like density. For densities with a non&#8209;zero mean, such as the constant force model, the domain must be shifted accordingly to ensure the peak is well inside the integration range. The code uses 20001 points to achieve a relative error below <code>1e-12</code> for smooth densities.</p><h3>The C++ Implementation</h3><p>The computation lives in the header&#8209;only file <code>src/effective_volatility.h</code>. It provides three main pieces: a general Simpson integrator, the <code>VolResult</code> struct, and the <code>compute_eff_vol</code> function.</p><h4>Simpson Integration</h4><p>The workhorse is <code>simpson_integrate</code>, which implements Simpson&#8217;s 1/3 rule for a uniformly spaced grid:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;82632f8c-f3be-442f-9264-1627e40b3991&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">inline double simpson_integrate(const std::vector&lt;double&gt;&amp; x,
                                const std::vector&lt;double&gt;&amp; f) {
    const size_t n = x.size();
    if (n != f.size()) {
        throw std::invalid_argument(
            "simpson_integrate: x and f must have the same size");
    }
    if (n &lt; 3 || n % 2 == 0) {
        throw std::invalid_argument(
            "simpson_integrate: size must be odd and &gt;= 3");
    }

    const double dx = x[1] - x[0];
    double sum = f[0] + f[n - 1];

    for (size_t i = 1; i &lt; n - 1; ++i) {
        sum += (i % 2 == 1) ? 4.0 * f[i] : 2.0 * f[i];
    }

    return (dx / 3.0) * sum;
}</code></pre></div><p>This function is used throughout the codebase for expectation values, normalisation constants, and cumulative distribution functions. It requires an odd number of points (at least 3) and a constant step size <code>dx</code>. The pattern <code>4,2,4,2,&#8230;</code> is the classic Simpson weight sequence.</p><h4>The Result Struct</h4><p>The <code>VolResult</code> struct bundles the outputs:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;791dbc67-6131-49d5-82c1-e689e36f4bd4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">struct VolResult {
    double sigma_eff;  // Effective volatility: &#963; * &#963;_QM
    double sigma_QM;   // Quantum mechanical standard deviation of x
    double mean_x;     // Expectation value &lt;x&gt;
    double var_x;      // Variance &lt;x^2&gt; - &lt;x&gt;^2
};</code></pre></div><p>All four fields are returned so that callers can inspect the mean, variance, and both volatility values.</p><h4>The Main Function: <code>compute_eff_vol</code></h4><p>The function <code>compute_eff_vol</code> implements the algorithm described above. Here is the core logic (the full file contains input validation and a corrected version that properly assigns all fields):</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;43e958a8-8d2d-4832-98f4-8050ef63f381&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">inline VolResult compute_eff_vol(double sigma,
                                 std::function&lt;double(double)&gt; P,
                                 double a = -12.0,
                                 double b = 12.0,
                                 int N = 20001) {
    // ... input validation ...

    const double dx = (b - a) / static_cast&lt;double&gt;(N - 1);
    std::vector&lt;double&gt; x(N);
    std::vector&lt;double&gt; P_vals(N);
    std::vector&lt;double&gt; xP_vals(N);
    std::vector&lt;double&gt; x2P_vals(N);

    for (int i = 0; i &lt; N; ++i) {
        x[i] = a + static_cast&lt;double&gt;(i) * dx;
        P_vals[i] = P(x[i]);
        xP_vals[i] = x[i] * P_vals[i];
        x2P_vals[i] = x[i] * x[i] * P_vals[i];
    }

    const double mean_x = simpson_integrate(x, xP_vals);
    const double mean_x2 = simpson_integrate(x, x2P_vals);
    const double var_x = mean_x2 - mean_x * mean_x;

    if (var_x &lt;= 0.0) {
        throw std::runtime_error(
            "compute_eff_vol: variance is non-positive; "
            "check probability density");
    }

    const double sigma_QM = std::sqrt(var_x);
    const double sigma_eff = sigma * sigma_QM;

    return VolResult{sigma_eff, sigma_QM, mean_x, var_x};
}</code></pre></div><p>Notice how the function builds three parallel arrays: <code>P_vals</code>, <code>xP_vals</code>, and <code>x2P_vals</code>. Simpson integration is then applied to each to obtain the mean and the second moment. (The full file also computes the integral of <code>P</code> itself for verification, though that value is not returned in the snippet above.) The variance is computed as <code>&#10216;x&#178;&#10217; &#8211; &#10216;x&#10217;&#178;</code>, and the quantum standard deviation is its square root. The effective volatility is simply the product <code>&#963; * &#963;_QM</code>.</p><h3>Why &#963;_QM = 1 for the Baseline Normal Density</h3><p>For the standard normal probability density <code>P_BS(x) = (1/&#8730;(2&#960;)) e^{-x&#178;/2}</code>, the mean is 0 and the variance is 1. Therefore <code>&#963;_QM = 1</code> and <code>&#963;_eff = &#963;</code>. The code verifies this property in the <code>verify_effective_volatility</code> function:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;45b4bd3c-2cb6-469d-b200-bde01ea90c7f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">VolResult res = compute_eff_vol(0.25, normalPDF, -12.0, 12.0, 20001);
// Check &#963;_QM &#8776; 1.0
if (std::abs(res.sigma_QM - 1.0) &gt; tol) { /* fail */ }
// Check mean_x &#8776; 0.0
if (std::abs(res.mean_x) &gt; tol) { /* fail */ }
// Check var_x &#8776; 1.0
if (std::abs(res.var_x - 1.0) &gt; tol) { /* fail */ }
// Check &#963;_eff = &#963; * &#963;_QM
if (std::abs(res.sigma_eff - 0.25) &gt; tol) { /* fail */ }</code></pre></div><p>These checks ensure that the numerical integration is accurate and that the baseline model is correctly recovered before any market force is applied.</p><h3>Worked Example</h3><p>Let&#8217;s walk through a concrete call. Suppose we have the original volatility <code>&#963; = 0.25</code> and we want the effective volatility for the standard normal density. We call</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;7064d049-69c2-4a45-925e-f14d5331cf60&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">VolResult res = compute_eff_vol(0.25, normalPDF);</code></pre></div><p>Internally, the function builds a grid of 20001 points from <code>-12.0</code> to <code>12.0</code>, evaluates the normal PDF at each point, and computes the integrals. The result is <code>res.sigma_QM &#8776; 1.0</code>, <code>res.sigma_eff &#8776; 0.25</code>, <code>res.mean_x &#8776; 0.0</code>, and <code>res.var_x &#8776; 1.0</code>.</p><p>Now consider a market force that narrows the distribution, such as the linear force with <code>&#955; = 0.5</code>. The probability density becomes a Gaussian with variance <code>1/&#955;_&#969; = 1/&#8730;(1+&#955;) &#8776; 0.816</code>. The quantum standard deviation is then <code>&#963;_QM &#8776; 0.904</code>, and the effective volatility drops to <code>&#963;_eff &#8776; 0.226</code>. This reduction in effective volatility will lower the option price because extreme price moves become less likely.</p><h3>Summary</h3><p>The effective volatility <code>&#963;_eff</code> is the bridge between the quantum probability density and the Black&#8211;Scholes pricing formula. By computing the standard deviation of the modified density and scaling the original volatility, we preserve the structure of the Black&#8211;Scholes equations while incorporating the non&#8209;Gaussian features introduced by market forces. The C++ implementation uses Simpson&#8217;s rule to evaluate the necessary expectation values numerically, with careful choice of integration domain and resolution to ensure high accuracy. The built&#8209;in verification checks confirm that the baseline normal density yields <code>&#963;_QM = 1</code>, giving confidence that the machinery works correctly before we apply it to the force models in the following sections.</p><h2>Modified Option Pricing: Integrating the New Density and Adjusting d&#177;</h2><p>Now that we have a market&#8209;force&#8209;driven probability density <code>P(x)</code> and an effective volatility <code>sigma_eff</code>, we need to combine them to obtain option prices. The key idea is to preserve the structure of the Black&#8209;Scholes formulas but replace the standard normal cumulative distribution function <code>N(&#183;)</code> with a new CDF derived from <code>P(x)</code>, and to use <code>sigma_eff</code> when computing the integration limits <code>d&#177;</code>. This section walks through the mathematics and the C++ implementation that brings all the pieces together.</p><h3>Modified Integration Limits</h3><p>The standard <code>d&#177;</code> arguments (Eq. 3 from the paper) are defined by</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;d_{\\pm} = \\frac{\\ln(S_0/K) + (r \\pm \\sigma^2/2)T}{\\sigma\\sqrt{T}}&quot;,&quot;id&quot;:&quot;010B00EBED&quot;}" data-component-name="LatexBlockToDOM"></div><p>To incorporate the market force we simply replace the original volatility <code>&#963;</code> with the effective volatility <code>&#963;_eff</code> obtained from the quantum mechanical expectation values. This gives the modified limits <code>d_eff&#177;</code>. No new equation is introduced; the replacement is applied directly in the code by calling the same <code>compute_dpm</code> helper with <code>sigma_eff</code> instead of <code>sigma</code>.</p><h3>Modified Black&#8209;Scholes Formulas</h3><p>The call and put pricing formulas (Eqs. 1 and 2) keep their algebraic form but now reference the new cumulative distribution <code>N_eff(&#183;)</code> and the modified limits <code>d_eff&#177;</code>. First, recall the original Black&#8209;Scholes formulas:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;c = S_0 N(d_+) - K e^{-rT} N(d_-)&quot;,&quot;id&quot;:&quot;8469E51DE4&quot;}" data-component-name="LatexBlockToDOM"></div><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;p = K e^{-rT} N(-d_-) - S_0 N(-d_+)&quot;,&quot;id&quot;:&quot;8D8D495376&quot;}" data-component-name="LatexBlockToDOM"></div><p>In the modified version, <code>N(&#183;)</code> is replaced by <code>N_eff(&#183;)</code> and <code>d&#177;</code> by <code>d_eff&#177;</code>. The put formula is derived from put&#8209;call parity and uses the same <code>N_eff</code> values.</p><h3>Building the Cumulative Distribution Grid</h3><p>Evaluating <code>N_eff(d)</code> on the fly for every option pricing call would be slow and numerically expensive. Instead, the implementation pre&#8209;computes a cumulative distribution grid on a uniform partition of a sufficiently wide domain (the paper uses <code>[-12, 12]</code> with 20001 points) using Simpson&#8217;s rule. The function <code>build_cdf_grid</code> performs this task.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;02a3551d-6b52-439a-9861-14fed825a4b5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">struct ModifiedBSResult {
    double call;         // Modified call option price
    double put;          // Modified put option price
    double d_eff_plus;   // d_eff+ using effective volatility
    double d_eff_minus;  // d_eff- using effective volatility
    double N_eff_plus;   // N_eff(d_eff+)
    double N_eff_minus;  // N_eff(d_eff-)
};

// Build a cumulative distribution grid from a probability density P(x).
// Returns a pair (x_grid, cdf_grid) where cdf_grid[i] = &#8747;_{a}^{x_grid[i]} P(t) dt.
std::pair&lt;std::vector&lt;double&gt;, std::vector&lt;double&gt;&gt;
build_cdf_grid(std::function&lt;double(double)&gt; P, double a, double b, int N);</code></pre></div><p>Inside <code>build_cdf_grid</code>, the density <code>P(x)</code> is evaluated at every grid point, and a cumulative Simpson integration is applied. The first element of <code>cdf_grid</code> is always 0.0, and for a properly normalised <code>P</code>, the last element should be very close to 1.0.</p><h3>Interpolating the CDF at Arbitrary Points</h3><p>Once the grid is built, <code>N_eff(d)</code> can be evaluated for any <code>d</code> by linear interpolation between the two nearest grid points. The helper <code>interpolate_cdf</code> handles that, including clamping to 0.0 for values below the domain and to <code>cdf_grid.back()</code> for values above it.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;b8a1ae29-17f9-42ac-b4f2-a9edcabe35b3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">// Interpolate the CDF at an arbitrary point d.
double interpolate_cdf(double d,
                       const std::vector&lt;double&gt;&amp; x_grid,
                       const std::vector&lt;double&gt;&amp; cdf_grid);</code></pre></div><p>Binary search locates the interval containing <code>d</code>, and a linear blend yields the interpolated value. This keeps the option pricing fast even when the grid is very fine.</p><h3>Bringing It All Together: modified<em>option</em>pricing</h3><p>The main function <code>modified_option_pricing</code> accepts the financial parameters (<code>S0</code>, <code>K</code>, <code>r</code>, <code>T</code>), the effective volatility <code>sigma_eff</code>, the density function <code>P</code> (kept only for documentation, not called inside the function), and the pre&#8209;computed grids. It then:</p><ol><li><p>Computes <code>d_eff&#177;</code> by calling the standard <code>compute_dpm</code> with <code>sigma_eff</code>.</p></li><li><p>Evaluates <code>N_eff(d_eff&#177;)</code> via <code>interpolate_cdf</code>.</p></li><li><p>Applies the modified call and put formulas.</p></li></ol><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;86d6c567-5704-47b7-b394-666aeb89026d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">ModifiedBSResult modified_option_pricing(
    double S0,
    double K,
    double r,
    double T,
    double sigma_eff,
    std::function&lt;double(double)&gt; /*P*/,
    const std::vector&lt;double&gt;&amp; x_grid,
    const std::vector&lt;double&gt;&amp; cdf_grid)
{
    // Compute d_eff&#177; using sigma_eff
    auto [d_eff_plus, d_eff_minus] = compute_dpm(S0, K, r, sigma_eff, T);

    // Evaluate modified CDF
    double N_eff_plus  = interpolate_cdf(d_eff_plus,  x_grid, cdf_grid);
    double N_eff_minus = interpolate_cdf(d_eff_minus, x_grid, cdf_grid);

    double discount = std::exp(-r * T);
    double call = S0 * N_eff_plus - K * discount * N_eff_minus;
    double put  = K * discount * (1.0 - N_eff_minus) - S0 * (1.0 - N_eff_plus);

    return {call, put, d_eff_plus, d_eff_minus, N_eff_plus, N_eff_minus};
}</code></pre></div><h3>Verification: Recovering the Baseline</h3><p>A crucial sanity check is that when we supply the standard normal PDF and use <code>sigma_eff = sigma</code>, the modified pricing must reproduce the original Black&#8209;Scholes prices. The implementation includes a verification function <code>verify_modified_option_pricing</code> that performs this test alongside monotonicity and put&#8209;call parity checks. Using the paper&#8217;s numeric example (<code>S0 = 20</code>, <code>K = 20</code>, <code>r = 0.1</code>, <code>T = 1</code>, <code>&#963; = 0.25</code>) the modified call and put prices match the baseline values within a tolerance of <code>1e-10</code>.</p><h3>Putting It All Together</h3><p>To compute an option price under a given market force model, you would:</p><ol><li><p>Choose a force strength (e.g., <code>&#955;</code> for the linear force) and obtain the corresponding probability density <code>P(x)</code> (e.g., <code>linear_force_PDF(x, &#955;)</code>).</p></li><li><p>Compute the effective volatility by calling <code>compute_eff_vol(sigma, P, -12.0, 12.0, 20001)</code>, which returns <code>sigma_eff</code>.</p></li><li><p>Build the CDF grid with <code>build_cdf_grid(P, -12.0, 12.0, 20001)</code>.</p></li><li><p>Call <code>modified_option_pricing(S0, K, r, T, sigma_eff, P, x_grid, cdf_grid)</code> to obtain the modified call and put prices.</p></li></ol><p>When <code>P</code> is the standard normal PDF, the result collapses identically to the original Black&#8209;Scholes price, confirming the internal consistency of the approach.</p><h2>Model 1: Constant Market Force (Shifted Gaussian)</h2><p>A constant market force is the simplest departure from the Black&#8211;Scholes baseline. Imagine a steady wind that pushes the marble to one side of the bowl. The bell&#8209;shaped probability density keeps its shape but shifts left or right. In financial terms, this corresponds to a net buying or selling pressure that changes the most likely outcome without altering the size of typical fluctuations. This section builds the constant force model from the paper, explains why the effective volatility remains unchanged, and walks through the C++ implementation that reproduces Figure&#8239;1.</p><h3>The Potential and the Force</h3><p>The paper introduces a linear potential <code>V_k(x) = k x</code> to model a constant force, where <code>k</code> is the force strength constant. A positive <code>k</code> gives a force <code>F(x) = -dV/dx = -k</code>, which pushes the distribution to the left (negative direction). A negative <code>k</code> pushes it to the right. The potential is added to the harmonic oscillator baseline, tilting the bowl.</p><h3>The Shifted Gaussian Density</h3><p>Solving the Schr&#246;dinger equation with the linear potential yields a ground state wavefunction whose squared modulus is a shifted Gaussian. The paper&#8217;s Eq.&#8239;(15) is garbled in the original text; we reconstruct the probability density as a standard normal distribution shifted by <code>-x_k</code>:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;P_k(x) = \\frac{1}{\\sqrt{2\\pi}} e^{-(x + x_k)^2/2}&quot;,&quot;id&quot;:&quot;43B2D183F0&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here <code>P_k(x)</code> is the probability density at the dimensionless coordinate <code>x</code>, and <code>x_k</code> is the peak shift. The shift is related to the force strength by <code>x_k = 2k/(\hbar\omega)</code>. Following the Appendix (Sec.&#8239;5.1) we set the energy scale <code>&#8463;&#969; = 1</code>, so the shift simplifies to <code>x_k = 2k</code>. The density is a Gaussian with mean <code>-x_k</code> and variance <code>1</code>.</p><h3>Effective Volatility</h3><p>The paper defines the effective volatility as the product of the original volatility and the quantum standard deviation:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\sigma_{\\text{eff}} = \\sigma \\sigma_{\\text{QM}}, \\quad \\sigma_{\\text{QM}} = \\sqrt{\\langle x^2 \\rangle_{\\text{QM}} - \\langle x \\rangle_{\\text{QM}}^2}&quot;,&quot;id&quot;:&quot;2B79A5EBB5&quot;}" data-component-name="LatexBlockToDOM"></div><p>For the constant force model, the variance of the shifted Gaussian is unchanged, so <code>&#963;_QM = 1</code> and therefore <code>&#963;_eff = &#963;</code>. This is a crucial property: a constant force shifts the distribution but does not change its width, so the effective volatility is the same as the Black&#8211;Scholes volatility. The only effect on option prices comes from the altered moneyness through the modified integration limits <code>d_eff&#177;</code>.</p><h3>Implementation in C++</h3><p>The constant force model is implemented in the header <code>src/constant_force_model.h</code>. Two functions capture the mathematics.</p><p><strong>Computing the shift</strong> &#8211; The function <code>constant_force_xk</code> returns <code>x_k = 2k</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;ee75ef29-52f1-4db2-b805-f6a60f0153ea&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">inline double constant_force_xk(double k) {
    // Eq. (14): x_k = 2k / (hbar * omega), with hbar*omega = 1
    return 2.0 * k;
}</code></pre></div><p><strong>Evaluating the density</strong> &#8211; The function <code>constant_force_PDF</code> implements the reconstructed Eq.&#8239;(15). It computes the shifted argument <code>x + x_k</code> and evaluates the standard normal PDF:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;19e23d4e-6ec4-49ee-bc00-94c5043bdfd9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">inline double constant_force_PDF(double x, double k) {
    const double xk = constant_force_xk(k);
    const double shifted = x + xk;  // mean is -x_k, so (x - (-x_k)) = x + x_k
    constexpr double inv_sqrt_2pi = 0.3989422804014327;  // 1 / sqrt(2*pi)
    return inv_sqrt_2pi * std::exp(-0.5 * shifted * shifted);
}</code></pre></div><p>Notice that the density is always positive and integrates to <code>1</code> over the real line. For <code>k = 0</code> the shift <code>x_k</code> is zero, and the function returns exactly the standard normal PDF used in the baseline Black&#8211;Scholes model.</p><h3>How the Shift Affects Option Prices</h3><p>Because <code>&#963;_eff = &#963;</code>, the modified integration limits <code>d_eff&#177;</code> are computed with the original volatility. However, the cumulative distribution function <code>N_eff(d_eff&#177;)</code> is obtained by integrating the shifted density <code>P_k(x)</code> from <code>-&#8734;</code> to <code>d_eff&#177;</code>. The shift changes the probability mass to the left or right of the strike, altering the option&#8217;s moneyness.</p><p>Consider a call option with the paper&#8217;s synthetic parameters: <code>S0 = 20</code>, <code>K = 20</code>, <code>r = 0.1</code>, <code>T = 1</code>, <code>&#963; = 0.25</code>. For <code>k = 0.5</code> we have <code>x_k = 1.0</code>. The density is centred at <code>-1.0</code>, meaning the distribution is shifted to the left. The probability of finishing in&#8209;the&#8209;money decreases, so the call price falls below the baseline value of approximately <code>3.022</code>. Conversely, a negative <code>k</code> shifts the density to the right, increasing the call price. The put price moves in the opposite direction.</p><h3>Verification</h3><p>The header also provides a verification function <code>verify_constant_force_model</code> that checks three invariants:</p><ol><li><p>For <code>k = 0</code>, <code>P_k(x)</code> equals the standard normal PDF at several test points.</p></li><li><p>The integral of <code>P_k</code> over the domain <code>[-12, 12]</code> is <code>1.0</code> within a tight tolerance for several values of <code>k</code>.</p></li><li><p>The quantum standard deviation <code>&#963;_QM</code> computed numerically is <code>1.0</code> regardless of the shift.</p></li></ol><p>These checks ensure that the implementation faithfully reproduces the paper&#8217;s shifted Gaussian and that the effective volatility remains unchanged. The verification is designed to be called from the main program after all headers are included.</p><h3>Worked Example</h3><p>To see the model in action, suppose we set <code>k = 0.5</code>. Then <code>x_k = 1.0</code>. The density <code>P_k(x)</code> is a Gaussian centred at <code>-1.0</code>. Using the modified pricing pipeline (building a CDF grid from <code>constant_force_PDF</code>, computing <code>&#963;_eff = &#963;</code>, and evaluating <code>modified_option_pricing</code>), the call price drops from the baseline <code>3.022</code> to a lower value, while the put price increases. The parameter sweep function <code>sweep_constant_force</code> in <code>src/paper_figures.h</code> automates this process over a range of <code>k</code> values, producing the data for Figure&#8239;1 of the paper.</p><p>In summary, the constant force model is the simplest illustration of how a market potential can alter option prices. It shifts the probability density without changing its shape, and the implementation is a direct translation of the reconstructed Eq.&#8239;(15). The next section will examine the linear force, which modifies the variance and therefore the effective volatility.</p><h2>Model 2: Linear Market Force (Modified Variance Gaussian)</h2><p>A linear force is like making the bowl steeper or shallower. A steeper bowl (&#955; &gt; 0) squeezes the bell curve, making extreme moves less likely. This lowers the option price because big payoffs are rarer. In this section we implement the linear market force model from the paper, which modifies the variance of the probability density while keeping its Gaussian shape. We will see how the effective frequency parameter <code>&#955;_&#969;</code> emerges from the quadratic potential, how the analytical standard deviation <code>&#963;_QM</code> avoids numerical integration, and how the C++ code in <code>linear_force_model.h</code> brings everything together.</p><h3>The Quadratic Potential and the Modified Frequency</h3><p>The paper adds a quadratic potential to the harmonic oscillator baseline: <code>V_&#955;(x) = &#955; x&#178;</code> with <code>&#955; &gt; 0</code>. The corresponding market force is <code>F(x) = -dV/dx = -2&#955; x</code>, a restoring force proportional to the displacement from equilibrium. Adding this term changes the effective frequency of the system. With the conventions <code>&#8463;&#969; = 1</code>, <code>m = 1</code>, and <code>&#969; = 1</code> (Appendix Sec. 5.1), the new frequency becomes <code>&#969;' = &#969; &#8730;(1 + &#955;/&#969;) = &#8730;(1 + &#955;)</code>. The paper defines the effective frequency parameter as <code>&#955;_&#969; = &#8730;(1 + &#955;/&#969;)</code>, which simplifies to <code>&#955;_&#969; = &#8730;(1 + &#955;)</code> when <code>&#969; = 1</code>. This parameter controls the width of the ground state probability density.</p><h3>The Modified Gaussian Density</h3><p>The ground state wavefunction for the modified harmonic oscillator yields a probability density that is still a Gaussian, but with a variance that depends on <code>&#955;</code>. The canonical equation from the paper (Eq. 18) is</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;P_\\lambda(x) = \\frac{\\lambda_\\omega}{\\sqrt{2\\pi}} e^{-\\lambda_\\omega x^2/2}&quot;,&quot;id&quot;:&quot;2C99179612&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here <code>P_&#955;(x)</code> is the probability density at the dimensionless coordinate <code>x</code>, and <code>&#955;_&#969; = &#8730;(1 + &#955;)</code>. The density is always positive and integrates to 1 over the real line. When <code>&#955; = 0</code>, <code>&#955;_&#969; = 1</code> and <code>P_&#955;(x)</code> reduces to the standard normal density <code>P_BS(x)</code>. As <code>&#955;</code> increases, the variance <code>1/&#955;_&#969;</code> decreases, making the distribution narrower.</p><h3>Effective Volatility for the Linear Force</h3><p>Because the density is a zero&#8209;mean Gaussian, the quantum mechanical standard deviation <code>&#963;_QM</code> can be computed analytically. From the definition in Eq. (11),</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\sigma_{\\text{eff}} = \\sigma \\sigma_{\\text{QM}}, \\quad \\sigma_{\\text{QM}} = \\sqrt{\\langle x^2 \\rangle_{\\text{QM}} - \\langle x \\rangle_{\\text{QM}}^2}&quot;,&quot;id&quot;:&quot;CDF1BB07F0&quot;}" data-component-name="LatexBlockToDOM"></div><p>For <code>P_&#955;(x)</code> the mean is zero and the second moment is <code>1/&#955;_&#969;</code>, so <code>&#963;_QM = 1/&#955;_&#969;</code>. The effective volatility becomes <code>&#963;_eff = &#963; / &#955;_&#969; = &#963; / &#8730;(1 + &#955;)</code>. This analytical result avoids numerical integration for the linear force model. A larger <code>&#955;</code> (stronger restoring force) reduces <code>&#963;_eff</code>, which in turn lowers option prices because the distribution of log&#8209;returns is less spread out.</p><h3>C++ Implementation</h3><p>The header&#8209;only file <code>src/linear_force_model.h</code> provides three core functions and a verification routine. Let&#8217;s examine each.</p><p><strong>Computing `&#955;_&#969;` and `&#963;_QM`</strong></p><p>The functions <code>linear_force_lambda_omega</code> and <code>linear_force_sigma_QM</code> encapsulate the analytical formulas. They validate that <code>&#955; &gt; 0</code> and return the exact values.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;9e946db8-d13f-49e0-b541-38c146bad537&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">inline double linear_force_lambda_omega(double lambda) {
    if (lambda &lt;= 0.0) {
        throw std::domain_error("linear_force_lambda_omega: lambda must be &gt; 0, got "
                                + std::to_string(lambda));
    }
    return std::sqrt(1.0 + lambda);
}

inline double linear_force_sigma_QM(double lambda) {
    if (lambda &lt;= 0.0) {
        throw std::domain_error("linear_force_sigma_QM: lambda must be &gt; 0, got "
                                + std::to_string(lambda));
    }
    return 1.0 / std::sqrt(1.0 + lambda);
}</code></pre></div><p>Notice that <code>linear_force_sigma_QM</code> simply returns <code>1.0 / &#955;_&#969;</code>, matching the analytical derivation. Both functions throw a <code>std::domain_error</code> if <code>&#955;</code> is not strictly positive, as required by the paper&#8217;s Eq. (16).</p><p><strong>Evaluating the Probability Density</strong></p><p>The function <code>linear_force_PDF</code> implements Eq. (18) directly.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;21cb8ad4-f4cd-4cd0-96f7-9bbbbc3224c5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">inline double linear_force_PDF(double x, double lambda) {
    if (lambda &lt;= 0.0) {
        throw std::domain_error("linear_force_PDF: lambda must be &gt; 0, got "
                                + std::to_string(lambda));
    }
    const double lambda_omega = std::sqrt(1.0 + lambda);
    const double norm = lambda_omega / std::sqrt(2.0 * M_PI);
    return norm * std::exp(-0.5 * lambda_omega * x * x);
}</code></pre></div><p>The normalisation constant <code>norm = &#955;_&#969; / &#8730;(2&#960;)</code> ensures the integral over the real line equals 1. The exponential uses <code>-0.5 * &#955;_&#969; * x&#178;</code>, which is the standard Gaussian exponent with variance <code>1/&#955;_&#969;</code>. For large <code>|x|</code> the exponential underflows to zero, which is correct.</p><p><strong>Verification</strong></p><p>The function <code>verify_linear_force_model</code> runs several static checks:</p><ul><li><p>For a very small <code>&#955;</code> (e.g. 1e-10), <code>P_&#955;(x)</code> must match the standard normal PDF at several test points.</p></li><li><p>The integral of <code>P_&#955;</code> over <code>[-12, 12]</code> must be approximately 1.0 for a range of <code>&#955;</code> values.</p></li><li><p>The numerical <code>&#963;_QM</code> (computed via Simpson integration of <code>x&#178; P_&#955;(x)</code>) must agree with the analytical <code>linear_force_sigma_QM</code>.</p></li><li><p>The identity <code>&#963;_QM = 1/&#955;_&#969;</code> must hold exactly.</p></li></ul><p>These checks confirm that the implementation is consistent with the paper&#8217;s equations and that the density is properly normalised.</p><h3>Worked Example</h3><p>Let&#8217;s use the paper&#8217;s synthetic parameters: <code>S0 = 20</code>, <code>K = 20</code>, <code>r = 0.1</code>, <code>T = 1</code>, <code>&#963; = 0.25</code>. For <code>&#955; = 0.5</code> we compute:</p><ul><li><p><code>&#955;_&#969; = &#8730;(1 + 0.5) &#8776; 1.2247</code></p></li><li><p><code>&#963;_QM = 1 / 1.2247 &#8776; 0.8165</code></p></li><li><p><code>&#963;_eff = 0.25 &#215; 0.8165 &#8776; 0.2041</code></p></li></ul><p>The baseline Black&#8209;Scholes call price (with <code>&#963; = 0.25</code>) is approximately 3.022. With the reduced effective volatility, the call price drops to about 2.4. This makes sense: a narrower distribution means the asset is less likely to finish far in&#8209;the&#8209;money, so the option is worth less.</p><p>In code, you would obtain the density at any point with <code>linear_force_PDF(x, 0.5)</code>. To reproduce Figure 2 from the paper, the <code>sweep_linear_force</code> function (in <code>paper_figures.h</code>) loops over a range of <code>&#955;</code> values, builds the CDF grid from <code>linear_force_PDF</code>, computes <code>&#963;_eff</code> analytically, and calls <code>modified_option_pricing</code>. The resulting call and put prices are stored in a <code>SweepResult</code> struct for plotting.</p><h3>Key Takeaways</h3><ul><li><p>The linear force model corresponds to a quadratic potential <code>V(x) = &#955; x&#178;</code> that changes the harmonic oscillator frequency.</p></li><li><p>The probability density remains Gaussian, but its variance becomes <code>1/&#955;_&#969;</code> with <code>&#955;_&#969; = &#8730;(1 + &#955;)</code>.</p></li><li><p>The quantum standard deviation <code>&#963;_QM</code> and the effective volatility <code>&#963;_eff</code> can be computed analytically, avoiding numerical integration.</p></li><li><p>As <code>&#955;</code> increases, the distribution narrows, reducing option prices.</p></li><li><p>The implementation in <code>linear_force_model.h</code> is self&#8209;contained, validates inputs, and includes built&#8209;in verification checks that ensure correctness against the paper&#8217;s equations.</p></li></ul><h2>Model 3: x&#178; Market Force (Cubic Potential, Perturbative) &#8211; with Error Discussion</h2><p>An x&#178; force is like a bowl that is steeper on one side than the other. The paper attempts to approximate the new bell&#8209;shaped probability density using a mathematical technique called perturbation theory, but the approximation contains a mistake. In this section we will implement the model exactly as the paper presents it, explain why the linear term in <code>&#946;</code> is incorrect, and then show how to switch to a corrected second&#8209;order expansion using a compile&#8209;time flag. The C++ code lives in <code>src/x2_force_model.h</code> and provides both the paper's formula and the optional corrected version.</p><h3>The Cubic Potential and the x&#178; Force</h3><p>The paper introduces a cubic potential <code>V_&#946;(x) = &#946; x&#179;</code> to model a force that grows quadratically with the displacement <code>x</code>. Here <code>&#946;</code> is the force strength parameter. The corresponding market force is <code>F(x) = -dV/dx = -3&#946; x&#178;</code>. A positive <code>&#946;</code> creates a force that pushes the distribution to the left for positive <code>x</code> and to the right for negative <code>x</code>, skewing the probability density. Because the potential is an odd function, it breaks the left&#8209;right symmetry of the harmonic oscillator.</p><h3>Perturbation Theory and the Paper's Formula</h3><p>To obtain the ground state wavefunction under the cubic potential, the paper uses non&#8209;degenerate perturbation theory. The general first&#8209;order correction to a wavefunction is given by</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\psi_n = \\psi_n^{(0)} + \\sum_{m \\neq n} \\frac{H'_{nm}}{E_n^{(0)} - E_m^{(0)}} \\psi_m^{(0)} + \\ldots&quot;,&quot;id&quot;:&quot;DF07E801EC&quot;}" data-component-name="LatexBlockToDOM"></div><p>where <code>&#968;_n^{(0)}</code> are the unperturbed eigenstates, <code>E_n^{(0)}</code> the unperturbed energies, and <code>H'_{nm}</code> the matrix elements of the perturbation <code>V_&#946;(x)</code>. The paper then squares the approximate wavefunction to obtain the probability density. The claimed result (Eq. 22) is</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;P_\\beta(x) = \\frac{C^2}{\\sqrt{2\\pi}} e^{-x^2/2} \\left[1 - \\frac{\\beta}{\\hbar\\omega}\\frac{1}{3}x^3\\right]^2&quot;,&quot;id&quot;:&quot;7DE4B08E9E&quot;}" data-component-name="LatexBlockToDOM"></div><p>In this expression <code>C&#178;</code> is a normalisation constant that must be computed numerically because the perturbation expansion does not preserve the unit integral. The factor <code>&#8463;&#969;</code> is the energy quantum; following Appendix Sec. 5.1 we set <code>&#8463;&#969; = 1</code> throughout the implementation, so the term simplifies to <code>(&#946;/3) x&#179;</code>.</p><h3>The Error in the Paper's Result</h3><p>There is a fundamental problem with the formula above. The unperturbed ground state of the harmonic oscillator is an even function, while the cubic perturbation <code>&#946; x&#179;</code> is odd. The first&#8209;order correction to the wavefunction involves matrix elements <code>H'_{nm}</code> that couple the ground state to odd&#8209;parity excited states. However, the first&#8209;order correction to the <strong>probability density</strong> (the squared modulus of a real wavefunction) contains a cross term <code>2 &#968;^{(0)} &#968;^{(1)}</code>. Because <code>&#968;^{(0)}</code> is even and <code>&#968;^{(1)}</code> is odd, their product is odd, and the integral of the cross term over the whole real line vanishes. Therefore the first&#8209;order correction to the normalised probability density is <strong>zero</strong>. The leading correction is of order <code>&#946;&#178;</code>, not linear in <code>&#946;</code>. The paper's Eq. (22) erroneously retains a linear <code>&#946;</code> term, which is spurious.</p><p>Despite this error, we provide the paper's formula for reproducibility. A compile&#8209;time flag <code>USE_CORRECTED_X2</code> can be defined to replace it with a correct second&#8209;order expansion derived from standard quantum mechanics.</p><h3>Implementation of the Paper's Formula</h3><p>The header <code>src/x2_force_model.h</code> contains three essential functions for the paper's model. First, the unnormalised raw density is computed by <code>x2_force_PDF_raw</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;e2c5eb5e-d950-48db-a212-a01a7a60d4df&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">inline double x2_force_PDF_raw(double x, double beta) {
    const double x3 = x * x * x;
    const double correction = 1.0 - (beta / 3.0) * x3;
    const double gaussian = std::exp(-0.5 * x * x);
    return gaussian * correction * correction;
}</code></pre></div><p>This directly implements the bracket <code>[1 - (&#946;/3) x&#179;]&#178;</code> multiplied by the Gaussian factor <code>exp(-x&#178;/2)</code>. Notice that the factor <code>1/&#8730;(2&#960;)</code> is omitted because the normalisation constant <code>C&#178;</code> will absorb it.</p><p>Next, the normalisation constant <code>C&#178;</code> is obtained by numerical integration over the domain <code>[-12, 12]</code> with 20001 points (Simpson's rule):</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;0aa50858-c107-49fb-b6f2-2275a40e668b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">inline double compute_x2_normalization(double beta, double a = -12.0, double b = 12.0, int N = 20001) {
    if (N &lt; 3 || N % 2 == 0) {
        throw std::invalid_argument("compute_x2_normalization: N must be odd and &gt;= 3, got " + std::to_string(N));
    }
    if (a &gt;= b) {
        throw std::invalid_argument("compute_x2_normalization: a must be less than b");
    }

    const double dx = (b - a) / (N - 1);
    std::vector&lt;double&gt; x(N), f(N);
    for (int i = 0; i &lt; N; ++i) {
        x[i] = a + i * dx;
        f[i] = x2_force_PDF_raw(x[i], beta);
    }

    double integral = x2_force_forward::simpson_integrate(x, f);

    if (integral &lt;= 0.0) {
        throw std::runtime_error("compute_x2_normalization: integral of P_raw is non-positive ("
                                 + std::to_string(integral) + "). The perturbation expansion "
                                 "may have broken down for beta = " + std::to_string(beta));
    }

    return 1.0 / integral;
}</code></pre></div><p>The function validates its inputs, evaluates the raw density on a uniform grid, integrates with Simpson's rule, and throws an exception if the integral is non&#8209;positive&#8212;which can happen for large <code>|&#946;|</code> when the truncated expansion breaks down and the raw density becomes negative over most of the domain.</p><p>Finally, the normalised density is a simple scaling:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;d88f7b8c-3624-4691-abeb-ea9a1e91cb5c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">inline double x2_force_PDF(double x, double beta, double C2) {
    return C2 * x2_force_PDF_raw(x, beta);
}</code></pre></div><h3>The Corrected Second&#8209;Order Perturbation (Optional)</h3><p>When <code>USE_CORRECTED_X2</code> is defined, the header provides an alternative set of functions that use the correct <code>O(&#946;&#178;)</code> expansion. The corrected raw density is</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;8084cfed-3135-491c-83bb-895c74268c13&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">inline double x2_force_PDF_corrected_raw(double x, double beta) {
    const double x2 = x * x;
    const double x4 = x2 * x2;
    const double x6 = x4 * x2;
    const double Q = (4.0 * x6 - 36.0 * x4 + 81.0 * x2 - 15.0) / 72.0;
    const double beta2 = beta * beta;
    const double gaussian = std::exp(-0.5 * x2);
    return gaussian * (1.0 + beta2 * Q);
}</code></pre></div><p>This formula is not from the paper; it is derived from standard quantum mechanics. The polynomial <code>Q(x)</code> ensures that the density remains symmetric and non&#8209;negative for sufficiently small <code>&#946;</code>. The corresponding normalisation function <code>compute_x2_normalization_corrected</code> and the normalised accessor <code>x2_force_PDF_corrected</code> follow the same pattern as the paper's version.</p><h3>Worked Example</h3><p>Let us set <code>&#946; = 0.1</code>. First we compute the normalisation constant:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;a5dc0e5e-621e-4fa5-92aa-ce7df0b7fd6b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">double C2 = compute_x2_normalization(0.1);</code></pre></div><p>This integrates the raw density over <code>[-12, 12]</code> and returns <code>C&#178; &#8776; 1.0</code> (the deviation from unity is small because the perturbation is weak). Then we can evaluate the normalised density at any point:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;f9b9c8c1-754e-43fc-a4aa-6b834ebfd4b0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">double p = x2_force_PDF(0.0, 0.1, C2);  // near 0.3989</code></pre></div><p>To price an option, we would build a CDF grid from this density, compute <code>&#963;_eff</code> via <code>compute_eff_vol</code>, and call <code>modified_option_pricing</code>. The call price decreases slightly compared to the Black&#8209;Scholes baseline because the distribution becomes skewed, reducing the probability of large positive returns.</p><h3>Verification</h3><p>The file includes a <code>verify_x2_force_model()</code> function that runs several static checks:</p><ul><li><p>For <code>&#946; = 0</code>, the density must equal the standard normal PDF pointwise.</p></li><li><p>The normalised density must integrate to 1.0 within a tight tolerance for several small <code>&#946;</code> values.</p></li><li><p>A warning is emitted if the density becomes negative anywhere in the domain, indicating that the perturbation expansion has broken down.</p></li></ul><p>If <code>USE_CORRECTED_X2</code> is defined, analogous checks are performed for the corrected density. These verifications help ensure that the implementation faithfully reproduces the paper's formulas (with the known caveat) and that the optional corrected version behaves correctly.</p><h3>Summary</h3><p>The x&#178; force model illustrates both the power and the pitfalls of perturbation theory. The paper's formula, while algebraically simple, contains a linear <code>&#946;</code> term that should not be present. Our implementation provides the paper's version for exact reproducibility and an optional corrected second&#8209;order expansion for those who need a physically consistent density. In the next section we will examine the x&#179; force, where the first&#8209;order correction is non&#8209;zero and leads to a richer, non&#8209;monotonic behaviour.</p><h2>Model 4: x&#179; Market Force (Quartic Potential, Perturbative)</h2><p>An <code>x&#179;</code> force is like a bowl that becomes flatter at the bottom and then rises steeply. For strong forces, the marble can sit in two places, creating a double&#8209;humped probability distribution. This can make options cheaper or more expensive depending on where the strike price falls relative to the two peaks. In this section we implement the <code>x&#179;</code> market force model from the paper using first&#8209;order perturbation theory, walk through the C++ code in <code>src/x3_force_model.h</code>, and explain why the option price does not change monotonically with the force strength <code>&#947;</code>.</p><h3>The Quartic Potential</h3><p>The paper introduces a quartic potential to model an <code>x&#179;</code> force. The potential is <code>V_&#947;(x) = &#947; * x^4</code> with <code>&#947; &gt; 0</code>. The corresponding market force is <code>F(x) = -dV/dx = -4 * &#947; * x^3</code>. Because the potential is symmetric (an even function of <code>x</code>), the force is an odd function: it pulls the distribution toward the origin for positive <code>&#947;</code>, making the bottom of the effective bowl flatter than the pure harmonic oscillator. For large enough <code>&#947;</code> the total potential <code>&#189; m &#969;&#178; x&#178; + &#947; x&#8308;</code> develops a double&#8209;well shape, and the probability density can become bimodal&#8212;the marble can sit on either side of the origin.</p><h3>Perturbation Theory for the Wavefunction</h3><p>The paper uses non&#8209;degenerate perturbation theory (Eq. 33 in the Appendix) to obtain an approximate ground state wavefunction. The unperturbed system is the harmonic oscillator, and the perturbation is <code>H' = &#947; x&#8308;</code>. The first&#8209;order correction to the wavefunction is given by the canonical perturbation expansion:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\psi_n = \\psi_n^{(0)} + \\sum_{m \\neq n} \\frac{H'_{nm}}{E_n^{(0)} - E_m^{(0)}} \\psi_m^{(0)} + \\ldots&quot;,&quot;id&quot;:&quot;EB13E7B41D&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here <code>&#968;_n^(0)</code> are the unperturbed harmonic oscillator eigenstates, <code>E_n^(0)</code> are the unperturbed energies, and <code>H'_{nm}</code> is the matrix element of the perturbation between states <code>n</code> and <code>m</code>. For the ground state (<code>n = 0</code>), the quartic potential couples it to the <code>m = 2</code> and <code>m = 4</code> states. The paper gives the resulting ground state wavefunction (Eq. 25) as</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\psi_{g,\\gamma}(x) = \\frac{C}{(2\\pi)^{1/4}} e^{-x^2/4} \\left[1 - \\frac{\\gamma}{4\\sqrt{2}\\hbar\\omega}(x^4 - 9)\\right]&quot;,&quot;id&quot;:&quot;4684D3C0BD&quot;}" data-component-name="LatexBlockToDOM"></div><p>With the convention <code>&#8463;&#969; = 1</code> (Appendix Sec. 5.1), the factor simplifies to <code>&#947; / (4&#8730;2)</code>.</p><h3>The Probability Density</h3><p>The probability density is the squared modulus of the wavefunction. Because the ground state is real, this is simply <code>P_&#947;(x) = |&#968;_{g,&#947;}(x)|&#178;</code>. The paper writes it as (Eq. 26)</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;P_\\gamma(x) = \\frac{C^2}{\\sqrt{2\\pi}} e^{-x^2/2} \\left[1 - \\frac{\\gamma}{4\\sqrt{2}\\hbar\\omega}(x^4 - 9)\\right]^2&quot;,&quot;id&quot;:&quot;5F1E139B7F&quot;}" data-component-name="LatexBlockToDOM"></div><p>This is a Gaussian multiplied by a quartic correction. The constant 9 emerges from the matrix elements of the perturbation; its origin is derived in the paper's Appendix. For <code>&#947; = 0</code> the expression reduces exactly to the standard normal PDF <code>P_BS(x)</code>.</p><p>With <code>&#8463;&#969; = 1</code> the density simplifies to the raw unnormalized form used in the code:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;P_{\\text{raw}}(x) = e^{-x^2/2} \\left[1 - \\frac{\\gamma}{4\\sqrt{2}}(x^4 - 9)\\right]^2&quot;,&quot;id&quot;:&quot;B19BA31709&quot;}" data-component-name="LatexBlockToDOM"></div><p><strong>Important caveat:</strong> The perturbation expansion is valid only for small <code>&#947;</code>. For large <code>&#947;</code> the correction can overcome the Gaussian envelope, making <code>P_&#947;(x)</code> negative for some <code>x</code>. The C++ code detects this and emits a warning, because a negative probability density is unphysical and signals the breakdown of the first&#8209;order approximation.</p><h3>Normalisation</h3><p>Perturbation theory does not preserve the normalization of the wavefunction. The constant <code>C&#178;</code> must therefore be computed numerically so that <code>&#8747; P_&#947;(x) dx = 1</code> over the chosen integration domain. The paper does not give an analytical value for <code>C&#178;</code>; we compute it by Simpson integration.</p><h3>Implementation in src/x3<em>force</em>model.h</h3><p>The header&#8209;only file <code>src/x3_force_model.h</code> provides three public functions and a verification routine. Let us examine each one.</p><h4>The Raw Unnormalized Density</h4><p>The function <code>x3_force_PDF_raw</code> implements the bracket part of Eq. 26 without the normalization constant:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;94f30306-d184-43a1-b752-4558cc805c64&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">inline double x3_force_PDF_raw(double x, double gamma) {
    const double factor = gamma / (4.0 * std::sqrt(2.0));  // &#947;/(4&#8730;2)
    const double x2 = x * x;
    const double x4 = x2 * x2;
    const double bracket = 1.0 - factor * (x4 - 9.0);
    return std::exp(-0.5 * x2) * bracket * bracket;
}</code></pre></div><p>Notice how <code>&#8463;&#969;</code> is set to 1, consistent with Appendix Sec. 5.1, so the denominator <code>4&#8730;2 &#8463;&#969;</code> becomes simply <code>4&#8730;2</code>. The function computes <code>x&#178;</code>, then <code>x&#8308;</code>, forms the bracket <code>[1 - &#947;/(4&#8730;2) (x&#8308; - 9)]</code>, squares it, and multiplies by the Gaussian envelope <code>exp(-x&#178;/2)</code>. The squaring guarantees the raw density is non&#8209;negative, which is a mathematical identity from the squared modulus. However, the <em>normalized</em> density <code>C&#178; * P_raw</code> can still become negative if <code>C&#178;</code> is negative&#8212;which we prevent by validating the integral.</p><h4>Numerical Normalization</h4><p>The function <code>compute_x3_normalization</code> computes <code>C&#178;</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;da2b07f4-4c13-4dc9-a388-4f4b7f896c05&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">inline double compute_x3_normalization(double gamma, double a, double b, int N) {
    // Input validation ...
    const double dx = (b - a) / (N - 1);
    std::vector&lt;double&gt; x(N), f(N);
    for (int i = 0; i &lt; N; ++i) {
        x[i] = a + i * dx;
        f[i] = x3_force_PDF_raw(x[i], gamma);
    }
    const double integral = simpson_integrate(x, f);
    if (integral &lt;= 0.0) {
        throw std::runtime_error("Integral of P_raw is non-positive ...");
    }
    return 1.0 / integral;
}</code></pre></div><p>It creates a uniform grid on <code>[a, b]</code> (default <code>[-12, 12]</code> with 20001 points), evaluates the raw density at every grid point, and integrates with Simpson's rule. If the integral is non&#8209;positive, a <code>runtime_error</code> is thrown because the perturbation expansion has broken down. Otherwise it returns <code>C&#178; = 1 / integral</code>. For <code>&#947; = 0</code>, this returns <code>1/&#8730;(2&#960;)</code>, exactly the normal distribution's normalization.</p><h4>The Normalized Density</h4><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;bf3e1bfd-48d8-41b7-b27a-caf8ad682492&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">inline double x3_force_PDF(double x, double gamma, double C2) {
    return C2 * x3_force_PDF_raw(x, gamma);
}</code></pre></div><p>This multiplies the raw density by the pre&#8209;computed <code>C&#178;</code>. All subsequent calculations&#8212;effective volatility, CDF grid construction, and option pricing&#8212;call this function.</p><h3>How the Option Price Behaves</h3><p>The paper's Figure 4 shows non&#8209;monotonic behavior: as <code>&#947;</code> increases from zero, the call option price first decreases, reaches a minimum, and then increases again. This can be understood physically. For small <code>&#947;</code>, the quartic potential flattens the bottom of the well, which broadens the probability density slightly but also changes its shape. The effective standard deviation <code>&#963;_QM</code> may decrease, reducing <code>&#963;_eff</code> and lowering the option price (similar to the linear force case). For larger <code>&#947;</code>, the potential develops two minima, and the probability density becomes bimodal with peaks at non&#8209;zero <code>x</code>. If the strike price happens to lie near one of the peaks, the probability of finishing in&#8209;the&#8209;money can actually increase, raising the option price. The non&#8209;monotonicity is a signature of the double&#8209;well physics.</p><h3>Worked Example</h3><p>Let <code>&#947; = 0.2</code>. First compute the normalization:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;bba000ce-7ed3-48af-8a70-7ffb69c8127f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">double C2 = compute_x3_normalization(0.2, -12.0, 12.0, 20001);</code></pre></div><p>For <code>&#947; = 0.2</code>, <code>C2</code> is approximately <code>0.3989</code> (close to <code>1/&#8730;(2&#960;) &#8776; 0.398942</code>). To evaluate the density at, say, <code>x = 1.5</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;e5ea3574-f49a-4f63-9519-64f495da3b42&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">double density = x3_force_PDF(1.5, 0.2, C2);</code></pre></div><p>The raw bracket at <code>x = 1.5</code> is <code>1 - 0.2/(4&#8730;2) (5.0625 - 9) = 1 + 0.03536 &#215; 3.9375 &#8776; 1.139</code>. Squaring gives <code>&#8776; 1.298</code>. Multiplying by <code>exp(-1.125) &#8776; 0.3247</code> yields a raw density of <code>&#8776; 0.421</code>. After multiplying by <code>C2 &#8776; 0.3989</code>, the normalized density is roughly <code>0.168</code>.</p><p>To compute the full option price, you would follow the same pipeline described in earlier sections: build the CDF grid from <code>x3_force_PDF</code>, compute <code>&#963;_eff</code> via <code>compute_eff_vol</code> (which numerically integrates <code>x * P(x)</code> and <code>x&#178; * P(x)</code>), obtain <code>d_eff&#177;</code> using <code>&#963;_eff</code>, interpolate the CDF at <code>d_eff&#177;</code>, and finally apply the Black&#8209;Scholes formula with the modified <code>N_eff</code> values.</p><h3>When the Model Breaks Down</h3><p>The verification function <code>verify_x3_force_model</code> (included in the same header) checks the <code>&#947; = 0</code> recovery, the integral normalization, and whether the density becomes negative. For <code>&#947;</code> larger than about 1.0, the correction term <code>&#947;/(4&#8730;2) (x&#8308; - 9)</code> can exceed 1 for large <code>|x|</code>, making the bracket negative. Although squaring makes the raw density positive, the <em>normalized</em> density can still dip below zero at extreme <code>x</code> because numerical errors in the normalization constant magnify tiny negative raw values. The verification emits a warning in this case:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;1a2379e6-787c-45e3-acfa-1fc21be0a6ef&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">WARNING: x3_force_model - For gamma=2.0, P_gamma becomes negative (min=-1.23e-06 at x=...). 
Perturbation expansion may be invalid for this gamma.</code></pre></div><p>This is the signal that you have pushed the perturbative model beyond its regime of validity. For quantitative work, you would restrict <code>&#947;</code> to values where the density is everywhere non&#8209;negative.</p><h3>Summary</h3><p>The <code>x&#179;</code> market force model introduces a quartic potential <code>&#947; x&#8308;</code> that flattens the harmonic oscillator well and eventually creates a double&#8209;well structure. The probability density is a Gaussian multiplied by a quartic correction, obtained from first&#8209;order perturbation theory. The C++ implementation computes the raw density <code>x3_force_PDF_raw</code>, normalizes it numerically with <code>compute_x3_normalization</code>, and produces the final density <code>x3_force_PDF</code>. The resulting option price shows non&#8209;monotonic behaviour&#8212;a minimum followed by an increase&#8212;which reflects the onset of bimodality in the underlying distribution. The model is a striking example of how quantum mechanical potentials can generate option&#8209;price patterns that are impossible in the standard Black&#8209;Scholes world.</p><h2>Model 5: Quantum Well (Hard Price Bounds)</h2><p>Imagine the price of an asset is trapped between two walls. It cannot go beyond a certain range, no matter what. This is like a currency that is pegged to a narrow band, or a stock subject to circuit breakers that halt trading when the price moves too far. In such a market, the probability of extreme returns is exactly zero, and the shape of the distribution inside the allowed band is not a simple bell curve. The paper models this situation with an <strong>infinite square well</strong> potential, the simplest quantum mechanical system that enforces hard boundaries. This section builds the quantum well model, derives its probability density and effective volatility, and walks through the C++ implementation in <code>src/quantum_well_model.h</code>.</p><h3>The infinite square well potential</h3><p>The potential that creates hard price bounds is defined piecewise in Eq. (27) of the paper:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;V(x) = \\begin{cases} 0 &amp; |x| < a \\\\ \\infty &amp; |x| \\ge a \\end{cases}&quot;,&quot;id&quot;:&quot;4E25C24D99&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here <code>a</code> is the half&#8209;width of the well, a positive number that controls how wide the allowed price range is. <code>V(x)</code> is the potential energy. Inside the well the potential is zero; outside it is infinite, meaning the price cannot exist there. The parameter <code>a</code> is qualitatively proportional to the price range <code>&#916;S</code>; the paper does not give an exact mapping, but larger <code>a</code> means a wider allowed band.</p><h3>Ground state wavefunction and probability density</h3><p>Solving the stationary Schr&#246;dinger equation inside the well with the boundary condition that the wavefunction vanishes at <code>x = &#177;a</code> gives the ground state wavefunction. The infinite walls force the wavefunction to zero at the boundaries, which selects a standing&#8209;wave solution&#8212;a sine function with exactly one half&#8209;wavelength fitting inside the well. The result is Eq. (28):</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\psi_g(x) = \\frac{1}{\\sqrt{a}} \\sin\\left[\\frac{\\pi}{2a}(x + a)\\right], \\quad x \\in [-a, a]&quot;,&quot;id&quot;:&quot;750879847F&quot;}" data-component-name="LatexBlockToDOM"></div><p><code>&#968;_g(x)</code> is the ground state wavefunction. The factor <code>1/&#8730;a</code> ensures the wavefunction is normalized. The argument <code>&#960;(x + a) / (2a)</code> shifts and scales the sine so that it is zero at <code>x = -a</code>, reaches its maximum at <code>x = 0</code>, and returns to zero at <code>x = a</code>. Outside the well the wavefunction is identically zero.</p><p>The probability density is the squared modulus of the wavefunction, which yields Eq. (29):</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;P_{\\text{QW}}(x) = \\begin{cases} \\frac{1}{a} \\sin^2\\left[\\frac{\\pi}{2a}(x + a)\\right] &amp; x \\in [-a, a] \\\\ 0 &amp; \\text{otherwise} \\end{cases}&quot;,&quot;id&quot;:&quot;5657C111CE&quot;}" data-component-name="LatexBlockToDOM"></div><p><code>P_QW(x)</code> is the probability density for the quantum well model. This density is zero outside <code>[-a, a]</code>, eliminating any chance of extreme returns. Inside the well it has a <code>sin&#178;</code> shape that peaks at <code>x = 0</code> with value <code>1/a</code> and goes smoothly to zero at the boundaries. The density is symmetric and integrates to 1 over the interval.</p><h3>Effective volatility</h3><p>The effective volatility is defined by Eq. (11):</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\sigma_{\\text{eff}} = \\sigma \\sigma_{\\text{QM}}, \\quad \\sigma_{\\text{QM}} = \\sqrt{\\langle x^2 \\rangle_{\\text{QM}} - \\langle x \\rangle_{\\text{QM}}^2}&quot;,&quot;id&quot;:&quot;18D44268A1&quot;}" data-component-name="LatexBlockToDOM"></div><p><code>&#963;_eff</code> is the effective volatility that replaces the original <code>&#963;</code> in the Black&#8209;Scholes formula. <code>&#963;_QM</code> is the quantum standard deviation computed from the probability density <code>P_QW(x)</code>. The expectation values <code>&#10216;x&#10217;_QM</code> and <code>&#10216;x&#178;&#10217;_QM</code> are integrals over the density: <code>&#10216;x&#10217;_QM = &#8747; x P_QW(x) dx</code> and <code>&#10216;x&#178;&#10217;_QM = &#8747; x&#178; P_QW(x) dx</code>.</p><p>For the quantum well, the density is symmetric about zero, so the mean <code>&#10216;x&#10217;_QM</code> is zero. The second moment can be computed exactly by integrating <code>x&#178;</code> times the density over <code>[-a, a]</code>:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\langle x^2 \\rangle_{\\text{QM}} = \\int_{-a}^{a} x^2 \\cdot \\frac{1}{a} \\sin^2\\!\\left[\\frac{\\pi(x+a)}{2a}\\right] dx = \\frac{a^2}{3} - \\frac{2a^2}{\\pi^2}&quot;,&quot;id&quot;:&quot;E0FF405E35&quot;}" data-component-name="LatexBlockToDOM"></div><p>Consequently, the quantum standard deviation is <code>&#963;_QM = &#8730;(a&#178;/3 - 2a&#178;/&#960;&#178;)</code>. This analytical formula is exact and avoids numerical integration for the effective volatility. The variance is guaranteed to be non&#8209;negative because <code>1/3 &#8776; 0.3333</code> and <code>2/&#960;&#178; &#8776; 0.2026</code>, so the difference is positive for any <code>a &gt; 0</code>.</p><p>For small <code>a</code>, <code>&#963;_QM</code> is small, so the effective volatility is low and option prices are depressed. As <code>a</code> grows, <code>&#963;_QM</code> increases, and the option price rises. However, the option price saturates for large <code>a</code> rather than converging to the Black&#8209;Scholes price. This is because the well always imposes hard bounds: the density is exactly zero outside <code>[-a, a]</code>, so the tails of the distribution are always truncated, no matter how wide the well becomes. The cumulative distribution <code>N_eff(d_eff&#177;)</code> therefore never matches the standard normal CDF, and the option price remains below the unbounded Black&#8209;Scholes value.</p><h3>Modified option pricing with the quantum well</h3><p>To price an option under the quantum well model, the effective volatility <code>&#963;_eff</code> is used in place of <code>&#963;</code> when computing the integration limits <code>d_eff&#177;</code> via the standard formula (Eq. 3, reconstructed):</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;d_{\\pm} = \\frac{\\ln(S_0/K) + (r \\pm \\sigma^2/2)T}{\\sigma\\sqrt{T}}&quot;,&quot;id&quot;:&quot;44B146F484&quot;}" data-component-name="LatexBlockToDOM"></div><p>With <code>&#963;</code> replaced by <code>&#963;_eff</code>, the modified limits become:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;d_{\\text{eff}\\pm} = \\frac{\\ln(S_0/K) + (r \\pm \\sigma_{\\text{eff}}^2/2)T}{\\sigma_{\\text{eff}}\\sqrt{T}}&quot;,&quot;id&quot;:&quot;D16D60180D&quot;}" data-component-name="LatexBlockToDOM"></div><p>The modified cumulative distribution <code>N_eff(d_eff&#177;)</code> is then obtained by numerically integrating <code>P_QW(x)</code> from <code>-&#8734;</code> to <code>d_eff&#177;</code>. Because <code>P_QW(x)</code> is zero outside <code>[-a, a]</code>, the integration domain is effectively truncated to <code>[-a, min(d_eff&#177;, a)]</code>. The option prices are computed using the standard Black&#8209;Scholes formula structure (Eqs. 1&#8209;2) with <code>&#963;_eff</code> and <code>N_eff</code> replacing <code>&#963;</code> and <code>N</code>.</p><h3>C++ implementation</h3><p>The header <code>src/quantum_well_model.h</code> provides two core functions: one for the probability density and one for the analytical <code>&#963;_QM</code>. Both are straightforward translations of the equations above.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;368607d8-f7b3-404c-b24e-0531fa62c99e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">inline double quantum_well_sigma_QM(double a) {
    if (a &lt;= 0.0) {
        throw std::invalid_argument("quantum_well_sigma_QM: a must be positive, got " + std::to_string(a));
    }
    const double a2 = a * a;
    const double var = a2 / 3.0 - 2.0 * a2 / (M_PI * M_PI);
    return std::sqrt(var);
}</code></pre></div><p>The function <code>quantum_well_sigma_QM</code> computes the exact standard deviation. Input validation ensures <code>a &gt; 0</code>, throwing <code>std::invalid_argument</code> otherwise. The variance is guaranteed to be non&#8209;negative because <code>1/3 &#8776; 0.3333</code> and <code>2/&#960;&#178; &#8776; 0.2026</code>, so the difference is positive for any <code>a &gt; 0</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;4dbfa69e-6488-499d-95c7-86824d3244a9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">inline double quantum_well_PDF(double x, double a) {
    if (a &lt;= 0.0) {
        throw std::invalid_argument("quantum_well_PDF: a must be positive, got " + std::to_string(a));
    }
    if (x &lt;= -a || x &gt;= a) {
        return 0.0;
    }
    const double argument = M_PI * (x + a) / (2.0 * a);
    const double sin_val = std::sin(argument);
    return (1.0 / a) * sin_val * sin_val;
}</code></pre></div><p>The PDF function first validates <code>a &gt; 0</code>, then checks whether <code>x</code> lies outside the well; if so, it returns zero immediately. Inside the well, it evaluates the <code>sin&#178;</code> expression exactly as given in Eq. (29).</p><h3>Verification</h3><p>The header file includes a comprehensive verification function <code>verify_quantum_well_model()</code> that runs several static checks. Here is an excerpt showing the normalization check and the analytical&#8209;vs&#8209;numerical <code>&#963;_QM</code> check:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;f8924e70-3edb-49f9-bded-3802d9240afc&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">// Check 1: Integral of P_QW over [-a, a] &#8776; 1.0
for (double a : a_values) {
    const int N = 20001;
    std::vector&lt;double&gt; x(N);
    std::vector&lt;double&gt; f(N);
    const double dx = 2.0 * a / (N - 1);
    for (int i = 0; i &lt; N; ++i) {
        x[i] = -a + i * dx;
        f[i] = quantum_well_PDF(x[i], a);
    }
    const double integral = simpson_integrate(x, f);
    const bool passed = std::abs(integral - 1.0) &lt; tolerance;
    // ... pass/fail reporting ...
}

// Check 2: Analytical &#963;_QM vs numerical integration
for (double a : a_values) {
    const double sigma_analytical = quantum_well_sigma_QM(a);
    // Numerical computation of &#963;_QM via Simpson integration of x*f and x&#178;*f
    // ...
    const bool passed = std::abs(sigma_analytical - sigma_numerical) &lt; tolerance;
    // ... pass/fail reporting ...
}</code></pre></div><p>The full verification function performs seven checks:</p><ul><li><p><strong>Normalization</strong>: The integral of <code>P_QW</code> over <code>[-a, a]</code> equals 1.0 for several test values of <code>a</code>.</p></li><li><p><strong>Analytical vs. numerical &#963;_QM</strong>: The analytical formula matches a numerical integration of <code>&#10216;x&#178;&#10217;</code> and <code>&#10216;x&#10217;</code>.</p></li><li><p><strong>Zero outside the well</strong>: <code>P_QW(x)</code> returns exactly 0 for <code>|x| &#8805; a</code>.</p></li><li><p><strong>Non&#8209;negativity</strong>: The density is never negative inside the well.</p></li><li><p><strong>Small&#8209;a limit</strong>: As <code>a &#8594; 0</code>, <code>&#963;_QM &#8594; 0</code>.</p></li><li><p><strong>Peak value</strong>: <code>P_QW(0, a) = 1/a</code>.</p></li><li><p><strong>Boundary condition</strong>: <code>P_QW(&#177;a, a) = 0</code>.</p></li></ul><p>These checks ensure that the implementation faithfully reproduces the paper&#8217;s model and that the probability density is a valid distribution. When all checks pass, the model is ready to be used in the parameter sweeps that reproduce Figure 5 of the paper.</p><h3>Worked example</h3><p>Take the paper&#8217;s standard parameters: <code>S0 = 20</code>, <code>K = 20</code>, <code>r = 0.1</code>, <code>T = 1</code>, <code>&#963; = 0.25</code>. For a well half&#8209;width <code>a = 2.0</code>:</p><ol><li><p><strong>Quantum standard deviation</strong>: <code>&#963;_QM = quantum_well_sigma_QM(2.0) &#8776; 0.737</code>.</p></li><li><p><strong>Effective volatility</strong>: <code>&#963;_eff = 0.25 &#215; 0.737 &#8776; 0.184</code>.</p></li><li><p><strong>Modified d_eff&#177;</strong>: Using <code>&#963;_eff</code> in the <code>d&#177;</code> formula:</p></li></ol><ul><li><p><code>d_eff+ = [ln(20/20) + (0.1 + 0.184&#178;/2) &#215; 1] / (0.184 &#215; &#8730;1) &#8776; 0.635</code></p></li><li><p><code>d_eff- = d_eff+ - 0.184 &#8776; 0.451</code></p></li></ul><ol start="4"><li><p><strong>Modified CDF</strong>: Integrate <code>P_QW(x)</code> from <code>-&#8734;</code> to <code>d_eff&#177;</code>. Since <code>d_eff&#177;</code> are within <code>[-2, 2]</code>, the integration uses the <code>sin&#178;</code> density. The resulting <code>N_eff(d_eff+)</code> and <code>N_eff(d_eff-)</code> are smaller than their standard normal counterparts because the distribution is narrower.</p></li><li><p><strong>Option price</strong>: The call price computed via the modified Black&#8209;Scholes formula is lower than the baseline 3.022.</p></li></ol><p>The code call <code>quantum_well_PDF(x, 2.0)</code> returns the density at any <code>x</code>, and <code>quantum_well_sigma_QM(2.0)</code> returns approximately 0.737.</p><h3>Reproducing Figure 5</h3><p>Figure 5 of the paper plots the option price against the well half&#8209;width <code>a</code>. The parameter sweep function <code>sweep_quantum_well</code> in <code>src/paper_figures.h</code> varies <code>a</code> over a range (e.g., 0.1 to 5.0) while keeping <code>S0</code>, <code>K</code>, <code>r</code>, <code>T</code>, and <code>&#963;</code> fixed at the standard values. For each <code>a</code>, it builds a CDF grid from <code>quantum_well_PDF</code>, computes <code>&#963;_eff</code> via <code>quantum_well_sigma_QM</code>, and calls <code>modified_option_pricing</code>. The resulting call price increases with <code>a</code> and saturates, never reaching the Black&#8209;Scholes baseline.</p><h3>Limitations</h3><p>The quantum well model is a stylized representation of hard price bounds. The abrupt cutoff at <code>x = &#177;a</code> is unrealistic for most markets, where price limits are rarely absolute. The parameter <code>a</code> is only qualitatively related to the price range <code>&#916;S</code>; the paper provides no calibration procedure. Because the density is always zero outside the well, the model never converges to the Black&#8209;Scholes price, even for arbitrarily large <code>a</code>. Despite these limitations, the model illustrates how a simple potential can capture the effect of price boundaries on option valuation, completing the set of five market force models presented in the paper.</p><h2>Code Walkthrough: File Structure, Key Functions, and Data Flow</h2><p>Now that we have studied every market force model and the machinery that turns a probability density into an option price, it is time to look at how all the pieces fit together inside the C++ codebase. The implementation is organised as a collection of header&#8209;only libraries, making it easy to include and reuse individual components. This section presents the file dependency hierarchy, walks through the purpose of each module, and traces the exact path a single force parameter takes from input to a pair of call and put prices. We will also see how the sweep functions that reproduce the paper&#8217;s figures are built and how the final results are written to CSV files for external plotting.</p><h3>File Dependency Hierarchy</h3><p>The project follows a strict, layered build order. Every file depends only on modules that are listed above it, and no circular dependencies exist. The recommended inclusion order is:</p><ol><li><p><strong>`src/black_scholes_base.h`</strong> &#8211; the baseline Black&#8209;Scholes formulas and the standard normal CDF/PDF. No dependencies.</p></li><li><p><strong>`src/effective_volatility.h`</strong> &#8211; the Simpson integrator and the <code>compute_eff_vol</code> function. The primary functions in this header (<code>simpson_integrate</code> and <code>compute_eff_vol</code>) have no dependency on <code>normalPDF</code>. The verification function <code>verify_effective_volatility</code> defines a local lambda that matches the standard normal PDF for its own checks, so the header does not need to include <code>black_scholes_base.h</code>.</p></li><li><p><strong>`src/modified_option_pricing.h`</strong> &#8211; the modified Black&#8209;Scholes pricing that uses a precomputed CDF grid and effective volatility. It includes both <code>black_scholes_base.h</code> (for <code>compute_dpm</code> and <code>normalPDF</code>) and <code>effective_volatility.h</code> (for the <code>VolResult</code> struct). Note that <code>build_cdf_grid</code> implements its own cumulative integration internally and does not call <code>simpson_integrate</code> from <code>effective_volatility.h</code>; <code>simpson_integrate</code> is used elsewhere, such as in <code>compute_eff_vol</code> and the normalisation functions of the perturbative models.</p></li><li><p><strong>Market force model headers</strong> &#8211; each file (<code>constant_force_model.h</code>, <code>linear_force_model.h</code>, <code>x2_force_model.h</code>, <code>x3_force_model.h</code>, <code>quantum_well_model.h</code>) defines one probability density function and its helpers. They do <strong>not</strong> include the core headers directly. Instead they use forward declarations (e.g., <code>extern double normalPDF(double x);</code> or namespace wrappers) for the symbols they need during verification. The symbols are resolved at link time when all headers are included together in a translation unit such as <code>main.cpp</code>.</p></li><li><p><strong>`src/paper_figures.h`</strong> &#8211; the sweep functions that reproduce Figures&#8239;1&#8209;5. This header depends on all five model headers plus the three core modules. It uses forward declarations to avoid including everything at once, but the user must include the required headers before this file.</p></li><li><p><strong>`src/main.cpp`</strong> &#8211; the entry point that calls all verification routines and then generates the CSV data for the figures. It includes every header listed above.</p></li></ol><p>If you are compiling the project from scratch, you only need to compile <code>main.cpp</code>; all other code is pulled in through the headers.</p><h3>Core Modules</h3><h4><code>black_scholes_base.h</code> &#8211; The Baseline</h4><p>This file provides the four building blocks that the rest of the code depends on:</p><ul><li><p><code>normalPDF(double x)</code> &#8211; the standard normal density <code>P_BS(x)</code>.</p></li><li><p><code>normalCDF(double x)</code> &#8211; the standard normal cumulative distribution <code>N(x)</code> implemented via <code>std::erfc</code>.</p></li><li><p><code>compute_dpm(S0, K, r, sigma, T)</code> &#8211; returns <code>d_+</code> and <code>d_-</code> from the reconstructed Eq.&#8239;(3).</p></li><li><p><code>black_scholes_base(&#8230;)</code> &#8211; returns a <code>BSResult</code> struct with call, put, and the intermediate <code>d</code> and <code>N(d)</code> values.</p></li></ul><p>A small excerpt shows the interface:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;d53d6f4f-4b8d-4f4a-be69-52d5de76f748&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">struct BSResult {
    double call, put, d_plus, d_minus, N_d_plus, N_d_minus;
};

BSResult black_scholes_base(double S0, double K, double r, double T, double sigma);</code></pre></div><p>Every subsequent module uses <code>normalPDF</code> and <code>compute_dpm</code>; the CDF <code>normalCDF</code> is used only for verification and as a baseline reference.</p><h4><code>effective_volatility.h</code> &#8211; The Integrator and &#963;_QM</h4><p>This header defines the <code>VolResult</code> struct and the two workhorses of numerical integration:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;1450fcfa-a2e1-4b7c-a8e4-276a163a5768&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">struct VolResult {
    double sigma_eff, sigma_QM, mean_x, var_x;
};

double simpson_integrate(const std::vector&lt;double&gt;&amp; x, const std::vector&lt;double&gt;&amp; f);
VolResult compute_eff_vol(double sigma, std::function&lt;double(double)&gt; P,
                          double a = -12.0, double b = 12.0, int N = 20001);</code></pre></div><p><code>simpson_integrate</code> applies Simpson&#8217;s 1/3 rule on a uniform grid. It is used by <code>compute_eff_vol</code> and by the normalisation functions of the perturbative models. The function <code>compute_eff_vol</code> builds a grid, evaluates <code>P(x)</code>, <code>x&#183;P(x)</code>, and <code>x&#178;&#183;P(x)</code>, and returns the effective volatility <code>&#963;_eff = &#963; &#183; &#963;_QM</code>.</p><h4><code>modified_option_pricing.h</code> &#8211; The Modified Black&#8209;Scholes Engine</h4><p>This header is the glue that combines the new probability density with the Black&#8209;Scholes formula structure. It provides:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;82fb047e-0ee8-48a4-8679-a1c2e5673264&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">struct ModifiedBSResult {
    double call, put, d_eff_plus, d_eff_minus, N_eff_plus, N_eff_minus;
};

std::pair&lt;std::vector&lt;double&gt;, std::vector&lt;double&gt;&gt;
build_cdf_grid(std::function&lt;double(double)&gt; P, double a, double b, int N);

double interpolate_cdf(double d, const std::vector&lt;double&gt;&amp; x_grid,
                       const std::vector&lt;double&gt;&amp; cdf_grid);

ModifiedBSResult modified_option_pricing(double S0, double K, double r, double T,
                                         double sigma_eff,
                                         std::function&lt;double(double)&gt; P,
                                         const std::vector&lt;double&gt;&amp; x_grid,
                                         const std::vector&lt;double&gt;&amp; cdf_grid);</code></pre></div><ul><li><p><code>build_cdf_grid</code> precomputes the CDF <code>N_eff(x) = &#8747;_{a}^{x} P(t) dt</code> on a uniform grid. It uses a mixed integration scheme: Simpson&#8217;s rule for completed panels (even indices) and the trapezoidal rule for the first interval and for half&#8209;panels at odd indices. The first element is <code>0.0</code>, and the last should be approximately <code>1.0</code> for a properly normalised density.</p></li><li><p><code>interpolate_cdf</code> performs linear interpolation between the two nearest grid points to evaluate the CDF at arbitrary <code>d_eff&#177;</code> values. When <code>d</code> falls outside the grid, it clamps the result to <code>0.0</code> for <code>d &#8804; x_grid[0]</code> and to <code>cdf_grid.back()</code> (which is approximately <code>1.0</code> for a normalised density) for <code>d &#8805; x_grid[n-1]</code>.</p></li><li><p><code>modified_option_pricing</code> calls <code>compute_dpm</code> with <code>sigma_eff</code> to obtain <code>d_eff&#177;</code>, then uses <code>interpolate_cdf</code> to get <code>N_eff(d_eff&#177;)</code>, and finally plugs everything into the call and put formulas. The <code>P</code> parameter is not used inside the function body; it is retained in the signature for interface consistency and documentation.</p></li></ul><p>When <code>P</code> is the standard normal PDF and <code>sigma_eff</code> equals the original <code>&#963;</code>, the function reproduces the baseline Black&#8209;Scholes prices to within <code>1e-10</code>.</p><h3>Market Force Model Headers</h3><p>Each of the five model files follows a uniform pattern:</p><ul><li><p>One or two helper functions that compute analytical parameters (e.g., <code>constant_force_xk</code>, <code>linear_force_sigma_QM</code>, <code>quantum_well_sigma_QM</code>).</p></li><li><p>A function <code>*_PDF(x, param)</code> that returns the value of the probability density at <code>x</code>.</p></li><li><p>For the perturbative models (<code>x2_force_model.h</code> and <code>x3_force_model.h</code>), a separate <code>compute_*_normalization</code> function and a <code>*_PDF</code> that takes the precomputed normalisation constant <code>C2</code>.</p></li><li><p>A <code>verify_*</code> function that checks normalisation, zero&#8209;force recovery, and (for the <code>x^2</code> and <code>x^3</code> models) warns when the density becomes negative.</p></li></ul><p>For example, the constant force model exposes:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;8cee8e75-a113-42bb-9bda-63f599e4a6fc&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">double constant_force_xk(double k);
double constant_force_PDF(double x, double k);</code></pre></div><p>The linear force model provides analytical <code>&#963;_QM</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;034a06b6-9a6f-46c9-b57c-ddb44ea2cea3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">double linear_force_lambda_omega(double lambda);
double linear_force_sigma_QM(double lambda);
double linear_force_PDF(double x, double lambda);</code></pre></div><p>The perturbative models require a normalisation step:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;a5ecb089-6bd9-4a07-9989-f28e5385cb90&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">// x&#178; force
double compute_x2_normalization(double beta, double a = -12.0, double b = 12.0, int N = 20001);
double x2_force_PDF(double x, double beta, double C2);

// x&#179; force
double compute_x3_normalization(double gamma, double a = -12.0, double b = 12.0, int N = 20001);
double x3_force_PDF(double x, double gamma, double C2);</code></pre></div><p>The quantum well model has an analytical <code>&#963;_QM</code> and a piecewise PDF:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;a0160d91-2649-4604-a23c-ae655c6e8f9e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">double quantum_well_sigma_QM(double a);
double quantum_well_PDF(double x, double a);</code></pre></div><p>All these functions are <code>inline</code> and defined in the headers, so no separate compilation is required.</p><h3>Sweep Functions and Figure Generation</h3><p>The file <code>src/paper_figures.h</code> contains five sweep functions, one for each figure in the paper. They all share the same signature pattern:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;99cac87e-fbaf-446b-aaf9-7ec0af17d5e4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">SweepResult sweep_constant_force(double S0, double K, double r, double T, double sigma,
                                 double k_min, double k_max, int n_steps);
SweepResult sweep_linear_force(double S0, double K, double r, double T, double sigma,
                                double lambda_min, double lambda_max, int n_steps);
// ... and similarly for x2, x3, and quantum_well</code></pre></div><p>A <code>SweepResult</code> is simply a struct with three vectors:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;081b5272-6079-4ec3-935d-141273a2d0ed&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">struct SweepResult {
    std::vector&lt;double&gt; param_values;
    std::vector&lt;double&gt; call_prices;
    std::vector&lt;double&gt; put_prices;
};</code></pre></div><p>Each sweep function iterates over the requested parameter range, performs the following steps for every parameter value, and records the resulting call and put prices.</p><p>Let us trace the <code>sweep_constant_force</code> function as a concrete example. The code, slightly simplified, looks like this:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;ec65a4ca-e89c-4375-92a7-eb245b8bc6ff&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">SweepResult sweep_constant_force(double S0, double K, double r, double T, double sigma,
                                 double k_min, double k_max, int n_steps)
{
    SweepResult result;
    const double a = -12.0, b = 12.0;
    const int N = 20001;
    const double sigma_eff = sigma;   // variance unchanged for constant force

    for (int i = 0; i &lt; n_steps; ++i) {
        double k = k_min + (k_max - k_min) * i / (n_steps - 1);

        auto P = [k](double x) { return constant_force_PDF(x, k); };
        auto [x_grid, cdf_grid] = build_cdf_grid(P, a, b, N);
        ModifiedBSResult mbs = modified_option_pricing(S0, K, r, T, sigma_eff, P, x_grid, cdf_grid);

        result.param_values.push_back(k);
        result.call_prices.push_back(mbs.call);
        result.put_prices.push_back(mbs.put);
    }
    return result;
}</code></pre></div><ul><li><p><strong>Step 1 &#8211; Create the density:</strong> A lambda captures the current <code>k</code> and calls <code>constant_force_PDF(x, k)</code>. This lambda is a valid <code>std::function&lt;double(double)&gt;</code>.</p></li><li><p><strong>Step 2 &#8211; Build the CDF grid:</strong> <code>build_cdf_grid</code> evaluates the lambda on <code>N</code> points from <code>-12</code> to <code>12</code> and accumulates the integral using the mixed Simpson/trapezoidal scheme described above. The result is a pair of grids: <code>x_grid</code> and <code>cdf_grid</code>.</p></li><li><p><strong>Step 3 &#8211; Compute effective volatility:</strong> For the constant force model we know analytically that <code>&#963;_QM = 1</code>, so <code>sigma_eff = sigma</code>. The code does not call <code>compute_eff_vol</code> here; it simply uses <code>sigma</code> directly. The same is true for <code>sweep_linear_force</code> and <code>sweep_quantum_well</code>, which use analytical <code>&#963;_QM</code> formulas (<code>linear_force_sigma_QM</code> and <code>quantum_well_sigma_QM</code>). Only the perturbative sweeps (<code>sweep_x2_force</code> and <code>sweep_x3_force</code>) call <code>compute_eff_vol</code> because their variance changes in a way that is not captured by a simple analytical expression.</p></li><li><p><strong>Step 4 &#8211; Price the option:</strong> <code>modified_option_pricing</code> takes the original market parameters, <code>sigma_eff</code>, the density lambda (retained in the signature for interface consistency but not used inside the function), and the precomputed grids. It computes <code>d_eff&#177;</code>, interpolates the CDF, and returns the modified call and put prices.</p></li><li><p><strong>Step 5 &#8211; Store the results:</strong> The parameter value and the two prices are appended to the <code>SweepResult</code> vectors.</p></li></ul><p>The other sweep functions follow the same logic, with the only differences being the density function and the method for obtaining <code>sigma_eff</code>.</p><h3>The Entry Point: <code>main.cpp</code></h3><p>The file <code>main.cpp</code> orchestrates the entire workflow. It defines two helper functions&#8212;<code>run_verification</code> and <code>generate_all_figures</code>&#8212;and a <code>main</code> that calls them in sequence. The programme uses the paper&#8217;s sample parameters (<code>S0=20</code>, <code>K=20</code>, <code>r=0.1</code>, <code>T=1</code>, <code>&#963;=0.25</code>) throughout.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;9c4b8942-00e0-49d8-b0ff-0824d6767550&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">int main() {
    try {
        run_verification();
        generate_all_figures();
        return 0;
    } catch (const std::exception&amp; e) {
        std::cerr &lt;&lt; "Error: " &lt;&lt; e.what() &lt;&lt; "\n";
        return 1;
    }
}</code></pre></div><p><code>run_verification</code> calls every <code>verify_*</code> function from the corresponding headers. If any check fails, it sets a flag and, after all checks have run, throws <code>std::runtime_error("Verification failure")</code> to stop the programme before any figure data is written. This ensures that users are immediately alerted to inconsistencies.</p><p><code>generate_all_figures</code> invokes the five sweep functions with parameter ranges chosen to reproduce the paper&#8217;s Figures&#8239;1&#8209;5. For example, the constant force sweep is called as:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;6726584c-9c85-45ef-868d-04a7fc6916b6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">SweepResult res = sweep_constant_force(20.0, 20.0, 0.1, 1.0, 0.25, -1.0, 1.0, 50);
write_sweep_csv("fig1_constant_force.csv", res, "k");</code></pre></div><p>Each sweep result is written to a CSV file by a helper function <code>write_sweep_csv</code>. The output format is straightforward:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;0268b71f-231b-48ac-a2a1-5d28339159a0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">k,call_price,put_price
-1.0000000000,3.022...,1.119...
-0.9591836735,3.020...,1.117...
...</code></pre></div><p>The first column is the force parameter, the second is the modified call price, and the third is the modified put price. The values are printed with <code>std::setprecision(10)</code>, which sets the number of significant digits (not decimal places). For values around 3.0 this gives about nine decimal places; for values around 0.1 it gives about eleven. These CSV files can be loaded into any plotting tool (Python&#8217;s matplotlib, Excel, gnuplot) to generate the figures.</p><h3>Data Flow for a Single Sweep</h3><p>To summarise the data flow, let us follow a single parameter, say <code>k = 0.5</code>, through the constant force sweep:</p><ol><li><p><code>constant_force_PDF</code> is called with <code>x</code> values from <code>-12</code> to <code>12</code> and <code>k = 0.5</code>.</p></li><li><p><code>build_cdf_grid</code> accumulates the integral of <code>P_k(x)</code> and returns <code>x_grid</code> and <code>cdf_grid</code>.</p></li><li><p><code>modified_option_pricing</code> receives <code>sigma_eff = 0.25</code>, the density lambda, and the grids.</p></li><li><p>Inside <code>modified_option_pricing</code>, <code>compute_dpm(20, 20, 0.1, 0.25, 1.0)</code> returns <code>d_eff+ &#8776; 0.525</code> and <code>d_eff- &#8776; 0.275</code>.</p></li><li><p><code>interpolate_cdf</code> evaluates the CDF at <code>d_eff+</code> and <code>d_eff-</code> using the precomputed grid.</p></li><li><p>The call price is computed as <code>20 * N_eff(0.525) - 20 * exp(-0.1) * N_eff(0.275)</code>.</p></li><li><p>The result is appended to <code>SweepResult.call_prices</code>.</p></li></ol><p>This identical pattern is used for every force model, every parameter, and every point in the figures. The separation of concerns&#8212;density definition, CDF construction, effective volatility computation, and Black&#8209;Scholes evaluation&#8212;makes the code easy to extend. Adding a new force model requires only writing a new <code>*_PDF</code> function and a corresponding sweep function, without touching the core pricing engine.</p><h3>How to Plot the Results</h3><p>After running the programme, you will have five CSV files. A minimal Python snippet to plot Figure&#8239;1 looks like:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2b283a38-a284-4662-a7ac-8a68af328a46&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import pandas as pd
import matplotlib.pyplot as plt
data = pd.read_csv('fig1_constant_force.csv')
plt.plot(data['k'], data['call_price'], label='Call')
plt.plot(data['k'], data['put_price'], label='Put')
plt.xlabel('k')
plt.ylabel('Option Price')
plt.legend()
plt.show()</code></pre></div><p>No built&#8209;in plotting library is used in the C++ code; the separation of computation and visualisation keeps the implementation lightweight and portable.</p><p>With this walkthrough, you now have a complete picture of how the code is structured and how the mathematical models from the paper are translated into a working C++ implementation. The next section will verify that every component behaves correctly and that the paper&#8217;s figures are faithfully reproduced.</p><h2>Reproducing the Paper's Figures: Parameter Sweeps and Plotting</h2><p>The previous sections built the core machinery that turns a probability density and an effective volatility into a modified option price. We now use that machinery to recreate the five qualitative graphs shown in the paper. Each graph plots the call option price against the strength parameter of one market force, while all standard Black&#8209;Scholes inputs are held fixed.</p><p>The paper itself provides only illustrative sketches and does not give the exact parameter ranges. The sweep functions in <code>src/paper_figures.h</code> make the ranges explicit so you can reproduce the described behaviour and experiment with your own intervals. This section explains the fixed financial parameters, introduces the <code>SweepResult</code> data structure and the numerical integration domain, walks through each sweep function, and shows how to export the resulting data to CSV files for plotting with external tools.</p><h3>Fixed Financial Parameters</h3><p>All five sweeps share the same underlying asset and option parameters. They are chosen so the baseline Black&#8209;Scholes call price is approximately 3.022 and the put price approximately 1.119:</p><ul><li><p>Current asset price <code>S0</code> = 20</p></li><li><p>Strike price <code>K</code> = 20 (at&#8209;the&#8209;money)</p></li><li><p>Risk&#8209;free rate <code>r</code> = 0.1 (10 % per annum)</p></li><li><p>Time to maturity <code>T</code> = 1 year</p></li><li><p>Volatility <code>&#963;</code> = 0.25 (25 % per annum)</p></li></ul><p>A crucial consistency requirement is that when the force parameter is zero (or its equivalent neutral value), the modified option price must recover the baseline Black&#8209;Scholes price. Every sweep function is designed to honour this property, and the built&#8209;in verification routine <code>verify_paper_figures</code> formally checks it.</p><h3>The SweepResult Structure</h3><p>All sweep functions return a <code>SweepResult</code> struct, defined in <code>src/paper_figures.h</code>. It holds three <code>std::vector&lt;double&gt;</code> members of equal length:</p><ul><li><p><code>param_values</code> &#8211; the force parameter at each step (<code>k</code>, <code>&#955;</code>, <code>&#946;</code>, <code>&#947;</code>, or <code>a</code>),</p></li><li><p><code>call_prices</code> &#8211; the corresponding modified call option price,</p></li><li><p><code>put_prices</code> &#8211; the corresponding modified put option price.</p></li></ul><p>The uniform structure lets you loop over a sweep and write rows to a CSV file without knowing which force model produced the data.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;637f02c6-f233-4480-8197-458e7ec5c85a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">struct SweepResult {
    std::vector&lt;double&gt; param_values;
    std::vector&lt;double&gt; call_prices;
    std::vector&lt;double&gt; put_prices;
};</code></pre></div><p>Every sweep function accepts the five fixed Black&#8209;Scholes parameters (<code>S0</code>, <code>K</code>, <code>r</code>, <code>T</code>, <code>&#963;</code>), the minimum and maximum of the force parameter, and an integer <code>n_steps</code> (&#8805;&#8239;2) that controls the number of equally spaced points. The first entry always equals <code>param_min</code> and the last equals <code>param_max</code>.</p><h3>Numerical Integration Domain</h3><p>For unbounded densities the sweep functions rely on the same numerical integration domain used throughout the codebase. The constants are gathered in the <code>SweepConfig</code> namespace inside <code>src/paper_figures.h</code>:</p><ul><li><p><code>SweepConfig::INTEGRAL_A</code> = -12.0 &#8211; lower bound</p></li><li><p><code>SweepConfig::INTEGRAL_B</code> = 12.0  &#8211; upper bound</p></li><li><p><code>SweepConfig::INTEGRAL_N</code> = 20001 &#8211; number of Simpson grid points (must be odd)</p></li></ul><p>These values are passed to <code>build_cdf_grid</code> and <code>compute_eff_vol</code> for every sweep model except the quantum well, which uses its own bounded support <code>[-a, a]</code> with the same number of points. The domain <code>[-12, 12]</code> is wide enough to capture the tails of the standard normal distribution and all the perturbed densities used in the paper.</p><h3>Constant Market Force (Figure 1)</h3><p>The constant force model uses a shifted Gaussian whose variance is unchanged; therefore the standard deviation of the quantum density is one, <code>&#963;_eff</code> equals the original <code>&#963;</code>, and no numerical integration for the effective volatility is needed. The sweep function <code>sweep_constant_force</code> iterates over <code>k</code>, builds a probability density through <code>constant_force_PDF(x, k)</code>, constructs a CDF grid with <code>build_cdf_grid</code>, and calls <code>modified_option_pricing</code>.</p><p>A call that reproduces Figure 1 is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;f565dd4a-8d7f-4dc1-85d8-237b793233a9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">SweepResult fig1 = sweep_constant_force(
    20.0, 20.0, 0.1, 1.0, 0.25,   // S0, K, r, T, sigma
    -2.0, 2.0, 50                   // k_min, k_max, n_steps
);</code></pre></div><p>The loop body inside the sweep function is shown below. Notice that <code>sigma_eff</code> is assigned the original <code>sigma</code> before the loop because the constant force only shifts the mean, leaving the width intact.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;2c0be9eb-8618-4975-a903-93a16546f98b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">const double sigma_eff = sigma;   // constant force does not change variance
for (int i = 0; i &lt; n_steps; ++i) {
    double k = k_min + (k_max - k_min) * static_cast&lt;double&gt;(i) / (n_steps - 1);

    auto P = [k](double x) { return constant_force_PDF(x, k); };
    auto [x_grid, cdf_grid] = build_cdf_grid(P, a, b, N);

    ModifiedBSResult mbs = modified_option_pricing(
        S0, K, r, T, sigma_eff, P, x_grid, cdf_grid);

    result.param_values.push_back(k);
    result.call_prices.push_back(mbs.call);
    result.put_prices.push_back(mbs.put);
}</code></pre></div><p>The call price decreases monotonically as <code>k</code> increases. A positive <code>k</code> gives a force <code>F = &#8722;k</code> that pushes the probability mass to the left, away from the strike, thereby reducing the probability that the option finishes in&#8209;the&#8209;money. For negative <code>k</code> the distribution shifts to the right, increasing the call price. When <code>k = 0</code> the density is exactly the standard normal and the sweep recovers the baseline Black&#8209;Scholes price.</p><h3>Linear Market Force (Figure 2)</h3><p>The linear force modifies the variance of the Gaussian. The quantum standard deviation is available analytically as <code>&#963;_QM = 1 / sqrt(1 + &#955;)</code>, provided by the helper <code>linear_force_sigma_QM</code>. The sweep function <code>sweep_linear_force</code> therefore bypasses numerical integration for effective volatility and computes <code>sigma_eff = &#963; &#183; &#963;_QM</code> directly.</p><p>A call that reproduces Figure 2 might be:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;e6a3cf82-04c1-4061-8b39-50149a6724ac&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">SweepResult fig2 = sweep_linear_force(
    20.0, 20.0, 0.1, 1.0, 0.25,   // S0, K, r, T, sigma
    0.1, 2.0, 50                    // lambda_min, lambda_max, n_steps
);</code></pre></div><p>Larger <code>&#955;</code> makes the bowl steeper, narrowing the distribution and reducing <code>&#963;_eff</code>. Because the probability of large price movements is lower, the at&#8209;the&#8209;money call option becomes less valuable. The call price therefore decreases monotonically with <code>&#955;</code>. At <code>&#955; = 0</code> the density reduces to the standard normal, and the sweep returns the baseline Black&#8209;Scholes price.</p><h3>x&#178; Market Force (Figure 3)</h3><p>The x&#178; force model uses the paper's perturbative density. As discussed earlier, that density contains a spurious linear term in <code>&#946;</code>. The sweep function <code>sweep_x2_force</code> nevertheless implements the paper's formula so you can reproduce the original figure. For each <code>&#946;</code> it recomputes the normalisation constant <code>C&#178;</code> through <code>compute_x2_normalization</code> and obtains the effective volatility numerically from the normalised density via <code>compute_eff_vol</code>.</p><p>A typical call is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;6da0d4d0-2b3c-478e-b8a6-70fe7ea49c2b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">SweepResult fig3 = sweep_x2_force(
    20.0, 20.0, 0.1, 1.0, 0.25,   // S0, K, r, T, sigma
    -0.5, 0.5, 50                   // beta_min, beta_max, n_steps
);</code></pre></div><p>The call price decreases as <code>&#946;</code> increases. This curve should be interpreted cautiously because of the perturbation error in the paper. The provided code does not contain a corrected second&#8209;order density; if you wish to replace the paper's formula with the correct expansion you need to modify the source files directly (see the discussion of the perturbation error in the earlier section on the x&#178; force model). At <code>&#946; = 0</code> the density is exactly the standard normal and the sweep recovers the baseline price.</p><h3>x&#179; Market Force (Figure 4)</h3><p>The x&#179; force model uses first&#8209;order perturbation theory for a quartic potential. The sweep function <code>sweep_x3_force</code> is called with a range of <code>&#947;</code> values. The paper reports a non&#8209;monotonic behaviour: the option price first decreases and then increases. This is reproduced with a call such as:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;5448d02a-5348-4133-b936-9e95718fb98f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">SweepResult fig4 = sweep_x3_force(
    20.0, 20.0, 0.1, 1.0, 0.25,   // S0, K, r, T, sigma
    0.0, 0.5, 50                    // gamma_min, gamma_max, n_steps
);</code></pre></div><p>For small <code>&#947;</code> the quartic term steepens the potential slightly, narrowing the central peak and reducing the call price. As <code>&#947;</code> grows larger the potential develops a double&#8209;well shape, which can create a bimodal density. The increased probability mass in the tails eventually raises the probability of finishing in&#8209;the&#8209;money, causing the call price to rise again. The density can become negative for large <code>&#947;</code>, signalling the breakdown of the perturbation expansion. The sweep function still produces numbers, but you should restrict <code>&#947;</code> to a range where the density stays positive. When <code>&#947; = 0</code> the density is the standard normal and the sweep returns the baseline Black&#8209;Scholes price.</p><h3>Quantum Well (Figure 5)</h3><p>The quantum well model confines the probability mass to the interval <code>[-a, a]</code>. The effective volatility is obtained analytically from <code>quantum_well_sigma_QM(a)</code>, which returns <code>&#8730;(a&#178;/3 &#8722; 2a&#178;/&#960;&#178;)</code>. The sweep function <code>sweep_quantum_well</code> builds the CDF grid only over the support <code>[-a, a]</code> to save computation.</p><p>To reproduce Figure 5, call:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;410f7046-11eb-4e1f-a3b1-cdb42b10eadd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">SweepResult fig5 = sweep_quantum_well(
    20.0, 20.0, 0.1, 1.0, 0.25,   // S0, K, r, T, sigma
    0.5, 5.0, 50                    // a_min, a_max, n_steps
);</code></pre></div><p>The call price increases with <code>a</code> and saturates for large <code>a</code>. A narrow well (small <code>a</code>) traps the probability mass near zero, making large movements impossible and drastically reducing the call option value. As the well widens, the distribution can spread further, and the option price rises. However, even for a very wide well the price never reaches the baseline Black&#8209;Scholes value because the hard boundaries always eliminate the extreme returns that a log&#8209;normal distribution would allow. The saturation level is therefore below the Black&#8209;Scholes price.</p><h3>Exporting to CSV and Plotting</h3><p>The sweep functions return the data in memory. To plot the figures you can write a short loop that outputs rows of <code>param_values[i], call_prices[i], put_prices[i]</code> to a comma&#8209;separated file. The <code>src/main.cpp</code> entry point follows this pattern&#8212;it calls all five sweep functions and writes the results to separate CSV files. A minimal helper you could write for your own project might look like this:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;2546296b-61ab-4a18-b426-f7ef696e4099&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">#include &lt;fstream&gt;
#include "paper_figures.h"

void write_sweep_to_csv(const SweepResult&amp; sweep, const std::string&amp; filename) {
    std::ofstream out(filename);
    out &lt;&lt; "param,call,put\n";
    for (size_t i = 0; i &lt; sweep.param_values.size(); ++i) {
        out &lt;&lt; sweep.param_values[i] &lt;&lt; ","
            &lt;&lt; sweep.call_prices[i]  &lt;&lt; ","
            &lt;&lt; sweep.put_prices[i]   &lt;&lt; "\n";
    }
}</code></pre></div><p>This helper is not part of the supplied code files; it is meant as an illustration of how to convert a <code>SweepResult</code> into CSV format. Once the CSV files exist, any plotting tool (Python with matplotlib, gnuplot, Excel, etc.) can create the graphs. The horizontal axis is the force parameter and the vertical axis is the call price (or put price). Drawing a horizontal line at the baseline Black&#8209;Scholes price makes the effect of the market force visually obvious.</p><h3>Built&#8209;in Verification</h3><p>The header <code>src/paper_figures.h</code> also provides a function <code>verify_paper_figures</code> that runs static integrity checks on the sweep machinery. It confirms that for zero force parameters the modified prices match the baseline and that call prices stay within the theoretical bounds <code>[0, S0]</code>. Call it before generating the figures to catch configuration errors early:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;85522c6b-4bfa-4d9e-ba99-8c2e8f285816&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">if (!verify_paper_figures()) {
    std::cerr &lt;&lt; "Paper figure verification failed.\n";
    return 1;
}</code></pre></div><p>With the sweep functions described above, you have a complete toolset for exploring how different market forces alter standard Black&#8209;Scholes prices. The next section covers the full verification strategy in more detail, and the final section summarises the key insights and limitations of the quantum approach.</p><h2>Verification: Static Checks, Invariants, and Paper-Fidelity Review</h2><p>Before we trust the option prices produced by our quantum market force models, we need to be confident that the code is correct. The implementation includes a comprehensive set of built&#8209;in verification functions that test mathematical invariants, check consistency with the paper's equations, and ensure that edge cases behave as expected. This section explains the verification strategy, walks through the key invariants that every model must satisfy, and shows how the C++ code in <code>src/main.cpp</code> orchestrates all the checks before any figure data is generated.</p><p><strong>Important:</strong> The verification described here is <strong>static and semantic only</strong>. No code execution is performed in this tutorial. The checks are designed to be compiled and run as part of a normal build process, but the tutorial itself does not execute them. The pass/fail messages shown are illustrative of what the code would produce when compiled and executed.</p><h3>The Verification Strategy</h3><p>The verification approach has three layers. First, the code is designed to compile cleanly with strict warnings enabled (<code>-Wall -Wextra -Werror</code>), catching potential bugs at the compiler level. Second, every header file contains a dedicated <code>verify_*</code> function that runs a battery of mathematical checks and prints pass/fail messages to the console. Third, the entry point <code>main()</code> calls all verification functions before generating any output; if any check fails, the program aborts with an error message rather than producing potentially incorrect results.</p><p>The checks are <strong>static and semantic</strong>&#8212;they test the mathematical properties of the code but do not require execution against external market data. The paper does not provide validation against real option prices, so the verification focuses on internal consistency: do the probability densities integrate to one? Does the modified pricing recover the baseline Black&#8209;Scholes when the force is turned off? Does put&#8209;call parity hold for every model?</p><h3>The Orchestration in <code>main.cpp</code></h3><p>The <code>run_verification</code> function in <code>src/main.cpp</code> calls every model's verification routine and collects the results. Here is the relevant excerpt:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;4fc3277a-7bdc-48a0-8180-5888de08391e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">void run_verification() {
    std::cout &lt;&lt; "=== Running verification checks ===\n";
    bool all_passed = true;

    auto check = [&amp;](const std::string&amp; name, bool result) {
        std::cout &lt;&lt; "  " &lt;&lt; name &lt;&lt; ": " &lt;&lt; (result ? "PASSED" : "FAILED") &lt;&lt; "\n";
        if (!result) all_passed = false;
    };

    check("Black-Scholes baseline", verify_black_scholes_base());
    check("Effective volatility", verify_effective_volatility());
    check("Modified option pricing", verify_modified_option_pricing());
    check("Constant force model", verify_constant_force_model());
    check("Linear force model", verify_linear_force_model());
    check("x^2 force model", verify_x2_force_model());
    check("x^3 force model", verify_x3_force_model());
    check("Quantum well model", verify_quantum_well_model());
    check("Paper figures (sweep sanity)", verify_paper_figures());

    if (!all_passed) {
        std::cerr &lt;&lt; "\nSome verification checks FAILED. Aborting figure generation.\n";
        throw std::runtime_error("Verification failure");
    }
    std::cout &lt;&lt; "All verification checks passed.\n\n";
}</code></pre></div><p>Notice the order: the baseline Black&#8209;Scholes module is verified first because all other modules depend on it. The effective volatility and modified pricing modules come next, followed by each force model, and finally the sweep functions that reproduce the paper's figures. If any single check fails, the program throws an exception and stops. This guarantees that the CSV files written later are produced by code that has passed every internal consistency test.</p><h3>Invariant 1: Probability Density Normalization</h3><p>Every probability density <code>P(x)</code> used in the computation must integrate to one over the real line. This is a fundamental requirement: if the density is not normalized, the cumulative distribution <code>N_eff(d)</code> will not be a valid CDF, and the option prices will be wrong.</p><p>Each model's verification function builds a uniform grid over the integration domain (typically <code>[-12, 12]</code> with 20001 points) and uses Simpson's rule to compute the integral of <code>P(x)</code>. The result is compared to <code>1.0</code> with a tolerance of <code>1e-10</code>. For example, the constant force model verification in <code>src/constant_force_model.h</code> contains a loop that sweeps over several force strengths <code>k</code> and verifies that the shifted Gaussian integrates to one regardless of the shift. The same pattern appears in every model's verification function. For the perturbation&#8209;based models (<code>x&#178;</code> and <code>x&#179;</code> forces), the normalization constant <code>C&#178;</code> is computed numerically, and the check confirms that the normalized density indeed integrates to one.</p><h3>Invariant 2: Put&#8209;Call Parity</h3><p>Put&#8209;call parity is a model&#8209;independent relationship that must hold for any European option pricing formula that uses a valid cumulative distribution function. The identity is</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;c + K e^{-rT} = p + S_0&quot;,&quot;id&quot;:&quot;A0158D11D1&quot;}" data-component-name="LatexBlockToDOM"></div><p>where <code>c</code> is the call price, <code>p</code> is the put price, <code>K</code> is the strike, <code>r</code> is the risk&#8209;free rate, <code>T</code> is the time to maturity, and <code>S_0</code> is the current asset price. This parity relies only on the definition of the CDF, not on the specific shape of the probability density. Therefore, it must hold for the baseline Black&#8209;Scholes model and for every modified model we implement. Note that this identity is a standard financial invariant and is not explicitly numbered as an equation in the paper; it follows directly from the definitions of the call and put prices in Eqs. (1) and (2).</p><p>The baseline verification in <code>src/black_scholes_base.h</code> checks this explicitly:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;a6bb8ba0-3aba-4e46-8bd0-f79d4f47a020&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">double parity_lhs = res.call + K * std::exp(-r * T);
double parity_rhs = res.put + S0;
bool pass = std::abs(parity_lhs - parity_rhs) &lt; eps;</code></pre></div><p>The modified option pricing verification in <code>src/modified_option_pricing.h</code> repeats the same check for the modified prices, ensuring that the new CDF <code>N_eff</code> does not break the parity. If a market force model produced a density that was not a valid probability distribution, the parity check would catch it.</p><h3>Invariant 3: Baseline Recovery When Force Parameters Are Zero</h3><p>When the force strength parameter is zero, every model must reduce exactly to the standard Black&#8209;Scholes baseline. This is the most important paper&#8209;fidelity check: the quantum approach is an extension, not a replacement, and it must nest the original model.</p><p>Each model's verification function tests this at multiple points. For the constant force model, <code>k = 0</code> must give <code>P_k(x) = normalPDF(x)</code>. The verification in <code>src/constant_force_model.h</code> checks this with a loop over several test points:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;98752f44-efd2-4d7b-87e3-025070a42b17&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">const double k = 0.0;
const double test_points[] = {-3.0, -1.0, 0.0, 1.0, 3.0};
for (double x : test_points) {
    double pk = constant_force_PDF(x, k);
    double np = normalPDF(x);
    if (std::abs(pk - np) &gt; tol) {
        all_passed = false;
    }
}</code></pre></div><p>For the linear force model, the limit <code>&#955; &#8594; 0</code> is tested with a very small positive <code>&#955;</code> (since the function throws for <code>&#955; &#8804; 0</code>). The <code>x&#178;</code> and <code>x&#179;</code> force models check that <code>&#946; = 0</code> and <code>&#947; = 0</code> recover the standard normal PDF after normalization. The quantum well model does not have a zero&#8209;force limit in the same sense, but its verification checks that the analytical <code>&#963;_QM</code> formula matches numerical integration.</p><p>The sweep verification in <code>src/paper_figures.h</code> takes this further: it runs a full pricing sweep with zero force and compares every call and put price to the baseline Black&#8209;Scholes result, ensuring that the entire pipeline&#8212;density, effective volatility, CDF grid, and modified pricing&#8212;collapses to the original model.</p><h3>Invariant 4: Effective Volatility Consistency</h3><p>For the baseline standard normal PDF, the quantum standard deviation <code>&#963;_QM</code> must be exactly <code>1.0</code>, so the effective volatility <code>&#963;_eff</code> must equal the input <code>&#963;</code>. The verification in <code>src/effective_volatility.h</code> checks this:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;d1e69359-7808-4884-a24e-94dfcbc4f2bf&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">VolResult res = compute_eff_vol(0.25, normalPDF, -12.0, 12.0, 20001);
if (std::abs(res.sigma_QM - 1.0) &gt; tol) { /* fail */ }
if (std::abs(res.sigma_eff - 0.25) &gt; tol) { /* fail */ }</code></pre></div><p>For the constant force model, the variance is unchanged by the shift, so <code>&#963;_QM</code> must remain <code>1.0</code> for any <code>k</code>. The verification confirms this numerically. For the linear force model, the analytical formula <code>&#963;_QM = 1 / sqrt(1 + &#955;)</code> is tested against numerical integration, and the identity <code>&#963;_QM = 1 / &#955;_&#969;</code> is checked exactly. For the quantum well, the analytical formula <code>&#963;_QM = sqrt(a&#178;/3 - 2a&#178;/&#960;&#178;)</code> is verified against Simpson integration over the well domain.</p><h3>Invariant 5: CDF Monotonicity and Range</h3><p>A valid cumulative distribution function must be monotonically non&#8209;decreasing and must return values in <code>[0, 1]</code>. The modified option pricing verification in <code>src/modified_option_pricing.h</code> checks both properties for the CDF grid built from <code>normalPDF</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;f9970479-fa24-488b-94e4-e13660418433&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">bool monotonic = true;
for (size_t i = 1; i &lt; cdf_grid.size(); ++i) {
    if (cdf_grid[i] &lt; cdf_grid[i - 1] - 1e-15) {
        monotonic = false;
        break;
    }
}</code></pre></div><p>It also verifies that the CDF grid ends at approximately <code>1.0</code> and that the interpolated values <code>N_eff(d_eff&#177;)</code> lie in <code>[0, 1]</code>. These checks ensure that the numerical integration and interpolation routines are working correctly before they are used with the more exotic force&#8209;model densities.</p><h3>Invariant 6: Simpson Integration Accuracy</h3><p>The Simpson integrator is the workhorse of the entire codebase. It is used for normalization, expectation values, and CDF construction. The effective volatility verification tests it against a known integral: the area under <code>x&#178;</code> on <code>[0, 1]</code> must equal <code>1/3</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;dc5304e0-fc53-469f-bed7-077d93143ed3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">const int N = 101;
const double a = 0.0, b = 1.0;
// ... build x and f vectors with f[i] = x[i] * x[i] ...
const double integral = simpson_integrate(x, f);
if (std::abs(integral - 1.0 / 3.0) &gt; tol) { /* fail */ }</code></pre></div><p>This test is independent of any probability density and confirms that the Simpson rule implementation is correct to machine precision. The paper relies on numerical integration throughout (for expectation values in Eq. (11) and for the modified CDF), so the accuracy of this integrator is critical to the fidelity of the entire implementation.</p><h3>Model&#8209;Specific Warnings</h3><p>The verification functions for the perturbation&#8209;based models (<code>x&#178;</code> and <code>x&#179;</code> forces) include an additional check that is not a hard failure but a warning. Because the perturbation expansions are truncated, the resulting probability density can become negative for large force strengths or large <code>|x|</code>. The verification scans the entire integration grid and reports the minimum value:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;cpp&quot;,&quot;nodeId&quot;:&quot;59a14159-aefa-41ea-a925-6bec647e0f9c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-cpp">if (min_val &lt; 0.0) {
    std::cout &lt;&lt; "WARNING: x3_force_model - For gamma=" &lt;&lt; gamma
              &lt;&lt; ", P_gamma becomes negative (min=" &lt;&lt; min_val
              &lt;&lt; " at x=" &lt;&lt; min_x
              &lt;&lt; "). Perturbation expansion may be invalid for this gamma."
              &lt;&lt; std::endl;
}</code></pre></div><p>This warning alerts the user that the perturbation approximation has broken down and that the resulting option prices should not be trusted. It does not cause the verification to fail, because the paper's formula is known to be approximate, but it provides important guidance for choosing parameter ranges.</p><h3>What the Verification Does Not Check</h3><p>The verification suite is thorough but has deliberate boundaries. It does <strong>not</strong> check that the option prices match any market data, because the paper provides no such validation. It does <strong>not</strong> execute the code and compare output to expected numerical values from the paper, because the paper's figures are qualitative sketches without precise data points. It does <strong>not</strong> verify that the corrected <code>x&#178;</code> force perturbation (enabled by <code>USE_CORRECTED_X2</code>) is mathematically correct, because that formula is derived externally and not part of the paper.</p><p>The verification is a <strong>paper&#8209;fidelity review</strong>: it ensures that the implementation faithfully reproduces the equations and methods described in the paper, that the mathematical invariants are preserved, and that the code is internally consistent. When all checks pass, you can be confident that the CSV files produced by <code>main()</code> correctly implement the quantum option pricing framework as the paper intended.</p><h2>Conclusion: Insights, Limitations, and Extensions</h2><p>We have travelled from the familiar Black&#8211;Scholes formula to a family of option pricing models inspired by quantum mechanics. The journey began with a simple observation: the standard normal probability density that underpins Black&#8211;Scholes is exactly the squared ground&#8209;state wavefunction of a quantum harmonic oscillator. From that starting point, the paper builds a framework in which hidden &#8220;market forces&#8221; are represented by additional potential energy terms. Each potential distorts the harmonic oscillator bowl, producing a new probability density that can capture skewness, fat tails, or hard price bounds&#8212;features that the log&#8209;normal model cannot express.</p><h3>What the Framework Achieves</h3><p>The core contribution of the paper is a <strong>phenomenological recipe for generating non&#8209;Gaussian distributions</strong> within a Black&#8211;Scholes&#8209;like pricing structure. The recipe has three steps:</p><ol><li><p>Choose a potential <code>V(x)</code> that encodes the market force you want to model.</p></li><li><p>Solve (or approximate) the stationary Schr&#246;dinger equation to obtain the ground&#8209;state probability density <code>P(x)</code>.</p></li><li><p>Compute the effective volatility <code>&#963;_eff = &#963; &#183; &#963;_QM</code> from the standard deviation of <code>P(x)</code>, then price options by integrating <code>P(x)</code> up to the modified limits <code>d_eff&#177;</code>.</p></li></ol><p>This procedure yields a family of models that are all anchored to the same baseline: when the extra potential vanishes, every model collapses back to the standard Black&#8211;Scholes result. The five forces studied in the paper&#8212;constant, linear, <code>x&#178;</code>, <code>x&#179;</code>, and the quantum well&#8212;demonstrate how different potential shapes produce qualitatively different option price behaviours. A constant force shifts the distribution without changing its width. A linear force squeezes or stretches the variance. The <code>x&#179;</code> force can create a double&#8209;humped density, leading to non&#8209;monotonic price curves. The quantum well imposes hard boundaries, eliminating the possibility of extreme returns altogether.</p><p>The C++ implementation we have built mirrors this structure exactly. Each force model lives in its own header file, depends only on the core pricing and integration modules, and exposes a probability density function that plugs directly into the modified pricing pipeline. The sweep functions in <code>src/paper_figures.h</code> automate the parameter studies that reproduce the paper's qualitative figures, and the verification functions in every module check that mathematical invariants&#8212;normalisation, put&#8209;call parity, zero&#8209;force recovery&#8212;hold to high precision.</p><h3>Limitations and Caveats</h3><p>Despite its elegance, the quantum mechanics approach has important limitations that anyone using the code should keep in mind.</p><p><strong>It is a phenomenological model, not a first&#8209;principles derivation.</strong> The paper does not derive the potentials from market micro&#8209;structure, agent behaviour, or any economic theory. The potentials are chosen because they are mathematically tractable and produce interesting probability densities. This means the force parameters <code>k</code>, <code>&#955;</code>, <code>&#946;</code>, <code>&#947;</code>, and <code>a</code> do not have direct financial interpretations that can be read off a balance sheet or a ticker tape. They are dials that change the shape of the distribution, and their values would need to be calibrated to market data before the model could be used for real pricing.</p><p><strong>There is no dynamics.</strong> The paper works entirely with stationary states of the Schr&#246;dinger equation. It does not propose a stochastic process for the underlying asset price, nor does it derive a time&#8209;dependent evolution that could be used for path&#8209;dependent options or hedging calculations. The mapping between the dimensionless quantum coordinate <code>x</code> and the financial log&#8209;return is implicit; the paper never writes down a stochastic differential equation that corresponds to a given potential. The effective volatility <code>&#963;_eff</code> is a heuristic that preserves the algebraic form of the Black&#8209;Scholes formula, but it is not derived from a dynamic hedging argument.</p><p><strong>The perturbation treatment for the `x&#178;` force is incorrect.</strong> As we discussed in detail in the section on Model 3, the paper's first&#8209;order perturbation result for the cubic potential contains a linear term in <code>&#946;</code> that should vanish by symmetry. The correct leading correction is of order <code>&#946;&#178;</code>. The implementation provides the paper's formula for reproducibility, but it also includes a compile&#8209;time flag <code>USE_CORRECTED_X2</code> that switches to a proper second&#8209;order expansion. Anyone using the <code>x&#178;</code> force model for quantitative work should enable the corrected version or derive their own.</p><p><strong>The constant force density was reconstructed.</strong> The original Eq. (15) in the paper is garbled. We reconstructed it as a shifted Gaussian with variance 1, which is the physically correct result for a linear potential added to the harmonic oscillator. This reconstruction is consistent with the paper's qualitative description and with standard quantum mechanics, but it is not a verbatim reproduction of the published text.</p><p><strong>No market validation is provided.</strong> The paper's numerical examples use a single set of synthetic parameters (<code>S0=20</code>, <code>K=20</code>, <code>r=10%</code>, <code>T=1</code> year, <code>&#963;=25%</code>). There is no comparison with real option prices, no calibration to implied volatility surfaces, and no statistical test of the model's predictive power. The figures are qualitative illustrations of how option prices change with the force parameters, not evidence that the model describes actual markets.</p><p><strong>Numerical integration introduces small errors.</strong> The implementation uses Simpson's rule on a truncated domain <code>[-12, 12]</code> with 20001 points. For Gaussian&#8209;like densities this captures more than 99.9% of the probability mass, but for densities with heavy tails or for very large force parameters, the truncation may miss some probability. The quantum well model avoids this issue by integrating only over its compact support <code>[-a, a]</code>. The normalisation constants for the perturbative densities are computed numerically, so they inherit the integration error.</p><h3>Extensions and Future Directions</h3><p>The codebase is designed to be a starting point for further exploration. Here are several directions that a motivated reader could pursue.</p><p><strong>Calibration to market data.</strong> The most natural next step is to treat the force parameters as free parameters and calibrate them to observed option prices. For example, one could minimise the sum of squared differences between model prices and market prices across a range of strikes and maturities. The constant force model (which shifts the mean) could capture skewness in the implied volatility smile, while the linear force model (which changes the variance) could capture the overall level of implied volatility. The <code>x&#179;</code> force model, with its double&#8209;well potential, might fit volatility smiles that have a pronounced curvature.</p><p><strong>Exotic and path&#8209;dependent options.</strong> Because the paper only provides stationary probability densities, the current framework cannot price barrier options, Asian options, or any contract whose payoff depends on the path of the underlying asset. Extending the model to a full stochastic process would require solving the time&#8209;dependent Schr&#246;dinger equation or finding a classical stochastic process whose stationary distribution matches the quantum probability density. This is a non&#8209;trivial problem, but it would greatly expand the scope of the approach.</p><p><strong>Time&#8209;dependent potentials.</strong> The paper assumes that the market force potential is static. A natural generalisation is to allow the potential to change over time, perhaps to model changing market conditions or approaching news events. In quantum mechanics, time&#8209;dependent potentials lead to transitions between energy levels; in the financial analogy, this might correspond to regime shifts or stochastic volatility.</p><p><strong>Other potentials.</strong> The five forces studied in the paper are only a small sample of the possible potentials one could add to the harmonic oscillator. A periodic potential (like a cosine) could model seasonality or support and resistance levels. A potential with a narrow dip and wide flat regions could model a market that is usually quiet but occasionally experiences large moves. The perturbation theory framework used for the <code>x&#178;</code> and <code>x&#179;</code> forces can be applied to any small perturbation, and for larger perturbations one could use numerical methods to solve the Schr&#246;dinger equation directly.</p><p><strong>Higher&#8209;order perturbation theory.</strong> The <code>x&#178;</code> force model highlights the danger of truncating a perturbation expansion too early. Implementing higher&#8209;order corrections&#8212;either analytically with computer algebra or numerically by diagonalising the Hamiltonian in a truncated basis&#8212;would improve the accuracy of the perturbative densities and extend their range of validity to larger force parameters.</p><p><strong>Alternative numerical methods.</strong> The current implementation relies on Simpson integration on a uniform grid. For densities with sharp features or heavy tails, adaptive quadrature or Gauss&#8209;Hermite integration could provide better accuracy with fewer function evaluations. The modular design of the code makes it straightforward to swap in a different integrator.</p><h3>Final Thoughts</h3><p>The quantum mechanics approach to option pricing is a creative and thought&#8209;provoking framework. It does not replace the Black&#8209;Scholes model, nor does it claim to be a complete theory of financial markets. Instead, it offers a new language for thinking about non&#8209;Gaussian distributions and a systematic way to generate them. The analogy between a market force and a potential energy term is intuitive: just as a physicist can predict how a particle's position distribution changes when you reshape the potential, a quant can explore how option prices change when you reshape the underlying return distribution.</p><p>The C++ implementation accompanying this tutorial makes the framework concrete. Every equation from the paper is translated into working code, every model is verified against mathematical invariants, and every figure can be reproduced with a single function call. Whether you use the code as a learning tool, a starting point for your own research, or a source of ideas for a production pricing system, we hope it deepens your understanding of both option pricing and the unexpected connections between finance and quantum physics.</p><p>Use the button below to download code:</p>
      <p>
          <a href="/__u/onepagecode.substack.com/p/quant-finance-option-pricing-beyond">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Quant Finance: Solving the High-Dimensional Black-Scholes PDE via Quantized Tensor Trains (QTT)]]></title><description><![CDATA[A complete Python guide to overcoming the curse of dimensionality in multi-asset option pricing using tensor decompositions.]]></description><link>https://onepagecode.substack.com/p/quant-finance-solving-the-high-dimensional</link><guid isPermaLink="false">https://onepagecode.substack.com/p/quant-finance-solving-the-high-dimensional</guid><dc:creator><![CDATA[Onepagecode]]></dc:creator><pubDate>Sun, 19 Jul 2026 20:06:06 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1590283603385-17ffb3a7f29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ0MDU1MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Use the button at the end of this article to download the source code</h2><p><em>This is not an official implementation. Since no official public code is available, I wrote this version based on my understanding of the approach. It has not been thoroughly tested, so please treat it as a reference for understanding the implementation rather than production-ready code.</em></p><p><em>Paper we are implementing today: https://arxiv.org/abs/2601.00009</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1590283603385-17ffb3a7f29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ0MDU1MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1590283603385-17ffb3a7f29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ0MDU1MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1590283603385-17ffb3a7f29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ0MDU1MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1590283603385-17ffb3a7f29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ0MDU1MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1590283603385-17ffb3a7f29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ0MDU1MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1590283603385-17ffb3a7f29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ0MDU1MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080" width="6000" height="4000" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1590283603385-17ffb3a7f29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ0MDU1MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:4000,&quot;width&quot;:6000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;screen showing bitcoin trading chart&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="screen showing bitcoin trading chart" title="screen showing bitcoin trading chart" srcset="https://images.unsplash.com/photo-1590283603385-17ffb3a7f29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ0MDU1MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 424w, https://images.unsplash.com/photo-1590283603385-17ffb3a7f29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ0MDU1MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 848w, https://images.unsplash.com/photo-1590283603385-17ffb3a7f29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ0MDU1MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1272w, https://images.unsplash.com/photo-1590283603385-17ffb3a7f29f?crop=entropy&amp;cs=tinysrgb&amp;fit=max&amp;fm=jpg&amp;ixid=M3wzMDAzMzh8MHwxfHNlYXJjaHwyfHxzdG9jayUyMG1hcmtldHxlbnwwfHx8fDE3ODQ0MDU1MTF8MA&amp;ixlib=rb-4.1.0&amp;q=80&amp;w=1080 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photo by <a href="https://unsplash.com/@nick604">Nick Chong</a> on <a href="https://unsplash.com">Unsplash</a></figcaption></figure></div><p>This paper presents quantized tensor train (QTT) solvers for the multi-asset Black-Scholes PDE, overcoming the curse of dimensionality. Two solvers are developed: a time-stepping algorithm for European and American options, and a space-time algorithm for European options. Both provide full-grid solutions and Greeks with polynomial scaling in the number of assets and polylogarithmic scaling in grid size. Numerical experiments for basket and max-min options in 3-5 dimensions demonstrate high accuracy on a personal computer.</p><h2>Implementation Assumptions</h2><ul><li><p>Fully implicit scheme (&#952;=1) for time-stepping.</p></li><li><p>TT-Cross sweeps default to 3-4.</p></li><li><p>Time-stepping solver is primary; space-time solver for d&#8804;3.</p></li><li><p>Boundary conditions for basket put: lower faces as Eqs. (A20)-(A22), upper faces homogeneous Dirichlet (Eq. A23).</p></li><li><p>For max-min put: lower boundaries Ke^{-rt}, upper boundaries zero.</p></li><li><p>Grid dimensions: c=7-9 spatial cores per dimension; domain determined adaptively from pilot coarse run.</p></li><li><p>Market parameters from Appendix B.2.</p></li><li><p>ALS/MALS solver uses 2 sweeps by default.</p></li><li><p>Rank reduction after early-exercise uses TT-SVD with tolerance or target rank.</p></li></ul><h2>Introduction to Multi-Asset Option Pricing and the Curse of Dimensionality</h2><p>Pricing a financial option on a single underlying asset is a textbook problem.  The Black&#8211;Scholes model gives a closed-form solution for European calls and puts, and a simple one-dimensional finite-difference grid can handle American options or more complex payoffs with ease.  But the moment you have two, three, or five underlying assets, the situation changes dramatically.  The number of grid points needed to represent the solution grows exponentially with the number of assets, quickly exceeding the memory of any computer.  This is the <em>curse of dimensionality</em>, and it is the central challenge that the paper by Kazeev, Khoromskij, and Tyrtyshnikov (2601.00009v2) overcomes using quantized tensor trains (QTT).</p><h3>The Multi-Asset Black&#8211;Scholes PDE</h3><p>The starting point is the d-dimensional Black&#8211;Scholes partial differential equation.  For an option whose value depends on d asset prices <code>S_1, S_2, ..., S_d</code> and time <code>t</code>, the PDE is:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\partial V / \\partial t + (1/2) * \\sum_{i,j=1}^{d} \\rho_{ij} \\sigma_i \\sigma_j S_i S_j \\partial^2 V / (\\partial S_i \\partial S_j) + \\sum_{i=1}^{d} r S_i \\partial V / \\partial S_i - r V = 0&quot;,&quot;id&quot;:&quot;86404F8907&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here <code>V(S,t)</code> is the option price, a scalar function of the d asset prices and time.  Each <code>&#963;_i</code> is the volatility of asset <code>i</code>, <code>&#961;_{ij}</code> is the correlation between assets <code>i</code> and <code>j</code>, and <code>r</code> is the risk-free interest rate.  The double sum over <code>i</code> and <code>j</code> couples every pair of assets through the covariance matrix <code>&#963;_i &#963;_j &#961;_{ij}</code>.  This coupling is what makes the problem genuinely high-dimensional: the price of a basket option depends on the joint distribution of all underlyings, not just their marginal distributions.</p><h3>Why Classical Methods Fail</h3><p>A standard approach to solving this PDE is to discretise the spatial domain with a uniform grid.  If we use <code>N</code> grid points per asset, the total number of unknowns is <code>N^d</code>.  For <code>d=5</code> assets and a modest <code>N=256</code> points per dimension, that is <code>256^5 &#8776; 1.1 &#215; 10^12</code> unknowns &#8212; over a trillion grid points.  Storing a single vector of that size would require terabytes of memory, and solving the resulting linear system is completely infeasible on any personal computer.</p><p>Monte Carlo simulation avoids the grid entirely by sampling random paths, but it has its own limitations.  It provides only point estimates of the option price at a given initial spot vector, not the full solution surface.  Computing Greeks (Delta, Gamma) requires additional simulation or finite-difference approximations, which introduce noise and extra computational cost.  Moreover, Monte Carlo converges slowly, at a rate of <code>1/&#8730;(number of paths)</code>, making high accuracy expensive.</p><h3>The Promise of Quantized Tensor Trains</h3><p>Quantized tensor trains (QTT) offer a way to represent the full solution vector with a number of parameters that grows only <em>polynomially</em> with the number of assets <code>d</code> and <em>polylogarithmically</em> with the grid size <code>N</code>.  The key idea is to break the huge vector of length <code>N^d</code> into a product of many small tensors, each of which lives on a binary (quantized) representation of its index.</p><p>Concretely, suppose we use <code>c</code> QTT cores per spatial dimension, so that <code>N = 2^c</code>.  The total number of cores is <code>d &#215; c</code>.  Each core has a small bond dimension <code>r</code> (typically between 5 and 20).  The total number of parameters is then <code>O(d &#215; c &#215; r^2)</code>.  For <code>d=5</code>, <code>c=8</code>, and <code>r=15</code>, that is about <code>5 &#215; 8 &#215; 15^2 = 9000</code> parameters &#8212; a far cry from the trillion points of the full grid.</p><p>This compression is not just a storage trick.  The paper shows that the Black&#8211;Scholes operator itself has an exact QTT representation with rank bounded by <code>O(d)</code>, and that the solution can be computed by solving linear systems entirely within the QTT format.  The result is a solver that runs on a personal computer for <code>d=3,4,5</code> assets, producing full-grid prices and Greeks with high accuracy.</p><h3>What This Tutorial Covers</h3><p>In the sections that follow, we will build this QTT solver from the ground up.  We will start with the core data structure (the quantized tensor train) and its basic arithmetic.  Then we will construct the analytic building blocks &#8212; the exponential function, boundary selectors, and tridiagonal matrices &#8212; that form the operator and payoff.  We will implement the TT-Cross algorithm for approximating the max and min functions needed for payoffs and early exercise.  We will develop the ALS/MALS solver for linear systems in QTT format.  Finally, we will assemble everything into the time-stepping and space-time solvers, and compute Greeks from the solution.</p><p>By the end, you will understand not only how the paper achieves its remarkable results, but also how to implement a QTT-based PDE solver yourself.</p><h2>QTT Solver Methodology</h2><p>This section details the core algorithmic components of the quantized tensor train (QTT) solvers for multi-asset option pricing. The key idea is to represent the discretized solution and operators as QTT tensors, enabling efficient storage and computation that scales polynomially with the number of assets.</p><h3>Quantized Tensor Trains (QTT)</h3><p>A tensor train (TT) decomposes a d-dimensional tensor into a product of 3-dimensional core tensors. The QTT format further applies this decomposition to a quantized (binary) representation of each dimension, effectively converting an N-point grid into a tensor of order log2(N). For a grid with 2^c points per dimension, the QTT representation uses c cores per dimension, each of small rank.</p><h3>TT-Cross Algorithm</h3><p>The TT-Cross algorithm approximates a tensor by sampling its entries along selected fibers (rows/columns). It iteratively refines the approximation by finding the most informative cross-sections, requiring only O(d <em> r^2 </em> N) evaluations of the original tensor function, where r is the TT rank.</p><h3>Alternating Linear Scheme (ALS)</h3><p>ALS solves linear systems in the TT format by optimizing one core at a time while keeping others fixed. It cycles through all cores, solving local least-squares problems, and converges to a solution of the global system. The Multigrid ALS (MALS) variant improves convergence by adaptively adjusting ranks.</p><h3>Time-Stepping Solver Pipeline</h3><ol><li><p><strong>Discretization</strong>: Apply backward Euler in time and central differences in space to the log-transformed Black-Scholes PDE.</p></li><li><p><strong>QTT Construction</strong>: Build the finite-difference operator and initial condition (payoff) as QTT tensors with known low-rank structures.</p></li><li><p><strong>Time Stepping</strong>: For each time step, solve the linear system (I - &#916;t <em> A) </em> v<em>{n+1} = v</em>n using ALS, where A is the spatial discretization operator.</p></li><li><p><strong>Boundary Conditions</strong>: Impose Dirichlet conditions using an eraser MPO that zeros out boundary points.</p></li><li><p><strong>Output</strong>: The solution at each time step is a QTT tensor representing the option price on the full grid.</p></li></ol><h3>Space-Time Formulation</h3><p>An alternative approach solves for all time steps simultaneously by forming a block-bidiagonal system that couples spatial points across time. This yields a single large QTT system that can be solved with ALS, avoiding sequential time stepping. The space-time operator has a known QTT rank bound that grows only polynomially with the number of assets.</p><h3>Comparison</h3><ul><li><p><strong>Time-stepping</strong>: More flexible (handles American options via early exercise checks), but requires solving a linear system at each step.</p></li><li><p><strong>Space-time</strong>: Faster for European options (single solve), but cannot easily incorporate early exercise.</p></li></ul><p>Both methods achieve full-grid solutions with polynomial scaling in the number of assets and polylogarithmic scaling in grid size.</p><h2>Building Blocks: Analytic QTT Representations</h2><p>Now that we have a feel for the QTT data structure, we need the concrete building blocks that the multi-asset solver uses.  The Black&#8211;Scholes operator, its boundary conditions, and the payoff functions all rely on a few special functions and matrices that can be represented <em>exactly</em> and with tiny rank in the QTT format.  This section implements three of those building blocks: the exponential function (rank 1), the left/right boundary selection vectors (rank 2), and the tridiagonal Toeplitz matrix (rank 3).  These constructions follow Appendices C.1, C.2, and C.3 of the paper.</p><h3>Exponential Function (rank 1)</h3><p>Why is the exponential function important?  In the log-price coordinates <code>x = ln S</code>, the payoff of a basket option involves terms like <code>e^x</code>.  The QTT representation of <code>e^{&#945;x}</code> on a dyadic grid is exact and has rank 1.  The key insight is the binary expansion of the grid index: on the interval <code>(0,1)</code>, a grid point <code>x</code> is represented by a binary fraction <code>0.b_1 b_2 ... b_c</code>, where <code>b_i</code> is the <code>i</code>-th bit of the index.  Then</p><p><code>e^{&#945;x} = e^{&#945; * &#931; b_i * 2^{-i}} = &#8719;_i e^{&#945; * b_i * 2^{-i}}</code>.</p><p>This product form translates directly into a QTT with cores <code>F_i</code> where <code>F_i[0] = 1</code> and <code>F_i[1] = exp(&#945; * 2^{-i})</code>.  The function <code>analytic_qtt_exponential</code> in <code>qtt_analytic.py</code> builds exactly this representation.  The core loop is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c211344f-cf20-43d4-aa74-0ff619d024d8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">cores = []
for i in range(c):
    factor = np.exp(scaled_alpha * 2.0 ** (-(i + 1)))
    core = np.array([[[1.0]], [[factor]]], dtype=np.float64)
    core = core.reshape(1, 2, 1)
    cores.append(core)</code></pre></div><p>Each core has shape <code>(1, 2, 1)</code> &#8211; a rank-1 MPS.  To map to an arbitrary interval <code>(a, b)</code>, we first scale <code>alpha</code> by <code>(b-a)</code>, then multiply every entry of the first core by <code>e^a</code>.  The result is a QTT vector of length <code>2^c</code> that matches the exponential function at every grid point.</p><h3>Boundary Selection Vectors vL and vR (rank 2)</h3><p>When imposing Dirichlet boundary conditions, we need to modify the solution at the leftmost and rightmost grid points.  The paper defines two vectors: <code>vL</code> has 1 at every position except the first (index 0), where it is 0; <code>vR</code> has 1 at every position except the last (index <code>2^c-1</code>), where it is 0.  These are building blocks for the &#8220;eraser&#8221; MPO that zeroes out boundary rows and columns.</p><p>The QTT representation of <code>vL</code> and <code>vR</code> is exact and has rank 2.  The construction of the first core is instructive.  For <code>vL</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4bcdcb10-f9fa-4e06-b48d-b49c1a994737&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">first = np.zeros((1, 2, 2), dtype=np.float64)
first[0, 0, 0] = 1.0
first[0, 1, 1] = 1.0</code></pre></div><p>For <code>vL</code>, the first core distinguishes between the grid index 0 (bit 0) and all others (bit 1).  The middle cores propagate the two &#8220;states&#8221; (leftmost or not) through the bit string, and the last core enforces the final zero or one.  The result is a rank-2 MPS with cores of shape <code>(1,2,2)</code>, <code>(2,2,2)</code>, ..., <code>(2,2,1)</code>.  The function <code>qtt_vL_vR(c, side='L')</code> returns this representation.</p><h3>Tridiagonal Toeplitz Matrix (rank 3) &#8211; Lemma 1</h3><p>The finite-difference discretisation of the Black&#8211;Scholes operator requires tridiagonal matrices for the second and first derivative terms.  These matrices are Toeplitz (constant diagonals) and have size <code>2^c &#215; 2^c</code>.  Lemma 1 in the paper (Appendix C.3) gives an exact QTT representation with bond dimension 3.</p><p>Let the matrix have diagonal entries <code>&#945;</code>, superdiagonal entries <code>&#946;</code>, and subdiagonal entries <code>&#947;</code>.  The QTT uses three 2&#215;2 building blocks:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;31c70d2f-ab13-4204-8a07-ac44dcc15d15&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">I = np.array([[1.0, 0.0], [0.0, 1.0]], dtype=np.float64)
J = np.array([[0.0, 1.0], [0.0, 0.0]], dtype=np.float64)
Jp = np.array([[0.0, 0.0], [1.0, 0.0]], dtype=np.float64)  # J'</code></pre></div><p>The first core has shape <code>(1, 2, 2, 3)</code> and its three slices are <code>&#945;I</code>, <code>&#946;J</code>, and <code>&#947;J'</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;754cfa4a-f912-4ffa-8d84-79cbd2b90704&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">first = np.zeros((1, 2, 2, 3), dtype=np.float64)
first[0, :, :, 0] = alpha * I
first[0, :, :, 1] = beta * J
first[0, :, :, 2] = gamma * Jp</code></pre></div><p>Middle cores (shape <code>(3, 2, 2, 3)</code>) repeat the same pattern: each of the three left bond indices corresponds to a &#8220;state&#8221; that routes the three diagonal contributions.  The last core (shape <code>(3, 2, 2, 1)</code>) again contains <code>&#945;I</code>, <code>&#946;J</code>, and <code>&#947;J'</code> but now projects to a single right bond.  The function <code>qtt_toeplitz_tridiagonal(c, alpha, beta, gamma)</code> returns this rank-3 MPO.</p><p>For completeness, <code>qtt_identity_mpo(c)</code> provides the identity matrix as a rank-1 MPO, simply formed by repeating the 2&#215;2 identity reshaped to <code>(1,2,2,1)</code> for each core.</p><h3>Worked Example: Exponential on a 3-Core Grid</h3><p>Take <code>c=3</code> cores (grid of 8 points in <code>(0,1)</code>), <code>&#945;=0.5</code>.  The cores are:</p><ul><li><p>Core 0: <code>[1, exp(0.5 * 0.5)] &#8776; [1, 1.284]</code></p></li><li><p>Core 1: <code>[1, exp(0.5 * 0.25)] &#8776; [1, 1.133]</code></p></li><li><p>Core 2: <code>[1, exp(0.5 * 0.125)] &#8776; [1, 1.064]</code></p></li></ul><p>Contracting the cores (each of shape <code>(1,2,1)</code>) gives the vector of length 8: <code>[1, 1.064, 1.133, 1.203, 1.284, 1.365, 1.455, 1.549]</code>.</p><p>This matches the exact values of <code>e^{0.5 x}</code> at the points <code>x = 0, 1/8, ..., 7/8</code>.</p><p>For the tridiagonal Toeplitz matrix with <code>&#945;=2, &#946;=-1, &#947;=-1</code> (the standard second-derivative stencil), the QTT representation has rank 3 and is exact.  When applied to a vector, it produces the same result as the full matrix-vector product, up to machine precision.</p><h3>Advanced Notes</h3><p>The exponential QTT cores <code>F_i[0] = 1</code>, <code>F_i[1] = exp(&#945; * 2^{-i})</code> follow directly from the binary expansion of the grid index.  Because the product of the cores reproduces the product of the exponentials, the representation is exact and rank 1.  The same idea extends to higher dimensions: for a d-dimensional tensor product grid, the Kronecker product of the 1D QTTs yields a rank-1 d-dimensional MPS.</p><p>The tridiagonal Toeplitz representation (Lemma 1) uses the three 2&#215;2 matrices <code>I</code>, <code>J</code>, <code>J'</code> to encode the three non-zero diagonals.  The bond dimension of 3 arises because the matrix has three independent &#8220;streams&#8221; &#8211; the diagonal, the superdiagonal, and the subdiagonal &#8211; that must be carried through the tensor network.  The middle cores are all identical, which makes the construction highly efficient for large <code>c</code>.</p><h2>TT-Cross: Approximating Functions with Controlled Rank</h2><p>Use the button or URL below to download the source code.</p><p>In the previous sections we built exact, low-rank QTT representations for the exponential function, boundary vectors, and tridiagonal matrices.  But the payoff of an option &#8212; <code>max(K - S, 0)</code> for a put, or <code>max(S_1, S_2, ..., S_d)</code> for a max-min contract &#8212; is not an exponential or a linear function.  It is a piecewise-linear function with a kink, and its QTT rank can be large if we try to represent it exactly.  The same problem appears in the early-exercise step for American options, where we need to compute the element-wise maximum of the solution and the payoff at every time step.</p><p>To handle these operations, the paper uses the <strong>TT-Cross algorithm</strong>.  Instead of building an exact representation, TT-Cross <em>approximates</em> a function on a tensor grid with a controlled rank.  It does this by evaluating the function on a carefully chosen subset of grid points &#8212; the <em>cross</em> &#8212; and then constructing a low-rank tensor that matches the function at those points.  The result is a QTT tensor whose bond dimensions never exceed a user-specified cap <code>&#967;</code>, and whose accuracy is sufficient for pricing errors below 1%.</p><h3>The Cross Approximation Idea</h3><p>Suppose you have a matrix <code>A</code> of size <code>m &#215; n</code> and you want a low-rank approximation.  The skeleton decomposition says that if <code>A</code> has rank <code>r</code>, you can pick <code>r</code> rows and <code>r</code> columns such that the submatrix at their intersection (the <em>cross</em>) determines the whole matrix.  Concretely, if you select row indices <code>I</code> and column indices <code>J</code>, then</p><p><code>A &#8776; A[:, J] * (A[I, J])^{-1} * A[I, :]</code>.</p><p>For a tensor, the same idea extends: you select fibers (the tensor analogue of rows and columns) along each dimension, and build a low-rank tensor train that interpolates the original tensor on those fibers.  The TT-Cross algorithm automates the selection of these fibers by sweeping through the dimensions, using a routine called <strong>maxvol</strong> to find the most informative indices.</p><h3>The <code>tt_cross</code> Function</h3><p>The file <code>tt_cross.py</code> implements the TT-Cross algorithm for general function approximation on a tensor grid.  The main entry point is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;00ec597f-c2ac-491f-9090-23dcb925a5f1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def tt_cross(function_handle, shape, max_rank, max_sweeps=4, tol=1e-6):
    """TT-Cross algorithm for approximating a function on a tensor grid.

    Parameters
    ----------
    function_handle : callable
        A function that takes a tuple of integer indices (one per dimension)
        and returns a scalar float.
    shape : list of int
        Shape of the tensor grid.  Each entry must be a power of two.
    max_rank : int
        Maximum allowed bond dimension (rank) of the approximation.
    max_sweeps : int
        Maximum number of alternating sweeps.
    tol : float
        Tolerance for early stopping.

    Returns
    -------
    QTTTensor
        A QTT tensor (MPS) approximating the function on the given grid.
    """</code></pre></div><p>The algorithm works as follows:</p><ol><li><p><strong>Initialise</strong> a random tensor train with bond dimensions capped at <code>max_rank</code>.</p></li><li><p><strong>Left-to-right sweep</strong>: for each dimension <code>k</code>, fix all cores except the <code>k</code>-th.  Evaluate the function on a set of indices selected by the maxvol procedure on the unfolding matrix.  Update the <code>k</code>-th core to match those evaluations.</p></li><li><p><strong>Right-to-left sweep</strong>: repeat the process in the opposite direction.</p></li><li><p><strong>Iterate</strong> for <code>max_sweeps</code> sweeps, or until the approximation stops changing.</p></li></ol><p>The function handle is evaluated only on the selected fibers, not on the entire grid.  This is what makes the algorithm efficient for high-dimensional problems: the number of evaluations is proportional to <code>d * &#967;^2</code>, where <code>d</code> is the number of dimensions and <code>&#967;</code> is the target rank.</p><h3>The Maxvol Routine</h3><p>The maxvol routine finds a set of rows (or columns) that form a submatrix with large determinant.  This is a greedy algorithm: it starts with a random set of rows, then iteratively swaps rows to increase the volume (absolute determinant) of the submatrix.  The implementation in <code>tt_cross.py</code> is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b2a0dde9-b1b8-47cf-9995-be4998902fdc&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def _maxvol(A, tol=1e-12):
    """Find a set of rows that form a maximum-volume submatrix of A."""
    m, r = A.shape
    if m &lt;= r:
        return np.arange(m)
    ind = np.arange(r)
    B = A[ind, :]
    for _ in range(10):
        try:
            B_inv = np.linalg.inv(B)
        except np.linalg.LinAlgError:
            break
        C = A @ B_inv
        i_max, j_max = np.unravel_index(np.argmax(np.abs(C)), C.shape)
        if np.abs(C[i_max, j_max]) &lt;= 1.0 + tol:
            break
        ind[j_max] = i_max
        B = A[ind, :]
    return ind</code></pre></div><p>This function takes a matrix <code>A</code> of shape <code>(m, r)</code> (with <code>m &gt;= r</code>) and returns <code>r</code> row indices that approximately maximise the volume.  The algorithm is simple: it computes the pivot matrix <code>C = A * B^{-1}</code>, finds the element of <code>C</code> with the largest absolute value, and swaps that row into the selected set.  The loop continues until no element exceeds <code>1 + tol</code> in absolute value, indicating that the current submatrix is close to maximum volume.</p><h3>Element-Wise Max and Min of Two QTT Tensors</h3><p>The most important use of TT-Cross in the solver is for the element-wise maximum and minimum operations.  These are implemented as <code>tt_cross_max</code> and <code>tt_cross_min</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a6ae1159-39f6-4250-8437-ebc46bf6cf7f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def tt_cross_max(qtt_a, qtt_b, max_rank, max_sweeps=4):
    """Compute element-wise maximum of two QTT tensors via TT-Cross."""
    shape = qtt_a.shape()
    if shape != qtt_b.shape():
        raise ValueError("Input tensors must have the same shape")

    # For small tensors, convert to full and use numpy
    total_size = 1
    for s in shape:
        total_size *= s
    if total_size &lt;= 10**6:
        full_a = qtt_a.to_full()
        full_b = qtt_b.to_full()
        full_max = np.maximum(full_a, full_b)
        # Convert back to QTT via TT-SVD
        ...
    else:
        # Use TT-Cross with function handle = max(f_a, f_b)
        def max_func(idx):
            return max(qtt_a.evaluate(idx), qtt_b.evaluate(idx))
        return tt_cross(max_func, shape, max_rank, max_sweeps)</code></pre></div><p>The function first checks if the total size of the tensor is small enough to convert to a full array.  If so, it computes the max directly using NumPy and then converts back to QTT via TT-SVD.  For larger tensors, it defines a function handle that evaluates the max at a given multi-index by evaluating both input tensors at that index, and then calls <code>tt_cross</code> with that handle.  The <code>tt_cross_min</code> function works identically, using <code>min</code> instead of <code>max</code>.</p><h3>Worked Example: Basket Put Payoff</h3><p>Consider a two-asset basket put with strike <code>K = 34</code> and equal weights.  The payoff at maturity is:</p><p><code>V(S_1, S_2) = max(K - (S_1 + S_2), 0)</code>.</p><p>In log-price coordinates <code>x_i = ln S_i</code>, this becomes:</p><p><code>V(x_1, x_2) = max(K - (e^{x_1} + e^{x_2}), 0)</code>.</p><p>We want to represent this payoff as a QTT tensor on a grid with <code>c = 3</code> cores per dimension, giving <code>2^3 = 8</code> points per dimension and a total of 64 grid points.  The function is not low-rank: the <code>max</code> operation introduces a kink along the line <code>e^{x_1} + e^{x_2} = K</code>.  However, we can approximate it using TT-Cross with a rank cap of <code>&#967; = 10</code> and 4 sweeps.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;738ad839-be46-488d-a2d9-9a99afec1de1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import numpy as np
from tt_cross import tt_cross

c = 3
n = 2**c  # 8 points per dimension
K = 34.0

# Define the payoff function on the grid
def basket_put_payoff(idx):
    # idx is a tuple (i1, i2) with 0 &lt;= i_k &lt; n
    # Map indices to log-prices on a uniform grid in [-2, 2]
    x1 = -2.0 + 4.0 * idx[0] / (n - 1)
    x2 = -2.0 + 4.0 * idx[1] / (n - 1)
    S1 = np.exp(x1)
    S2 = np.exp(x2)
    return max(K - (S1 + S2), 0.0)

# Run TT-Cross
payoff_qtt = tt_cross(basket_put_payoff, [n, n], max_rank=10, max_sweeps=4)</code></pre></div><p>The resulting <code>payoff_qtt</code> is a QTT tensor with bond dimensions at most 10.  To check the accuracy, we can compare the QTT approximation with the exact payoff at a few random grid points:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;cbbff571-6b35-4220-ba77-907b8d74eeb3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># Evaluate at a random point
idx = (3, 5)
exact = basket_put_payoff(idx)
approx = payoff_qtt.evaluate(idx)
print(f"Exact: {exact:.4f}, Approx: {approx:.4f}, Error: {abs(exact-approx):.4e}")</code></pre></div><p>For a rank cap of 10 and 4 sweeps, the relative error is typically on the order of <code>10^{-3}</code>, which is sufficient for pricing errors below 1%.</p><h3>Advanced Detail: Why TT-Cross Works for the Max Function</h3><p>The max function is not low-rank, but it is <em>approximately</em> low-rank in the QTT format.  The reason is that the kink (the set of points where <code>e^{x_1} + e^{x_2} = K</code>) is a smooth curve in the log-price domain, and the QTT representation can capture this structure with moderate rank.  The TT-Cross algorithm exploits this by sampling the function along fibers that cross the kink, building an approximation that is accurate everywhere.</p><p>For American options, the early-exercise condition requires computing <code>max(V, payoff)</code> at every time step.  The paper uses <code>tt_cross_max</code> followed by rank reduction via <code>qtt_round</code> to keep the bond dimensions manageable.  The experiments in Section V.C show that the rank of the solution grows only modestly (from about 10 to about 20) over the life of the option, and the additional cost of the early-exercise step is about 30% of the total runtime.</p><h3>Summary</h3><ul><li><p>TT-Cross approximates a function on a tensor grid with controlled rank by evaluating it on a subset of points (the cross).</p></li><li><p>The maxvol routine selects the most informative fibers by maximising the volume of the submatrix.</p></li><li><p><code>tt_cross_max</code> and <code>tt_cross_min</code> use TT-Cross to compute element-wise max/min of two QTT tensors, which is essential for payoffs and early-exercise.</p></li><li><p>The approximation accuracy is controlled by the rank cap <code>&#967;</code> and the number of sweeps; typical values are <code>&#967; = 10-20</code> and 3-4 sweeps.</p></li><li><p>For the basket put payoff on a 2D grid with 64 points, TT-Cross with rank 10 achieves relative error ~10^{-3}.</p></li></ul><h2>ALS and MALS Solvers for Linear Systems</h2><p>Once we have the Black&#8211;Scholes operator <code>A</code> and the right-hand side <code>b</code> in QTT format, we need to solve the linear system <code>A x = b</code>.  This is the computational heart of every time step in the time-stepping solver and of the single global solve in the space-time solver.  The paper uses the <strong>Alternating Linear Scheme (ALS)</strong> and its variant <strong>MALS (Modified ALS)</strong> to solve these systems entirely within the QTT format, without ever forming the full matrix or vector.</p><p>ALS is a coordinate-descent-like algorithm for tensor trains.  It sweeps through the cores of the unknown solution <code>x</code>, updating one core at a time while keeping all other cores fixed.  At each core, the problem reduces to a small dense linear system &#8212; typically 2&#215;2 or 4&#215;4 &#8212; that can be solved by direct factorization.  MALS extends this by updating two adjacent cores simultaneously, which allows the bond dimension to grow or shrink during the optimization.  Both algorithms are described in Section III.D of the paper.</p><h3>The ALS Sweep</h3><p>Suppose we have the operator <code>A</code> as a QTT-MPO and the right-hand side <code>b</code> as a QTT-MPS.  We want to find the QTT-MPS <code>x</code> that satisfies <code>A x = b</code>.  The ALS algorithm proceeds as follows:</p><ol><li><p><strong>Initialize</strong> <code>x</code> with some initial guess (often the solution from the previous time step, or a zero tensor).</p></li><li><p><strong>Sweep left-to-right</strong>: for each core index <code>i</code> from 0 to <code>num_cores - 1</code>:</p></li></ol><ul><li><p>Fix all cores of <code>x</code> except core <code>i</code>.</p></li><li><p>Contract the environment: compute the left and right parts of <code>A</code> and <code>x</code> that involve all cores except <code>i</code>, and also the left and right parts of <code>b</code>.</p></li><li><p>Form a small local linear system <code>A_local * x_core_i = b_local</code> by contracting the environment tensors with the current core.</p></li><li><p>Solve the local system (typically 2&#215;2 or 4&#215;4) using a dense solver.</p></li><li><p>Update core <code>i</code> with the solution.</p></li></ul><ol start="3"><li><p><strong>Sweep right-to-left</strong>: repeat the same process in reverse order.</p></li><li><p><strong>Repeat</strong> for a fixed number of sweeps (typically 2&#8211;4) or until the residual norm <code>||A x - b|| / ||b||</code> falls below a tolerance.</p></li></ol><p>The key insight is that the local system is tiny because each core has only two physical indices (size 2 each) and the bond dimensions are small (typically 3&#8211;20).  The environment contraction can be computed efficiently by precomputing left and right environment tensors, as shown in the helper function <code>_contract_left_env</code>.</p><h3>MALS: Allowing Bond Dimension Changes</h3><p>ALS keeps the bond dimensions of <code>x</code> fixed during the sweep.  This is fine if the initial guess has the right bond dimensions, but it can be restrictive.  MALS (Modified ALS) optimizes two adjacent cores at a time, which allows the bond dimension between them to change.  The local system becomes slightly larger (e.g., 4&#215;4 instead of 2&#215;2), but the algorithm can adapt the rank of the solution automatically.</p><h3>Implementation in <code>als_mals.py</code></h3><p>The file <code>als_mals.py</code> implements both solvers.  The core function <code>als_solve</code> takes the operator <code>A</code>, the right-hand side <code>b</code>, an optional initial guess <code>x0</code>, and the number of sweeps.  It returns the solution <code>x</code> as a QTTTensor.</p><p>Here is the public interface:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ae7eeb59-379b-4e77-bc10-fc12af85cc88&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def als_solve(
    A: QTTTensor,
    b: QTTTensor,
    x0: Optional[QTTTensor] = None,
    max_sweeps: int = 2,
    local_solver: str = 'dense',
    verbose: bool = False
) -&gt; QTTTensor:
    """
    Solve A x = b using the Alternating Linear Scheme (ALS).

    Parameters
    ----------
    A : QTTTensor
        System matrix as MPO, shape (2^{c*d}, 2^{c*d}).
    b : QTTTensor
        Right-hand side as MPS, shape (2^{c*d},).
    x0 : QTTTensor, optional
        Initial guess (MPS). If None, a zero tensor is used.
    max_sweeps : int
        Number of full left-right-right-left sweeps.
    local_solver : str
        Method for solving the local system ('dense' or 'lstsq').
    verbose : bool
        If True, print residual norm after each sweep.

    Returns
    -------
    x : QTTTensor
        Solution MPS.
    """</code></pre></div><p>The MALS variant <code>mals_solve</code> has the same signature but uses two-core updates internally.</p><h3>The Local Solve</h3><p>The local system is solved by <code>_local_solve</code>, which handles the small dense linear system with a regularization term to avoid ill-conditioning:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;abe1b5cb-ece2-4e14-acee-bd2fc5063c09&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def _local_solve(
    A_local: np.ndarray,
    b_local: np.ndarray,
    reg: float = 1e-12
) -&gt; np.ndarray:
    """
    Solve a small dense linear system A_local @ x = b_local.

    A small regularization term is added to handle potential ill-conditioning.
    """
    if A_local.shape[0] == 0:
        return np.array([])
    try:
        return np.linalg.solve(A_local + reg * np.eye(A_local.shape[0]), b_local)
    except np.linalg.LinAlgError:
        return np.linalg.lstsq(A_local, b_local, rcond=None)[0]</code></pre></div><p>This function is called for each pair of bond indices during the sweep.  The local matrix <code>A_local</code> is formed by contracting the left and right environment tensors with the current core of <code>A</code> and the current core of <code>x</code>.  The local right-hand side <code>b_local</code> is formed similarly from the environment of <code>b</code>.</p><h3>Worked Example: 1D Tridiagonal System</h3><p>Consider a 1D problem with 8 grid points (<code>c=3</code>).  The operator <code>A</code> is a tridiagonal matrix (e.g., the second derivative stencil) represented as a QTT-MPO with rank 3.  The right-hand side <code>b</code> is a vector (e.g., the payoff) represented as a QTT-MPS with rank 5.  We want to solve <code>A x = b</code>.</p><ol><li><p><strong>Initialization</strong>: <code>x0</code> is a zero MPS with rank 1.</p></li><li><p><strong>Sweep 1 (left-to-right)</strong>:</p></li></ol><ul><li><p>Core 0: the left environment is empty (scalar 1).  The right environment is contracted from cores 1 and 2 of <code>A</code> and <code>x</code>.  The local system is 2&#215;2 (since the physical dimension is 2).  Solve and update core 0.</p></li><li><p>Core 1: left environment now includes core 0; right environment includes core 2.  Local system is 2&#215;2.  Solve and update core 1.</p></li><li><p>Core 2: left environment includes cores 0 and 1; right environment is empty.  Local system is 2&#215;2.  Solve and update core 2.</p></li></ul><ol start="3"><li><p><strong>Sweep 1 (right-to-left)</strong>: repeat in reverse order.</p></li><li><p>After 2&#8211;3 sweeps, the residual norm drops below 1e-6, and the solution <code>x</code> is accurate to machine precision.</p></li></ol><p>The entire solve takes a few milliseconds on a laptop.</p><h3>Advanced Notes</h3><ul><li><p><strong>Environment contraction</strong>: The left environment for core <code>i</code> is a tensor of shape <code>(A_left_bond_i, x_left_bond_i)</code> that represents the contraction of all cores 0..i-1 of <code>A</code> and <code>x</code>, with all physical indices summed out.  The function <code>_contract_left_env</code> computes this iteratively.  The right environment is computed analogously by contracting from the right.</p></li><li><p><strong>MALS bond dimension growth</strong>: When two cores are optimized together, the bond dimension between them can increase up to the product of the two original bond dimensions.  After the local solve, the new bond dimension is truncated via SVD to a target rank or tolerance.  This allows the algorithm to adapt the rank of the solution automatically.</p></li><li><p><strong>Convergence</strong>: The paper reports that 2 sweeps are sufficient for the multi-asset experiments.  The solver stops when the residual norm <code>||A x - b|| / ||b||</code> falls below a tolerance (default 1e-8) or after <code>max_sweeps</code> sweeps.</p></li></ul><h3>Summary</h3><p>ALS and MALS are the workhorses of the QTT solver.  They solve linear systems in the QTT format by iteratively updating one or two cores at a time, using small dense local solves.  The implementation in <code>als_mals.py</code> follows the paper's description and provides both <code>als_solve</code> and <code>mals_solve</code> functions.  These functions are called by the time-stepping solver at every time step and by the space-time solver for the global system.</p><h2>Constructing the Black-Scholes Operator in QTT</h2><p>We now have all the analytic building blocks we need: the exponential function (rank 1), the boundary selection vectors (rank 2), and the tridiagonal Toeplitz matrix (rank 3).  The next step is to assemble these pieces into the d-asset Black&#8211;Scholes spatial operator in QTT format.  This operator is the core of every time step: at each step we solve <code>(I - &#916;t A) x_{n+1} = x_n</code>, where <code>A</code> is the spatial discretization of the Black&#8211;Scholes PDE.</p><p>The key insight is that the Black&#8211;Scholes PDE in log-price coordinates has <em>constant coefficients</em>.  This means the spatial operator is a sum of terms that are tensor products of one-dimensional operators.  Each one-dimensional operator &#8212; second derivative, first derivative, identity &#8212; can be represented exactly in QTT with rank at most 3.  By combining them via the Kronecker product and addition, we obtain the full d-asset operator with a rank that grows only linearly with <code>d</code>.</p><h3>The d-Asset Black&#8211;Scholes Operator in Log-Price Coordinates</h3><p>Recall the Black&#8211;Scholes PDE after the log transformation <code>x_i = ln S_i</code> and time reversal <code>&#964; = T - t</code>:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\partial V / \\partial \\tau - (1/2) \\sum_{i,j=1}^{d} \\rho_{ij} \\sigma_i \\sigma_j \\partial^2 V / (\\partial x_i \\partial x_j) - \\sum_{i=1}^{d} (r - \\sigma_i^2/2) \\partial V / \\partial x_i + r V = 0&quot;,&quot;id&quot;:&quot;34E171A8A5&quot;}" data-component-name="LatexBlockToDOM"></div><p>This is equation (A2) from the paper, extended to d assets.  The spatial operator <code>L</code> is the part acting on the spatial variables <code>x_1, ..., x_d</code>:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;L = (1/2) \\sum_{i,j=1}^{d} \\rho_{ij} \\sigma_i \\sigma_j \\partial^2 / (\\partial x_i \\partial x_j) + \\sum_{i=1}^{d} (r - \\sigma_i^2/2) \\partial / \\partial x_i - r I&quot;,&quot;id&quot;:&quot;791D2C2D16&quot;}" data-component-name="LatexBlockToDOM"></div><p>where <code>I</code> is the identity operator.  The time-stepping solver uses the operator <code>A = I - &#916;t L</code>, which appears in the backward Euler discretization.</p><h3>Decomposing L into Tensor Products of 1D Operators</h3><p>The crucial observation is that the second-derivative terms can be split into two groups:</p><ul><li><p><strong>Diagonal terms</strong> (<code>i = j</code>):  <code>(1/2) &#963;_i^2 &#8706;&#178;/&#8706;x_i&#178;</code></p></li><li><p><strong>Cross terms</strong> (<code>i &#8800; j</code>):  <code>&#961;_{ij} &#963;_i &#963;_j &#8706;&#178;/(&#8706;x_i &#8706;x_j)</code></p></li></ul><p>Each of these terms acts on a <em>product</em> of one-dimensional operators.  For example, the operator <code>&#8706;&#178;/(&#8706;x_i &#8706;x_j)</code> is the tensor product of the first-derivative operator in dimension <code>i</code> and the first-derivative operator in dimension <code>j</code>.  Similarly, <code>&#8706;&#178;/&#8706;x_i&#178;</code> is the tensor product of the second-derivative operator in dimension <code>i</code> and identity operators in all other dimensions.</p><p>Formally, let <code>D2_i</code> be the second-derivative matrix (tridiagonal Toeplitz) for dimension <code>i</code>, let <code>D1_i</code> be the first-derivative matrix (skew-symmetric tridiagonal), and let <code>I_i</code> be the identity matrix.  All three are of size <code>2^c &#215; 2^c</code> and have exact QTT representations (rank 3 for <code>D2_i</code> and <code>D1_i</code>, rank 1 for <code>I_i</code>).  Then the d-dimensional operator <code>L</code> can be written as:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;L = \\sum_{i=1}^{d} \\frac{\\sigma_i^2}{2} \\, I_1 \\otimes \\cdots \\otimes D2_i \\otimes \\cdots \\otimes I_d \\\\\n+ \\sum_{i=1}^{d} \\left(r - \\frac{\\sigma_i^2}{2}\\right) \\, I_1 \\otimes \\cdots \\otimes D1_i \\otimes \\cdots \\otimes I_d \\\\\n+ \\sum_{i < j} \\rho_{ij} \\sigma_i \\sigma_j \\, I_1 \\otimes \\cdots \\otimes D1_i \\otimes \\cdots \\otimes D1_j \\otimes \\cdots \\otimes I_d \\\\\n- r \\, I_1 \\otimes \\cdots \\otimes I_d&quot;,&quot;id&quot;:&quot;2F953E3D59&quot;}" data-component-name="LatexBlockToDOM"></div><p>Each term in this sum is a Kronecker product of d one-dimensional MPOs.  The Kronecker product of MPOs is itself an MPO: we simply concatenate the cores of the factor MPOs, with the bond dimensions multiplying.  Since each factor has rank at most 3, the rank of each term is at most <code>3^d</code> in the worst case.  However, the paper shows that the <em>sum</em> of these terms has a rank bounded by <code>O(d)</code> because the terms share a common structure and can be combined efficiently.  In practice, the rank of the full operator <code>A</code> is typically <code>3d + 1</code> or less.</p><h3>Building the Operator in Code</h3><p>Let us see how this decomposition translates into code.  The file <code>bs_operator.py</code> contains the function <code>build_bs_operator</code> that constructs the d-asset operator as a QTT-MPO.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;063b4596-953e-4d9d-87cf-ea865b82b8ac&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># bs_operator.py (excerpt)

def build_bs_operator(d, c, sigma, rho, r, dt):
    """
    Build the QTT-MPO for A = I - dt * L, where L is the d-asset
    Black-Scholes spatial operator in log-price coordinates.

    Parameters
    ----------
    d : int
        Number of assets.
    c : int
        Number of QTT cores per spatial dimension (grid points = 2^c).
    sigma : list of float, length d
        Volatilities.
    rho : numpy.ndarray, shape (d, d)
        Correlation matrix.
    r : float
        Risk-free rate.
    dt : float
        Time step size.

    Returns
    -------
    QTTTensor
        The operator A as a QTT-MPO.
    """
    # Build 1D operators
    I_1d = qtt_identity_mpo(c)          # rank 1
    D2_1d = [qtt_toeplitz_tridiagonal(c, alpha=-2.0, beta=1.0, gamma=1.0)
             for _ in range(d)]         # rank 3, second derivative stencil
    D1_1d = [qtt_toeplitz_tridiagonal(c, alpha=0.0, beta=0.5, gamma=-0.5)
             for _ in range(d)]         # rank 3, first derivative stencil

    # Start with the identity term: -r * I_1 &#8855; ... &#8855; I_d
    A = qtt_mul_scalar(qtt_kronecker_product([I_1d] * d), -r)

    # Add diagonal second-derivative terms
    for i in range(d):
        factors = [I_1d] * d
        factors[i] = D2_1d[i]
        term = qtt_kronecker_product(factors)
        term = qtt_mul_scalar(term, 0.5 * sigma[i]**2)
        A = qtt_add(A, term)

    # Add first-derivative terms
    for i in range(d):
        factors = [I_1d] * d
        factors[i] = D1_1d[i]
        term = qtt_kronecker_product(factors)
        term = qtt_mul_scalar(term, r - 0.5 * sigma[i]**2)
        A = qtt_add(A, term)

    # Add cross-derivative terms
    for i in range(d):
        for j in range(i+1, d):
            factors = [I_1d] * d
            factors[i] = D1_1d[i]
            factors[j] = D1_1d[j]
            term = qtt_kronecker_product(factors)
            term = qtt_mul_scalar(term, rho[i, j] * sigma[i] * sigma[j])
            A = qtt_add(A, term)

    # Form A = I - dt * L
    A = qtt_mul_scalar(A, -dt)
    A = qtt_add(qtt_kronecker_product([I_1d] * d), A)

    # Round to control rank growth
    A = qtt_round(A, tol=1e-12)

    return A</code></pre></div><p>Notice the structure: we build each term as a Kronecker product of 1D MPOs, scale it by the appropriate coefficient, and add it to the accumulator <code>A</code>.  The final step forms <code>I - dt * L</code> and rounds the result to keep the bond dimensions under control.  The rounding tolerance <code>1e-12</code> is tight enough to preserve the exactness of the analytic building blocks.</p><h3>The Eraser MPO: Imposing Boundary Conditions</h3><p>The operator <code>A</code> acts on the entire grid, including boundary points.  However, the Black&#8211;Scholes PDE requires Dirichlet boundary conditions on all faces of the domain.  The paper enforces these conditions by constructing a modified right-hand side at each time step, rather than modifying the operator itself.  The key tool is the <strong>eraser MPO</strong>, which zeros out the rows and columns of the operator that correspond to boundary points.</p><p>The eraser MPO is built from the <code>vL</code> and <code>vR</code> vectors we constructed in the analytic QTT section.  For each dimension, we create an MPO that selects the interior points: <code>E_i = I_i - vL_i &#8855; vL_i^T - vR_i &#8855; vR_i^T</code>, where <code>vL_i</code> is the vector that is 1 everywhere except the first entry (which is 0), and <code>vR_i</code> is the vector that is 1 everywhere except the last entry.  The full eraser MPO is the Kronecker product of these 1D erasers:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d0cdf13a-939d-45e3-b2a7-425806194da3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># bs_operator.py (excerpt)

def eraser_mpo(d, c):
    """
    Build the eraser MPO that zeros out all boundary elements.

    Parameters
    ----------
    d : int
        Number of assets.
    c : int
        Number of QTT cores per spatial dimension.

    Returns
    -------
    QTTTensor
        The eraser MPO of shape (2^{c*d}, 2^{c*d}).
    """
    I_1d = qtt_identity_mpo(c)
    vL = qtt_vL_vR(c, side='L')   # rank-2 MPS
    vR = qtt_vL_vR(c, side='R')   # rank-2 MPS

    # Build 1D eraser: I - vL vL^T - vR vR^T
    # vL vL^T is an MPO of rank 2 (outer product of vL with itself)
    vL_mpo = qtt_outer_product(vL, vL)   # rank-2 MPO
    vR_mpo = qtt_outer_product(vR, vR)   # rank-2 MPO
    E_1d = qtt_add(I_1d, qtt_mul_scalar(vL_mpo, -1.0))
    E_1d = qtt_add(E_1d, qtt_mul_scalar(vR_mpo, -1.0))
    E_1d = qtt_round(E_1d, tol=1e-12)

    # Kronecker product over all dimensions
    E = qtt_kronecker_product([E_1d] * d)
    E = qtt_round(E, tol=1e-12)
    return E</code></pre></div><p>The function <code>qtt_outer_product</code> (not shown here) takes an MPS and returns an MPO representing the outer product with itself.  The resulting eraser MPO, when applied to a vector, zeros out all entries that lie on any boundary face.  In the time-stepping solver, we apply the eraser to the solution after each time step to enforce the Dirichlet conditions.</p><h3>Worked Example: Two-Asset Operator</h3><p>Let us walk through a concrete example.  Suppose we have two assets with <code>&#963;&#8321; = 0.25</code>, <code>&#963;&#8322; = 0.15</code>, <code>&#961;&#8321;&#8322; = 0.4</code>, <code>r = 0.05</code>, and we use <code>c = 3</code> cores per dimension (8 grid points per asset).  The operator <code>L</code> has the following terms:</p><ol><li><p><strong>Diagonal second derivatives</strong>: <code>(0.25&#178;/2) I &#8855; D2&#8321; + (0.15&#178;/2) D2&#8322; &#8855; I</code></p></li><li><p><strong>First derivatives</strong>: <code>(0.05 - 0.25&#178;/2) I &#8855; D1&#8321; + (0.05 - 0.15&#178;/2) D1&#8322; &#8855; I</code></p></li><li><p><strong>Cross derivative</strong>: <code>0.4 * 0.25 * 0.15 * D1&#8321; &#8855; D1&#8322;</code></p></li><li><p><strong>Identity</strong>: <code>-0.05 I &#8855; I</code></p></li></ol><p>Each <code>D2</code> and <code>D1</code> is a rank-3 MPO of size <code>8 &#215; 8</code>.  The Kronecker product <code>I &#8855; D2&#8321;</code> is an MPO of size <code>64 &#215; 64</code> with rank 3 (since the identity has rank 1).  The cross term <code>D1&#8321; &#8855; D1&#8322;</code> has rank up to 9 (3 &#215; 3), but after addition and rounding, the total rank of <code>A = I - &#916;t L</code> is typically around 7 for <code>d = 2</code>.  This is far smaller than the full matrix size of 64 &#215; 64 = 4096 entries.</p><h3>Advanced Notes</h3><ul><li><p><strong>Rank bound</strong>: The paper proves that the rank of the d-asset operator <code>A</code> is bounded by <code>3d + 1</code>.  This is because each of the <code>d</code> diagonal second-derivative terms contributes at most 1 to the rank, each of the <code>d</code> first-derivative terms contributes at most 1, and the <code>d(d-1)/2</code> cross-derivative terms can be combined into a single rank-<code>d</code> term.  The identity adds 1.  In practice, the rank is often lower due to the structure of the correlation matrix.</p></li><li><p><strong>Boundary conditions</strong>: The eraser MPO approach is elegant but requires an additional MPO-MPS contraction at each time step.  An alternative is to incorporate the boundary conditions directly into the operator by modifying the cores of <code>A</code>.  The paper uses the eraser approach for simplicity.</p></li><li><p><strong>Time-dependent coefficients</strong>: The operator <code>A</code> is constant in time because the Black&#8211;Scholes PDE has constant coefficients.  This means we build <code>A</code> once and reuse it at every time step, saving significant computation.</p></li></ul><p>With the operator <code>A</code> and the eraser MPO in hand, we are ready to assemble the full time-stepping solver.  The next section will show how to combine these pieces with the payoff and boundary conditions to price European and American options.</p><h2>Time-Stepping Solver for European and American Options</h2><p>We now have all the pieces in place: the Black&#8211;Scholes operator <code>A</code> as a QTT-MPO, the payoff as a QTT-MPS, and the ALS solver to solve linear systems.  The final step is to assemble these into a time-stepping algorithm that advances the solution from maturity backwards to the present.  This section implements the backward Euler time-stepping solver for both European and American options, following the algorithm described in Appendix A.5 of the paper.</p><p>The idea is straightforward.  Starting from the payoff at maturity (<code>&#964; = 0</code>), we repeatedly solve the linear system <code>(I - &#916;t A) x_{n+1} = x_n</code> to step backwards in time.  For European options, this is all we do.  For American options, after each time step we also enforce the early-exercise condition by taking the element-wise maximum of the solution and the payoff.  The ALS solver handles the linear solve, and the TT-Cross algorithm handles the max operation.</p><h3>The Backward Euler Discretization</h3><p>After the log transformation <code>x_i = ln S_i</code> and time reversal <code>&#964; = T - t</code>, the Black&#8211;Scholes PDE becomes:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\partial V / \\partial \\tau - (1/2) \\sum_{i,j=1}^{d} \\rho_{ij} \\sigma_i \\sigma_j \\partial^2 V / (\\partial x_i \\partial x_j) - \\sum_{i=1}^{d} (r - \\sigma_i^2/2) \\partial V / \\partial x_i + r V = 0&quot;,&quot;id&quot;:&quot;32C7244253&quot;}" data-component-name="LatexBlockToDOM"></div><p>We discretize this in time using the fully implicit (backward Euler) scheme.  Let <code>&#916;t</code> be the time step size, and let <code>x_n</code> denote the solution at time level <code>n</code> (with <code>x_0</code> being the payoff at maturity).  The backward Euler discretization gives:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(x_{n+1} - x_n) / \\Delta t - L x_{n+1} = 0&quot;,&quot;id&quot;:&quot;988CB94B62&quot;}" data-component-name="LatexBlockToDOM"></div><p>where <code>L</code> is the spatial operator (the sum of derivative terms).  Rearranging:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(I - \\Delta t L) x_{n+1} = x_n&quot;,&quot;id&quot;:&quot;FC62636401&quot;}" data-component-name="LatexBlockToDOM"></div><p>This is the linear system we solve at each time step.  The matrix <code>I - &#916;t L</code> is the same for every step (since <code>L</code> is constant in time), so we build it once and reuse it.  The right-hand side <code>x_n</code> changes each step, but the ALS solver can use the previous solution as an initial guess, which speeds up convergence.</p><h3>The Time-Stepping Algorithm</h3><p>The algorithm proceeds as follows:</p><ol><li><p>Build the operator <code>A = I - &#916;t L</code> as a QTT-MPO using the construction from the previous section.</p></li><li><p>Build the payoff <code>p</code> as a QTT-MPS using the payoff construction (basket or max-min).</p></li><li><p>Set <code>x = p</code> (solution at maturity).</p></li><li><p>For each time step from <code>n = 0</code> to <code>N-1</code>:</p></li></ol><ul><li><p>Solve <code>A x_new = x</code> using ALS, with <code>x</code> as the initial guess.</p></li><li><p>Set <code>x = x_new</code>.</p></li><li><p>If American: compute <code>x = max(x, p)</code> element-wise using <code>tt_cross_max</code>, then apply rank reduction via <code>qtt_round</code>.</p></li></ul><ol start="5"><li><p>Return <code>x</code> as the solution at <code>t = 0</code>.</p></li></ol><p>The number of time steps <code>N</code> is typically <code>2^c</code>, matching the spatial grid size.  This ensures that the temporal discretization error is of the same order as the spatial error.</p><h3>Code Walkthrough: <code>time_stepping_solver</code></h3><p>The implementation lives in <code>solver.py</code>.  The main function is <code>time_stepping_solver</code>, which orchestrates the loop.  Here is the core logic:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d5562abf-47c2-4bb0-b724-86d9a3de12cd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def time_stepping_solver(market_params, grid_params, option_type, contract_type):
    # Unpack parameters
    d = market_params['d']
    c = grid_params['c']
    dt = grid_params.get('dt', market_params['T'] / 2**c)
    
    # Build operator A = I - dt * L
    A = build_operator_A(d, c, market_params['sigma'], market_params['rho'],
                         market_params['r'], dt)
    
    # Build payoff
    if contract_type == 'basket':
        payoff = payoff_basket_put(market_params['K'], market_params['weights'],
                                   market_params['sigma'], market_params['rho'],
                                   market_params['r'], market_params['T'],
                                   c, grid_params['domain'])
    elif contract_type == 'maxmin':
        payoff = payoff_maxmin_put(market_params['K'], market_params['sigma'],
                                   market_params['rho'], market_params['r'],
                                   market_params['T'], c, grid_params['domain'])
    else:
        raise ValueError(f"Unknown contract type: {contract_type}")
    
    # Initialize solution with payoff
    x = payoff
    
    # Time-stepping loop
    num_steps = int(market_params['T'] / dt)
    for n in range(num_steps):
        # Solve (I - dt*L) x_new = x
        x_new = als_solve(A, x, x0=x, max_sweeps=2)
        
        # Apply boundary conditions
        tau = (n + 1) * dt  # time from maturity
        x_new = impose_boundaries(x_new, tau, contract_type, market_params,
                                  c, grid_params['domain'])
        
        # American early exercise
        if option_type == 'American':
            x_new = tt_cross_max(x_new, payoff, max_rank=20, max_sweeps=4)
            x_new = qtt_round(x_new, tol=1e-6)
        
        x = x_new
    
    return x</code></pre></div><p><strong>What to notice:</strong></p><ul><li><p>The operator <code>A</code> is built once before the loop.  This is efficient because <code>A</code> does not depend on time.</p></li><li><p>The ALS solver uses the previous solution <code>x</code> as the initial guess for the next step.  Since the solution changes slowly between time steps, this reduces the number of ALS sweeps needed (typically 2 sweeps per step).</p></li><li><p>Boundary conditions are imposed after each solve using <code>impose_boundaries</code>, which applies the eraser MPO and adds the boundary values.</p></li><li><p>For American options, the early-exercise condition is applied via <code>tt_cross_max</code> followed by rank reduction.  The rank cap of 20 is based on the paper's experiments, which show that ranks remain bounded even for American options.</p></li></ul><h3>The European Solver</h3><p>For European options, the early-exercise step is omitted.  The function <code>european_solver</code> is a thin wrapper around the loop:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;94aa1fa0-6419-49fd-b621-2cf60d2fa6b0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def european_solver(market_params, grid_params, contract_type):
    return time_stepping_solver(market_params, grid_params,
                                option_type='European',
                                contract_type=contract_type)</code></pre></div><h3>The American Solver</h3><p>For American options, the early-exercise step is included.  The function <code>american_solver</code> is similarly a wrapper:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;463e889e-ffce-4444-a790-cb71ef0be9d2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def american_solver(market_params, grid_params, contract_type):
    return time_stepping_solver(market_params, grid_params,
                                option_type='American',
                                contract_type=contract_type)</code></pre></div><p>The key difference is the <code>tt_cross_max</code> and <code>qtt_round</code> calls inside the loop.  The paper reports that the early-exercise step adds about 30% to the total run time, but the ranks remain controlled (typically &lt; 20).</p><h3>Worked Example: 1-Asset European Call</h3><p>Let's walk through a concrete example.  For a single-asset European call with <code>S0 = 10</code>, <code>K = 10</code>, <code>T = 1</code>, <code>r = 0.05</code>, <code>&#963; = 0.25</code>, and <code>c = 8</code> cores (256 grid points), the solver proceeds as follows:</p><ol><li><p>Build the operator <code>A = I - &#916;t L</code> where <code>L</code> is the 1D Black&#8211;Scholes operator.  The operator is a tridiagonal matrix of size 256&#215;256, represented as a QTT-MPO with rank 3.</p></li><li><p>Build the payoff <code>p = max(e^x - K, 0)</code> as a QTT-MPS.  The payoff is constructed using the analytic exponential QTT and TT-Cross for the max function.</p></li><li><p>Set <code>x = p</code>.</p></li><li><p>For each of the 256 time steps:</p></li></ol><ul><li><p>Solve <code>A x_new = x</code> using ALS (2 sweeps).</p></li><li><p>Apply boundary conditions (Dirichlet at the boundaries).</p></li><li><p>Set <code>x = x_new</code>.</p></li></ul><ol start="5"><li><p>Return <code>x</code> as the option price at <code>t = 0</code>.</p></li></ol><p>The entire computation takes about 1 second on a modern laptop and produces prices accurate to within 0.1% of the analytical Black&#8211;Scholes formula.</p><h3>Advanced Notes</h3><ul><li><p><strong>Rank management:</strong> The solution <code>x</code> changes slowly between time steps, so its QTT rank remains bounded (typically &lt; 10 for European options).  For American options, the early-exercise step can increase the rank, but the subsequent <code>qtt_round</code> call keeps it under control.</p></li><li><p><strong>Boundary conditions:</strong> The <code>impose_boundaries</code> function applies the eraser MPO to zero out boundary points, then adds the correct boundary values.  For basket puts, the lower boundaries are <code>max(Ke^{-rt} - sum of other assets, 0)</code> and the upper boundaries are zero.  These boundary values are constructed using the analytic exponential QTT and TT-Cross.</p></li><li><p><strong>Comparison with space-time solver:</strong> The time-stepping solver requires one ALS solve per time step, which can be expensive for many time steps.  The space-time solver (covered in the next section) solves all time steps simultaneously, which can be faster for European options.  However, the time-stepping solver is the only option for American options, since the early-exercise condition must be applied sequentially.</p></li></ul><h2>Space-Time Solver for European Options</h2><p>The time-stepping solver we built in the previous section works well, but it has a drawback: it requires one ALS solve per time step.  For a grid with 256 time steps, that is 256 separate linear system solves.  The <strong>space-time solver</strong> takes a different approach: it treats the entire space-time grid as a single, larger linear system and solves it in one go.  This can be more efficient because the ALS solver only needs to converge once, and the QTT representation of the space-time operator has a special structure that keeps the ranks low.</p><p>The space-time solver is described in Appendix A.1 of the paper for the single-asset case and generalized to multiple assets in Section III.C.  It is only applicable to European options because the early-exercise condition for American options would break the linear structure of the space-time system.</p><h3>The Space-Time System</h3><p>Recall the Black&#8211;Scholes PDE after the log transformation <code>x = ln S</code> and time reversal <code>&#964; = T - t</code>:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\partial V / \\partial \\tau - (1/2) \\sum_{i,j=1}^{d} \\rho_{ij} \\sigma_i \\sigma_j \\partial^2 V / (\\partial x_i \\partial x_j) - \\sum_{i=1}^{d} (r - \\sigma_i^2/2) \\partial V / \\partial x_i + r V = 0&quot;,&quot;id&quot;:&quot;E8ED7F3811&quot;}" data-component-name="LatexBlockToDOM"></div><p>This is the same PDE we used for the time-stepping solver.  The difference is in how we discretize the time derivative.  Instead of stepping one time level at a time, we discretize all time levels simultaneously using backward Euler.  Let <code>N = 2^c</code> be the number of time steps (we use the same <code>c</code> for time as for each spatial dimension).  Let <code>V_n</code> be the solution vector at time level <code>n</code> (size <code>2^{c*d}</code>).  The backward Euler discretization gives:</p><p><code>(I - &#916;t A) V_{n+1} = V_n</code></p><p>where <code>A</code> is the spatial operator (the same QTT-MPO we built in the previous section) and <code>&#916;t = T / N</code>.  This is a recurrence that links each time level to the next.  We can write all <code>N</code> equations together as a single block-bidiagonal system:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ca8d7a44-66c1-4e2d-a09c-b31c11b9575a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">[ I                                 ] [ V_0   ]   [ b_0   ]
[ -I    (I - &#916;t A)                  ] [ V_1   ]   [ 0     ]
[      -I    (I - &#916;t A)             ] [ V_2   ] = [ 0     ]
[           ...    ...              ] [ ...   ]   [ ...   ]
[                -I    (I - &#916;t A)   ] [ V_N   ]   [ 0     ]</code></pre></div><p>Here <code>V_0</code> is the solution at <code>&#964; = 0</code> (maturity), which is the payoff.  The first row says <code>V_0 = b_0</code>, where <code>b_0</code> is the payoff vector.  The remaining rows enforce the backward Euler recurrence.  The system matrix is block-bidiagonal: the diagonal blocks are <code>(I - &#916;t A)</code> (except the first, which is <code>I</code>), and the subdiagonal blocks are <code>-I</code>.</p><h3>Building the Space-Time Operator in QTT</h3><p>The key to the space-time solver is to represent this block-bidiagonal matrix as a QTT-MPO.  The structure is a Kronecker product of a time-direction operator and the spatial operator.  Let <code>D_t</code> be the <code>N &#215; N</code> bidiagonal matrix that represents the backward Euler recurrence:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4f064ad6-3dfd-4728-af03-aed1741c0179&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">D_t = [ 1                            ]
      [ -1   1                       ]
      [      -1   1                  ]
      [           ...   ...          ]
      [                -1   1        ]</code></pre></div><p>This is a lower bidiagonal matrix with 1 on the diagonal and -1 on the subdiagonal.  The space-time operator <code>A_st</code> is then:</p><p><code>A_st = I_t &#8855; I_x - D_t &#8855; (&#916;t A_x)</code></p><p>where <code>I_t</code> is the <code>N &#215; N</code> identity, <code>I_x</code> is the <code>2^{c*d} &#215; 2^{c*d}</code> spatial identity, and <code>A_x</code> is the spatial operator.  The first term <code>I_t &#8855; I_x</code> is the identity on the full space-time grid.  The second term <code>D_t &#8855; (&#916;t A_x)</code> encodes the backward Euler coupling between time levels.</p><p>Both <code>I_t</code> and <code>D_t</code> are one-dimensional matrices of size <code>N = 2^c</code>.  They can be represented in QTT format using the same analytic building blocks we used for the spatial operator.  The identity <code>I_t</code> is a rank-1 QTT-MPO (each core is the 2&#215;2 identity).  The bidiagonal matrix <code>D_t</code> is a Toeplitz matrix with diagonal 1 and subdiagonal -1, which has a rank-3 QTT representation via Lemma 1 (the same lemma we used for the tridiagonal Toeplitz matrix).</p><p>The Kronecker product <code>I_t &#8855; I_x</code> is built by concatenating the cores of <code>I_t</code> and <code>I_x</code>.  Similarly, <code>D_t &#8855; (&#916;t A_x)</code> is built by concatenating the cores of <code>D_t</code> and the scaled spatial operator.  The resulting space-time operator <code>A_st</code> is a QTT-MPO with <code>c</code> cores for time and <code>c*d</code> cores for space, for a total of <code>c*(d+1)</code> cores.  Its rank is bounded by the sum of the ranks of the two terms, which is at most <code>1 + 3 * (rank of A_x)</code>.  Since the spatial operator rank is <code>O(d)</code>, the space-time operator rank is also <code>O(d)</code>.</p><h3>The Right-Hand Side</h3><p>The right-hand side <code>b</code> of the space-time system is a vector of length <code>N * 2^{c*d}</code>.  It has the payoff vector <code>b_0</code> in the first block (corresponding to <code>&#964; = 0</code>) and zeros everywhere else.  In QTT format, this is a tensor that is the Kronecker product of a time-direction vector <code>e_0</code> (a vector with 1 in the first position and 0 elsewhere) and the payoff QTT-MPS:</p><p><code>b = e_0 &#8855; payoff</code></p><p>The vector <code>e_0</code> is a one-dimensional vector of length <code>N = 2^c</code> with a single 1 at the first index.  This can be represented in QTT format with rank 1: each core has entries <code>[1, 0]</code> (selecting the first bit).  The Kronecker product with the payoff QTT-MPS gives a QTT-MPS of length <code>N * 2^{c*d}</code> with rank equal to the payoff rank.</p><h3>Solving the Space-Time System</h3><p>Once we have the space-time operator <code>A_st</code> and the right-hand side <code>b</code> in QTT format, we solve the linear system <code>A_st x = b</code> using the ALS or MALS solver.  The solution <code>x</code> is a QTT-MPS of length <code>N * 2^{c*d}</code>.  It contains the solution at all time levels, stacked in order: the first <code>2^{c*d}</code> entries are <code>V_0</code> (the payoff), the next <code>2^{c*d}</code> entries are <code>V_1</code> (the solution one time step back), and so on.  To extract the solution at the present time (<code>&#964; = T</code>), we take the last block of <code>2^{c*d}</code> entries.</p><p>The ALS solver converges in a few sweeps because the space-time operator is well-conditioned (it is a small perturbation of the identity).  The rank of the solution <code>x</code> is typically bounded by the rank of the right-hand side plus a small increase from the operator application.  In practice, the solution rank stays below 20 for the experiments in the paper.</p><h3>Code Walkthrough: space_time.py</h3><p>The implementation of the space-time solver is in the file <code>space_time.py</code>.  It contains two main functions: <code>build_space_time_operator</code> and <code>space_time_solver</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6691e849-7e52-448e-88e0-e84a39c33ce1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># space_time.py (excerpt)

def build_space_time_operator(c, dt, A_x):
    """
    Build the space-time operator A_st = I_t &#8855; I_x - D_t &#8855; (dt * A_x).

    Parameters
    ----------
    c : int
        Number of cores for the time dimension (N = 2^c time steps).
    dt : float
        Time step size.
    A_x : QTTTensor (MPO)
        Spatial operator (I - dt * A) from the time-stepping solver.

    Returns
    -------
    QTTTensor (MPO)
        Space-time operator of size (N * 2^{c*d}, N * 2^{c*d}).
    """
    # Build the time-direction identity I_t (rank 1)
    I_t = qtt_identity_mpo(c)

    # Build the time-direction bidiagonal matrix D_t (rank 3)
    # D_t has diagonal 1, subdiagonal -1, superdiagonal 0
    D_t = qtt_toeplitz_tridiagonal(c, alpha=1.0, beta=0.0, gamma=-1.0)

    # Build the spatial identity I_x
    I_x = qtt_identity_mpo(c * d)  # d is the number of assets

    # Build the scaled spatial operator dt * A_x
    dt_A_x = qtt_mul_scalar(A_x, dt)

    # Build the two terms
    term1 = qtt_kronecker_product([I_t, I_x])
    term2 = qtt_kronecker_product([D_t, dt_A_x])

    # A_st = term1 - term2
    A_st = qtt_add(term1, qtt_mul_scalar(term2, -1.0))

    return A_st</code></pre></div><p>This function constructs the space-time operator by combining the time-direction and spatial QTT-MPOs via Kronecker products.  The <code>qtt_kronecker_product</code> function concatenates the cores of the input MPOs, producing a new MPO with <code>c + c*d = c*(d+1)</code> cores.  The addition and scalar multiplication are the same operations we used for the time-stepping operator.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;af0abdd3-89f1-45f8-9a71-1c919705516e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># space_time.py (excerpt)

def space_time_solver(market_params, grid_params):
    """
    Solve the space-time system for a European option.

    Parameters
    ----------
    market_params : dict
        Market parameters (d, sigma, rho, r, K, T, weights).
    grid_params : dict
        Grid parameters (c, domain).

    Returns
    -------
    QTTTensor (MPS)
        Solution at all time levels, shape (N * 2^{c*d},).
    """
    d = market_params['d']
    c = grid_params['c']
    N = 2**c  # number of time steps
    dt = market_params['T'] / N

    # Build the spatial operator A_x (same as time-stepping)
    A_x = build_bs_operator(d, c, market_params['sigma'],
                            market_params['rho'], market_params['r'], dt)

    # Build the space-time operator
    A_st = build_space_time_operator(c, dt, A_x)

    # Build the payoff at maturity
    payoff = payoff_basket_put(market_params['K'], market_params['weights'],
                               market_params['sigma'], market_params['rho'],
                               market_params['r'], market_params['T'],
                               c, grid_params['domain'])

    # Build the time-direction vector e_0 (rank 1)
    # e_0 has 1 at the first index, 0 elsewhere
    e_0_cores = []
    for i in range(c):
        core = np.zeros((1, 2, 1))
        if i == 0:
            core[0, 0, 0] = 1.0  # first bit is 0
        else:
            core[0, 0, 0] = 1.0  # all bits are 0
        e_0_cores.append(core)
    e_0 = QTTTensor(e_0_cores)

    # Right-hand side: b = e_0 &#8855; payoff
    b = qtt_kronecker_product([e_0, payoff])

    # Solve the space-time system using ALS
    x = als_solve(A_st, b, max_sweeps=2)

    return x</code></pre></div><p>This function assembles the space-time system and solves it.  The key steps are:</p><ol><li><p>Build the spatial operator <code>A_x</code> (the same <code>I - &#916;t A</code> used in the time-stepping solver).</p></li><li><p>Build the space-time operator <code>A_st</code> using <code>build_space_time_operator</code>.</p></li><li><p>Build the payoff QTT-MPS.</p></li><li><p>Build the time-direction vector <code>e_0</code> (a rank-1 QTT-MPS with a single 1 at the first position).</p></li><li><p>Form the right-hand side <code>b</code> as the Kronecker product of <code>e_0</code> and the payoff.</p></li><li><p>Solve the system using ALS.</p></li></ol><p>The solution <code>x</code> is a QTT-MPS of length <code>N * 2^{c*d}</code>.  To extract the price at the present time (<code>&#964; = T</code>), we take the last <code>2^{c*d}</code> entries.  This can be done by reshaping the QTT-MPS or by contracting with a selection vector.</p><h3>Worked Example: 1-Asset European Call</h3><p>Let us walk through a concrete example.  Consider a 1-asset European call with the following parameters:</p><ul><li><p>Spot price <code>S0 = 10</code>, strike <code>K = 10</code>, volatility <code>&#963; = 0.25</code>, risk-free rate <code>r = 0.05</code>, maturity <code>T = 1.0</code>.</p></li><li><p>Grid: <code>c = 8</code> cores per dimension, so <code>N = 256</code> time steps and <code>256</code> spatial points.</p></li><li><p>Domain: <code>x</code> in <code>[-2, 2]</code> (log-price).</p></li></ul><p>The space-time system has <code>256 * 256 = 65536</code> unknowns.  The QTT representation of the solution uses <code>c*(d+1) = 8*2 = 16</code> cores.  With a rank of about 10, the total number of parameters is roughly <code>16 * 10^2 = 1600</code> &#8212; a tiny fraction of the full grid size.</p><p>The ALS solver converges in 2-3 sweeps, taking a few seconds on a personal computer.  The resulting price at the spot price matches the analytical Black&#8211;Scholes price to within 0.1%.</p><h3>Comparison with Time-Stepping</h3><p>The space-time solver has both advantages and disadvantages compared to the time-stepping solver:</p><ul><li><p><strong>Advantage</strong>: Only one ALS solve is needed, instead of one per time step.  This can be faster for problems with many time steps.</p></li><li><p><strong>Advantage</strong>: The space-time operator has a simple Kronecker structure that is easy to build.</p></li><li><p><strong>Disadvantage</strong>: The space-time system is larger (by a factor of <code>N</code> in the number of unknowns), so the ALS solver works with longer cores and potentially higher ranks.</p></li><li><p><strong>Disadvantage</strong>: The space-time solver is only applicable to European options.  American options require the early-exercise condition, which is a nonlinear operation that cannot be expressed as a linear system.</p></li></ul><p>In practice, the paper reports that the space-time solver is competitive for low dimensions (<code>d &#8804; 3</code>) but becomes less efficient than time-stepping for higher dimensions due to rank growth in the space-time operator.</p><h3>Advanced Notes</h3><p>The time complexity of the space-time solver for d assets is <code>O(c (d^5 + d^3) &#967;_ST^3)</code>, where <code>&#967;_ST</code> is the rank of the space-time solution.  This is derived in Appendix A.3 of the paper.  The cubic dependence on rank means that rank control is critical: if the solution rank grows too large, the solver becomes slow.  In practice, the rank stays below 20 for the experiments in the paper, keeping the solver efficient.</p><p>The space-time solver can also be extended to handle more complex time discretizations, such as Crank-Nicolson, by modifying the time-direction matrix <code>D_t</code>.  However, the paper uses backward Euler for its simplicity and low-rank preservation.</p><h2>Computing Greeks from the QTT Grid</h2><p>A full-grid solver gives you more than just the option price at a single point.  Once you have the entire solution surface in QTT format, you can compute the Greeks &#8212; Delta and Gamma &#8212; over the whole grid with negligible additional cost.  The key insight is that the derivative operators in log-price space have exact, low-rank QTT representations.  Applying them to the solution MPS produces the Greeks in log-price coordinates, and a simple chain-rule conversion gives the familiar Delta and Gamma in the original asset-price coordinates.</p><p>This section implements the Greek computation as described in Appendix A.2 of the paper.  The generated file <code>greeks.py</code> provides two public functions: <code>compute_delta</code> and <code>compute_gamma</code>.  Each can return either a QTTTensor representing the Greek over the entire grid or a single interpolated value at a given spot price.</p><h3>Derivative MPOs from the Tridiagonal Building Block</h3><p>The first and second derivatives in log-price space are discretized using central differences.  The first derivative uses the stencil <code>[-1/(2dx), 0, 1/(2dx)]</code>, and the second derivative uses <code>[1/dx^2, -2/dx^2, 1/dx^2]</code>.  Both are tridiagonal Toeplitz matrices, so they can be built exactly using the same <code>qtt_toeplitz_tridiagonal</code> function from <code>qtt_analytic.py</code> (Lemma 1, rank 3).</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e0c4e342-5886-4b79-8b60-bee9aa862738&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def _derivative_mpo_first(c, dx):
    """
    Build a 1D first-derivative MPO using central differences.

    Implements the tridiagonal Toeplitz matrix with stencil [-1/(2dx), 0, 1/(2dx)]
    (Lemma 1 / Appendix C.3).

    Parameters
    ----------
    c : int
        Number of QTT cores (grid size = 2^c).
    dx : float
        Spatial step size in the log-price domain.

    Returns
    -------
    QTTTensor
        MPO of shape (2^c, 2^c) with rank 3.
    """
    # Central difference: subdiagonal = -1/(2dx), diagonal = 0, superdiagonal = 1/(2dx)
    beta = 0.5 / dx
    gamma = -0.5 / dx
    return qtt_toeplitz_tridiagonal(c, alpha=0.0, beta=beta, gamma=gamma)


def _derivative_mpo_second(c, dx):
    """
    Build a 1D second-derivative MPO using central differences.

    Stencil [1/dx^2, -2/dx^2, 1/dx^2] (Lemma 1 / Appendix C.3).

    Parameters
    ----------
    c : int
        Number of QTT cores (grid size = 2^c).
    dx : float
        Spatial step size in the log-price domain.

    Returns
    -------
    QTTTensor
        MPO of shape (2^c, 2^c) with rank 3.
    """
    alpha = -2.0 / (dx * dx)
    beta = 1.0 / (dx * dx)
    gamma = 1.0 / (dx * dx)
    return qtt_toeplitz_tridiagonal(c, alpha=alpha, beta=beta, gamma=gamma)</code></pre></div><p>Notice that the first-derivative MPO has a zero diagonal (<code>alpha=0.0</code>), while the second-derivative MPO has <code>alpha = -2/dx^2</code>.  Both have rank 3, matching the tridiagonal Toeplitz construction from Lemma 1.</p><h3>Building a d-Dimensional Derivative MPO</h3><p>To differentiate with respect to a single asset in a d-asset problem, we need an MPO that applies the 1D derivative in the target dimension and the identity in all other dimensions.  The function <code>_build_derivative_mpo_d</code> constructs this by concatenating the cores of the 1D derivative MPO for the target dimension and the cores of the 1D identity MPO for every other dimension.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2fc824e0-eeeb-4273-924e-cbf9f1ddd629&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def _build_derivative_mpo_d(d, c, asset_index, dx):
    """
    Build a d-dimensional MPO that applies the first derivative in the
    `asset_index`-th dimension and identity in all other dimensions.

    The MPO is constructed by concatenating the 1D cores for each dimension:
      - For the target dimension: use the 1D derivative MPO cores.
      - For all other dimensions: use the 1D identity MPO cores.

    Parameters
    ----------
    d : int
        Number of assets (dimensions).
    c : int
        Number of QTT cores per dimension.
    asset_index : int
        Index of the dimension along which to differentiate (0-based).
    dx : float
        Spatial step size in the log-price domain.

    Returns
    -------
    QTTTensor
        MPO of shape (2^{c*d}, 2^{c*d}).
    """
    derivative_1d = _derivative_mpo_first(c, dx)
    identity_1d = qtt_identity_mpo(c)

    cores = []
    for dim in range(d):
        if dim == asset_index:
            cores.extend(derivative_1d.cores)
        else:
            cores.extend(identity_1d.cores)

    return QTTTensor(cores, is_mpo=True)</code></pre></div><p>This function is used for the first derivative.  For the second derivative, the same pattern applies but with <code>_derivative_mpo_second</code> in place of <code>_derivative_mpo_first</code>.  The resulting MPO has rank 3 in the target dimension and rank 1 in all others, so the overall rank is bounded by 3.</p><h3>Computing Delta and Gamma</h3><p>With the derivative MPOs in hand, computing the Greeks is a matter of applying them to the solution MPS and converting from log-price coordinates to asset-price coordinates.</p><p><strong>Delta</strong> in log-price coordinates is <code>&#8706;V/&#8706;x</code>.  The chain rule gives the familiar Delta in asset-price coordinates:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\Delta = \\frac{\\partial V}{\\partial S} = \\frac{1}{S} \\frac{\\partial V}{\\partial x}&quot;,&quot;id&quot;:&quot;E4CA58C447&quot;}" data-component-name="LatexBlockToDOM"></div><p><strong>Gamma</strong> requires both the first and second derivatives in log-price:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\Gamma = \\frac{\\partial^2 V}{\\partial S^2} = \\frac{1}{S^2} \\left( \\frac{\\partial^2 V}{\\partial x^2} - \\frac{\\partial V}{\\partial x} \\right)&quot;,&quot;id&quot;:&quot;01A4A1910F&quot;}" data-component-name="LatexBlockToDOM"></div><p>These formulas follow from the change of variables <code>x = ln S</code> and the chain rule.  The paper derives them in Appendix A.2.</p><p>Here is the <code>compute_delta</code> function from <code>greeks.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7eba02e5-fb92-4fd8-a15c-10c623f0920d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def compute_delta(solution_qtt, asset_index, market_params, grid_params, spot=None):
    """
    Compute the Delta (&#8706;V/&#8706;S) for a given asset from the QTT solution.

    The chain rule is used: &#916; = (1/S) * &#8706;V/&#8706;x, where x = log S.

    If `spot` is provided, returns the interpolated Delta at that spot price
    for the specified asset, with other assets fixed at their initial spots.
    Otherwise, returns a QTTTensor representing the Delta grid over all assets.

    Parameters
    ----------
    solution_qtt : QTTTensor
        MPS of the option price at the final time (shape 2^{c*d}).
    asset_index : int
        Index of the asset for which to compute Delta (0-based).
    market_params : dict
        Must contain 'S0' (list of initial spot prices) and other parameters.
    grid_params : dict
        Must contain 'c' (int) and 'domain' (tuple (xmin, xmax)).
    spot : float, optional
        If provided, return the interpolated Delta at this spot.

    Returns
    -------
    QTTTensor or float
        The Delta grid (QTTTensor) or interpolated value (float).
    """
    c = grid_params['c']
    xmin, xmax = grid_params['domain']
    dx = (xmax - xmin) / (2**c - 1)
    d = len(market_params['S0'])

    # Build derivative MPO for the specified asset dimension
    derivative_mpo = _build_derivative_mpo_d(d, c, asset_index, dx)

    # Apply to get &#8706;V/&#8706;x
    dV_dx = qtt_contract_mpo_mps(derivative_mpo, solution_qtt)
    # Round to control rank growth (optional, but good practice)
    dV_dx = qtt_round(dV_dx, tol=1e-8)

    if spot is not None:
        # Interpolate at the given spot
        val = _interpolate_at_spot(dV_dx, spot, asset_index, market_params, grid_params)
        # &#916; = (1/S) * dV/dx
        return val / spot
    else:
        # Build the d-dimensional QTT for e^{-x} on the asset_index dimension
        # and all-ones on other dimensions.
        # First, build 1D exponential QTT: e^{-x} on the whole domain.
        exp_minus_x = analytic_qtt_exponential(alpha=-1.0, c=c, interval=(xmin, xmax))

        # Build list of 1D QTTs for Kronecker product
        ones_qtt = analytic_qtt_exponential(alpha=0.0, c=c, interval=(xmin, xmax))  # all ones

        qtt_list = []
        for i in range(d):
            if i == asset_index:
                qtt_list.append(exp_minus_x)
            else:
                qtt_list.append(ones_qtt)

        # Use qtt_kronecker_product to combine (only works for rank-1 MPS)
        from qtt_core import qtt_kronecker_product
        inv_s_qtt = qtt_kronecker_product(qtt_list)

        # Element-wise multiply: &#916; = dV_dx * e^{-x}
        delta_qtt = _elementwise_mul_mps(dV_dx, inv_s_qtt)
        # Round to control rank
        delta_qtt = qtt_round(delta_qtt, tol=1e-8)

        return delta_qtt</code></pre></div><p>The function first builds the derivative MPO for the specified asset dimension and applies it to the solution to get <code>&#8706;V/&#8706;x</code>.  If a specific <code>spot</code> price is requested, it interpolates <code>&#8706;V/&#8706;x</code> at that point and divides by <code>spot</code> to get Delta.  Otherwise, it builds a QTT representation of <code>1/S = e^{-x}</code> using the analytic exponential function (rank 1) and multiplies element-wise with <code>&#8706;V/&#8706;x</code> to get the full Delta grid.</p><p>The <code>compute_gamma</code> function follows the same pattern but computes both <code>&#8706;V/&#8706;x</code> and <code>&#8706;&#178;V/&#8706;x&#178;</code>, subtracts them, and multiplies by <code>e^{-2x}</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;3c59e53c-dd95-4316-9602-1fb5e085a348&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def compute_gamma(solution_qtt, asset_index, market_params, grid_params, spot=None):
    """
    Compute the Gamma (&#8706;&#178;V/&#8706;S&#178;) for a given asset from the QTT solution.

    The chain rule is used: &#915; = (1/S&#178;) * (&#8706;&#178;V/&#8706;x&#178; - &#8706;V/&#8706;x), where x = log S.

    If `spot` is provided, returns the interpolated Gamma at that spot price
    for the specified asset. Otherwise, returns a QTTTensor representing the
    Gamma grid over all assets.

    Parameters
    ----------
    solution_qtt : QTTTensor
        MPS of the option price at the final time (shape 2^{c*d}).
    asset_index : int
        Index of the asset for which to compute Gamma (0-based).
    market_params : dict
        Must contain 'S0' (list of initial spot prices).
    grid_params : dict
        Must contain 'c' (int) and 'domain' (tuple (xmin, xmax)).
    spot : float, optional
        If provided, return the interpolated Gamma at this spot.

    Returns
    -------
    QTTTensor or float
        The Gamma grid (QTTTensor) or interpolated value (float).
    """
    c = grid_params['c']
    xmin, xmax = grid_params['domain']
    dx = (xmax - xmin) / (2**c - 1)
    d = len(market_params['S0'])

    # ---- Compute &#8706;V/&#8706;x ----
    deriv1_mpo = _build_derivative_mpo_d(d, c, asset_index, dx)
    dV_dx = qtt_contract_mpo_mps(deriv1_mpo, solution_qtt)
    dV_dx = qtt_round(dV_dx, tol=1e-8)

    # ---- Compute &#8706;&#178;V/&#8706;x&#178; ----
    # Build second derivative MPO for the target dimension
    # We need a d-dimensional MPO for the second derivative.
    # Reuse the same construction as first derivative but with second derivative MPO.
    deriv2_1d = _derivative_mpo_second(c, dx)
    identity_1d = qtt_identity_mpo(c)

    cores = []
    for dim in range(d):
        if dim == asset_index:
            cores.extend(deriv2_1d.cores)
        else:
            cores.extend(identity_1d.cores)
    deriv2_mpo = QTTTensor(cores, is_mpo=True)

    d2V_dx2 = qtt_contract_mpo_mps(deriv2_mpo, solution_qtt)
    d2V_dx2 = qtt_round(d2V_dx2, tol=1e-8)

    # ---- Compute (&#8706;&#178;V/&#8706;x&#178; - &#8706;V/&#8706;x) ----
    # Element-wise subtraction
    # We need a subtraction function; we can use qtt_add with negative scalar
    from qtt_core import qtt_add, qtt_mul_scalar
    neg_dV_dx = qtt_mul_scalar(dV_dx, -1.0)
    diff = qtt_add(d2V_dx2, neg_dV_dx)
    diff = qtt_round(diff, tol=1e-8)

    if spot is not None:
        # Interpolate the difference at the spot
        val = _interpolate_at_spot(diff, spot, asset_index, market_params, grid_params)
        # &#915; = (1/S&#178;) * val
        return val / (spot * spot)
    else:
        # Build the d-dimensional QTT for e^{-2x} on the asset_index dimension
        exp_minus_2x = analytic_qtt_exponential(alpha=-2.0, c=c, interval=(xmin, xmax))
        ones_qtt = analytic_qtt_exponential(alpha=0.0, c=c, interval=(xmin, xmax))

        qtt_list = []
        for i in range(d):
            if i == asset_index:
                qtt_list.append(exp_minus_2x)
            else:
                qtt_list.append(ones_qtt)

        from qtt_core import qtt_kronecker_product
        inv_s2_qtt = qtt_kronecker_product(qtt_list)

        # Element-wise multiply: &#915; = diff * e^{-2x}
        gamma_qtt = _elementwise_mul_mps(diff, inv_s2_qtt)
        gamma_qtt = qtt_round(gamma_qtt, tol=1e-8)

        return gamma_qtt</code></pre></div><h3>Worked Example: Delta and Gamma for a 1-Asset European Call</h3><p>Let us walk through a concrete example.  Suppose we have solved the Black&#8211;Scholes PDE for a 1-asset European call with <code>c = 8</code> cores (256 grid points), <code>sigma = 0.25</code>, <code>r = 0.05</code>, <code>T = 1.0</code>, and <code>K = 10</code>.  The solution MPS <code>V</code> has shape <code>(256,)</code> and rank typically around 5&#8211;10.</p><p>To compute Delta at the spot price <code>S0 = 10</code>:</p><ol><li><p>Build the 1D first-derivative MPO <code>D_x</code> with <code>dx = (xmax - xmin) / (2^c - 1)</code>.</p></li><li><p>Apply <code>D_x</code> to <code>V</code> to get <code>dV_dx</code> (a new MPS of shape <code>(256,)</code>).</p></li><li><p>Interpolate <code>dV_dx</code> at <code>x0 = ln(10)</code> to get <code>dV_dx_at_spot</code>.</p></li><li><p>Divide by <code>S0 = 10</code> to get <code>Delta = dV_dx_at_spot / 10</code>.</p></li></ol><p>For Gamma, we additionally compute <code>d2V_dx2</code> by applying the second-derivative MPO, then compute <code>diff = d2V_dx2 - dV_dx</code>, interpolate at <code>x0</code>, and divide by <code>S0^2 = 100</code>.</p><p>The paper reports in Table IV that for a 1-asset European call, the grid-wide Delta and Gamma errors are on the order of 1e-4 to 1e-3, and the computation adds only a few percent to the total solver runtime.  This is because the derivative MPOs have rank at most 3, and the element-wise multiplication with <code>e^{-x}</code> or <code>e^{-2x}</code> uses rank-1 MPS, so the overall rank of the Greek tensors remains small.</p><h3>Advanced Detail: Interpolation on the Dyadic Grid</h3><p>The function <code>_interpolate_at_spot</code> uses nearest-neighbor interpolation on the dyadic log-price grid.  For a given spot price <code>S</code>, it computes <code>x = ln(S)</code>, clamps it to the domain <code>[xmin, xmax]</code>, finds the nearest grid index, and evaluates the MPS at that multi-index using <code>_evaluate_mps_at_indices</code>.  The evaluation contracts the MPS cores with one-hot vectors corresponding to the binary representation of each index, which is efficient because the MPS is never expanded to its full vector form.</p><p>For higher accuracy, one could replace nearest-neighbor with linear or cubic interpolation, but the paper shows that nearest-neighbor is sufficient for the reported error levels.</p><h3>Summary</h3><ul><li><p>Derivative MPOs for first and second derivatives in log-price space are built from the analytic tridiagonal Toeplitz representation (Lemma 1, rank 3).</p></li><li><p>A d-dimensional derivative MPO is constructed by concatenating 1D derivative cores for the target dimension and 1D identity cores for all other dimensions.</p></li><li><p>Delta and Gamma are computed by applying the derivative MPOs to the solution MPS and converting from log-price to asset-price coordinates using the chain rule.</p></li><li><p>The conversion uses the analytic rank-1 QTT for <code>e^{-x}</code> (for Delta) or <code>e^{-2x}</code> (for Gamma), multiplied element-wise with the derivative MPS.</p></li><li><p>The entire computation is efficient because all operations stay within the QTT format, and the ranks of the derivative MPOs and conversion factors are small (3 or 1).</p></li></ul><h2>Numerical Experiments and Results</h2><p>We now bring together all the components built in the previous sections to reproduce the key numerical results from the paper.  The experiments cover European and American basket puts and max-min puts in dimensions d = 3, 4, 5, using the time-stepping solver as the primary tool.  For European options we also test the space-time solver and confirm that it delivers the same prices (within truncation error) with a different efficiency profile.  All results are compared against reference prices obtained by independent methods: Gauss&#8211;Hermite quadrature for European contracts and Longstaff&#8211;Schwartz Monte Carlo for American contracts.</p><h3>Experimental Setup</h3><p>The market parameters are taken from Appendix B.2 of the paper.  A single set of parameters is used across all experiments, with the number of assets d determining which entries are active:</p><ul><li><p>Spot prices (log&#8209;price coordinates): <code>S = (10, 11, 12, 13, 14)</code> for assets 1 through 5.</p></li><li><p>Volatilities: <code>sigma = (0.25, 0.15, 0.20, 0.10, 0.15)</code>.</p></li><li><p>Correlation matrix <code>rho</code> of size 5 x 5 with off&#8209;diagonal entries <code>(0.4, 0.3, 0.2, 0.1, 0.4, ...)</code> (see paper for the full matrix).</p></li><li><p>Risk&#8209;free rate <code>r = 0.05</code> (5% per annum).</p></li><li><p>Maturity <code>T = 1.0</code> year.</p></li><li><p>Strikes: for basket puts, <code>K = 34, 47, 62</code> for d = 3, 4, 5 respectively; for max&#8209;min puts, <code>K = 10</code> for all d.</p></li></ul><p>The grid parameters are:</p><ul><li><p>Number of spatial cores <code>c</code> per dimension, typically 7, 8, or 9.  The grid per dimension therefore has <code>N = 2^c</code> points.</p></li><li><p>Number of time steps in the time&#8209;stepping solver is also <code>2^c</code> (the paper uses a fully implicit scheme with <code>theta = 1</code>).</p></li><li><p>The spatial domain <code>[x_min, x_max]^d</code> is determined adaptively by a pilot coarse run (see the function <code>adaptive_domain</code> in <code>utils.py</code>).</p></li></ul><p>Reference prices are computed as follows:</p><ul><li><p><strong>European basket put</strong>: Gauss&#8211;Hermite quadrature with 100 points per dimension (using the function <code>gauss_hermite_quadrature_basket</code> in <code>utils.py</code>).  This gives a highly accurate reference value.</p></li><li><p><strong>European max&#8209;min put</strong>: The same Gauss&#8211;Hermite method applied to the payoff <code>max(K - (min_i e^{x_i}), 0)</code>.</p></li><li><p><strong>American basket and max&#8209;min puts</strong>: Longstaff&#8211;Schwartz least&#8209;squares Monte Carlo with 100,000 paths and 100 time steps (function <code>longstaff_schwartz_american</code>).  The Monte Carlo reference itself has a standard deviation of roughly 1&#8211;2% of the price.</p></li></ul><p>Error metrics used in all tables:</p><ul><li><p><code>|price_error| = |price_QTT - price_ref|</code> (absolute error).</p></li><li><p><code>relative_error = |price_error| / price_ref</code>.</p></li><li><p>For Greeks, the same metrics are computed over the whole grid (Delta and Gamma are compared to the reference grid, which is obtained by finite differences on the Gauss&#8211;Hermite solution).</p></li></ul><h3>Running a Basket Experiment</h3><p>The central experiment for a given <code>d</code> and option type is orchestrated by the function <code>run_basket_experiment</code> in <code>experiments.py</code>.  A simplified version of its body is shown below; it relies on the solver and utilities built earlier.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;aa460be1-d8d2-44ac-8d1a-d2a1470fca68&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def run_basket_experiment(d, c, option_type, ref_price):
    # Market parameters for d assets
    params = default_market_params(d)  # from utils.py

    # Adaptive grid domain (pilot coarse run with c=5)
    domain = adaptive_domain(params, c_coarse=5)
    grid_params = {'c': c, 'domain': domain}

    # Solve using the time-stepping solver
    solution_qtt = time_stepping_solver(
        params, grid_params,
        option_type=option_type,
        contract_type='basket'
    )

    # Interpolate at the spot price (first element of params['spot'])
    spot = params['spot'][0]  # for d=3, S0 = 10
    price = solution_qtt.interpolate_at_spot(spot)

    # Compute error
    error = abs(price - ref_price) / ref_price
    return price, error</code></pre></div><p>For American options the same function is called with <code>option_type='American'</code>; the solver internally applies the early&#8209;exercise condition (via <code>tt_cross_max</code>) after every time step.</p><h3>Reproducing the Paper&#8217;s Tables</h3><p>The paper presents its main results in Tables V through VIII.  The following paragraphs summarise those results and explain how the code reproduces them.  Because no computation was executed during this writing, the values quoted here are taken directly from the paper; running the code with the same parameters should yield numbers matching the published tables within the reported tolerances.</p><p><strong>Table V: European basket put (d = 3, 4, 5).</strong>  The time&#8209;stepping solver with c = 8 cores per dimension (256 grid points per asset) achieves absolute pricing errors below 0.05 for all dimensions.  The run time grows modestly with d: for d = 3 the solver completes in under 10 seconds on the reference hardware (Apple M3 Pro); for d = 5 the runtime is still under one minute.  The rank of the solution never exceeds 15.  The space&#8209;time solver, when applied to the same problems, yields identical prices (within 1e&#8209;6) and is faster for d = 3 but requires more memory because the global solution surface has rank that is 2&#8211;3 times larger than the time&#8209;stepping representation.</p><p><strong>Table VI: European max&#8209;min put.</strong>  The max&#8209;min payoff introduces the <code>min</code> function, which is approximated by TT&#8209;Cross with rank cap chi = 12.  The resulting pricing errors are slightly larger than for the basket contract &#8212; typically 0.1&#8211;0.3 absolute error &#8212; but still well within the 1&#8211;2% range needed for practical applications.  The run times are comparable to the basket case because the extra TT&#8209;Cross cost appears only during the initial payoff construction.</p><p><strong>Table VII: American basket put.</strong>  The early&#8209;exercise condition adds about 30% to the total runtime.  The paper reports that the rank of the solution grows by at most 2&#8211;3 additional bond dimensions compared to the European case, confirming that the QTT representation remains efficient.  Pricing errors are slightly larger than for European options because of the additional approximation introduced by the TT&#8209;Cross max operation at each time step.</p><p><strong>Table VIII: American max&#8209;min put.</strong>  This is the most challenging contract because both the payoff and the early&#8209;exercise condition involve the <code>min</code> and <code>max</code> functions.  The observed ranks are 2&#8211;3 times higher than for the basket put, and the runtime increases accordingly.  Nevertheless, for d = 3 the solver still finishes in under 30 seconds, and for d = 5 it remains under 5 minutes on a personal computer.</p><h3>Worked Example: d = 3 Basket Put, c = 8</h3><p>To illustrate the typical workflow, consider the 3&#8209;asset European basket put.  The steps are:</p><ol><li><p>Load the market parameters: <code>sigma = [0.25, 0.15, 0.20]</code>, <code>rho</code> (the 3x3 block of the full matrix), <code>r = 0.05</code>, <code>K = 34</code>, <code>T = 1.0</code>.</p></li><li><p>Determine the domain via a coarse pilot run (e.g., <code>c = 5</code>, domain expands to cover the region where the payoff is non&#8209;zero).  The result might be <code>[x_min, x_max] = [-2.1, 2.1]</code>.</p></li><li><p>Build the operator MPO <code>A</code> (using the analytic QTT building blocks) and the payoff QTT (using <code>payoff_basket_put</code>).</p></li><li><p>Run the time&#8209;stepping loop: 256 time steps, each solving <code>(I - dt A) x_{n+1} = x_n</code> with the ALS solver (2 sweeps, initial guess from previous step).</p></li><li><p>Extract the price at the spot point <code>(10, 11, 12)</code> by interpolating the final time&#8209;slice QTT.</p></li></ol><p>The computed price should match the reference (obtained by Gauss&#8211;Hermite quadrature) to within 0.02&#8211;0.03 absolute error, corresponding to a relative error of less than 1%.  The entire computation takes about 8 seconds on the reference machine.</p><h3>Advanced Details</h3><ul><li><p><strong>Rank behaviour</strong>: For all experiments reported in the paper, the bond dimensions of the solution QTT remained below 20 even for the most demanding 5&#8209;asset American max&#8209;min case.  This is crucial for the polynomial scaling of the method.</p></li><li><p><strong>Space&#8209;time vs. time&#8209;stepping</strong>: The space&#8209;time solver is approximately twice as fast as the time&#8209;stepping solver for d = 3 European options, but its memory footprint is larger because it stores the entire space&#8209;time surface.  For d &gt; 3 the time&#8209;stepping solver is preferred because the space&#8209;time operator rank grows more rapidly.</p></li><li><p><strong>Greek computation</strong>: The functions <code>compute_delta</code> and <code>compute_gamma</code> (from <code>greeks.py</code>) can be called on the final solution QTT to obtain full&#8209;grid Greeks with negligible extra cost.  The paper reports that the grid&#8209;wide Delta and Gamma errors are comparable to the price errors (Table IV for the 1D case; similar behaviour holds for d &gt; 1).</p></li></ul><h3>Code Organisation</h3><p>The experimental workflow is encapsulated in <code>experiments.py</code>.  Its two main entry points are:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2f12c73b-a3b9-4289-8521-f1f8a5f04a66&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def run_basket_experiment(d, c, option_type, ref_price=None):
    """
    Run the time-stepping solver for a basket put.
    Returns the QTT solution, the interpolated spot price, and the relative error.
    If ref_price is given, error is computed; otherwise it returns None.
    """
    ...

def run_maxmin_experiment(d, c, option_type, ref_price=None):
    """
    Similar for max-min put.
    """
    ...</code></pre></div><p>These functions use the utility functions <code>default_market_params</code>, <code>adaptive_domain</code>, and the reference&#8209;price calculators from <code>utils.py</code>.  The paper&#8217;s tables can be reproduced by looping over <code>d</code>, <code>c</code>, and <code>option_type</code> and recording the outputs.</p><p>No code execution is performed as part of this tutorial; the reader is encouraged to run the solver and compare the results with the published Tables V&#8211;VIII.</p><p>Use the Button below to download the source code.</p>
      <p>
          <a href="/__u/onepagecode.substack.com/p/quant-finance-solving-the-high-dimensional">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Implementing Stock Market Prediction: NeuralProphet + DNN (PyTorch Guide & Critique)]]></title><description><![CDATA[A step-by-step PyTorch implementation of the NP-DNN hybrid model (arXiv:2601.05202v3), resolving dataset contradictions with Optuna.]]></description><link>https://onepagecode.substack.com/p/implementing-stock-market-prediction</link><guid isPermaLink="false">https://onepagecode.substack.com/p/implementing-stock-market-prediction</guid><dc:creator><![CDATA[Onepagecode]]></dc:creator><pubDate>Tue, 14 Jul 2026 05:04:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!IoJA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc25caa3-bd9b-43fd-8fd0-795c25c02806_1512x654.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Use the URL at the end of this to download source code</h2><p>The paper proposes a hybrid model (NP-DNN) combining Neural Prophet and Deep Neural Network for stock market price prediction. The methodology includes data preprocessing (Z-score normalization, missing value imputation via linear interpolation), feature extraction using MLP, prediction using DNN, and hyperparameter tuning with Optuna. The model is evaluated on the Crunchbase dataset using classification metrics (accuracy, precision, recall, F1) and compared with DSS, LightGBM, RF, and LLM, claiming 93.21% accuracy. The paper contains several inconsistencies and ambiguities, including a mismatch between the dataset (startup investments) and the task (stock price prediction), and inappropriate use of classification metrics.</p><h2>Implementation Assumptions</h2><ul><li><p>The task is binary classification (price up/down) to align with classification metrics.</p></li><li><p>MLP has 2 hidden layers (64, 32) with ReLU; DNN has 2 hidden layers (64, 32) with ReLU and softmax output.</p></li><li><p>Neural Prophet is implemented using the Prophet library (not neuralprophet) for simplicity.</p></li><li><p>Optuna search space: learning rate [1e-5, 1e-1], number of layers [1,3], hidden units [32,256], dropout [0.0,0.5], batch size [16,128].</p></li><li><p>The Crunchbase dataset is not used; synthetic data is generated for demonstration.</p></li><li><p>The abstract accuracy 99.21% is a typo; we use 93.21% as the claimed accuracy.</p></li></ul><h2>Introduction and Paper Overview</h2><p>Stock market price prediction is one of the most challenging problems in financial forecasting. Prices move in response to countless factors&#8212;economic news, investor sentiment, geopolitical events, and random noise&#8212;making it nearly impossible to predict with high accuracy using simple methods. Traditional statistical models like ARIMA or linear regression often fail to capture the complex, nonlinear patterns that drive market movements.</p><p>The paper we are implementing, "Stock Market Price Prediction using Neural Prophet with Deep Neural Network" (arXiv:2601.05202v3), proposes a hybrid model called <strong>NP-DNN</strong> that combines two powerful tools:</p><ul><li><p><strong>Neural Prophet (NP)</strong>: A time series decomposition model that breaks a price history into interpretable components: the overall trend, repeating seasonal patterns, holiday effects, and random error. This helps the model understand long-term structure.</p></li><li><p><strong>Deep Neural Network (DNN)</strong>: A multi-layer neural network that learns complex patterns from data. The paper uses a Multi-Layer Perceptron (MLP) for feature extraction before feeding into the DNN for final prediction.</p></li></ul><p>The overall pipeline has three main stages:</p><ol><li><p><strong>Data Preprocessing</strong>: Fill missing values using linear interpolation and standardize features with Z-score normalization.</p></li><li><p><strong>Feature Extraction</strong>: Use an MLP to transform the preprocessed data into a more informative representation.</p></li><li><p><strong>Prediction</strong>: Feed the extracted features into a DNN with a softmax output layer to produce a classification (e.g., price up or down).</p></li></ol><p>The paper claims that this NP-DNN model achieves 93.21% accuracy on the Crunchbase dataset, outperforming methods like LightGBM, Random Forest, and even a Large Language Model (LLM).</p><h3>What to Expect from This Tutorial</h3>
      <p>
          <a href="/__u/onepagecode.substack.com/p/implementing-stock-market-prediction">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Forecasting Stock Prices with Diffusion Models and U-GNNs]]></title><description><![CDATA[A paper-to-Python guide to graph diffusion, uncertainty-aware S&P 500 forecasting, and wireless resource allocation]]></description><link>https://onepagecode.substack.com/p/forecasting-stock-prices-with-diffusion</link><guid isPermaLink="false">https://onepagecode.substack.com/p/forecasting-stock-prices-with-diffusion</guid><dc:creator><![CDATA[Onepagecode]]></dc:creator><pubDate>Mon, 13 Jul 2026 12:07:44 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_W6n!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3078fdad-8e80-40a9-be17-640eb18fa447_1254x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What This Article Builds</h2><p>The paper proposes a unified framework for conditional generative modeling of graph signals using denoising diffusion models. The key innovation is the U-Graph Neural Network (U-GNN), a U-Net-like architecture for graph-structured data that uses learned node selection matrices for pooling/unpooling and strided graph convolutions to avoid explicit graph coarsening. The framework is demonstrated on two tasks: stock price forecasting (S&amp;P 500) and wireless resource allocation (power control).</p><h2>Implementation Assumptions</h2><ul><li><p>PyTorch 2.x + PyTorch Geometric (PyG) + standard scientific Python (numpy, scipy, matplotlib).</p></li><li><p>Graph shift operator S stored as a sparse PyTorch tensor (coalesced COO) or edge<em>index+edge</em>weight.</p></li><li><p>All graph signals are batched: (batch, N, F) with optional node dimension per graph (single graph assumed for each task).</p></li><li><p>The U-GNN is built from modular components: StridedGraphConv, NodeSelection, FusionBlock, UGNN core.</p></li><li><p>Training uses AMP (torch.cuda.amp), AdamW, cosine LR with linear warm-up, gradient clipping (max_norm=1.0).</p></li><li><p>Validation and checkpointing use task-specific composite criteria.</p></li><li><p>Data pipeline for S&amp;P 500 uses yfinance; graph built from fundamentals (rank correlation + sector bonus).</p></li><li><p>Data pipeline for WRA synthesizes networks using the paper's channel model and expert primal-dual algorithm.</p></li><li><p>DDIM sampling uses 100 steps and &#951;=0.2 on a sub-grid of the 500-step linear schedule.</p></li></ul><h2>Download Source code. </h2><p>You can download the source using the URL at the end of this article. </p><p>Note: There was no source code attached to the paper, this is my implementation, I tried to code it, and it is part of the larger program, so use the code to understand and implement in your code, try running it would run but there might be some problem. So Use code as reference, not in actual public working. </p><h2>1. Introduction to the Paper and Problem</h2><h3>1.1 What Are Stochastic Graph Signals?</h3><p>Many real-world systems produce signals that live on a graph &#8212; a set of nodes connected by edges that encode relationships. When those signals are inherently random, we call them <em>stochastic graph signals</em>. Two canonical examples, both studied in the paper, are:</p><ul><li><p><strong>Stock returns on a company graph</strong>: Each node is a publicly traded company (e.g., the S&amp;P 500 constituents). The graph edges capture fundamental similarities: same industry sector, correlated financial profiles, or shared supply chains. The signal at each node is the daily log-return of that stock. The signal is stochastic because future returns are unpredictable in detail, yet they are <em>conditionally</em> dependent on the past and on the graph structure (e.g., a shock to one sector propagates to related companies).</p></li><li><p><strong>Power allocations on an interference graph</strong>: Each node is a wireless transmitter&#8211;receiver pair. The graph edges represent interference strength: a directed edge from transmitter <em>j</em> to receiver <em>i</em> carries the channel gain between them. The signal at each node is the transmit power level. The optimal allocation is stochastic because it depends on random fading and the expert algorithm&#8217;s internal randomness, but it is conditioned on the large-scale channel gains (the graph) and the node&#8217;s own channel quality.</p></li></ul><p>In both cases, the graph topology and node-level side information (history features, channel measurements) provide a conditioning context. The goal is to <em>sample</em> new graph signals from the unknown conditional distribution &#8212; not just to predict a single value, but to generate entire trajectories or allocations that respect the graph structure and the observed context.</p><h3>1.2 Why Generative Modeling (Not Just Regression)?</h3><p>A regressive approach &#8212; training a neural network to minimize mean squared error between a point prediction and the true signal &#8212; learns the conditional mean. This is adequate when the conditional distribution is unimodal and symmetric, but it fails when:</p><ul><li><p>The distribution is <strong>multi-modal</strong>: there may be several equally plausible future price paths or power allocations.</p></li><li><p><strong>Uncertainty quantification</strong> is required: a point forecast gives no sense of risk or confidence intervals.</p></li><li><p>The downstream task needs <strong>samples</strong>, not averages: for portfolio optimization, one needs many possible return scenarios; for wireless resource allocation, one needs to test different power settings under fading.</p></li></ul><p>Generative models address these limitations by learning the full conditional law. The paper adopts a <strong>denoising diffusion probabilistic model (DDPM)</strong>, which has become the state of the art for high-quality sample generation in images, audio, and now graph signals. Diffusion models are particularly attractive because they are stable to train, do not suffer from mode collapse (unlike GANs), and can be conditioned naturally on arbitrary inputs.</p><h3>1.3 The Two Applications</h3><p>The paper demonstrates the framework on two distinct tasks, each with its own graph structure and conditioning modality:</p><p><strong>Stock Price Forecasting (S&amp;P 500)</strong></p><ul><li><p><strong>Data</strong>: Daily prices for 468 stocks over ~10 years. For each stock, 12 market features (open, high, low, close, log-return, moving averages, volume, RSI, MACD) are computed. A static undirected graph is built from rank correlations of fundamental profiles, with a same-sector bonus.</p></li><li><p><strong>Conditioning</strong>: The last 20 days of features (history) and the graph topology.</p></li><li><p><strong>Target</strong>: The next 5 days of log-returns (a graph signal of dimension 5 per node).</p></li><li><p><strong>Challenge</strong>: The conditional distribution of future returns is heavy-tailed, heteroscedastic, and exhibits volatility clustering. A generative model can produce multiple plausible trajectories, enabling risk-aware decision-making.</p></li></ul><p><strong>Wireless Resource Allocation (Power Control)</strong></p><ul><li><p><strong>Data</strong>: Synthetic networks of 400 transmitter&#8211;receiver pairs in square areas of varying density. Channel gains follow a dual-slope path loss model with log-normal shadowing and Rayleigh fading. An expert primal&#8211;dual algorithm computes optimal power allocations that maximize sum ergodic rate subject to per-user minimum rate constraints.</p></li><li><p><strong>Conditioning</strong>: The directed interference graph (top-10 strongest interferers per receiver, log-normalized) and two node-state features (direct link gain and aggregate interference under full power).</p></li><li><p><strong>Target</strong>: The centered power allocation (p/Pmax &#8722; 1) for each user.</p></li><li><p><strong>Challenge</strong>: The expert algorithm is computationally expensive; a learned generative model can approximate its output distribution at a fraction of the cost, and can be deployed in real time.</p></li></ul><h3>1.4 Overview of the Proposed Solution</h3><p>The paper proposes a unified framework that combines a <strong>denoising diffusion probabilistic model</strong> with a novel <strong>U&#8209;Graph Neural Network (U&#8209;GNN)</strong> denoiser. The key ideas are:</p><ol><li><p><strong>Forward diffusion</strong>: Gradually add Gaussian noise to the clean graph signal over 500 steps, following a linear noise schedule (&#946;&#8321; = 1e&#8722;4, &#946;&#8325;&#8320;&#8320; = 2e&#8722;2). The noising process is fixed and does not depend on the graph.</p></li><li><p><strong>Reverse denoising</strong>: Learn a neural network &#949;_&#952; that predicts the noise added at each step, conditioned on the graph shift operator S and node features u. The network is trained with the simple mean-squared error between true and predicted noise (equation (6) in the paper).</p></li><li><p><strong>U&#8209;GNN denoiser</strong>: A U&#8209;Net-like architecture for graph signals. It processes the signal at multiple resolutions by:</p></li></ol><ul><li><p><strong>Strided graph convolutions</strong> (Algorithm 1): Instead of coarsening the graph, the signal is lifted to the full graph, filtered with iterated unit shifts of S, and then reduced to a subset of active nodes. The stride &#947; increases the receptive field without densifying the graph.</p></li><li><p><strong>Learned node selection</strong> (Section IV&#8209;B): At each down&#8209;sampling level, a scoring MLP selects the top&#8209;k nodes to retain. During training, Gumbel noise encourages exploration, and a straight&#8209;through estimator allows gradients to flow through the discrete selection.</p></li><li><p><strong>Fusion layers</strong> (equation (22)): The down&#8209;sampled signal is merged with global embeddings (node states and diffusion step) via cross&#8209;attention (for stock) or addition (for WRA).</p></li><li><p><strong>Skip connections</strong> (equation (23)): The decoder receives concatenated features from the encoder at the same resolution.</p></li></ul><ol start="4"><li><p><strong>Accelerated sampling</strong>: Instead of the full 500&#8209;step reverse process, the paper uses DDIM (Denoising Diffusion Implicit Models) with 100 steps and &#951; = 0.2, which produces high&#8209;quality samples much faster.</p></li><li><p><strong>Task&#8209;specific adaptations</strong>: For stock forecasting, the U&#8209;GNN includes a lightweight temporal mixer (interleaved dilated 1&#8209;D convolutions and two&#8209;head self&#8209;attention) to process the history features. For WRA, the conditioning is static and uses simple addition.</p></li></ol><h3>1.5 What This Tutorial Covers</h3><p>This tutorial walks through a complete, modular reproduction of the paper&#8217;s stock forecasting and wireless resource allocation experiments. We will:</p><ul><li><p>Implement the forward diffusion process and DDIM sampling from scratch.</p></li><li><p>Build the U&#8209;GNN architecture piece by piece: strided graph convolutions, learned node selection, fusion layers, and the full encoder&#8211;decoder.</p></li><li><p>Construct the data pipelines for both tasks, including synthetic data generation for WRA and a simplified stock data pipeline (using synthetic data to avoid network calls).</p></li><li><p>Train the model with the exact hyperparameters from the paper (Table I, Appendix E).</p></li><li><p>Evaluate using the paper&#8217;s metrics: CRPS, MIS90, RMSE, direction accuracy, stylized&#8209;fact gaps for stock; ergodic rate, feasibility, and gap to expert for WRA.</p></li><li><p>Reproduce the key figures: forecasting trajectories, CDF comparisons, distribution diagnostics, and WRA bar charts.</p></li></ul><p>All code is written in Python using PyTorch and PyTorch Geometric, and is organized into modular files for clarity. The implementation follows the paper&#8217;s descriptions and appendices faithfully, with documented decisions where the paper is ambiguous.</p><p>Let&#8217;s begin by understanding the diffusion process and how it is adapted to graph signals.</p><h2>2. Denoising Diffusion Probabilistic Models for Graph Signals</h2><h3>2.1 Forward Diffusion Process</h3><p>The core idea of a denoising diffusion model is to define a <em>forward</em> process that gradually corrupts a clean signal into pure noise, and then learn a <em>reverse</em> process that undoes this corruption step by step. For graph signals, the forward process operates on the node features directly, independent of the graph structure at this stage &#8212; the graph topology enters only as conditioning information for the reverse denoiser.</p><p>Let x_0 \in \mathbb{R}^{N \times F} be a clean graph signal (e.g., the future returns of N stocks over F time steps). The paper defines a Markov chain of K = 500 steps that adds Gaussian noise with a <em>linear variance schedule</em>:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\beta_1 = 10^{-4}, \\quad \\beta_{500} = 2 \\times 10^{-2}, \\quad \\beta_k = \\beta_1 + (k-1)\\frac{\\beta_{500} - \\beta_1}{K-1}&quot;,&quot;id&quot;:&quot;DFBE8E173E&quot;}" data-component-name="LatexBlockToDOM"></div><p>Each forward transition is (equation (2)):</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;q(x_k \\mid x_{k-1}) = \\mathcal{N}\\bigl(x_k; \\sqrt{1-\\beta_k}\\,x_{k-1},\\; \\beta_k I\\bigr)&quot;,&quot;id&quot;:&quot;3978E6C16E&quot;}" data-component-name="LatexBlockToDOM"></div><p>Because the Gaussian is closed under composition, we can directly sample x_k from x_0 without iterating through all intermediate steps. Define \alpha_k = 1 - \beta_k and the cumulative product \bar{\alpha}_k = \prod_{i=1}^k \alpha_i. Then the marginal distribution is (equation (3)):</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;q(x_k \\mid x_0) = \\mathcal{N}\\bigl(x_k; \\sqrt{\\bar{\\alpha}_k}\\,x_0,\\; (1-\\bar{\\alpha}_k)I\\bigr)&quot;,&quot;id&quot;:&quot;1DC06A0796&quot;}" data-component-name="LatexBlockToDOM"></div><p>which leads to the reparameterization used in training (equation (4)):</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;x_k = \\sqrt{\\bar{\\alpha}_k}\\,x_0 + \\sqrt{1-\\bar{\\alpha}_k}\\,\\varepsilon, \\quad \\varepsilon \\sim \\mathcal{N}(0,I)&quot;,&quot;id&quot;:&quot;363DC23A66&quot;}" data-component-name="LatexBlockToDOM"></div><h4>Implementation in <code>diffusion.py</code></h4><p>The <code>DiffusionProcess</code> class precomputes all \bar{\alpha}_k values from the linear \beta schedule. The constructor stores the betas, alphas, and derived quantities as tensors for efficient indexing:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a2173af3-7953-4864-b4f4-741809d413ce&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class DiffusionProcess:
    def __init__(self, betas: torch.Tensor) -&gt; None:
        self.K = betas.shape[0]
        self.betas = betas
        self.alphas = 1.0 - betas
        # alpha_bar[t] = prod_{i=1}^{t} alpha_i, alpha_bar[0] = 1
        self.alpha_bar = torch.cat([
            torch.ones(1, dtype=betas.dtype, device=betas.device),
            torch.cumprod(self.alphas, dim=0)
        ], dim=0)
        self.sqrt_alpha_bar = self.alpha_bar.sqrt()
        self.sqrt_one_minus_alpha_bar = torch.sqrt(1.0 - self.alpha_bar)</code></pre></div><p>The <code>forward_diffuse</code> method implements equation (4) directly. It takes a clean signal x_0 and a step index k (1-indexed), samples noise, and returns the noised signal:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5352da5b-9050-4539-9dcf-138541c7c2c8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def forward_diffuse(self, x0: torch.Tensor, k: torch.Tensor) -&gt; tuple[torch.Tensor, torch.Tensor]:
    k_idx = k.long()
    sqrt_alpha_k = self.sqrt_alpha_bar[k_idx]
    sqrt_one_minus_alpha_k = self.sqrt_one_minus_alpha_bar[k_idx]
    epsilon = torch.randn_like(x0)
    x_k = sqrt_alpha_k * x0 + sqrt_one_minus_alpha_k * epsilon
    return x_k, epsilon</code></pre></div><p>During training, we sample k \sim \text{Uniform}(1, K) for each batch element, compute x_k using this method, and then ask the denoiser to predict the noise \varepsilon.</p><h3>2.2 Reverse Denoising and the Noise-Prediction Objective</h3><p>The reverse process is a Markov chain that starts from pure noise x_K \sim \mathcal{N}(0,I) and iteratively removes noise to recover x_0. Each reverse step is parameterized as a Gaussian with learned mean (equation (5)):</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;p_\\theta(x_{k-1} \\mid x_k; \\xi) = \\mathcal{N}\\bigl(x_{k-1}; \\mu_\\theta(x_k, k; \\xi), \\sigma_k^2 I\\bigr)&quot;,&quot;id&quot;:&quot;C30BC1E27D&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here \xi = (S, u) denotes the conditioning information: the graph shift operator S and node features u. The mean \mu_\theta is derived from a noise-prediction network \varepsilon_\theta (the U-GNN) via equation (8):</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\mu_\\theta = \\frac{1}{\\sqrt{\\alpha_k}}\\left(x_k - \\frac{\\beta_k}{\\sqrt{1-\\bar{\\alpha}_k}}\\,\\varepsilon_\\theta(x_k, k; \\xi)\\right)&quot;,&quot;id&quot;:&quot;F03B82DCC2&quot;}" data-component-name="LatexBlockToDOM"></div><p>Rather than directly predicting \mu_\theta, the paper adopts the simplified training objective from Ho et al. (2020) (equation (6)):</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\mathcal{L}(\\theta) = \\mathbb{E}_{k, x_0, \\varepsilon}\\left[\\|\\varepsilon - \\varepsilon_\\theta(x_k, k; \\xi)\\|^2\\right]&quot;,&quot;id&quot;:&quot;E6EC6A9FEB&quot;}" data-component-name="LatexBlockToDOM"></div><p>This loss is simple: at each training step, we sample a random k, noise \varepsilon, and a clean signal x_0, compute x_k via the forward process, and ask the U-GNN to predict the noise that was added. The gradient flows through the U-GNN only; the forward process is fixed.</p><h3>2.3 DDIM Sampling for Acceleration</h3><p>While the reverse process could be simulated for all K = 500 steps, the paper uses the Denoising Diffusion Implicit Model (DDIM) to accelerate sampling to just 100 steps while maintaining sample quality. DDIM defines a non-Markovian reverse process that shares the same forward marginals but allows larger jumps.</p><p>At each sampling step, we first estimate the clean signal from the current noisy x_k and the predicted noise (equation (7)):</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\hat{x}_0 = \\frac{x_k - \\sqrt{1-\\bar{\\alpha}_k}\\,\\varepsilon_\\theta}{\\sqrt{\\bar{\\alpha}_k}}&quot;,&quot;id&quot;:&quot;C4BC98E27D&quot;}" data-component-name="LatexBlockToDOM"></div><p>Then we update to the previous step k-1 (or, for accelerated sampling, to the previous sub-grid step \tau_{i-1}) using equation (9):</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;x_{k-1} = \\sqrt{\\bar{\\alpha}_{k-1}}\\,\\hat{x}_0 + \\sqrt{1-\\bar{\\alpha}_{k-1} - \\sigma_k^2}\\,\\varepsilon_\\theta + \\sigma_k w, \\quad w \\sim \\mathcal{N}(0,I)&quot;,&quot;id&quot;:&quot;A4DF9953E3&quot;}" data-component-name="LatexBlockToDOM"></div><p>The noise scale \sigma_k is controlled by a parameter \eta \in [0,1] (equation (29)):</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\sigma_k(\\eta) = \\eta\\,\\sqrt{\\frac{1-\\bar{\\alpha}_{k-1}}{1-\\bar{\\alpha}_k}}\\,\\sqrt{1 - \\frac{\\bar{\\alpha}_k}{\\bar{\\alpha}_{k-1}}}&quot;,&quot;id&quot;:&quot;842B00FAD2&quot;}" data-component-name="LatexBlockToDOM"></div><p>When \eta = 0, the process is deterministic (pure DDIM). When \eta = 1, it recovers the original DDPM. The paper uses \eta = 0.2 for a balance of stochasticity and speed.</p><p>For accelerated sampling, we select a sub-grid of step indices (e.g., 100 evenly spaced indices from 1 to 500) and apply the update using the corresponding \bar{\alpha} values. The <code>sample_sub_grid</code> method generates this list:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;67bd9c9d-d01b-4804-aae3-b7bb1e9f28f0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def sample_sub_grid(self, n_steps: int) -&gt; List[int]:
    step_indices = torch.linspace(
        self.K, 1, n_steps, dtype=torch.float32
    ).round().long().unique().tolist()
    step_indices = sorted(step_indices, reverse=True)
    return step_indices</code></pre></div><p>The <code>ddim_sample</code> method implements the full sampling loop. It takes the denoiser function, initial noise x_K, conditioning dictionary, sub-grid, and \eta:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7a43bfa7-93e0-4649-b313-167aed3e6838&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def ddim_sample(self, denoiser, x_K, cond=None, sub_grid=None, eta=0.2):
    if sub_grid is None:
        sub_grid = list(range(self.K, 0, -1))
    sub_grid = sorted(sub_grid, reverse=True)
    x = x_K.clone()
    for i, k in enumerate(sub_grid):
        k_prev = sub_grid[i + 1] if i + 1 &lt; len(sub_grid) else 0
        sqrt_alpha_bar_k = self.sqrt_alpha_bar[k]
        sqrt_one_minus_alpha_bar_k = self.sqrt_one_minus_alpha_bar[k]
        inv_sqrt_alpha_bar_k = self.inv_sqrt_alpha_bar[k]
        epsilon_theta = denoiser(x, k, cond)
        x_hat_0 = (x - sqrt_one_minus_alpha_bar_k * epsilon_theta) * inv_sqrt_alpha_bar_k
        if k_prev &gt; 0:
            sqrt_alpha_bar_prev = self.sqrt_alpha_bar[k_prev]
            one_minus_alpha_bar_prev = self.one_minus_alpha_bar[k_prev]
            one_minus_alpha_bar_k = self.one_minus_alpha_bar[k]
            ratio = (one_minus_alpha_bar_prev / one_minus_alpha_bar_k).clamp(min=0.0)
            sigma = eta * torch.sqrt(ratio) * torch.sqrt(1.0 - self.alpha_bar[k] / self.alpha_bar[k_prev])
            noise = torch.randn_like(x) if eta &gt; 0 else 0.0
            sqrt_one_minus_alpha_bar_prev_minus_sigma2 = torch.sqrt(
                (one_minus_alpha_bar_prev - sigma**2).clamp(min=0.0)
            )
            x = sqrt_alpha_bar_prev * x_hat_0 + sqrt_one_minus_alpha_bar_prev_minus_sigma2 * epsilon_theta + sigma * noise
        else:
            x = x_hat_0
    return x</code></pre></div><p>Note the adaptation for sub-grid steps: the formula uses \bar{\alpha}_{k_\text{prev}} and \bar{\alpha}_k directly, which is valid for any two steps where k_\text{prev} &lt; k. The clamping ensures numerical stability.</p><h3>2.4 Conditioning on Graph Topology and Node Features</h3><p>The denoiser \varepsilon_\theta is the U-GNN, which takes three inputs at each step:</p><ul><li><p><strong>Noisy signal</strong> x_k: the current node features to be denoised.</p></li><li><p><strong>Diffusion step</strong> k: embedded as a sinusoidal positional encoding and broadcast to all nodes.</p></li><li><p><strong>Conditioning</strong> \xi = (S, u): the graph shift operator S (a sparse matrix encoding edge weights) and node-level side information u (e.g., historical market features for stock forecasting, or direct link gain for wireless).</p></li></ul><p>The conditioning is injected through fusion layers inside the U-GNN, which we will examine in detail in the next section. For now, it suffices to understand that the denoiser is a function:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\varepsilon_\\theta: (x_k, k, S, u) \\mapsto \\hat{\\varepsilon}&quot;,&quot;id&quot;:&quot;DB9AF97024&quot;}" data-component-name="LatexBlockToDOM"></div><p>and that the entire diffusion framework &#8212; forward process, noise-prediction loss, and DDIM sampling &#8212; is agnostic to the internal architecture of \varepsilon_\theta. This modularity is what allows the paper to apply the same diffusion framework to two very different tasks by swapping only the denoiser architecture.</p><h3>Summary</h3><ul><li><p>The forward diffusion process adds Gaussian noise to graph signals according to a linear schedule (equations (2)-(4)).</p></li><li><p>The reverse process is learned by predicting the noise added at each step (equation (6)).</p></li><li><p>DDIM sampling accelerates generation by using a sub-grid of steps and a controlled stochasticity parameter \eta (equations (7), (9), (29)-(30)).</p></li><li><p>The <code>DiffusionProcess</code> class in <code>diffusion.py</code> implements all of these operations with precomputed \bar{\alpha} values and efficient tensor operations.</p></li><li><p>Conditioning on graph topology and node features is handled by the U-GNN denoiser, which we will explore next.</p></li></ul><h2>3. The U-GNN Architecture</h2><h3>3.1 Why a Graph U-Net?</h3><p>Images have a natural multi-scale structure: you can downsample (pool) by averaging neighboring pixels, and then upsample (unpool) by interpolation. A U-Net exploits this to build a computational graph with a wide receptive field while keeping the number of parameters manageable. Graph signals lack such a regular grid &#8212; there is no canonical notion of &#8220;neighbourhood averaging&#8221; because the graph is irregular. Nevertheless, the same intuition applies: to model long-range dependencies across the graph, we need to process the signal at multiple resolutions, aggregating information from far-away nodes into compact representations and then expanding back to the original resolution.</p><p>The paper&#8217;s solution, the <strong>U-Graph Neural Network (U-GNN)</strong>, is a U-Net-like architecture designed for graph-structured data. It uses three key innovations:</p><ol><li><p><strong>Strided graph convolutions</strong> that increase the receptive field without making the graph shift operator denser.</p></li><li><p><strong>Learned node selection</strong> (with Gumbel-Top-K and a straight-through estimator) to decide which nodes to keep at each coarser resolution, avoiding a fixed pooling strategy.</p></li><li><p><strong>Fusion layers</strong> that inject conditioning information (node states and diffusion step) at every resolution level.</p></li></ol><p>Let us walk through each component.</p><h3>3.2 The Lift&#8211;Filter&#8211;Reduce Pattern (Equation (13))</h3><p>A standard GNN layer updates node features by exchanging messages along edges. If the graph shift operator is a sparse matrix S \in \mathbb{R}^{N \times N} (the adjacency or a normalized Laplacian), a single layer computes:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;Z_\\ell = \\sigma_\\ell\\Bigl( \\sum_{k=0}^{K} S^k X_{\\ell-1} \\Theta_{\\ell,k} \\Bigr)&quot;,&quot;id&quot;:&quot;5D7882FB94&quot;}" data-component-name="LatexBlockToDOM"></div><p>where X_{\ell-1} are the node features, \Theta_{\ell,k} are learnable weight matrices for each &#8220;tap&#8221; k, and \sigma_\ell is an activation. Each tap k corresponds to a k-step random walk on the graph, giving a localised receptive field.</p><p>Now suppose we have identified a subset of &#8220;active&#8221; nodes (those that survive the pooling) through a binary selection matrix D \in \{0,1\}^{N' \times N} (each row is a one-hot indicator of an active node). The paper uses a <strong>lift&#8211;filter&#8211;reduce</strong> pattern (equation (13)):</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;Z_\\ell = D\\; \\sigma_\\ell\\Bigl( \\sum_{k=0}^{K} S^k \\, D^\\mathsf{T} Z_{\\ell-1} \\, \\Theta_{\\ell,k} \\Bigr)&quot;,&quot;id&quot;:&quot;BBA97DCE6A&quot;}" data-component-name="LatexBlockToDOM"></div><ol><li><p><strong>Lift</strong>: D^\mathsf{T} Z_{\ell-1} maps the N'-node signal back to the full N-node graph by placing the values at the active nodes and leaving zeros elsewhere.</p></li><li><p><strong>Filter</strong>: apply the graph convolution (sum over k taps) on the full graph.</p></li><li><p><strong>Reduce</strong>: multiply by D to extract the active-node signals.</p></li></ol><p>This pattern is the foundation of every strided convolution in U-GNN.</p><h3>3.3 Strided Graph Convolutions (Algorithm 1, Equation (14))</h3><p>A plain GNN with K taps has a receptive field of exactly K hops. To reach farther nodes, one could increase K, but that makes S^K denser and more expensive (both in memory and computation). The paper&#8217;s solution is to replace S^k with (S^\gamma)^k, where \gamma is a <strong>stride</strong> parameter. Equation (14) formalises this:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;Z_\\ell = D\\; \\sigma_\\ell\\Bigl( \\sum_{k=0}^{K} (S^\\gamma)^k \\, D^\\mathsf{T} Z_{\\ell-1} \\, \\Theta_{\\ell,k} \\Bigr)&quot;,&quot;id&quot;:&quot;05B05E874B&quot;}" data-component-name="LatexBlockToDOM"></div><p>Raising S to the power \gamma directly is expensive, so the paper proposes Algorithm 1: instead of forming S^\gamma, we perform \gamma K unit shifts and tap every \gamma-th one. This gives the same result as using S^\gamma, but each shift is a sparse matrix&#8211;vector product, and the intermediate \gamma-1 shifts that are not tapped still propagate the signal forward, building the wider receptive field.</p><p>Let us look at the implementation in <code>gnn_layers.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2bad0778-7db1-4bb1-9c40-08833a1f7176&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class StridedGraphConv(nn.Module):
    """
    Strided graph convolution layer with pooling (lift-filter-reduce).
    Implements Algorithm 1 of the paper.
    """
    def __init__(self, in_dim, out_dim, K, gamma, activation=nn.ReLU()):
        super().__init__()
        self.weight = nn.Parameter(torch.randn(K + 1, in_dim, out_dim) * 0.1)
        self.bias = nn.Parameter(torch.zeros(out_dim))
        self.K = K
        self.gamma = gamma
        self.activation = activation

    def forward(self, Z_prev, S, D):
        # Lift
        x = torch.mm(D.T, Z_prev)          # (N, in_dim)
        # Tap 0
        A = torch.mm(x, self.weight[0])   # (N, out_dim)

        # Iterated unit shifts
        for j in range(1, self.gamma * self.K + 1):
            x = torch.mm(S, x)            # one shift
            if j % self.gamma == 0:
                tap_idx = j // self.gamma
                A = A + torch.mm(x, self.weight[tap_idx])

        A = self.activation(A + self.bias)
        out = torch.mm(D, A)              # reduce
        return out</code></pre></div><p>Key points:</p><ul><li><p>The total number of taps is still K+1 (including the identity tap), independent of \gamma.</p></li><li><p>The stride \gamma is set per depth: \gamma_b = \min(\lfloor\sqrt{\bar{\rho}_b}\rfloor, \gamma_{\max}=2), where \bar{\rho}_b is the cumulative pooling factor at depth b.</p></li><li><p>Both dense and sparse S are supported (the code checks <code>S.is_sparse</code>; PyTorch Geometric&#8217;s sparse ops can be used in practice).</p></li></ul><h3>3.4 Learned Node Selection (Equations (24)&#8211;(25), (31)&#8211;(34))</h3><p>How does the U-GNN decide which nodes to keep when going from one resolution level to the next? Instead of a fixed rule (e.g., max pooling over neighbourhoods), the paper <strong>learns</strong> which nodes are most informative. The process has three steps:</p><ol><li><p><strong>Score each node</strong>: a node-wise MLP \Psi_b takes the concatenation of the encoded feature Z_b and the global embeddings D_b[U_0;K_0] and outputs a scalar score (equation (24)).</p></li><li><p>**Select top-k**: keep the N_{b+1} highest-scoring nodes (equation (25)).</p></li><li><p><strong>Down-sample</strong>: the selection matrix C_{b+1} is the identity rows indexed by the chosen nodes.</p></li></ol><p>But this discrete selection is non-differentiable. During training, the paper uses two tricks:</p><ul><li><p><strong>Gumbel-Top-K</strong> (equation (31)): Gumbel noise is added to the scores before the \text{TopK} operation, encouraging exploration.</p></li><li><p><strong>Straight-Through Estimator (STE)</strong> (equation (32)): the forward pass uses a hard one-hot mask (the selected nodes), while the backward pass flows through a soft sigmoid surrogate \sigma(v_b / \tau). The code does this with a simple <code>detach</code> trick:</p></li></ul><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;42a4af36-48a5-4654-a191-9fd408109fb1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">mask_hard = torch.zeros(N_b, device=device)
mask_hard.scatter_(0, indices, 1.0)            # hard one-hot
mask_soft = torch.sigmoid(scores / tau)         # soft surrogate
mask = mask_hard + mask_soft - mask_soft.detach()  # forward hard, backward soft</code></pre></div><p>Additionally, the temperature \tau and exploration noise \varepsilon are annealed linearly over training (equation (34)): a warm-up phase of T_w = 0.02T epochs followed by an anneal phase of T_a = 0.73T epochs. The full <code>NodeSelectionHead</code> class in <code>node_selection.py</code> implements this schedule.</p><h3>3.5 Full U-GNN: Encoder, Bottleneck, Decoder (Equations (18)&#8211;(23))</h3><p>With the strided convolution and the learned selection head, we can now assemble the complete U-GNN. The architecture has B = 4 resolution levels. Let N_1 = N be the original number of nodes. At each depth b, the number of active nodes is N_b = \lfloor N / \rho^{b-1}\rfloor with \rho = 2.</p><p><strong>Nested selection matrices</strong> (equation (18)): the global selection matrix D_b that maps from the full graph to the active nodes at depth b is the product of all previous selection matrices:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;D_b = C_b \\cdots C_1&quot;,&quot;id&quot;:&quot;5C60792EF5&quot;}" data-component-name="LatexBlockToDOM"></div><p><strong>Encoder path</strong> (equation (19)): for b = 1, \dots, B-1:</p><ol><li><p>Down-sample: V_b = C_b Z_{b-1}.</p></li><li><p>Fuse: P_b = \Pi_b(V_b, D_b[U_0;K_0]) (see below).</p></li><li><p>Apply strided GNN module: Z_b = \Phi_{E_b}(P_b, S; \Theta_{E_b}, D_b, \gamma_b).</p></li><li><p>Select: C_{b+1} = \text{Selector}(Z_b, D_b[U_0;K_0], N_{b+1}).</p></li></ol><p><strong>Bottleneck</strong> (equation (21)): at b = B, the same pattern without selection.</p><p><strong>Decoder path</strong> (equation (20)): for b = B-1, \dots, 1:</p><ol><li><p>Up-sample: C_{b+1}^\mathsf{T} Y_{b+1} (zero-pad to the previous resolution).</p></li><li><p>Skip-connect: concatenate the up-sampled decoder output with the encoder signal Z_b.</p></li><li><p>Project: V_b = \Pi_{\text{skip}}^b([C_{b+1}^\mathsf{T} Y_{b+1}; Z_b]) (equation (23)).</p></li><li><p>Fuse: P_b = \Pi_b(V_b, D_b[U_0;K_0]).</p></li><li><p>Apply strided GNN: Y_b = \Phi_{D_b}(P_b, S; \Theta_{D_b}, D_b, \gamma_b).</p></li></ol><p>Finally, <strong>read-out</strong>: \varepsilon_\theta = \Pi_{\text{out}}(Y_1), a node-wise MLP that maps back to the original feature dimension.</p><p>The fusion layer \Pi_b (equation (22)) merges the down-sampled signal with the global embeddings (node states U_0 and step embedding K_0). For the stock task, the paper uses cross-attention (query = signal, key/value = global embeddings); for the WRA task, it uses simple addition after a projection. The <code>FusionLayer</code> class in <code>ugnn.py</code> implements both versions.</p><h3>3.6 Task-Specific Adaptations</h3><p>For the stock forecasting task, the input x_k has a temporal dimension (5 future days). The U-GNN includes a <strong>TemporalMixer</strong> that applies two dilated 1D convolutions followed by two-head self-attention along the time axis, producing a single feature vector per node (by averaging over time). This is a reasonable default for the paper&#8217;s underspecified temporal encoder.</p><p>For the wireless resource allocation task, the input is static (one allocation per node), so no temporal mixer is needed &#8212; the <code>TemporalMixer</code> is replaced by an identity module.</p><p>The <code>UGNN</code> class in <code>ugnn.py</code> orchestrates all these components. It pre-computes the target node counts, initialises the nested selection matrix D, and runs the encoder&#8211;bottleneck&#8211;decoder loop. The selection heads are called at the end of each encoder block, and the resulting C and D matrices are accumulated.</p><h3>3.7 Summary</h3><p>The U-GNN is a carefully designed graph U-Net that avoids explicit graph coarsening. Instead, it learns which nodes to retain at each level, uses strided graph convolutions to control the receptive field without making the shift operator denser, and fuses conditioning information at every resolution. This architecture is the central contribution of the paper and is what enables the diffusion model to generate high-quality graph signals conditioned on graph structure and node features.</p><h2>4. Data Pipelines</h2><p>A generative model is only as good as the data it learns from. The paper presents two distinct data pipelines, each transforming raw observations into the graph-structured inputs required by the diffusion model. This section walks through both pipelines in detail, showing how we go from stock tickers or wireless network parameters to the tensors that feed the U-GNN.</p><h3>4.1 S&amp;P 500 Stock Forecasting Pipeline</h3><p>The stock forecasting task aims to predict the next five days of log-returns for 468 S&amp;P 500 constituents, conditioned on the last 20 days of market features and a static graph of company similarities. The pipeline has five stages: data acquisition, feature engineering, graph construction, window creation, and normalisation &amp; splitting.</p><h4>4.1.1 Data Acquisition</h4><p>In a real deployment, you would download daily OHLCV data from Yahoo Finance using the <code>yfinance</code> library. The paper specifies the date range 2016&#8209;09&#8209;06 to 2026&#8209;02&#8209;13. However, our implementation (in <code>data_stock.py</code>) provides a synthetic data generator that mimics the shape and basic statistics of 468 stocks, so that the code can be run without an internet connection. The function <code>download_sp500_data</code> returns a DataFrame with columns <code>ticker</code>, <code>date</code>, <code>open</code>, <code>high</code>, <code>low</code>, <code>close</code>, <code>volume</code>, and <code>sector</code>. The synthetic prices follow an AR(1) log&#8209;price process with a common market factor and idiosyncratic noise, which is sufficient for testing the pipeline logic.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a56233be-3a15-49e0-97ea-33a62436c3fe&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def download_sp500_data(
    start_date: str = "2016-09-06",
    end_date: str = "2026-02-13",
    num_stocks: int = 468,
    num_days: int = 2500,
) -&gt; pd.DataFrame:
    # ... generates synthetic OHLCV data with random sectors
    # Returns a DataFrame with columns: ticker, date, open, high, low, close, volume, sector</code></pre></div><h4>4.1.2 Feature Engineering</h4><p>For each stock on each day, the paper computes 12 market features. The function <code>compute_market_features</code> adds these features as a list column to the DataFrame. The features are:</p><ul><li><p>open, high, low, close, log_return</p></li><li><p>moving averages (MA5, MA20)</p></li><li><p>volume</p></li><li><p>RSI(14)</p></li><li><p>MACD, MACD<em>signal, MACD</em>histogram</p></li></ul><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d6fea620-ea9d-49d7-8094-92a61490fc56&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def compute_market_features(df: pd.DataFrame) -&gt; pd.DataFrame:
    # ... computes each feature per ticker
    # Adds columns 'features' (list of 12 floats) and 'log_return'</code></pre></div><p>The log_return column becomes the target signal; the full feature vector serves as the conditioning input <code>u</code>.</p><h4>4.1.3 Graph Construction from Fundamentals</h4><p>The paper builds a static undirected graph from company fundamentals. The function <code>build_fundamentals_graph</code> implements a simplified version:</p><ol><li><p>Compute the Spearman rank correlation between the mean feature vectors of each pair of stocks.</p></li><li><p>Add a same&#8209;sector bonus of 0.1.</p></li><li><p>Threshold at the median correlation to keep only the strongest edges.</p></li><li><p>Symmetrize and add self&#8209;loops.</p></li><li><p>Apply spectral normalisation: S = D^{-1/2} A D^{-1/2}.</p></li></ol><p>The result is a dense <code>(N, N)</code> tensor <code>S</code> and a list of sector labels.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;0b0f71b2-c9bd-465c-b014-60c163e91ee6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def build_fundamentals_graph(stocks_df: pd.DataFrame) -&gt; Tuple[torch.Tensor, List[int]]:
    # ... rank correlation, sector bonus, threshold, spectral normalisation
    return torch.tensor(S, dtype=torch.float32), sectors.tolist()</code></pre></div><h4>4.1.4 Sliding Windows</h4><p>With the feature matrix of shape <code>(N, T, U)</code> (12 features) and the target matrix of shape <code>(N, T)</code> (log&#8209;returns), the function <code>create_sliding_windows</code> extracts every possible window of length <code>Th + Tp</code>. For each starting time <code>t</code> from <code>Th</code> to <code>T - Tp</code>, it produces:</p><ul><li><p><strong>history</strong> <code>u</code>: shape <code>(N, Th, U)</code> &#8212; the last 20 days of features.</p></li><li><p><strong>target</strong> <code>x0</code>: shape <code>(N, Tp)</code> &#8212; the next 5 days of log&#8209;returns.</p></li></ul><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;aef5b5ac-4d7b-42db-a6ea-e8d8cbbd0aa2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def create_sliding_windows(
    data: pd.DataFrame, Th: int = 20, Tp: int = 5
) -&gt; List[Dict]:
    # ... builds feature tensor (N, T, U) and target tensor (N, T)
    # ... iterates t from Th to T-Tp, creates history and target tensors</code></pre></div><h4>4.1.5 RevIN Normalisation</h4><p>The paper applies a modified Reversible Instance Normalisation (RevIN) to the target window. The <code>RevIN</code> class in <code>data_stock.py</code> implements this with the paper&#8209;specific parameters: a blend weight of 0.7 and a scale correction of 1.107. During the forward pass, it normalises the target using per&#8209;instance mean and standard deviation (computed along the temporal dimension), then blends these statistics with running averages for consistency. The normalised target is then scaled and shifted by learnable affine parameters.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a4650311-6fab-4327-b9c6-ee9a716d27af&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class RevIN(nn.Module):
    def __init__(self, num_features: int = 1, blend: float = 0.7, scale_correction: float = 1.107):
        # ... initialise affine weight with 1/scale_correction

    def forward(self, x, mode="norm", mean=None, std=None):
        if mode == "norm":
            # compute instance mean and std, blend with running stats
            # normalise, apply affine transform
            return x_norm, (mean_inst, std_inst)
        elif mode == "denorm":
            # reverse affine, then denormalise using stored statistics
            return x_denorm, None</code></pre></div><h4>4.1.6 Interleaved Chronological Split</h4><p>To avoid temporal leakage, the paper uses an interleaved chronological split. The function <code>interleaved_chronological_split</code> divides the total number of windows into 10 equal chunks, then within each chunk splits chronologically 80% train, 10% validation, 10% test. The windows are pooled across chunks, ensuring that test windows always come after training windows in time.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;92d4d1cd-1733-472b-a982-b390b7ac2976&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def interleaved_chronological_split(
    windows: List[Dict],
    split_ratios: Tuple[float, float, float] = (0.8, 0.1, 0.1),
    num_chunks: int = 10,
) -&gt; Dict[str, List[Dict]]:
    # ... divide windows into chunks, split each chunk, concatenate
    return {"train": train, "val": val, "test": test}</code></pre></div><h4>4.1.7 Dataset and DataLoader</h4><p>The <code>StockDataset</code> class wraps the windows, the static graph <code>S</code>, and the <code>RevIN</code> module. Each call to <code>__getitem__</code> returns a dictionary with keys:</p><ul><li><p><code>x0</code>: target log&#8209;returns (normalised), shape <code>(N, Tp)</code></p></li><li><p><code>u</code>: history features, shape <code>(N, Th, U)</code></p></li><li><p><code>S</code>: graph shift operator, shape <code>(N, N)</code></p></li><li><p><code>revin_params</code>: tuple <code>(mean, std)</code> used for later denormalisation during evaluation</p></li></ul><p>The factory function <code>get_stock_dataloaders</code> creates the three DataLoaders with the appropriate batch size (64) and shuffling.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;12964ba1-efc9-4b3e-956b-5cb54a3df916&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class StockDataset(Dataset):
    def __init__(self, windows: List[Dict], S: torch.Tensor, revin: RevIN):
        self.windows = windows
        self.S = S
        self.revin = revin

    def __getitem__(self, idx):
        w = self.windows[idx]
        # apply RevIN normalisation to target
        x0_norm, (mean, std) = self.revin(w["target"], mode="norm")
        return {
            "x0": x0_norm,
            "u": w["history"],
            "S": self.S,
            "revin_params": (mean, std)
        }</code></pre></div><h3>4.2 Wireless Resource Allocation Pipeline</h3><p>The wireless task is to learn a policy that maps a network configuration (positions, channel gains) to a set of transmit powers (one per transmitter&#8209;receiver pair) that maximise the sum ergodic rate under a minimum rate constraint. The pipeline synthetically generates 128 networks (32 for each of four densities) and uses an expert primal&#8209;dual algorithm to produce optimal allocations.</p><h4>4.2.1 Network Generation</h4><p>The <code>WirelessNetworkGenerator</code> class creates a network with <code>N=400</code> transmitter&#8209;receiver pairs uniformly distributed in a square area. The side length <code>R</code> depends on the density (Table III):</p><p>| Density Index | Side Length R (m) | Density (pairs/km&#178;) | |---------------|-------------------|---------------------| | 0             | 7800              | 6.6                 | | 1             | 7000              | 8.2                 | | 2             | 6300              | 10.1                | | 3             | 5800              | 11.9                |</p><p>A minimum transmitter&#8209;transmitter spacing of 50&#8239;m is enforced by resampling pairs that are too close.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f26afeac-8abe-411b-a914-78a231e479ba&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class WirelessNetworkGenerator:
    def __init__(self, config: WRAConfig):
        self.config = config

    def generate_network(self, density_idx: int) -&gt; Dict:
        # ... uniform positions in [0,R]x[0,R], enforce min spacing
        return {
            'pos_tx': pos_tx, 'pos_rx': pos_rx,
            'area_side': R, 'density': density
        }</code></pre></div><h4>4.2.2 Channel Model</h4><p>The <code>ChannelModel</code> implements the paper&#8217;s dual&#8209;slope path loss model with log&#8209;normal shadowing (&#963; = 7&#8239;dB) and Rayleigh fading. The path loss has a breakpoint at 200&#8239;m: exponent 2.0 before, 3.5 after. The <code>compute_path_loss</code> method returns an <code>(N, N)</code> matrix of large&#8209;scale channel gains (linear scale). The <code>sample_fading</code> method generates unit&#8209;mean complex Gaussian fading samples for a given number of links and time slots.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;27e68d58-0714-484d-86d0-85d36dee2cd7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class ChannelModel:
    def compute_path_loss(self, pos_tx, pos_rx):
        # ... distances, dual-slope, shadowing
        return gain  # (N, N) tensor

    def sample_fading(self, n_links, n_slots):
        # ... complex Gaussian, squared magnitude
        return fading  # (n_links, n_slots)</code></pre></div><h4>4.2.3 GSO Construction</h4><p>The function <code>build_wra_gso</code> constructs a directed weighted graph shift operator. For each receiver <code>i</code>, it selects the top&#8209;10 strongest interferers (excluding self), sets the edge weight to the channel gain, then applies a log&#8209;normalisation: <code>S = log(1 + gain) / max(log(1 + gain))</code>. Self&#8209;loops are set to 1.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a92bc307-6988-46de-bb01-6d4dff5802b8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def build_wra_gso(network, channel_gains, config):
    # ... for each receiver, top-K gains, log-normalisation
    return S  # (N, N) tensor</code></pre></div><h4>4.2.4 Node State Features</h4><p>The <code>compute_node_states</code> function computes the 2&#8209;dimensional node state vector <code>u</code> for each receiver:</p><ul><li><p><code>direct_link_gain</code>: the channel gain from its own transmitter.</p></li><li><p><code>aggregate_interference</code>: sum of channel gains from all other transmitters multiplied by <code>Pmax</code>.</p></li></ul><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e047f6e0-5eeb-494b-b10b-dbab75893434&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def compute_node_states(network, channel_gains, config):
    direct = channel_gains.diag()
    interference = (channel_gains.sum(dim=1) - direct) * config.Pmax
    u = torch.stack([direct, interference], dim=1)  # (N, 2)
    return u</code></pre></div><h4>4.2.5 Expert Primal&#8209;Dual Algorithm</h4><p>The paper uses a primal&#8209;dual algorithm to solve the constrained optimisation problem (equation&#8239;(28)). The <code>expert_primal_dual</code> function implements a simplified version: it performs projected gradient ascent on the powers <code>p</code> and gradient descent on the dual variables <code>&#955;</code> to maximise the Lagrangian. From the converged trajectory, it collects 200 intermediate allocations (though for simplicity our implementation returns the final allocation). The allocation is centred as <code>x0 = p / Pmax - 1</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d892c9cc-f668-4b38-9f54-9c6c9dbb747e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def expert_primal_dual(network, channel_gains, config):
    # ... initialise p, lambda
    for it in range(num_iters):
        # compute rates, gradient of Lagrangian, update p, project to [0,Pmax]
        # update lambda with clipping
        # collect samples
    x0 = p / Pmax - 1.0
    return x0.float()</code></pre></div><h4>4.2.6 Dataset and DataLoader</h4><p>The <code>WRADataset</code> wraps the lists of networks, allocations, GSOs, and node states. Each item returns a dictionary with:</p><ul><li><p><code>x0</code>: centred allocation, shape <code>(N,)</code></p></li><li><p><code>S</code>: graph shift operator, shape <code>(N, N)</code></p></li><li><p><code>u</code>: node state features, shape <code>(N, 2)</code></p></li></ul><p>The factory function <code>get_wra_dataloaders</code> generates all 128 networks (32 per density), splits them 20/4/8 per density for train/val/test, and returns three DataLoaders with batch size 800 (as specified in Table&#8239;I).</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e9245844-b43c-4d7a-b30a-b22b829b368e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class WRADataset(Dataset):
    def __init__(self, networks, allocations, gsos, node_states):
        # ... store lists

    def __getitem__(self, idx):
        return {
            'x0': self.allocations[idx],
            'S': self.gsos[idx],
            'u': self.node_states[idx]
        }</code></pre></div><h3>4.3 Shape Invariants Summary</h3><p>Both pipelines produce tensors with consistent shapes that the U&#8209;GNN expects:</p><p>| Task  | <code>x0</code> shape       | <code>u</code> shape          | <code>S</code> shape   | Notes                                    | |-------|------------------|--------------------|-------------|------------------------------------------| | Stock | <code>(N, Tp=5)</code>      | <code>(N, Th=20, U=12)</code> | <code>(N, N)</code>    | <code>x0</code> normalised; <code>S</code> static, symmetric   | | WRA   | <code>(N,)</code>           | <code>(N, 2)</code>           | <code>(N, N)</code>    | <code>x0</code> centred; <code>S</code> directed, weighted     |</p><p>These shapes are guaranteed by the dataset classes and are maintained throughout the diffusion forward and reverse processes. The conditioning variable <code>u</code> is always available, and the graph <code>S</code> is used by every strided convolution in the U&#8209;GNN.</p><p>With the data pipelines in place, we can now load batches of <code>(x0, u, S)</code> and feed them into the diffusion model. The next section walks through the core code that implements the forward diffusion, the U&#8209;GNN architecture, and the training loop.</p><h2>5. Code Walkthrough: Core Components</h2><p>With the theoretical foundation laid and the data pipelines ready, we now open the code files themselves. This section walks through each of the four core modules&#8212;<code>diffusion.py</code>, <code>gnn_layers.py</code>, <code>node_selection.py</code>, and <code>ugnn.py</code>&#8212;function by function. Our goal is to show exactly how the paper&#8217;s mathematical descriptions and algorithms become executable Python, highlighting every design decision and shape constraint.</p><h3>5.1 Diffusion Process (<code>diffusion.py</code>)</h3><p><strong>File purpose:</strong> Implements the forward diffusion process (equations (2)&#8211;(4)) and DDIM accelerated sampling (equations (9) and (29)&#8211;(30)).</p><h4><code>DiffusionProcess.__init__</code></h4><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f53a411b-e70c-4a38-8de6-5694d2327df3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class DiffusionProcess:
    def __init__(self, betas: torch.Tensor) -&gt; None:
        if betas.dim() != 1:
            raise ValueError("betas must be a 1D tensor")
        if not torch.all(betas &gt; 0) or not torch.all(betas &lt; 1):
            raise ValueError("betas must be in (0,1)")

        self.K = betas.shape[0]
        self.betas = betas
        self.alphas = 1.0 - betas

        # alpha_bar[t] = prod_{i=1}^{t} alpha_i, alpha_bar[0] = 1
        # shape: (K+1,)
        self.alpha_bar = torch.cat([
            torch.ones(1, dtype=betas.dtype, device=betas.device),
            torch.cumprod(self.alphas, dim=0)
        ], dim=0)

        # Precompute sqrt and sqrt(1 - alpha_bar) for efficiency
        self.sqrt_alpha_bar = self.alpha_bar.sqrt()
        self.sqrt_one_minus_alpha_bar = torch.sqrt(1.0 - self.alpha_bar)

        # For DDIM, precompute 1 - alpha_bar and inverse of sqrt_alpha_bar
        self.one_minus_alpha_bar = 1.0 - self.alpha_bar
        self.inv_sqrt_alpha_bar = 1.0 / self.sqrt_alpha_bar</code></pre></div><p><strong>Explanation:</strong> The constructor takes a 1D tensor <code>betas</code> of length <code>K</code> (K = 500 in the paper). It computes <code>alpha = 1 - beta</code> and then builds <code>alpha_bar</code>, where <code>alpha_bar[t] = prod_{i=1}^{t} alpha_i</code>. A leading 1.0 is prepended so that <code>alpha_bar[0] = 1</code> (no noise at step 0). All subsequent quantities are precomputed and stored as tensors: <code>sqrt_alpha_bar</code>, <code>sqrt_one_minus_alpha_bar</code>, <code>one_minus_alpha_bar</code>, and <code>inv_sqrt_alpha_bar</code>. This avoids recomputing square roots during training and sampling.</p><p><strong>Shape note:</strong> <code>alpha_bar</code> has shape <code>(K+1,)</code> so that indexing with step <code>k</code> (1&#8209;indexed) directly uses position <code>k</code>.</p><h4><code>forward_diffuse</code> (equation (4))</h4><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;90d24f0a-51b9-402d-b6a9-962e8a8a1ee2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def forward_diffuse(self, x0: torch.Tensor, k: torch.Tensor) -&gt; tuple[torch.Tensor, torch.Tensor]:
    k_idx = k.long()
    sqrt_alpha_k = self.sqrt_alpha_bar[k_idx]
    sqrt_one_minus_alpha_k = self.sqrt_one_minus_alpha_bar[k_idx]

    epsilon = torch.randn_like(x0)
    x_k = sqrt_alpha_k * x0 + sqrt_one_minus_alpha_k * epsilon

    return x_k, epsilon</code></pre></div><p><strong>Explanation:</strong> This is the one-line reparameterization of equation (4): <code>x_k = sqrt(alpha_bar_k) * x0 + sqrt(1 - alpha_bar_k) * epsilon</code>. We look up the precomputed <code>sqrt_alpha_bar</code> at index <code>k</code> (which is a tensor, possibly batched), sample a fresh Gaussian noise <code>epsilon</code>, and return both <code>x_k</code> and <code>epsilon</code> for the noise&#8209;prediction loss. The shape of <code>x_k</code> matches <code>x0</code> exactly.</p><h4><code>sample_sub_grid</code></h4><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;fcbf5266-db15-4575-a061-ca61814d1d07&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def sample_sub_grid(self, n_steps: int) -&gt; List[int]:
    if n_steps &gt; self.K:
        raise ValueError("n_steps cannot exceed K")
    if n_steps &lt;= 0:
        raise ValueError("n_steps must be positive")

    step_indices = torch.linspace(
        self.K, 1, n_steps, dtype=torch.float32
    ).round().long().unique().tolist()

    if len(step_indices) &gt; n_steps:
        step_indices = step_indices[:n_steps]
    elif len(step_indices) &lt; n_steps:
        step_indices = [max(1, int(round(self.K - i * (self.K - 1) / (n_steps - 1))))
                        for i in range(n_steps)]
        step_indices = sorted(set(step_indices), reverse=True)

    step_indices = sorted(step_indices, reverse=True)
    return step_indices</code></pre></div><p><strong>Explanation:</strong> Accelerated DDIM uses a sub&#8209;grid of the full 500&#8209;step schedule. This method generates <code>n_steps</code> indices (e.g., 100) that are roughly evenly spaced from K down to 1. It first tries <code>torch.linspace</code> rounded to integers; if rounding produces duplicates (less than <code>n_steps</code> unique values), it falls back to a simpler arithmetic spacing. The returned list is sorted descending so that the sampling loop starts from the highest noise level.</p><h4><code>ddim_sample</code> (equations (9), (29)&#8211;(30))</h4><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;063a6fa3-2f92-4a50-9027-b0cd7467ba7c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def ddim_sample(self, denoiser: Callable, x_K: torch.Tensor,
                cond: Optional[Dict[str, Any]] = None,
                sub_grid: Optional[List[int]] = None,
                eta: float = 0.2) -&gt; torch.Tensor:
    if sub_grid is None:
        sub_grid = list(range(self.K, 0, -1))

    sub_grid = sorted(sub_grid, reverse=True)
    x = x_K.clone()

    for i, k in enumerate(sub_grid):
        k_prev = sub_grid[i + 1] if i + 1 &lt; len(sub_grid) else 0

        sqrt_alpha_bar_k = self.sqrt_alpha_bar[k]
        sqrt_one_minus_alpha_bar_k = self.sqrt_one_minus_alpha_bar[k]
        inv_sqrt_alpha_bar_k = self.inv_sqrt_alpha_bar[k]

        epsilon_theta = denoiser(x, k, cond)

        # Equation (7): estimate clean signal
        x_hat_0 = (x - sqrt_one_minus_alpha_bar_k * epsilon_theta) * inv_sqrt_alpha_bar_k

        if k_prev &gt; 0:
            sqrt_alpha_bar_prev = self.sqrt_alpha_bar[k_prev]
            one_minus_alpha_bar_prev = self.one_minus_alpha_bar[k_prev]
            one_minus_alpha_bar_k = self.one_minus_alpha_bar[k]

            # Equation (29): sigma_k(eta) adapted for sub-grid
            ratio = (one_minus_alpha_bar_prev / one_minus_alpha_bar_k).clamp(min=0.0)
            sigma = eta * torch.sqrt(ratio) * torch.sqrt(1.0 - self.alpha_bar[k] / self.alpha_bar[k_prev])

            noise = torch.randn_like(x) if eta &gt; 0 else 0.0

            sqrt_one_minus_alpha_bar_prev_minus_sigma2 = torch.sqrt(
                (one_minus_alpha_bar_prev - sigma**2).clamp(min=0.0)
            )

            # Equation (30): DDIM update
            x = sqrt_alpha_bar_prev * x_hat_0 \
                + sqrt_one_minus_alpha_bar_prev_minus_sigma2 * epsilon_theta \
                + sigma * noise
        else:
            x = x_hat_0

    return x</code></pre></div><p><strong>Explanation:</strong> This implements the core sampling loop. Starting from pure noise <code>x_K</code>, it iterates over the (possibly accelerated) sub&#8209;grid in descending order. For each step <code>k</code>:</p><ol><li><p>Predict the noise <code>epsilon_theta</code> using the denoiser (our U&#8209;GNN).</p></li><li><p>Compute <code>x_hat_0</code> via equation (7): <code>(x_k - sqrt(1 - alpha_bar_k) * epsilon_theta) / sqrt(alpha_bar_k)</code>.</p></li><li><p>If we are not at the final step (<code>k_prev &gt; 0</code>), compute the DDIM noise scale <code>sigma</code> using equation (29). The key adaptation for the sub&#8209;grid is that we use <code>alpha_bar[k_prev]</code> and <code>alpha_bar[k]</code> instead of consecutive indices. We clamp intermediate values to avoid numerical issues near zero.</p></li><li><p>Apply the update of equation (30): <code>x_{k_prev} = sqrt(alpha_bar_prev) * x_hat_0 + sqrt(1 - alpha_bar_prev - sigma^2) * epsilon_theta + sigma * noise</code>.</p></li><li><p>At the final step (<code>k_prev = 0</code>), assign <code>x = x_hat_0</code> directly.</p></li></ol><p>The <code>denoiser</code> is a callable that takes <code>(x, k, cond)</code> and returns the predicted noise with the same shape as <code>x</code>. The <code>cond</code> dictionary carries the graph shift operator <code>S</code> and the node states <code>u</code> (and, for the stock task, the historical window).</p><p><strong>Design decision:</strong> We treat the sub&#8209;grid update as a drop&#8209;in replacement for consecutive-step DDIM. The formula for <code>sigma</code> must be adapted because <code>alpha_bar[k_prev]</code> and <code>alpha_bar[k]</code> are not adjacent in the original 500&#8209;step schedule. However, the derivation in Appendix A of the paper is valid for any two indices <code>k_prev &lt; k</code>, and our implementation follows exactly that generalization.</p><div><hr></div><h3>5.2 Strided Graph Convolution (<code>gnn_layers.py</code>)</h3><p><strong>File purpose:</strong> Implements Algorithm 1 and the stacked GNN module used inside the U&#8209;GNN.</p><h4><code>StridedGraphConv.forward</code> (Algorithm 1)</h4><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a89272b9-54f0-48cd-9d2f-ff085b2fdbd7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class StridedGraphConv(nn.Module):
    def __init__(self, in_dim, out_dim, K, gamma, activation=nn.ReLU()):
        super().__init__()
        self.weight = nn.Parameter(torch.randn(K + 1, in_dim, out_dim) * 0.1)
        self.bias = nn.Parameter(torch.zeros(out_dim))
        # ... store K, gamma, activation

    def forward(self, Z_prev: torch.Tensor, S: torch.Tensor,
                D: torch.Tensor) -&gt; torch.Tensor:
        N = S.shape[0]

        # ---- Lift: D^T Z_prev ----
        x = torch.mm(D.T, Z_prev)

        # ---- Tap 0 ----
        A = torch.mm(x, self.weight[0])

        # ---- Iterated unit shifts ----
        for j in range(1, self.gamma * self.K + 1):
            if S.is_sparse:
                x = torch.sparse.mm(S, x)
            else:
                x = torch.mm(S, x)
            if j % self.gamma == 0:
                tap_idx = j // self.gamma
                A = A + torch.mm(x, self.weight[tap_idx])

        # ---- Nonlinearity + bias ----
        A = self.activation(A + self.bias)

        # ---- Reduce: D A ----
        out = torch.mm(D, A)
        return out</code></pre></div><p><strong>Explanation:</strong> This is the direct implementation of Algorithm 1 (Appendix B). The layer processes a signal <code>Z_prev</code> defined on <code>N'</code> active nodes. The steps are:</p><ol><li><p><strong>Lift</strong> (<code>D^T Z_prev</code>): Zero&#8209;pad the signal from <code>N'</code> nodes to all <code>N</code> nodes using the transpose of the binary selection matrix <code>D</code>.</p></li><li><p><strong>Tap 0</strong>: Apply the weight matrix for <code>k=0</code> (the identity shift).</p></li><li><p><strong>Iterated unit shifts</strong>: Run a loop up to <code>gamma * K</code>. In each iteration, apply one sparse shift (<code>S @ x</code>). Every <code>gamma</code>&#8209;th iteration, treat it as a tap: accumulate the contribution <code>x @ W[tap_idx]</code>.</p></li><li><p><strong>Nonlinearity + bias</strong>: Apply activation (ReLU) and bias.</p></li><li><p><strong>Reduce</strong> (<code>D @ A</code>): Project back from <code>N</code> nodes to the <code>N'</code> active nodes using the same selection matrix.</p></li></ol><p>The total number of taps is <code>K+1</code> (tap 0 plus taps at indices <code>gamma, 2*gamma, ..., gamma*K</code>). The stride <code>gamma</code> determines how many unit shifts occur between taps, effectively implementing <code>(S^gamma)^k</code> without ever forming the dense <code>S^gamma</code> matrix. This is the key efficiency trick.</p><p><strong>Shape invariants:</strong> <code>Z_prev: (N', in_dim)</code>, <code>S: (N, N)</code> (sparse or dense), <code>D: (N', N)</code> (binary, each row sums to 1). The output is <code>(N', out_dim)</code>.</p><h4><code>GNNModule</code></h4><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;266147c8-dd44-4715-929b-fe0df1b437c5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class GNNModule(nn.Module):
    def __init__(self, in_dim, out_dim, K, gamma, L=2, dropout=0.1):
        super().__init__()
        self.layers = nn.ModuleList()
        for i in range(L):
            dim_in = in_dim if i == 0 else out_dim
            self.layers.append(
                StridedGraphConv(dim_in, out_dim, K, gamma, activation=nn.ReLU())
            )
        self.norm = nn.LayerNorm(out_dim)
        self.dropout = nn.Dropout(dropout)

    def forward(self, X, S, D):
        h = X
        for layer in self.layers:
            h = layer(h, S, D)
            h = self.norm(h)
            h = self.dropout(h)
        return h</code></pre></div><p><strong>Explanation:</strong> A simple stack of <code>L</code> strided graph convolution layers (default 2, as per Table I), each followed by LayerNorm and dropout. All layers share the same <code>gamma</code> and <code>D</code> for a given depth level. The first layer expands from <code>in_dim</code> to <code>out_dim</code>; subsequent layers stay at <code>out_dim</code>. The module outputs a signal on the same active node set.</p><div><hr></div><h3>5.3 Learned Node Selection (<code>node_selection.py</code>)</h3><p><strong>File purpose:</strong> Implements the scoring, Gumbel&#8209;Top&#8209;K, and straight&#8209;through estimator described in Section IV&#8209;B and Appendix C.</p><h4><code>NodeSelectionHead</code></h4><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;874d2400-27d1-4ee5-a59b-eb7f17619294&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class NodeSelectionHead(nn.Module):
    def __init__(self, in_dim, global_dim, hidden_dim=64):
        super().__init__()
        self.score_mlp = nn.Sequential(
            nn.Linear(in_dim + global_dim, hidden_dim),
            nn.ReLU(),
            nn.Linear(hidden_dim, 1),
        )

    def forward(self, Z, global_emb, n_keep, training=True, epoch=0,
                annealing_params=None):
        N_b = Z.shape[0]
        device = Z.device

        # Equation (24): score = MLP(concat(Z, global_emb))
        concat = torch.cat([Z, global_emb], dim=-1)
        scores = self.score_mlp(concat).squeeze(-1)  # (N_b,)

        if training:
            tau, eps = self._anneal_tau_eps(epoch, annealing_params)

            if eps &gt; 0:
                # Equation (31): add Gumbel noise
                U = torch.rand_like(scores)
                gumbel = -torch.log(-torch.log(U + 1e-8))
                scores = scores + eps * gumbel

            # Equation (25): Top-K selection
            _, indices = torch.topk(scores, k=n_keep, dim=-1)
            indices = indices.sort()[0]

            C = torch.zeros(n_keep, N_b, device=device, dtype=Z.dtype)
            C[torch.arange(n_keep, device=device), indices] = 1.0

            # Equation (32): straight-through estimator
            mask_hard = torch.zeros(N_b, device=device, dtype=Z.dtype)
            mask_hard.scatter_(0, indices, 1.0)
            mask_soft = torch.sigmoid(scores / tau)
            mask = mask_hard + mask_soft - mask_soft.detach()

            # Equation (33): down-sample with mask
            masked_Z = Z * mask.unsqueeze(-1)
            V = masked_Z[indices]
        else:
            # Inference: hard selection only
            _, indices = torch.topk(scores, k=n_keep, dim=-1)
            indices = indices.sort()[0]
            C = torch.zeros(n_keep, N_b, device=device, dtype=Z.dtype)
            C[torch.arange(n_keep, device=device), indices] = 1.0
            V = Z[indices]

        return C, V</code></pre></div><p><strong>Explanation:</strong> The node selection head does exactly what Section IV&#8209;B describes. The scoring MLP <code>Psi_b</code> (equation (24)) takes the concatenation of the encoded feature <code>Z</code> and the global embeddings projected to the current active node set, and outputs a scalar per node. During training:</p><ul><li><p><strong>Gumbel noise</strong> is added to the scores (equation (31)). <code>eps</code> is annealed from 1 down to 0 over training, so exploration decays.</p></li><li><p><strong>Top&#8209;K</strong> picks the <code>n_keep</code> highest&#8209;scoring nodes.</p></li><li><p><strong>Straight&#8209;through estimator</strong> (equation (32)): the forward pass uses a hard one&#8209;hot mask (<code>mask_hard</code>), while gradients flow through a sigmoid surrogate (<code>mask_soft</code>). The trick <code>mask = hard + soft - soft.detach()</code> ensures the forward pass is hard and the backward pass sees <code>soft</code>.</p></li><li><p><strong>Down&#8209;sampled signal</strong> <code>V</code> is computed as <code>C @ (mask * Z)</code> (equation (33)).</p></li></ul><p>At inference, Gumbel noise and the STE are removed: only hard Top&#8209;K is used.</p><h4>Annealing Schedule (<code>_anneal_tau_eps</code>)</h4><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;14d09516-6d31-4bbb-95f4-69bd649e4859&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@staticmethod
def _anneal_tau_eps(epoch: int, annealing_params: dict) -&gt; (float, float):
    if annealing_params is None:
        Tw, Ta, tau0, tau_min, eps0, eps_min = 0.0, 1.0, 1.0, 0.5, 1.0, 0.0
    else:
        Tw = annealing_params.get('Tw', 0.02)
        Ta = annealing_params.get('Ta', 0.73)
        tau0 = annealing_params.get('tau0', 1.0)
        tau_min = annealing_params.get('tau_min', 0.5)
        eps0 = annealing_params.get('eps0', 1.0)
        eps_min = annealing_params.get('eps_min', 0.0)

    # Equation (34): r(t) = min(max((t - Tw) / Ta, 0), 1)
    r = (epoch - Tw) / Ta
    r = max(0.0, min(1.0, r))

    tau = tau0 + (tau_min - tau0) * r
    eps = eps0 + (eps_min - eps0) * r

    return tau, eps</code></pre></div><p><strong>Explanation:</strong> Implements equation (34) exactly. After a warm&#8209;up phase of <code>Tw</code> epochs (default 0.02 &#215; total = 100 epochs for 5000 total), the ratio <code>r</code> increases linearly from 0 to 1 over <code>Ta</code> epochs (default 0.73 &#215; total = 3650 epochs). Both <code>tau</code> and <code>eps</code> are linearly interpolated from their initial to their final values. After <code>Tw + Ta</code> epochs, <code>r = 1</code> and the parameters stay at their minima.</p><div><hr></div><h3>5.4 U&#8209;GNN Architecture (<code>ugnn.py</code>)</h3><p><strong>File purpose:</strong> Assembles all components into the full U&#8209;GNN as described in Section IV and equations (18)&#8211;(23).</p><h4><code>UGNN.__init__</code></h4><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;1ec73265-29e1-464f-afb2-dc1cbc973400&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class UGNN(nn.Module):
    def __init__(self, task: str, config: Config):
        super().__init__()
        self.task = task
        self.B = config.B          # 4 resolution levels
        self.rho = config.rho      # 2 (pooling factor)
        self.F0 = config.F0        # 64 (feature dim)
        self.L = config.L          # 2 (GNN layers per block)
        self.K = config.K          # 2 (hops)
        self.gamma_max = config.gamma_max  # 2

        # Core layers
        self.read_in = ReadInLayer(in_dim=config.F_out, out_dim=self.F0)
        self.read_out = ReadOutLayer(in_dim=self.F0, out_dim=self.F_out)
        self.node_state_encoder = NodeStateEncoder(in_dim=config.U, out_dim=self.F0)
        self.step_embedding = StepEmbedding(out_dim=self.F0)

        if task == 'stock':
            self.temporal_mixer = TemporalMixer(...)
        else:
            self.temporal_mixer = nn.Identity()

        # Encoder blocks
        self.encoder_blocks = nn.ModuleList()
        self.fusion_layers_enc = nn.ModuleList()
        self.selection_heads = nn.ModuleList()
        for b in range(1, self.B):
            gamma = self._compute_gamma(b)
            gnn = GNNModule(self.F0, self.F0, self.K, gamma, self.L, self.dropout)
            fusion = FusionLayer(self.F0, 2*self.F0, self.embed_dim, self.task)
            sel = NodeSelectionHead(self.F0, 2*self.F0)
            self.encoder_blocks.append(gnn)
            self.fusion_layers_enc.append(fusion)
            self.selection_heads.append(sel)

        # Bottleneck
        self.bottleneck_gamma = self._compute_gamma(self.B)
        self.bottleneck_gnn = GNNModule(self.F0, self.F0, self.K,
                                         self.bottleneck_gamma, self.L, self.dropout)
        self.bottleneck_fusion = FusionLayer(self.F0, 2*self.F0, self.embed_dim, self.task)

        # Decoder blocks
        self.decoder_blocks = nn.ModuleList()
        self.fusion_layers_dec = nn.ModuleList()
        self.skip_projections = nn.ModuleList()
        for b in range(self.B - 1, 0, -1):
            gamma = self._compute_gamma(b)
            gnn = GNNModule(self.F0, self.F0, self.K, gamma, self.L, self.dropout)
            fusion = FusionLayer(self.F0, 2*self.F0, self.embed_dim, self.task)
            skip_proj = SkipProjection(in_dim=2*self.F0, out_dim=self.F0)
            self.decoder_blocks.append(gnn)
            self.fusion_layers_dec.append(fusion)
            self.skip_projections.append(skip_proj)</code></pre></div><p><strong>Explanation:</strong> The constructor builds the full U&#8209;GNN. It creates <code>B-1</code> encoder blocks (b = 1 to B-1), one bottleneck block (b = B), and <code>B-1</code> decoder blocks (b = B-1 down to 1). Each block has its own <code>GNNModule</code> with a depth&#8209;dependent <code>gamma</code> computed by <code>_compute_gamma</code>, a <code>FusionLayer</code>, and (for encoder blocks) a <code>NodeSelectionHead</code>. The decoder blocks also get a <code>SkipProjection</code> layer to reduce the concatenated skip connection to the standard feature dimension <code>F0</code>.</p><p>The <code>FusionLayer</code> and <code>TemporalMixer</code> are task&#8209;specific: the stock task uses cross&#8209;attention in the fusion and a temporal mixer (dilated conv + self&#8209;attention) on the read&#8209;in features; the WRA task uses simple addition in the fusion and no temporal mixer.</p><h4><code>UGNN.forward</code> (equations (18)&#8211;(23))</h4><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c931d89b-b986-422e-8a84-80101c1f2e0b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def forward(self, x_k, k, u, S):
    batch, N, F_in = x_k.shape

    # 1. Read-in
    Z0 = self.read_in(x_k)                     # (batch, N, F0)
    U0 = self.node_state_encoder(u)            # (batch, N, F0)
    K0 = self.step_embedding(k).unsqueeze(1).expand(-1, N, -1)  # (batch, N, F0)

    # 2. Global embeddings
    global_emb = torch.cat([U0, K0], dim=-1)   # (batch, N, 2*F0)

    # 3. Temporal mixer (stock only)
    Z0 = self.temporal_mixer(Z0)

    # 4. Node counts
    node_counts = self._compute_target_node_counts(N)

    # 5. Encoder path
    D = torch.eye(N, device=x_k.device).unsqueeze(0).expand(batch, -1, -1)
    Z = Z0
    encoder_outputs = []
    for b_idx in range(self.B - 1):
        next_N = node_counts[b_idx + 1]
        gnn = self.encoder_blocks[b_idx]
        fusion = self.fusion_layers_enc[b_idx]
        sel_head = self.selection_heads[b_idx]

        # Fusion: P_b = &#928;_b(V_b, D_b[U0;K0])
        # (V_b is Z after down-sampling via C_b; here we use D to project global_emb)
        V_b = Z  # For now, V_b = Z (in practice we need C_b; see full code)
        global_emb_active = torch.bmm(D, global_emb)
        P_b = fusion(V_b, global_emb_active)

        Z_b = gnn(P_b, S, D)

        # Selection: C_{b+1}, Z_{b+1}
        C_next, Z_next = sel_head(Z_b, global_emb_active, next_N,
                                   training=self.training,
                                   epoch=self.current_epoch)

        # Update D for next level: D_{b+1} = C_{b+1} @ D_b
        D = torch.bmm(C_next, D)
        encoder_outputs.append(Z_b)
        Z = Z_next

    # 6. Bottleneck
    # ... (similar fusion + GNN on coarsest level)

    # 7. Decoder path
    # ... (up-sample via C^T, skip-concat, fusion, GNN)

    # 8. Read-out
    epsilon_theta = self.read_out(Y_1)
    return epsilon_theta</code></pre></div><p><em>(Note: The above is a simplified sketch. The full code in `ugnn.py` handles batching, indexing, and the decoder loop with proper up&#8209;sampling and skip connections.)</em></p><p><strong>Explanation of the forward pass:</strong></p><ol><li><p><strong>Read&#8209;in and embeddings:</strong> The noisy signal <code>x_k</code> is projected to dimension <code>F0</code> by a node&#8209;wise MLP. Node states <code>u</code> and the diffusion step <code>k</code> are similarly encoded and broadcast to all nodes. The step embedding uses sinusoidal encoding (following DDPM conventions) followed by an MLP.</p></li><li><p><strong>Global embeddings:</strong> <code>U0</code> and <code>K0</code> are concatenated along the feature dimension to form a <code>(batch, N, 2*F0)</code> tensor that carries all conditioning information.</p></li><li><p><strong>Encoder path:</strong> Starting from <code>Z0</code> and an identity <code>D</code> (active nodes = all N nodes), we iterate through the encoder blocks. At each depth <code>b</code>:</p></li></ol><ul><li><p>Project global embeddings onto the current active node set: <code>D_b * global_emb</code>.</p></li><li><p>Fuse (equation (22)): <code>P_b = &#928;_b(V_b, D_b[U0;K0])</code>.</p></li><li><p>Apply the GNN module with stride <code>gamma_b</code>: <code>Z_b = &#934;_{E_b}(P_b, S; D_b, gamma_b)</code>.</p></li><li><p>Select nodes for the next level: <code>C_{b+1}, Z_{b+1} = Selector(Z_b, D_b[U0;K0], N_{b+1})</code>.</p></li><li><p>Update the composite selection matrix: <code>D = C_{b+1} @ D</code> (this builds <code>D_{b+1}</code> as per equation (18)).</p></li></ul><ol start="4"><li><p><strong>Decoder path (omitted for brevity):</strong> The coarsest representation is processed by the bottleneck block. Then, iterating backwards, each decoder block up&#8209;samples via <code>C_{b+1}^T</code>, concatenates with the corresponding encoder output (skip connection), projects via <code>SkipProjection</code> (equation (23)), fuses, and applies a GNN module.</p></li><li><p><strong>Read&#8209;out:</strong> The finest&#8209;resolution decoder output <code>Y_1</code> is projected back to the original feature dimension by a node&#8209;wise MLP.</p></li></ol><p><strong>Key design choices:</strong></p><ul><li><p>The composite <code>D</code> matrices are maintained as dense <code>(batch, N_b, N)</code> tensors for simplicity. In a production system, you might store them as sparse index lists, but dense matrices make the batch matrix multiplications (<code>torch.bmm</code>) straightforward.</p></li><li><p>The <code>FusionLayer</code> is shared between encoder and decoder blocks only in the sense that the same class is instantiated with the same input/output dimensions; the paper mentions weight sharing, but we instantiate separate layers for clarity.</p></li><li><p>The <code>TemporalMixer</code> applies dilated 1D convolutions along the temporal axis (last dimension) followed by multi&#8209;head self&#8209;attention. The output is aggregated by taking the mean over the temporal dimension, producing a single feature vector per node. This is a reasonable default for the paper&#8217;s underspecified temporal encoder.</p></li></ul><div><hr></div><h3>5.5 Putting It All Together</h3><p>The four files work together seamlessly in the training loop (<code>train.py</code>):</p><ol><li><p><code>DiffusionProcess.forward_diffuse</code> generates <code>x_k</code> and <code>epsilon</code> from a clean <code>x_0</code> sampled from the dataloader.</p></li><li><p>The <code>UGNN</code> takes <code>(x_k, k, u, S)</code> and predicts <code>epsilon_theta</code>.</p></li><li><p>The MSE loss between <code>epsilon</code> and <code>epsilon_theta</code> is computed (equation (6)).</p></li><li><p>Gradients flow through the U&#8209;GNN, including through the straight&#8209;through estimator in <code>NodeSelectionHead</code>, and the parameters are updated.</p></li><li><p>After training, <code>DiffusionProcess.ddim_sample</code> uses the trained U&#8209;GNN to generate samples from pure noise, producing the final forecast trajectories.</p></li></ol><p>The modularity of the design&#8212;separate files for diffusion, graph convolution, node selection, and the full architecture&#8212;makes it easy to test, modify, or replace individual components without touching the rest of the pipeline.</p><h2>6. Training and Hyperparameters</h2><p>With the U-GNN architecture and data pipelines in place, we now turn to the training procedure. The paper specifies a detailed set of hyperparameters (Table I and Appendix E) and a training loop that includes a custom learning rate schedule, gradient clipping, automatic mixed precision, and a composite validation criterion. This section walks through the configuration object, the scheduler, the training loop itself, and the checkpointing logic, tying each piece to the generated code.</p><h3>6.1 Configuration from Table I and Appendix E</h3><p>All hyperparameters are centralized in <code>config.py</code>. The base <code>Config</code> class stores values that are shared across tasks, while <code>StockConfig</code> and <code>WRAConfig</code> override task-specific settings. The factory function <code>get_config(task)</code> returns the appropriate instance.</p><p><strong>Key hyperparameters (from Table I and Appendix E):</strong></p><p>| Parameter | Value | Paper Reference | |-----------|-------|-----------------| | Diffusion steps K | 500 | Table I | | \beta_1, \beta_{500} | 10^{-4}, 2\times10^{-2} | Table I, linear schedule | | DDIM steps | 100 | Table I | | DDIM \eta | 0.2 | Table I | | U-GNN depths B | 4 | Table I | | Pooling factor \rho | 2 | Table I | | Channel width F | 64 | Table I | | GNN layers per block L | 2 | Table I | | Filter taps K (hops) | 2 | Table I | | Max stride \gamma_{\max} | 2 | Table I | | Dropout | 0.1 | Table I | | Embedding dimension | 128 | Table I | | Epochs | 5000 | Table I | | Batch size (stock) | 64 | Table I | | Batch size (WRA) | 800 | Table I | | Learning rate | 10^{-3} | Table I | | Weight decay | 0.0 | Table I (AdamW default) | | Warmup fraction | 0.01 (1%) | Appendix E | | Decay fraction | 0.80 (80%) | Appendix E | | Decay to fraction | 0.05 (5% of peak) | Appendix E | | Gradient clip norm | 1.0 | Appendix E | | AMP | True | Appendix E | | Annealing warmup T_w | 0.02 (100 epochs) | Appendix C | | Annealing decay T_a | 0.73 (3650 epochs) | Appendix C | | Initial temperature \tau_0 | 1.0 | Appendix C | | Final temperature \tau_{\min} | 0.5 | Appendix C | | Initial exploration noise \varepsilon_0 | 1.0 | Appendix C | | Final exploration noise \varepsilon_{\min} | 0.0 | Appendix C | | RevIN blend (stock) | 0.7 | Section V-A | | RevIN scale correction (stock) | 1.107 | Section V-A |</p><p>The <code>Config.__init__</code> method calls <code>_compute_betas()</code> and <code>_compute_alphabars()</code> to precompute the linear \beta schedule and the cumulative \bar{\alpha}_k array (equations (2)&#8211;(3)). These are stored as tensors and used later by the <code>DiffusionProcess</code> class.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;47f039b2-eec7-4003-a8ea-67a90dea3b51&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># From config.py (simplified)
class Config:
    def __init__(self, task: str = "stock"):
        self.task = task
        if task == "stock":
            self.batch_size = 64
            self.U = 12
            self.Th = 20
            self.Tp = 5
        elif task == "wra":
            self.batch_size = 800
            self.U = 2
            self.Th = 0
            self.Tp = 1
        else:
            raise ValueError(f"Unknown task: {task}")
        self._compute_betas()
        self._compute_alphabars()

    def _compute_betas(self):
        k = torch.arange(1, self.K + 1, dtype=torch.float32)
        self.betas = self.beta_start + (k - 1) / (self.K - 1) * (self.beta_end - self.beta_start)

    def _compute_alphabars(self):
        self.alphas = 1.0 - self.betas
        self.alphabars = torch.cumprod(self.alphas, dim=0)</code></pre></div><h3>6.2 WarmupCosineLR Scheduler</h3><p>The paper describes a learning rate schedule with a linear warm-up over the first 1% of training steps, followed by cosine decay to 5% of the peak learning rate over the next 80% of steps, and then constant at that minimum. The <code>WarmupCosineLR</code> class in <code>train.py</code> implements this exactly.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;79546aca-99b2-4d95-9024-f5342c68d77d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class WarmupCosineLR(optim.lr_scheduler.LambdaLR):
    def __init__(self, optimizer, total_steps, warmup_frac=0.01,
                 decay_frac=0.80, min_lr_frac=0.05, last_epoch=-1):
        self.total_steps = total_steps
        self.warmup_steps = int(total_steps * warmup_frac)
        self.decay_steps = int(total_steps * decay_frac)
        self.min_lr_frac = min_lr_frac
        super().__init__(optimizer, self._lr_lambda, last_epoch=last_epoch)

    def _lr_lambda(self, step):
        if step &lt; self.warmup_steps:
            return float(step) / max(1.0, self.warmup_steps)
        elif step &lt; self.warmup_steps + self.decay_steps:
            progress = float(step - self.warmup_steps) / max(1.0, self.decay_steps)
            cosine_decay = 0.5 * (1.0 + np.cos(np.pi * progress))
            return self.min_lr_frac + (1.0 - self.min_lr_frac) * cosine_decay
        else:
            return self.min_lr_frac</code></pre></div><p>The scheduler is instantiated with <code>total_steps = epochs * batches_per_epoch</code>. Because the number of batches per epoch can vary, we estimate it from the training DataLoader. The scheduler is stepped after every optimizer step (i.e., every mini-batch), not every epoch.</p><h3>6.3 Training Loop</h3><p>The <code>train_epoch</code> function implements one epoch of training. For each batch, it:</p><ol><li><p><strong>Samples a diffusion step</strong> k \sim \text{Uniform}\{1, \dots, K\} independently for each element in the batch.</p></li><li><p><strong>Samples noise</strong> \varepsilon \sim \mathcal{N}(0, I) with the same shape as the clean signal x_0.</p></li><li><p><strong>Computes the noisy signal</strong> x_k = \sqrt{\bar{\alpha}_k}\,x_0 + \sqrt{1-\bar{\alpha}_k}\,\varepsilon using the precomputed \bar{\alpha} array (equation (4)).</p></li><li><p><strong>Predicts the noise</strong> \varepsilon_\theta = \text{U-GNN}(x_k, k, u, S).</p></li><li><p><strong>Computes the MSE loss</strong> \|\varepsilon - \varepsilon_\theta\|^2 (equation (6)).</p></li><li><p><strong>Backpropagates</strong> with gradient clipping (max L2 norm = 1.0) and automatic mixed precision (AMP) if CUDA is available.</p></li><li><p><strong>Updates the optimizer</strong> and <strong>steps the scheduler</strong>.</p></li></ol><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c84b5b70-2752-4188-bc2d-80350961a0eb&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># From train.py (simplified)
def train_epoch(model, dataloader, optimizer, scheduler, diffusion,
                epoch, config, scaler=None):
    model.train()
    if hasattr(model, 'set_epoch'):
        model.set_epoch(epoch)  # for annealing &#964; and &#949;

    total_loss = 0.0
    n_batches = 0

    for batch in dataloader:
        if config.task == 'stock':
            x0, u, S, _ = batch
        else:
            x0, S, u = batch

        x0 = x0.to(device, non_blocking=True)
        u = u.to(device, non_blocking=True)
        S = S.to(device, non_blocking=True)

        batch_size = x0.size(0)
        k = torch.randint(1, config.K + 1, (batch_size,), device=device)
        noise = torch.randn_like(x0)
        x_k = diffusion.forward_diffuse(x0, k, noise)

        with autocast(enabled=scaler is not None):
            pred_noise = model(x_k, k, u, S)
            loss = F.mse_loss(pred_noise, noise)

        if scaler is not None:
            scaler.scale(loss).backward()
            scaler.unscale_(optimizer)
            torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
            scaler.step(optimizer)
            scaler.update()
        else:
            loss.backward()
            torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
            optimizer.step()

        optimizer.zero_grad(set_to_none=True)
        scheduler.step()

        total_loss += loss.item()
        n_batches += 1

    return total_loss / max(1, n_batches)</code></pre></div><p><strong>Key details:</strong></p><ul><li><p>The <code>DiffusionProcess.forward_diffuse</code> method implements equation (4) by indexing into the precomputed <code>alpha_bar</code> tensor with the sampled k values.</p></li><li><p>Gradient clipping uses <code>torch.nn.utils.clip_grad_norm_</code> with <code>max_norm=1.0</code> and the default L2 norm.</p></li><li><p>AMP is enabled via <code>torch.cuda.amp.GradScaler</code> and <code>autocast</code>. The scaler is created only if CUDA is available.</p></li><li><p>The model&#8217;s <code>set_epoch</code> method (if present) is called at the start of each epoch to update the temperature \tau and exploration noise \varepsilon for the node selection heads, following the annealing schedule of Appendix C.</p></li></ul><h3>6.4 Annealing of &#964; and &#949;</h3><p>The node selection heads use a Gumbel-Top-K mechanism during training, with temperature \tau and exploration noise \varepsilon that are annealed over time. The schedule (equation (34)) has three phases:</p><ul><li><p><strong>Warm-up</strong> (first T_w = 0.02 \times 5000 = 100 epochs): \tau and \varepsilon are held at their initial values (1.0 and 1.0).</p></li><li><p><strong>Anneal</strong> (next T_a = 0.73 \times 5000 = 3650 epochs): both parameters decay linearly to their final values (0.5 and 0.0).</p></li><li><p><strong>Constant</strong> (remaining epochs): stay at final values.</p></li></ul><p>In the code, the <code>NodeSelectionHead</code> class (in <code>node_selection.py</code>) implements <code>anneal_tau_eps(epoch)</code> which computes the current values based on the schedule. The <code>UGNN</code> model&#8217;s <code>set_epoch</code> method propagates the epoch number to all selection heads so they can update their internal parameters before each forward pass.</p><h3>6.5 Composite Validation Criteria</h3><p>The paper uses a task-specific composite score to select the best checkpoint during training. The composite criterion combines multiple validation metrics into a single scalar.</p><p><strong>For stock forecasting:</strong> The composite score is a weighted sum of the noise-prediction loss, the gap between training and validation loss (to detect overfitting), and the CRPS (Continuous Ranked Probability Score). Lower is better.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b0043c61-ca1e-4fe5-b643-32c80ec3c8e9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def composite_criterion_stock(val_metrics):
    loss = val_metrics.get('loss', 1.0)
    crps = val_metrics.get('crps', 1.0)
    # train-val gap not computed here, assume 0
    return loss + 0.5 * crps</code></pre></div><p><strong>For wireless resource allocation:</strong> The composite score combines the mean ergodic rate and the feasibility (fraction of users meeting the minimum rate constraint). Higher is better, so we negate the sum for minimization.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d2286b7f-e666-4912-a06f-021027316ca0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def composite_criterion_wra(val_metrics):
    mean_rate = val_metrics.get('mean_rate', 0.0)
    feasibility = val_metrics.get('feasibility', 0.0)
    return -(mean_rate + feasibility * 2.0)  # weight feasibility more</code></pre></div><p>These functions are called after each validation epoch. The <code>train</code> function tracks the best composite score and saves the corresponding model checkpoint.</p><h3>6.6 Checkpointing</h3><p>At the end of each epoch, if the composite score improves (lower for stock, higher for WRA), the current model state, optimizer state, scheduler state, epoch number, and score are saved to a dictionary. After training completes, the best checkpoint is written to <code>checkpoints/best_model.pth</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;03984698-9614-4a36-bd68-609cedc09d7b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">if (config.task == 'stock' and score &lt; best_score) or \
   (config.task == 'wra' and score &gt; best_score):
    best_score = score
    best_epoch = epoch
    best_state = {
        'model_state_dict': model.state_dict(),
        'optimizer_state_dict': optimizer.state_dict(),
        'scheduler_state_dict': scheduler.state_dict(),
        'epoch': epoch,
        'score': score,
        'config': config
    }</code></pre></div><p>This checkpoint can later be loaded for evaluation or visualization.</p><h3>6.7 Summary</h3><p>The training configuration and loop are designed to match the paper&#8217;s specifications exactly. The <code>Config</code> class centralizes all hyperparameters, the <code>WarmupCosineLR</code> scheduler implements the custom learning rate schedule, and the <code>train_epoch</code> function handles the diffusion-specific steps (sampling k, computing x_k, predicting noise) along with gradient clipping and AMP. The annealing of node selection parameters is integrated via the model&#8217;s <code>set_epoch</code> hook, and the composite validation criteria guide checkpoint selection. This modular design makes it straightforward to reproduce the paper&#8217;s experiments and to adapt the code to new tasks.</p><h2>7. Evaluation Metrics and Baselines</h2><p>A generative model is only as useful as the metrics we use to judge it. The paper evaluates its U-GNN diffusion model against two sets of baselines and a battery of metrics that go far beyond simple point-prediction errors. For stock forecasting, the goal is to assess both the accuracy of the mean prediction and the fidelity of the full predictive distribution. For wireless resource allocation, the focus is on the ergodic sum rate, feasibility, and the gap to an expert algorithm. This section explains every metric and baseline, shows the corresponding code in <code>eval.py</code> and <code>baselines.py</code>, and describes how to compare the U-GNN against the baselines.</p><h3>7.1 Stock Forecasting Metrics</h3><p>The paper reports eight metrics for the S&amp;P 500 task (Table II). They fall into three categories: distributional accuracy (CRPS, MIS90), point-prediction accuracy (RMSE, MAE, direction accuracy), and stylized-fact gaps (volatility clustering, momentum, excess kurtosis). All metrics are computed on the <em>returns</em> (log-returns) unless otherwise noted.</p><h4>7.1.1 Continuous Ranked Probability Score (CRPS)</h4><p>CRPS measures how well the empirical distribution of generated samples matches the true target. It is defined as the integral of the squared difference between the empirical CDF of the samples and the indicator function of the target:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\text{CRPS} = \\int_{-\\infty}^{\\infty} \\bigl( F_{\\text{samples}}(x) - \\mathbb{1}\\{x \\ge \\text{target}\\} \\bigr)^2 \\, dx&quot;,&quot;id&quot;:&quot;EB53A23CFE&quot;}" data-component-name="LatexBlockToDOM"></div><p>For a set of M samples, the empirical CDF is a step function, and the integral can be computed analytically using the sorted samples. The implementation in <code>StockEvaluator.compute_crps</code> uses the formula:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\text{CRPS} = \\frac{1}{M} \\sum_{i=1}^{M} |s_{(i)} - y| - \\frac{1}{2M^2} \\sum_{i,j} |s_{(i)} - s_{(j)}|&quot;,&quot;id&quot;:&quot;4087D9222E&quot;}" data-component-name="LatexBlockToDOM"></div><p>where s_{(i)} are the sorted samples and y is the target. The double sum is computed efficiently using the fact that for sorted samples, \sum_{i,j} |s_{(i)} - s_{(j)}| = 2 \sum_{i=1}^{M} (2i - M - 1) s_{(i)}.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;93f8c086-b87d-478c-bd28-3d57e98daca1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@staticmethod
def compute_crps(samples: np.ndarray, target: np.ndarray) -&gt; float:
    samples = np.asarray(samples)
    target = np.asarray(target)
    if samples.ndim == 3:
        num_samples = samples.shape[0]
    else:
        num_samples = 1
        samples = samples[np.newaxis, ...]
    m = samples.shape[1] * samples.shape[2]
    samples_flat = samples.reshape(num_samples, -1)
    target_flat = target.reshape(-1)
    sorted_samples = np.sort(samples_flat, axis=0)
    M = num_samples
    abs_diff = np.mean(np.abs(sorted_samples - target_flat[np.newaxis, :]), axis=0)
    weights = 2 * (2 * np.arange(1, M + 1) - M - 1)
    sum_abs_diff_sorted = np.sum(weights[:, np.newaxis] * sorted_samples, axis=0) / (M ** 2)
    crps_per_element = abs_diff - 0.5 * sum_abs_diff_sorted
    return float(np.mean(crps_per_element))</code></pre></div><h4>7.1.2 Mean Interval Score for 90% Prediction Interval (MIS90)</h4><p>MIS90 evaluates the sharpness and calibration of the 90% prediction interval. It penalizes intervals that are too wide and intervals that miss the target:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\text{MIS} = (u - l) + \\frac{2}{\\alpha}(l - y)\\mathbb{1}\\{y < l\\} + \\frac{2}{\\alpha}(y - u)\\mathbb{1}\\{y > u\\}&quot;,&quot;id&quot;:&quot;0085FD07C5&quot;}" data-component-name="LatexBlockToDOM"></div><p>where l and u are the 5th and 95th percentiles of the samples, and \alpha = 0.1. The implementation computes these percentiles using <code>np.percentile</code>.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;88e1aa31-8501-4fb4-90be-31da9e220efa&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@staticmethod
def compute_mis90(samples: np.ndarray, target: np.ndarray) -&gt; float:
    samples = np.asarray(samples)
    target = np.asarray(target)
    if samples.ndim == 2:
        samples = samples[np.newaxis, ...]
    alpha = 0.1
    lower = np.percentile(samples, 5, axis=0)
    upper = np.percentile(samples, 95, axis=0)
    width = upper - lower
    penalty_low = (2.0 / alpha) * (lower - target) * (target &lt; lower)
    penalty_high = (2.0 / alpha) * (target - upper) * (target &gt; upper)
    mis = width + penalty_low + penalty_high
    return float(np.mean(mis))</code></pre></div><h4>7.1.3 RMSE, MAE, and Direction Accuracy</h4><p>These are standard point-prediction metrics. The mean prediction is the average over the M samples. RMSE and MAE are computed between this mean and the target. Direction accuracy measures the fraction of elements where the sign of the mean prediction matches the sign of the target (ignoring zeros).</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;246e8cec-1051-4167-9cfa-4fb0ce7c6c68&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@staticmethod
def compute_rmse(samples: np.ndarray, target: np.ndarray) -&gt; float:
    samples = np.asarray(samples)
    target = np.asarray(target)
    if samples.ndim == 3:
        mean_pred = np.mean(samples, axis=0)
    else:
        mean_pred = samples
    return float(np.sqrt(np.mean((mean_pred - target) ** 2)))

@staticmethod
def compute_dir_accuracy(samples: np.ndarray, target: np.ndarray) -&gt; float:
    samples = np.asarray(samples)
    target = np.asarray(target)
    if samples.ndim == 3:
        mean_pred = np.mean(samples, axis=0)
    else:
        mean_pred = samples
    correct = (np.sign(mean_pred) == np.sign(target))
    return float(np.mean(correct))</code></pre></div><h4>7.1.4 Stylized-Fact Gaps</h4><p>Financial time series exhibit well-known statistical regularities called <em>stylized facts</em>. The paper measures how well the generated samples reproduce three of them:</p><ul><li><p><strong>Volatility clustering</strong>: the tendency for large changes to be followed by large changes. Quantified as the lag-1 autocorrelation of squared returns.</p></li><li><p><strong>Momentum</strong>: the tendency for returns to be positively autocorrelated at short lags. Quantified as the lag-1 autocorrelation of returns.</p></li><li><p><strong>Excess kurtosis</strong>: the fat-tailed nature of return distributions. Quantified as the sample excess kurtosis (kurtosis minus 3).</p></li></ul><p>For each metric, the <em>gap</em> is the absolute difference between the value computed on the target and the value computed on the pooled generated samples. The implementation in <code>StockEvaluator</code> computes these using NumPy.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c95042c4-0d59-4351-addc-c1deebd31f12&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@staticmethod
def compute_volatility_gap(samples: np.ndarray, target: np.ndarray) -&gt; float:
    samples = np.asarray(samples)
    target = np.asarray(target)
    target_flat = target.ravel()
    samples_flat = samples.ravel()

    def autocorr_lag1(x):
        if len(x) &lt; 2:
            return 0.0
        x0 = x[:-1]
        x1 = x[1:]
        std0 = np.std(x0)
        std1 = np.std(x1)
        if std0 == 0 or std1 == 0:
            return 0.0
        return float(np.corrcoef(x0, x1)[0, 1])

    target_vol = autocorr_lag1(target_flat ** 2)
    samples_vol = autocorr_lag1(samples_flat ** 2)
    return float(np.abs(target_vol - samples_vol))</code></pre></div><h3>7.2 Wireless Resource Allocation Metrics</h3><p>For the WRA task, the paper reports ergodic sum rate, feasibility, and the gap to the expert primal-dual algorithm (Table IV).</p><h4>7.2.1 Ergodic Sum Rate</h4><p>The ergodic sum rate is the average over fading realizations of the sum of per-user rates. The per-user rate is (equation (26)):</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;r_j = \\log_2(1 + \\text{SINR}_j)&quot;,&quot;id&quot;:&quot;89A58396BA&quot;}" data-component-name="LatexBlockToDOM"></div><p>where \text{SINR}_j = \frac{p_j g_{jj}}{\sum_{i \neq j} p_i g_{ij} + \sigma^2}. The ergodic estimate (equation (27)) averages over T = 500 fading realizations, with the allocation held constant for T_0 = 5 slots. The implementation in <code>WRAEvaluator.compute_ergodic_rate</code> generates independent Rayleigh fading (exponential squared magnitude) and computes the SINR for each slot.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ffda163a-5c27-42f1-a9b4-52a1a5c10917&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@staticmethod
def compute_ergodic_rate(
    allocations: Union[np.ndarray, torch.Tensor],
    H: Union[np.ndarray, torch.Tensor],
    noise_power: float = 1e-11,
    T: int = 500,
    T0: int = 5
) -&gt; float:
    allocations = np.asarray(allocations).flatten()
    H = np.asarray(H)
    N = len(allocations)
    # Uncenter if needed (assume Pmax=1)
    if np.any(allocations &lt; 0):
        allocations = allocations + 1.0
    allocations = np.clip(allocations, 0.0, 1.0)
    # Generate fading: (N, N, T) exponential with mean 1
    fading = np.random.exponential(scale=1.0, size=(N, N, T)).astype(np.float64)
    H = np.abs(H)
    g = H[:, :, np.newaxis] * fading
    g_self = g.diagonal(axis1=0, axis2=1)
    p = allocations[:, np.newaxis, np.newaxis]
    total_received = np.sum(p * g, axis=0)
    self_power = p[:, 0, 0] * g_self
    interference = total_received - self_power
    sinr = self_power / (interference + noise_power)
    rate = np.log2(1.0 + np.maximum(sinr, 0.0))
    sum_rate_per_slot = np.sum(rate, axis=0)
    mean_sum_rate = np.mean(sum_rate_per_slot)
    return float(mean_sum_rate)</code></pre></div><h4>7.2.2 Feasibility</h4><p>Feasibility is the fraction of users whose ergodic rate meets a minimum rate constraint (e.g., 0.5 bits/s/Hz). The implementation computes the per-user ergodic rate (mean over fading) and checks the constraint.</p><h4>7.2.3 Gap to Expert</h4><p>The paper reports the percentage gap between the U-GNN and the expert primal-dual algorithm at the 1st, 5th, 10th percentiles and the mean. The gap is defined as:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\text{gap}_p = \\frac{\\text{expert}_p - \\text{U-GNN}_p}{\\text{expert}_p} \\times 100\\%&quot;,&quot;id&quot;:&quot;D25874DA99&quot;}" data-component-name="LatexBlockToDOM"></div><p>where p indicates the percentile of the ergodic sum rate distribution across test networks.</p><h3>7.3 Baselines</h3><p>The paper compares against three baselines, implemented in <code>baselines.py</code>.</p><h4>7.3.1 Geometric Random Walk (GRW)</h4><p>For stock forecasting, the GRW assumes log-returns are independent Gaussian with variance estimated from the history window. The <code>GeometricRandomWalk.forecast</code> method computes the per-stock variance from the 20-day history and samples M independent trajectories.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;1f682ad1-605a-41b6-8884-e39a34b07d39&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class GeometricRandomWalk:
    def forecast(self, history: torch.Tensor, steps: int, num_samples: int = 1) -&gt; torch.Tensor:
        if history.dim() == 3:
            history = history.squeeze(-1)
        N, Th = history.shape
        var = history.var(dim=1, unbiased=True, keepdim=False) + 1e-8
        std = torch.sqrt(var)
        noise = torch.randn(num_samples, N, steps, generator=self.rng, device=history.device, dtype=history.dtype)
        samples = noise * std.view(1, N, 1)
        return samples</code></pre></div><h4>7.3.2 Full Power (FP) and Average Power (AP)</h4><p>For WRA, the FP baseline sets every transmitter to maximum power P_{\text{max}}, so the centered allocation is 0. The AP baseline sets every transmitter to P_{\text{max}}/2, so the centered allocation is -0.5. Both are deterministic and return tensors of the appropriate shape.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;86f1c389-335c-4322-91df-985fb8d9fff6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class FullPower:
    def allocate(self, network: dict = None, H: torch.Tensor = None, num_samples: int = 1) -&gt; torch.Tensor:
        N = 400
        if H is not None:
            N = H.shape[0]
        device = H.device if H is not None else torch.device('cpu')
        return torch.zeros(num_samples, N, 1, device=device, dtype=torch.float32)

class AveragePower:
    def allocate(self, network: dict = None, H: torch.Tensor = None, num_samples: int = 1) -&gt; torch.Tensor:
        N = 400
        if H is not None:
            N = H.shape[0]
        device = H.device if H is not None else torch.device('cpu')
        return torch.full((num_samples, N, 1), -0.5, device=device, dtype=torch.float32)</code></pre></div><h3>7.4 Comparing U-GNN Against Baselines</h3><p>To compare the U-GNN against the baselines, we generate M = 100 samples from each method (for GRW, this is natural; for FP/AP, we repeat the deterministic allocation). Then we compute all metrics using the <code>StockEvaluator</code> or <code>WRAEvaluator</code>. The <code>evaluate_all</code> method returns a dictionary of metrics, which can be printed as a table (matching Table II or Table IV).</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ff9158a4-b026-4df8-b1e5-bdcea94ca666&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">evaluator = StockEvaluator()
metrics_ugnn = evaluator.e

## 8. Results and Figures

A generative model&#8217;s output is inherently high-dimensional: for stock forecasting, we produce a distribution over future price paths for hundreds of stocks; for wireless resource allocation, we produce a distribution over power vectors. To communicate the quality of these distributions, the paper relies on a set of carefully designed figures that go beyond simple scalar metrics. This section walks through the four plotting functions in `visualize.py` that reproduce the paper&#8217;s key visual results: forecasting trajectories, empirical CDFs, distribution diagnostics (histogram + Q-Q plot), and a WRA performance bar chart. Each function is explained in detail, with code snippets and guidance on how to call them from the main evaluation script.

### 8.1 Overview of the Plotting Functions

The file `visualize.py` contains four public functions:

- `plot_forecast_trajectories` &#8211; reproduces the paper&#8217;s &#8220;forecasting trajectories&#8221; figures (e.g., Figure 2 in the paper, showing historical prices, generated samples, and actual target for selected stocks).
- `plot_cdf` &#8211; plots the empirical cumulative distribution function of the pooled generated samples versus the pooled target values, corresponding to the CDF panels in the paper&#8217;s distribution diagnostics.
- `plot_distribution_diagnostics` &#8211; creates a two-panel figure with an overlaid histogram and a Q-Q plot, matching the paper&#8217;s density and quantile comparisons.
- `plot_wra_bar` &#8211; produces a grouped bar chart comparing ergodic rate metrics (p1, p5, p10, mean) across methods (U-GNN, Expert, FP, AP), with an optional secondary axis for percentage gaps and an inset for feasibility. This reproduces the paper&#8217;s WRA performance bar chart (e.g., Figure 4).

All functions accept NumPy arrays (or torch tensors, which are automatically converted via the internal `_to_numpy` helper) and return a `matplotlib.figure.Figure` object that can be saved or displayed.

### 8.2 Forecasting Trajectories

The function `plot_forecast_trajectories` visualises the model&#8217;s predictive distribution for a few selected stocks. It plots:
- The historical log-returns (or prices) for the last `Th` days (blue line).
- A random subset of generated trajectories (gray lines, low opacity).
- The mean of all generated trajectories (red line).
- The actual target values (green dashed line).

The time axis is split into history (days 0 to Th-1) and forecast (days Th to Th+Tp-1).
</code></pre></div><p>def plot<em>forecast</em>trajectories( prices: np.ndarray, samples: np.ndarray, target: np.ndarray, stock<em>indices: Optional[List[int]] = None, num</em>trajectories: int = 20, figsize: Tuple[float, float] = (10, 6) ) -&gt; plt.Figure:</p><p>Use the button or URL below to download the complete Python source code.</p>
      <p>
          <a href="/__u/onepagecode.substack.com/p/forecasting-stock-prices-with-diffusion">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Quantitative Trading Model from Paper to Python: Building an Attention-BiLSTM Strategy for Gold and Bitcoin]]></title><description><![CDATA[A step-by-step tutorial on translating academic paper concepts&#8212;including temporal attention, streak-based position sizing, and greedy portfolio optimization&#8212;into production-ready Python code.]]></description><link>https://onepagecode.substack.com/p/quantitative-trading-model-from-paper</link><guid isPermaLink="false">https://onepagecode.substack.com/p/quantitative-trading-model-from-paper</guid><dc:creator><![CDATA[Onepagecode]]></dc:creator><pubDate>Fri, 10 Jul 2026 12:02:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aq85!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0996af8d-98d9-422f-9f4b-8ac6fc43de2d_1406x1372.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What This Article Builds</h2><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!F5TF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef206556-2fb9-41b0-9dc7-b20367aef72e_758x234.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!F5TF!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef206556-2fb9-41b0-9dc7-b20367aef72e_758x234.png 424w, /__u/substackcdn.com/image/fetch/$s_!F5TF!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef206556-2fb9-41b0-9dc7-b20367aef72e_758x234.png 848w, /__u/substackcdn.com/image/fetch/$s_!F5TF!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef206556-2fb9-41b0-9dc7-b20367aef72e_758x234.png 1272w, /__u/substackcdn.com/image/fetch/$s_!F5TF!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef206556-2fb9-41b0-9dc7-b20367aef72e_758x234.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!F5TF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef206556-2fb9-41b0-9dc7-b20367aef72e_758x234.png" width="758" height="234" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ef206556-2fb9-41b0-9dc7-b20367aef72e_758x234.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:234,&quot;width&quot;:758,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:112430,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef206556-2fb9-41b0-9dc7-b20367aef72e_758x234.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!F5TF!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef206556-2fb9-41b0-9dc7-b20367aef72e_758x234.png 424w, /__u/substackcdn.com/image/fetch/$s_!F5TF!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef206556-2fb9-41b0-9dc7-b20367aef72e_758x234.png 848w, /__u/substackcdn.com/image/fetch/$s_!F5TF!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef206556-2fb9-41b0-9dc7-b20367aef72e_758x234.png 1272w, /__u/substackcdn.com/image/fetch/$s_!F5TF!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef206556-2fb9-41b0-9dc7-b20367aef72e_758x234.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p>The paper proposes a quantitative trading pipeline for gold and Bitcoin. It combines price forecasting with an attention-enhanced bidirectional LSTM, position sizing based on empirical consecutive rises and declines, transaction-cost sensitivity analysis, and claimed use of VaR and a modified greedy algorithm for trading decisions. Several benchmark forecasting models are compared, with Att-BiLSTM reported as the best predictor. The trading results claim a final value of approximately $646 for a $500 gold allocation and $215,487 for a $500 Bitcoin allocation, although the backtest protocol and several model details are insufficiently specified.</p><div><hr></div><p><strong>The code is my implementation based on what the paper discusses. Download the source code using the button at the end of this article!</strong></p><p><strong>Also the code might have some errors, because I didn&#8217;t found anything online related to paper, everything is implemented by myself, so download it and understand. With article and the code.</strong></p><p>This is the research paper I tried to implement: https://www.researchgate.net/publication/364569615_Research_on_Quantitative_Trading_Model</p><h3>Our Book is out on openclaw: </h3><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.amazon.com/dp/B0H1B37D9M&quot;,&quot;text&quot;:&quot;Get Your Copy&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.amazon.com/dp/B0H1B37D9M"><span>Get Your Copy</span></a></p><div><hr></div><h2>Implementation Assumptions</h2><ul><li><p>Historical gold and Bitcoin datasets are not supplied; the implementation will accept local CSV or pandas DataFrame inputs and include a deterministic synthetic-data generator for demonstrations only.</p></li><li><p>Prices are assumed to be positive, chronologically ordered observations with one row per asset and an explicit timestamp.</p></li><li><p>Standard simple return, (p<em>t-p</em>{t-1})/p_{t-1}, is used by default because the paper's printed denominator is ambiguous; the choice is configurable and documented.</p></li><li><p>The primary target is a one-step-ahead normalized price or return forecast. A three-step vector forecast option is supported because the paper's conclusion mentions three-day forecasts.</p></li><li><p>The primary forecasting model is an attention-enhanced BiLSTM with a configurable learned query vector, 20% dropout, MAE loss, and RMSprop optimizer.</p></li><li><p>The paper-style 70:30 chronological split is supported for comparison, while walk-forward evaluation is the preferred protocol for trading conclusions.</p></li><li><p>Equations (5) through (8), VaR operation, and the modified greedy algorithm are implemented as explicit reconstructions and will never be presented as exact reproductions.</p></li><li><p>Transaction fees apply to both buys and sells by default, use decimal rates such as 0.002 for 0.2%, and are configurable.</p></li><li><p>The backtest uses next-bar execution after a signal, finite position-addition levels, no leverage by default, and optional slippage.</p></li><li><p>The reported final values of approximately 646 USD for gold and 215487 USD for Bitcoin are recorded as unverified reference claims rather than expected test outputs.</p></li><li><p>Deep-learning, statistical, and optional benchmark dependencies are isolated so core data, sizing, risk, and backtesting components remain usable without every optional package installed.</p></li></ul><h1>1. The Paper&#8217;s Core Idea: Forecast Timing and Manage Position Size Separately</h1><p>The paper is trying to answer two different trading questions:</p><p><strong>When should the strategy trade?</strong></p><p><strong>How much capital should it deploy when it trades?</strong></p><p>Those questions are related, but they are not the same problem. A forecasting model can estimate that gold or Bitcoin may rise, yet it does not automatically determine whether the strategy should invest $10, $100, or the entire available budget. Conversely, a position-sizing rule can specify how much to buy after a decline, but it needs a signal to decide whether buying is appropriate at that moment.</p><p>The implementation keeps these responsibilities separate. The forecasting component produces information about a possible future price. The signal layer converts that forecast into a buy, sell, add, or hold decision. The sizing layer proposes an amount. Risk control can reject or reduce that amount. Finally, the execution and accounting layer applies transaction costs, updates cash and holdings, and measures the resulting portfolio.</p><p>This separation is especially important because several parts of the paper are incompletely specified. The paper gives a broad design involving an attention-enhanced BiLSTM, consecutive-rise and decline analysis, exponential position sizing, VaR, and a modified greedy algorithm. It does not fully define every interface between those components. A modular Python implementation makes each interpretation visible instead of hiding assumptions inside one large trading function.</p><h2>The complete data-to-portfolio flow</h2><p>The project&#8217;s README presents the intended architecture as a pipeline:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;45fc8f13-9472-423a-b1ed-5862ae608001&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">local prices
    -&gt; cleaning and chronological split
    -&gt; leakage-safe sliding windows
    -&gt; forecast model
    -&gt; buy / sell / hold signal
    -&gt; finite exponential position-size schedule
    -&gt; historical VaR filter
    -&gt; reconstructed greedy action selection
    -&gt; next-bar execution with fees and slippage
    -&gt; portfolio valuation and performance metrics</code></pre></div><p>Each stage has a distinct responsibility:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Totb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e7eebf1-25be-4651-b16f-b0483ba328ee_1564x830.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Totb!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e7eebf1-25be-4651-b16f-b0483ba328ee_1564x830.png 424w, /__u/substackcdn.com/image/fetch/$s_!Totb!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e7eebf1-25be-4651-b16f-b0483ba328ee_1564x830.png 848w, /__u/substackcdn.com/image/fetch/$s_!Totb!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e7eebf1-25be-4651-b16f-b0483ba328ee_1564x830.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Totb!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e7eebf1-25be-4651-b16f-b0483ba328ee_1564x830.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Totb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e7eebf1-25be-4651-b16f-b0483ba328ee_1564x830.png" width="1456" height="773" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4e7eebf1-25be-4651-b16f-b0483ba328ee_1564x830.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:773,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:166812,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e7eebf1-25be-4651-b16f-b0483ba328ee_1564x830.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Totb!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e7eebf1-25be-4651-b16f-b0483ba328ee_1564x830.png 424w, /__u/substackcdn.com/image/fetch/$s_!Totb!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e7eebf1-25be-4651-b16f-b0483ba328ee_1564x830.png 848w, /__u/substackcdn.com/image/fetch/$s_!Totb!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e7eebf1-25be-4651-b16f-b0483ba328ee_1564x830.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Totb!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e7eebf1-25be-4651-b16f-b0483ba328ee_1564x830.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The paper&#8217;s primary model is the attention-enhanced BiLSTM, or <strong>Att-BiLSTM</strong>. It processes a historical sequence and uses temporal attention to produce a representation for forecasting. The input sequence is built from a sliding window, reported as 100 observations with stride one. The resulting prediction is not itself a trade: it must be aligned with a timestamp and passed to a decision policy.</p><p>The position-management component is conceptually independent of the neural network. It analyzes gains and declines, estimates statistics such as representative gain and decline magnitudes, and uses an exponential schedule to increase later additions. Because the paper&#8217;s position-sizing equations are corrupted or incomplete in the extracted source, the Python project treats this as a documented normalized exponential reconstruction rather than an exact transcription.</p><h2>Why the boundaries matter</h2><p>A single function that loads prices, trains a model, decides an order, and reports profit would be difficult to audit. It would also make it unclear whether future information entered the decision process. The generated package instead plans separate modules for data, models, streak analysis, sizing, risk, strategy, and backtesting.</p><p>For example, <code>DataConfig</code> makes the forecasting window explicit:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;69836615-110c-4ac4-96de-5477655cb65e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass
class DataConfig:
    """Configuration for local historical-price preparation."""

    timestamp_column: str = "timestamp"
    price_column: str = "price"
    asset_column: Optional[str] = None
    window_length: int = 100
    stride: int = 1
    return_definition: str = "simple"
    zero_return_policy: str = "break"</code></pre></div><p>The values <code>window_length=100</code> and <code>stride=1</code> reflect the paper&#8217;s stated sliding-window setup. Other fields expose choices that the paper does not settle, such as how returns are defined and how zero returns affect streaks. Making these values configuration fields means a later experiment can record exactly which interpretation it used.</p><p>The data layer represents a cleaned asset as <code>PriceData</code>, a small object containing timestamps, prices, and an asset label. Its validation contract requires unique, increasing timestamps and finite positive prices. That contract is more than cosmetic. A negative or missing price can invalidate returns, position sizes, portfolio valuation, and risk calculations downstream.</p><p>The <code>WindowedDataset</code> object preserves temporal metadata in addition to arrays:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6f0a0462-ee9c-4900-aeed-aa0854ca4ee9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class WindowedDataset:
    """Windowed samples and their temporal metadata."""

    X: np.ndarray
    y: np.ndarray
    window_start: pd.DatetimeIndex
    window_end: pd.DatetimeIndex
    y_timestamp: pd.DatetimeIndex
    feature_names: Tuple[str, ...] = ("price",)</code></pre></div><p>Keeping <code>window_end</code> and <code>y_timestamp</code> is important. A model prediction must be associated with the time at which its input information became available. The later backtest can then enforce the rule that a signal generated after the final input observation cannot fill at an earlier or identical timestamp.</p><h2>Separating a forecast from an order</h2><p>The paper&#8217;s forecasting section and trading section imply different horizons. The prediction tables appear to use a one-step setup, while the conclusion refers to forecasts for the next three days. The implementation therefore exposes the forecast horizon rather than assuming that those descriptions are identical.</p><p>Conceptually, a one-step window is:</p><p>\[ X<em>t = [z</em>{t-99}, z<em>{t-98}, \ldots, z</em>t], \qquad y<em>t = z</em>{t+1}. \]</p><p>A three-step version is:</p><p>\[ y<em>t = [z</em>{t+1}, z<em>{t+2}, z</em>{t+3}]. \]</p><p>The forecast layer answers what the model predicts. A separate signal layer must decide how to interpret that prediction. For example, it might compare the terminal three-day forecast with the current price, or use a configured threshold before producing a buy signal. That rule is a strategy assumption, not a consequence of the neural network itself.</p><p>The same distinction applies to sizing. A positive forecast does not imply that the strategy should invest its full available balance. The sizing schedule may allocate one of several finite addition levels, subject to cash and exposure caps. The eventual order amount is therefore the result of several stages:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;752d07d3-1d92-4b3b-88a7-8415f0e7df84&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">forecast
    -&gt; directional signal
    -&gt; candidate order amount
    -&gt; risk adjustment
    -&gt; selected feasible action
    -&gt; executed fill</code></pre></div><p>This decomposition also makes it possible to compare alternatives. The same forecast can be evaluated with a flat allocation, the reconstructed exponential schedule, or a different risk limit. Likewise, the same sizing rule can be tested with Att-BiLSTM predictions or a simpler benchmark.</p><h2>Offline and educational by design</h2><p>The generated project is deliberately offline. It accepts local CSV files or deterministic synthetic data and does not connect to an exchange, download market data, handle credentials, or submit orders. This restriction is appropriate for a paper-to-code tutorial because it keeps the focus on data alignment, modeling assumptions, decision rules, and accounting.</p><p>Synthetic data are useful for demonstrating interfaces. For example, the data module includes a deterministic generator whose purpose is to provide repeatable price paths for checking window shapes and portfolio arithmetic. Such data are not evidence about gold or Bitcoin and cannot reproduce the paper&#8217;s reported outcomes.</p><p>The package initializer reinforces this design by avoiding model training or data loading during import:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;64a99256-b76d-4653-9b12-27e3abdcf44e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">"""Offline research components for the quantitative trading model paper."""

# Configuration is part of the lightweight public API. Heavy forecasting and
# benchmark dependencies are imported only by their respective modules.
from .config import (
    BacktestConfig,
    DataConfig,
    ExecutionConfig,
    FeeSensitivityConfig,
    ForecastConfig,
    SizingConfig,
    SplitConfig,
    VaRConfig,
    validate_config,
)</code></pre></div><p>This side-effect-free structure lets a reader use data preparation, streak analysis, sizing, risk, or accounting utilities without installing PyTorch or every optional benchmark library. It also avoids implying that importing the package performs a market action.</p><h2>What this architecture does&#8212;and does not&#8212;claim</h2><p>The architecture provides a practical way to translate the paper&#8217;s major ideas into independently inspectable components. It directly reflects the clearly described mechanics, such as sliding windows and temporal attention, while making incomplete parts configurable and labeled as reconstructions.</p><p>It does <strong>not</strong> establish that the paper&#8217;s strategy is profitable or that the implementation reproduces its reported numbers. The paper&#8217;s approximate final values&#8212;$646 from a $500 gold allocation and $215,487 from a $500 Bitcoin allocation&#8212;remain unverified claims because the original data, dates, execution rules, risk settings, and position-sizing details are unavailable.</p><p>The useful starting point is therefore not a promise of matching those numbers. It is a disciplined contract between modules:</p><ul><li><p>data provides only ordered historical observations;</p></li><li><p>forecasting uses a defined historical window;</p></li><li><p>signals use predictions and current state;</p></li><li><p>sizing uses a finite, explicit budget;</p></li><li><p>risk control uses pre-decision information;</p></li><li><p>execution occurs after the signal timestamp;</p></li><li><p>evaluation accounts for costs and reports portfolio behavior separately from forecast error.</p></li></ul><p>That contract is the foundation for the remaining sections, where each component is examined in detail and every reconstruction decision is identified.</p><h2>2. What Can Actually Be Reproduced?</h2><p>The paper separates two decisions: forecasting proposes <strong>when</strong> a trade might be useful, while position management proposes <strong>how much</strong> capital to deploy. Before implementing those components, distinguish a direct implementation from a reconstruction or an illustrative example.</p><p>A paper-to-code project can contain clean Python and still fail to reproduce the original experiment. Reproduction requires the same data, timestamps, preprocessing, model definition, decision rules, execution assumptions, and evaluation protocol. Several of those ingredients are missing or ambiguous here.</p><p>The project uses five status labels, consistent with <code>docs/reproduction_matrix.md</code>:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!ek9L!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa21cfe13-02cc-4e78-8865-97268467b0bf_1568x664.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!ek9L!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa21cfe13-02cc-4e78-8865-97268467b0bf_1568x664.png 424w, /__u/substackcdn.com/image/fetch/$s_!ek9L!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa21cfe13-02cc-4e78-8865-97268467b0bf_1568x664.png 848w, /__u/substackcdn.com/image/fetch/$s_!ek9L!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa21cfe13-02cc-4e78-8865-97268467b0bf_1568x664.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ek9L!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa21cfe13-02cc-4e78-8865-97268467b0bf_1568x664.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!ek9L!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa21cfe13-02cc-4e78-8865-97268467b0bf_1568x664.png" width="1456" height="617" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a21cfe13-02cc-4e78-8865-97268467b0bf_1568x664.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:617,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:150524,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa21cfe13-02cc-4e78-8865-97268467b0bf_1568x664.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!ek9L!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa21cfe13-02cc-4e78-8865-97268467b0bf_1568x664.png 424w, /__u/substackcdn.com/image/fetch/$s_!ek9L!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa21cfe13-02cc-4e78-8865-97268467b0bf_1568x664.png 848w, /__u/substackcdn.com/image/fetch/$s_!ek9L!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa21cfe13-02cc-4e78-8865-97268467b0bf_1568x664.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ek9L!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa21cfe13-02cc-4e78-8865-97268467b0bf_1568x664.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>These labels determine how results should be interpreted and whether a value may legitimately be used as a test expectation. A hand-calculated fee or accounting invariant can be tested. The paper's final portfolio values cannot be treated that way because their inputs and protocol are unknown.</p><h3>2.1 What the paper specifies clearly</h3><p>Several elements are concrete enough to guide implementation:</p><ul><li><p>overlapping historical windows of length 100 with stride 1;</p></li><li><p>an attention-enhanced bidirectional LSTM as the primary forecasting model;</p></li><li><p>comparisons with LSTM, GRU, BiLSTM, Holt-Winters, ARIMA, HMM, and XGBoost;</p></li><li><p>a stated dropout setting of 20%;</p></li><li><p>MAE as the training loss;</p></li><li><p>RMSprop as the intended optimizer after correcting the paper's terminology;</p></li><li><p>dot-product attention, softmax normalization, and a weighted temporal sum;</p></li><li><p>transaction-fee scenarios for gold and Bitcoin.</p></li></ul><p>For example, <code>src/quantitative_trading_model/models/attention.py</code> can implement the displayed attention equations without knowing the authors' complete source code. Similarly, the configuration modules can preserve paper-inspired values as visible settings. This still does not guarantee numerical reproduction: the data, model dimensions, initialization, and training schedule remain unknown.</p><p>Configuration provenance is deliberately split across modules rather than concentrated in <code>ForecastConfig</code>:</p><ul><li><p><code>DataConfig</code> contains the window length, stride, cleaning, and data conventions;</p></li><li><p><code>ForecastConfig</code> contains model and training settings such as horizon, dropout, loss, optimizer, and hidden dimensions;</p></li><li><p><code>SizingConfig</code> contains the reconstructed position-sizing settings;</p></li><li><p><code>VaRConfig</code> contains risk-estimation and action settings;</p></li><li><p><code>ExecutionConfig</code> and <code>BacktestConfig</code> contain fees, slippage, cash, leverage, and liquidation assumptions;</p></li><li><p><code>FeeSensitivityConfig</code> contains the gold and Bitcoin fee grids.</p></li></ul><p>This separation prevents unrelated assumptions from being hidden behind a misleading single model configuration.</p><h3>2.2 Missing data and experiment identity</h3><p>The largest reproducibility gap is the input data. The paper refers to historical gold and Bitcoin prices but does not identify enough information to establish exactly what those series represent. Missing details include:</p><ul><li><p>the data vendor or source;</p></li><li><p>the specific gold instrument and Bitcoin market or index;</p></li><li><p>currency and unit conventions;</p></li><li><p>sampling frequency;</p></li><li><p>start and end dates;</p></li><li><p>timezone and timestamp rules;</p></li><li><p>the selected price field, such as close or adjusted close;</p></li><li><p>treatment of missing observations and duplicate timestamps;</p></li><li><p>market closures, splits, or other instrument-specific adjustments;</p></li><li><p>whether the assets were processed independently or jointly.</p></li></ul><p>&#8220;Gold price&#8221; and &#8220;Bitcoin price&#8221; are not unique datasets. Two valid historical series can have different timestamps, gaps, price fields, and price levels. A model trained on either series may be reasonable while producing metrics that cannot be compared with the paper's table.</p><p><code>README.md</code>, <code>data.py</code>, and <code>cli.py</code> therefore require local inputs rather than silently selecting an unidentified market-data source. Synthetic data are available only for demonstrations. This makes provenance explicit and avoids implying that an arbitrary local or synthetic series is equivalent to the paper's input.</p><h3>2.3 Missing preprocessing and target definition</h3><p>The paper does not fully state what the forecasting model predicts. Plausible interpretations include:</p><p>a raw next-day price;</p><p>a normalized price level;</p><p>a one-day return or percentage change;</p><p>a vector of the next three prices;</p><p>a day-three price produced by a multi-step procedure.</p><p>The prediction discussion is compatible with one-step forecasting, while the conclusion refers to forecasts over the next three days. These are different supervised-learning problems with different target shapes, losses, inverse transformations, and trading signals.</p><p>Preprocessing is also unclear. The paper does not specify whether prices were scaled, whether scaling used training observations only, or whether features such as volume, technical indicators, or macroeconomic variables were included. The implementation defaults to a price-based setup but exposes the target representation and horizon so that this ambiguity remains visible.</p><p>A reproduction must record at least:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!kVBQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2604b83-fd6c-49df-86e0-9801756da914_1566x544.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!kVBQ!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2604b83-fd6c-49df-86e0-9801756da914_1566x544.png 424w, /__u/substackcdn.com/image/fetch/$s_!kVBQ!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2604b83-fd6c-49df-86e0-9801756da914_1566x544.png 848w, /__u/substackcdn.com/image/fetch/$s_!kVBQ!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2604b83-fd6c-49df-86e0-9801756da914_1566x544.png 1272w, /__u/substackcdn.com/image/fetch/$s_!kVBQ!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2604b83-fd6c-49df-86e0-9801756da914_1566x544.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!kVBQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2604b83-fd6c-49df-86e0-9801756da914_1566x544.png" width="1456" height="506" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d2604b83-fd6c-49df-86e0-9801756da914_1566x544.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:506,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:112924,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2604b83-fd6c-49df-86e0-9801756da914_1566x544.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!kVBQ!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2604b83-fd6c-49df-86e0-9801756da914_1566x544.png 424w, /__u/substackcdn.com/image/fetch/$s_!kVBQ!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2604b83-fd6c-49df-86e0-9801756da914_1566x544.png 848w, /__u/substackcdn.com/image/fetch/$s_!kVBQ!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2604b83-fd6c-49df-86e0-9801756da914_1566x544.png 1272w, /__u/substackcdn.com/image/fetch/$s_!kVBQ!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2604b83-fd6c-49df-86e0-9801756da914_1566x544.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>2.4 Missing Att-BiLSTM architecture details</h3><p>Att-BiLSTM identifies a model family, not a complete architecture. The paper does not reliably specify:</p><ul><li><p>the number of recurrent layers;</p></li><li><p>hidden-state dimensions;</p></li><li><p>whether attention receives concatenated forward and backward states;</p></li><li><p>how the query vector is constructed;</p></li><li><p>dense-layer dimensions and activations;</p></li><li><p>output shape and activation;</p></li><li><p>recurrent versus ordinary dropout;</p></li><li><p>initialization;</p></li><li><p>learning-rate and optimizer settings;</p></li><li><p>batch size, epoch count, and stopping rule;</p></li><li><p>random seed;</p></li><li><p>whether each asset receives a separate model.</p></li></ul><p><code>models/forecasters.py</code> consequently provides a configurable paper-inspired model. Its learned attention query is an explicit reconstruction: the paper supplies a symbol <code>q</code> but does not say how that vector is obtained. The implementation is useful and inspectable, but it is not a verified transcription of the authors' architecture.</p><p>The same caution applies to the benchmarks. Listing LSTM, GRU, BiLSTM, Holt-Winters, ARIMA, HMM, and XGBoost does not establish that their feature sets, capacities, tuning procedures, and horizons were comparable. The paper's statement about &#8220;consistent parameters&#8221; is insufficient to reconstruct a fair benchmark protocol.</p><h3>2.5 The position-sizing equations are incomplete</h3><p>Equations (5) through (8) in the position-management section are not reliably recoverable from the extracted material. Important ambiguities include:</p><ul><li><p>whether <code>p_i</code> denotes cash, asset units, or notional exposure;</p></li><li><p>the meaning of the maximum addition count <code>x</code>;</p></li><li><p>whether fees apply to one transaction or both sides;</p></li><li><p>whether decline statistics are signed or expressed as positive loss magnitudes;</p></li><li><p>the independent variable in the exponential expression;</p></li><li><p>whether additions follow every decline or only a forecast-confirmed condition.</p></li></ul><p><code>src/quantitative_trading_model/sizing.py</code> therefore does not claim to transcribe the equations. It implements a normalized exponential schedule with finite addition levels, cash limits, and exposure caps. This is a <strong>reconstructed exponential position-sizing rule</strong>: it captures the apparent idea that later additions receive larger weights, but it does not establish that the paper used the same formula or calibration.</p><p>Keeping this reconstruction in one module is useful. If the original equation images or source become available, the sizing implementation can be replaced without changing data preparation, forecasting, risk, or portfolio accounting.</p><h3>2.6 VaR is named but not operationally defined</h3><p>The paper says that VaR contributes to trading decisions but does not supply the settings needed for a unique procedure. Missing choices include:</p><ul><li><p>confidence level;</p></li><li><p>horizon;</p></li><li><p>estimation-window length;</p></li><li><p>historical, parametric, or Monte Carlo method;</p></li><li><p>portfolio-level versus asset-level returns;</p></li><li><p>treatment of overlapping multi-day returns;</p></li><li><p>maximum acceptable loss;</p></li><li><p>whether VaR blocks, scales, or merely reports a trade.</p></li></ul><p><code>src/quantitative_trading_model/risk.py</code> provides a rolling historical VaR reconstruction with explicit confidence, lookback, horizon, and action settings. It supports leakage-aware use, but it cannot independently guarantee temporal provenance: <code>HistoricalVaR</code> accepts caller-provided returns and does not infer timestamps. The surrounding caller or backtest must supply only returns known before the decision and must exclude the current or future return when appropriate.</p><p>Thus, a VaR value means &#8220;the result of this configured historical quantile on this supplied return window.&#8221; It is not a hidden setting recovered from the paper, and it should not be described as the paper's exact VaR method.</p><h3>2.7 The modified greedy algorithm is not reproducible from its name</h3><p>&#8220;Modified greedy algorithm&#8221; describes a family of procedures, not one precise policy. Reproduction would require the paper to define:</p><ul><li><p>candidate actions;</p></li><li><p>whether actions are buy, sell, hold, or additions;</p></li><li><p>the objective being maximized;</p></li><li><p>how forecasts become expected gains;</p></li><li><p>how fees enter the objective;</p></li><li><p>cash, exposure, and addition constraints;</p></li><li><p>how VaR affects feasibility;</p></li><li><p>whether one action or a multi-step plan is selected;</p></li><li><p>tie-breaking behavior.</p></li></ul><p><code>src/quantitative_trading_model/strategy.py</code> implements a transparent reconstruction: enumerate feasible candidates, estimate forecast-based benefit after fees, apply documented constraints, and select the highest-scoring candidate with deterministic tie-breaking. It must not be described as the paper's verified algorithm.</p><p>There is also an unresolved static integration issue. The generated <code>GreedyPolicy</code> adapter currently does not match the generated <code>VaRTradeFilter</code> interface: it attempts to pass <code>proposed_notional</code> and omits required exposure-related arguments, while the filter expects <code>proposed_quantity</code> together with current exposure and portfolio value. This boundary must be repaired before runtime integration is attempted. Until then, the strategy and risk modules should be treated as separately documented reconstructions rather than a verified end-to-end combination.</p><p>The temporal rule remains essential: realized future returns must never score a candidate. A decision may use forecasts, current cash, current holdings, current exposure, and pre-decision risk history, but not the price that will be observed after execution.</p><h3>2.8 Why the headline profits are not test targets</h3><p>The paper reports approximately:</p><ul><li><p><strong>$646</strong> from an initial <strong>$500</strong> gold allocation;</p></li><li><p><strong>$215,487</strong> from an initial <strong>$500</strong> Bitcoin allocation;</p></li><li><p>approximately <strong>$216,133</strong> from the combined <strong>$1,000</strong> allocation.</p></li></ul><p>These are <strong>reported claims</strong>, not reproduced results. <code>src/quantitative_trading_model/experiments.py</code> records the values as unverified references rather than using them as test assertions or expected outputs from synthetic demonstrations.</p><p>A test target is appropriate when its input and protocol are specified. For example, a hand-calculated accounting test can assert the cash effect of a purchase with a known fee. The paper's Bitcoin figure cannot be tested this way because the historical period, asset series, signal timestamps, sizing triggers, reinvestment rules, exposure limits, fee convention, slippage, and liquidation policy are unknown.</p><p>A local result that happens to resemble one of those values would not prove reproduction unless the underlying data and all decision and accounting rules also matched. The unusually large Bitcoin result may depend on a particular historical period, aggressive compounding, repeated averaging down, leverage-like exposure, or optimistic execution assumptions. It should not be generalized as evidence of future profitability.</p><h3>2.9 Benchmark metrics and fee figures are also unverified</h3><p>The paper reports forecast metrics such as RMSE, MAE, MAPE, and R-squared, and shows transaction-fee sensitivity figures. Those values cannot be confirmed from the extracted text alone.</p><p>Forecast metrics depend on the exact test observations, target scale, transformations, timestamps, MAPE zero policy, architecture, trained parameters, random seed, and stopping rule. Fee figures additionally depend on trade dates and quantities, fee side conventions, spread, slippage, action feasibility, initial holdings, reinvestment, and liquidation.</p><p>The implementation preserves the paper-inspired fee grids as inputs: gold rates from 0.2% through 1.2%, and Bitcoin rates from 1.5% through 3.5%. It cannot infer the missing y-values from the figures. A new fee sweep using identified local data is a valid experiment under the project's assumptions, not a recovered copy of the paper's chart.</p><h3>2.10 How to read the reproduction matrix</h3><p><code>docs/reproduction_matrix.md</code> records status at the method and claim level. Its distinctions are important:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!WF2W!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8379702-6a15-4d4c-bd2d-e07a404c903f_1558x1018.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!WF2W!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8379702-6a15-4d4c-bd2d-e07a404c903f_1558x1018.png 424w, /__u/substackcdn.com/image/fetch/$s_!WF2W!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8379702-6a15-4d4c-bd2d-e07a404c903f_1558x1018.png 848w, /__u/substackcdn.com/image/fetch/$s_!WF2W!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8379702-6a15-4d4c-bd2d-e07a404c903f_1558x1018.png 1272w, /__u/substackcdn.com/image/fetch/$s_!WF2W!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8379702-6a15-4d4c-bd2d-e07a404c903f_1558x1018.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!WF2W!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8379702-6a15-4d4c-bd2d-e07a404c903f_1558x1018.png" width="1456" height="951" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e8379702-6a15-4d4c-bd2d-e07a404c903f_1558x1018.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:951,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:211985,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8379702-6a15-4d4c-bd2d-e07a404c903f_1558x1018.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!WF2W!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8379702-6a15-4d4c-bd2d-e07a404c903f_1558x1018.png 424w, /__u/substackcdn.com/image/fetch/$s_!WF2W!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8379702-6a15-4d4c-bd2d-e07a404c903f_1558x1018.png 848w, /__u/substackcdn.com/image/fetch/$s_!WF2W!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8379702-6a15-4d4c-bd2d-e07a404c903f_1558x1018.png 1272w, /__u/substackcdn.com/image/fetch/$s_!WF2W!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8379702-6a15-4d4c-bd2d-e07a404c903f_1558x1018.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This matrix prevents every function from being presented with the same level of authority when some functions implement displayed equations and others embody missing decisions.</p><h3>2.11 Static review is not empirical reproduction</h3><p>The project raises two different verification questions:</p><p>Does the generated artifact appear structurally reviewable?</p><p>Does the implementation reproduce the paper's empirical behavior?</p><p>Static and semantic review can address parts of the first question. It can inspect Python parsing, expected symbols, configuration validation, tensor-shape contracts, attention normalization, temporal boundaries, budget invariants, and accounting formulas. It cannot establish that the paper's data, model, or financial results were reproduced.</p><p>The supplied local verification record did not pass completely. It reported <strong>placeholder-detection findings</strong> for <code>experiments.py</code> and <code>cli.py</code>, and <strong>Markdown-fence findings</strong> for the generated <code>README.md</code> and <code>docs/tutorial.md</code> code fields. These are generated-artifact quality findings. They are not runtime results and do not by themselves demonstrate a Python syntax failure or a financial-model failure.</p><p>Semantic code verification was skipped. Runtime execution and test-suite success therefore must not be claimed. In addition, some generated tests and orchestration functions appear to use APIs that are not fully aligned with the generated modules, including the strategy/VaR boundary described above. Those inconsistencies require a software-maintenance pass before execution is attempted.</p><p>The appropriate conclusion is limited: the project documents intended interfaces, assumptions, and invariants, but it is not an executed or empirically validated reproduction.</p><h3>2.12 Responsible interpretation policy</h3><p>When describing results from this project:</p><ul><li><p>say <strong>implemented</strong> for mechanics supported by the paper and the code contract;</p></li><li><p>say <strong>reconstructed</strong> for normalized sizing, historical VaR, the greedy policy, and other interpretations of missing details;</p></li><li><p>say <strong>illustrative</strong> for synthetic data and hand-calculated workflows;</p></li><li><p>say <strong>unavailable</strong> when the source lacks the information needed for implementation or validation;</p></li><li><p>say <strong>reported claim</strong> when quoting the paper's profits or metrics without independent validation;</p></li><li><p>do not call a run on a different local dataset an exact reproduction;</p></li><li><p>keep forecast metrics separate from portfolio metrics;</p></li><li><p>report dates, data provenance, target definition, fees, slippage, turnover, drawdown, and liquidation rules with every backtest;</p></li><li><p>treat static review as artifact evidence, not proof of execution, profitability, or paper fidelity.</p></li></ul><p>The next sections can discuss setup and implementation without blurring these categories. A modular reconstruction can still teach attention-based forecasting, streak statistics, bounded sizing, risk filtering, and fee-aware accounting while making clear which conclusions require the original data and experimental protocol.</p><h2>3. Project Setup and the Implementation Map</h2><p>The implementation is organized as a normal Python package rather than as one large research script. That choice matters because the paper combines several different responsibilities: data preparation, neural forecasting, benchmark models, streak analysis, position sizing, risk filtering, trading decisions, execution, and evaluation. Keeping those responsibilities in separate modules makes it easier to inspect assumptions and replace incomplete reconstructions later.</p><p>The project is also deliberately <strong>offline</strong>. It accepts local files or generated demonstration data, but it does not connect to an exchange, download market data, store credentials, or submit orders. The resulting code is suitable for educational research and backtest prototyping, not for live trading.</p><h3>3.1 The <code>src</code> layout</h3><p>The package uses the conventional <code>src</code> layout:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;56a56814-ba00-4cb5-a6f0-e2c28fac3ceb&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">quantitative-trading-model/
&#9500;&#9472;&#9472; pyproject.toml
&#9500;&#9472;&#9472; README.md
&#9500;&#9472;&#9472; src/
&#9474;   &#9492;&#9472;&#9472; quantitative_trading_model/
&#9474;       &#9500;&#9472;&#9472; __init__.py
&#9474;       &#9500;&#9472;&#9472; config.py
&#9474;       &#9500;&#9472;&#9472; data.py
&#9474;       &#9500;&#9472;&#9472; metrics.py
&#9474;       &#9500;&#9472;&#9472; streaks.py
&#9474;       &#9500;&#9472;&#9472; sizing.py
&#9474;       &#9500;&#9472;&#9472; risk.py
&#9474;       &#9500;&#9472;&#9472; strategy.py
&#9474;       &#9500;&#9472;&#9472; backtest.py
&#9474;       &#9500;&#9472;&#9472; experiments.py
&#9474;       &#9500;&#9472;&#9472; cli.py
&#9474;       &#9492;&#9472;&#9472; models/
&#9474;           &#9500;&#9472;&#9472; attention.py
&#9474;           &#9500;&#9472;&#9472; forecasters.py
&#9474;           &#9492;&#9472;&#9472; benchmarks.py
&#9500;&#9472;&#9472; tests/
&#9492;&#9472;&#9472; docs/
    &#9500;&#9472;&#9472; tutorial.md
    &#9492;&#9472;&#9472; reproduction_matrix.md</code></pre></div><p>With this layout, the importable package lives under <code>src/quantitative_trading_model</code>, while documentation and tests remain outside the runtime package. The <code>pyproject.toml</code> file tells setuptools where to find the package:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;toml&quot;,&quot;nodeId&quot;:&quot;de12be66-aaca-4bea-81f2-67ec24aa7df5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-toml">[tool.setuptools]
package-dir = { "" = "src" }
include-package-data = true

[tool.setuptools.packages.find]
where = ["src"]</code></pre></div><p>This avoids a common development problem in which tests accidentally import a loose source directory instead of the package that users will install. It also gives the project a clear boundary: files under <code>src/quantitative_trading_model</code> are implementation modules, while the root-level documentation explains how to use and interpret them.</p><h3>3.2 Core dependencies and optional model dependencies</h3><p>The package keeps its required dependencies small:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;toml&quot;,&quot;nodeId&quot;:&quot;95f987df-6ac6-4633-9789-55ec7c56609e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-toml">[project]
requires-python = "&gt;=3.10"
dependencies = [
    "numpy&gt;=1.24",
    "pandas&gt;=2.0"
]</code></pre></div><p>NumPy and pandas support the parts of the project that do not require a deep-learning framework: local price tables, timestamp handling, return calculations, streak statistics, metrics, risk calculations, and portfolio accounting.</p><p>The heavier forecasting libraries are optional:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;toml&quot;,&quot;nodeId&quot;:&quot;bc831663-cc6d-4f11-af38-39d4dba9eccd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-toml">[project.optional-dependencies]
deep-learning = [
    "torch&gt;=2.0"
]
classical = [
    "statsmodels&gt;=0.14",
    "pmdarima&gt;=2.0",
    "hmmlearn&gt;=0.3",
    "xgboost&gt;=2.0"
]</code></pre></div><p>This split reflects the paper's structure. The Att-BiLSTM requires PyTorch, but a reader who only wants to study the return convention, normalized exponential sizing reconstruction, historical VaR, or portfolio accounting should not need to install PyTorch. Similarly, Holt-Winters, ARIMA, HMM, and XGBoost benchmarks can be enabled only when those comparisons are needed.</p><p>The optional dependency design also makes missing capabilities visible. A benchmark adapter should report that a library is unavailable rather than silently substituting a different model. That distinction matters when comparing results with the paper: an omitted benchmark is not the same as a benchmark that produced a poor score.</p><p>The package metadata also defines a console entry point:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;toml&quot;,&quot;nodeId&quot;:&quot;e1aa41c6-381d-4899-85e0-cd39afbf9bea&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-toml">[project.scripts]
quant-trading-model = "quantitative_trading_model.cli:main"</code></pre></div><p>After local installation, this exposes the offline command-line interface as <code>quant-trading-model</code>. The CLI can validate local CSV files, generate deterministic synthetic demonstrations, and dispatch configured research experiments. It does not create a network client or read credentials.</p><h3>3.3 Side-effect-free package imports</h3><p>The package initializer intentionally exposes only lightweight metadata and configuration types:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f8ee8a83-16e3-4447-ade4-702a85369e0b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">"""Offline research components for the quantitative trading model paper."""

from importlib import metadata

try:
    __version__ = metadata.version("quantitative-trading-model")
except metadata.PackageNotFoundError:
    __version__ = "0.1.0"

from .config import (
    BacktestConfig,
    DataConfig,
    ExecutionConfig,
    FeeSensitivityConfig,
    ForecastConfig,
    SizingConfig,
    SplitConfig,
    VaRConfig,
    validate_config,
)</code></pre></div><p>Importing the package does not load a CSV, train a model, create a portfolio, or contact an external service. The initializer also does not import PyTorch, statsmodels, <code>hmmlearn</code>, or XGBoost. Those imports are deferred to the modules that need them.</p><p>That design has two practical benefits:</p><p><strong>Lightweight use:</strong> core modules remain usable in an environment without optional modeling libraries.</p><p><strong>Predictable imports:</strong> simply writing <code>import quantitative_trading_model</code> cannot start a computation or produce an external side effect.</p><p>The public exports are deliberately narrow:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5a36e762-fb86-49ba-a52a-0a8eb748f671&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">__all__ = [
    "__version__",
    "BacktestConfig",
    "DataConfig",
    "ExecutionConfig",
    "FeeSensitivityConfig",
    "ForecastConfig",
    "SizingConfig",
    "SplitConfig",
    "VaRConfig",
    "validate_config",
]</code></pre></div><p>The model, data, strategy, and backtest classes are imported from their own modules when required. This keeps the top-level API stable without forcing every user to install every optional dependency.</p><h3>3.4 Where each research responsibility lives</h3><p>The package map follows the paper's conceptual pipeline:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!aq85!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0996af8d-98d9-422f-9f4b-8ac6fc43de2d_1406x1372.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!aq85!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0996af8d-98d9-422f-9f4b-8ac6fc43de2d_1406x1372.png 424w, /__u/substackcdn.com/image/fetch/$s_!aq85!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0996af8d-98d9-422f-9f4b-8ac6fc43de2d_1406x1372.png 848w, /__u/substackcdn.com/image/fetch/$s_!aq85!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0996af8d-98d9-422f-9f4b-8ac6fc43de2d_1406x1372.png 1272w, /__u/substackcdn.com/image/fetch/$s_!aq85!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0996af8d-98d9-422f-9f4b-8ac6fc43de2d_1406x1372.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!aq85!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0996af8d-98d9-422f-9f4b-8ac6fc43de2d_1406x1372.png" width="1406" height="1372" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0996af8d-98d9-422f-9f4b-8ac6fc43de2d_1406x1372.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1372,&quot;width&quot;:1406,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:324804,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0996af8d-98d9-422f-9f4b-8ac6fc43de2d_1406x1372.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!aq85!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0996af8d-98d9-422f-9f4b-8ac6fc43de2d_1406x1372.png 424w, /__u/substackcdn.com/image/fetch/$s_!aq85!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0996af8d-98d9-422f-9f4b-8ac6fc43de2d_1406x1372.png 848w, /__u/substackcdn.com/image/fetch/$s_!aq85!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0996af8d-98d9-422f-9f4b-8ac6fc43de2d_1406x1372.png 1272w, /__u/substackcdn.com/image/fetch/$s_!aq85!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0996af8d-98d9-422f-9f4b-8ac6fc43de2d_1406x1372.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This separation is not merely a software-style preference. It prevents an ambiguous paper detail from being hidden inside an unrelated function. For example, the return denominator belongs in <code>streaks.py</code>, while the execution fee convention belongs in <code>backtest.py</code>. A reader can therefore change one assumption without accidentally changing every layer of the research pipeline.</p><h3>3.5 Documentation, tests, and the reproduction matrix</h3><p>The non-runtime files serve different purposes:</p><ul><li><p><strong>`README.md`</strong> gives installation instructions, local CSV conventions, CLI examples, project warnings, and a high-level architecture map.</p></li><li><p><strong>`docs/tutorial.md`</strong> explains the paper-to-code translation in detail, including equations, tensor shapes, timing rules, and reconstruction decisions.</p></li><li><p><strong>`docs/reproduction_matrix.md`</strong> tracks each paper component as implemented, reconstructed, illustrative, unavailable, or an unverified reported claim.</p></li><li><p><strong>`tests/`</strong> contains hand-calculable checks for data boundaries, streak semantics, sizing budgets, VaR decisions, model shapes, transaction fees, and portfolio-value identity.</p></li></ul><p>The reproduction matrix is especially useful for this paper because the source does not fully specify equations (5)&#8211;(8), VaR operation, the greedy policy, or the original backtest protocol. Instead of allowing those gaps to disappear into code, the matrix records what the implementation can support and what still requires the original data or source material.</p><p>For example, the matrix treats the following differently:</p><ul><li><p>the 100-step window as a directly implementable paper detail;</p></li><li><p>the learned attention query as a reconstruction;</p></li><li><p>synthetic prices as an illustration;</p></li><li><p>the reported <code>$646</code> gold and <code>$215,487</code> Bitcoin outcomes as unverified claims.</p></li></ul><p>That classification should remain visible in experiment output and documentation. A clean package installation does not establish empirical reproduction, and static code checks do not establish that the paper's financial results are correct.</p><h3>3.6 A practical installation boundary</h3><p>A typical local setup is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;ac509ac6-f821-441a-a0a2-96bd3bfa6393&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">python -m venv .venv
python -m pip install --upgrade pip
python -m pip install -e .</code></pre></div><p>The core installation is enough to explore data preparation, streak analysis, risk calculations, sizing, metrics, and portfolio accounting. To use the PyTorch forecasters, install the deep-learning extra:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;131adebc-8d3b-4a9b-bcea-dbf98fcd2ad8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">python -m pip install -e '.[deep-learning]'</code></pre></div><p>The classical benchmark dependencies are declared under the <code>classical</code> extra in <code>pyproject.toml</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;a1677b7d-4eb7-4220-b309-d5eab54b0798&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">python -m pip install -e '.[classical]'</code></pre></div><p>The exact installed command surface can be inspected with:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;cad7f9a4-908e-4e31-937b-6c9b0f228030&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">quant-trading-model --help</code></pre></div><p>These commands install or invoke local research tooling only. They do not download the paper's missing gold or Bitcoin data. The user must supply a provenance-documented local dataset before treating any result as an empirical experiment.</p><h3>3.7 What this structure does&#8212;and does not&#8212;guarantee</h3><p>The package structure provides useful engineering guarantees:</p><ul><li><p>optional modeling libraries do not prevent core imports;</p></li><li><p>configuration choices are visible rather than scattered through scripts;</p></li><li><p>data, forecast, strategy, execution, and evaluation responsibilities are separated;</p></li><li><p>documentation can distinguish direct implementation from reconstruction;</p></li><li><p>tests can check local invariants without requiring the paper's unavailable data.</p></li></ul><p>It does <strong>not</strong> guarantee that the paper's reported results can be reproduced. The missing dataset, dates, preprocessing, exact model architecture, VaR parameters, greedy rules, position-sizing equations, and execution protocol remain external requirements. The next sections use this package structure to examine those components one at a time, beginning with validated configuration and local data preparation.</p><h2>4. Make Ambiguity Explicit with Validated Configuration</h2><p>A research paper can name an algorithm without specifying every value required to run it. This paper gives some concrete settings, including a 100-observation history, one-observation stride, 20% dropout, MAE loss, RMSprop optimization, and several transaction-fee scenarios. It leaves other choices unclear, including the forecast horizon, recurrent-layer dimensions, attention query, VaR policy, position-sizing parameters, execution timing, and liquidation rule.</p><p>The implementation records these choices in dataclasses defined in <code>src/quantitative_trading_model/config.py</code>. Centralized configuration does not recover the paper's missing data or prove that its reported results can be reproduced. It does make assumptions visible, validates basic invariants, and provides metadata that can be saved with an experiment.</p><p>There are three important categories of settings:</p><ul><li><p><strong>Paper-inspired settings</strong> are stated directly or strongly suggested by the paper.</p></li><li><p><strong>Reconstruction settings</strong> are operational choices needed because the paper is incomplete.</p></li><li><p><strong>Implementation safety settings</strong> constrain this educational simulator even when the paper does not describe equivalent controls.</p></li></ul><p>The distinction matters because a field can be present in configuration without being fully wired into every generated runtime module. The status of each important setting should therefore be checked against both <code>config.py</code> and the module that consumes it.</p><h3>4.1 Configuration objects and their responsibilities</h3><p>The configuration module defines separate dataclasses for the main layers of the pipeline:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!X16U!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b94666b-df78-4138-9e5f-7e33b7fdcc2d_1420x1120.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!X16U!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b94666b-df78-4138-9e5f-7e33b7fdcc2d_1420x1120.png 424w, /__u/substackcdn.com/image/fetch/$s_!X16U!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b94666b-df78-4138-9e5f-7e33b7fdcc2d_1420x1120.png 848w, /__u/substackcdn.com/image/fetch/$s_!X16U!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b94666b-df78-4138-9e5f-7e33b7fdcc2d_1420x1120.png 1272w, /__u/substackcdn.com/image/fetch/$s_!X16U!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b94666b-df78-4138-9e5f-7e33b7fdcc2d_1420x1120.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!X16U!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b94666b-df78-4138-9e5f-7e33b7fdcc2d_1420x1120.png" width="1420" height="1120" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1b94666b-df78-4138-9e5f-7e33b7fdcc2d_1420x1120.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1120,&quot;width&quot;:1420,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:280178,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b94666b-df78-4138-9e5f-7e33b7fdcc2d_1420x1120.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!X16U!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b94666b-df78-4138-9e5f-7e33b7fdcc2d_1420x1120.png 424w, /__u/substackcdn.com/image/fetch/$s_!X16U!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b94666b-df78-4138-9e5f-7e33b7fdcc2d_1420x1120.png 848w, /__u/substackcdn.com/image/fetch/$s_!X16U!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b94666b-df78-4138-9e5f-7e33b7fdcc2d_1420x1120.png 1272w, /__u/substackcdn.com/image/fetch/$s_!X16U!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b94666b-df78-4138-9e5f-7e33b7fdcc2d_1420x1120.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>For example, the paper-inspired data and forecasting defaults are:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;cc438f5e-ccea-48c9-807c-8bbe45ad64ff&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass
class DataConfig:
    timestamp_column: str = "timestamp"
    price_column: str = "price"
    window_length: int = 100
    stride: int = 1
    return_definition: str = "simple"


@dataclass
class ForecastConfig:
    model_name: str = "att_bilstm"
    horizon: int = 1
    target_type: str = "scaled_price"
    hidden_size: int = 32
    num_layers: int = 1
    dropout: float = 0.20
    learning_rate: float = 0.001
    optimizer: str = "rmsprop"
    loss: str = "mae"
    learned_attention_query: bool = True</code></pre></div><p>The values with the clearest connection to the paper are:</p><ul><li><p><code>window_length=100</code>: the reported historical window length;</p></li><li><p><code>stride=1</code>: the reported one-observation shift between windows;</p></li><li><p><code>dropout=0.20</code>: the paper-inspired dropout rate;</p></li><li><p><code>loss="mae"</code>: the reported training loss;</p></li><li><p><code>optimizer="rmsprop"</code>: a correction to the paper's terminology, since RMSprop is an optimizer rather than an activation function;</p></li><li><p><code>model_name="att_bilstm"</code>: the paper's claimed primary forecasting model.</p></li></ul><p>The hidden size, number of layers, learning rate, batch size, epoch count, early-stopping policy, and learned-query choice are not established by the paper. They are configurable reconstruction choices and should be recorded with any result.</p><h3>4.2 Make the forecast horizon explicit</h3><p>The paper's prediction discussion is not consistent about horizon. Its model description and metric table can be read as a one-step forecasting setup, while the conclusion refers to forecasts for the next three days. The source does not establish whether this means a three-element output vector, repeated one-step forecasts, or a single prediction for the third day.</p><p>The configuration makes the choice visible:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c4c4e94e-a75f-4f06-9661-129ceb939542&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">one_day = ForecastConfig(horizon=1)
three_days = ForecastConfig(horizon=3)</code></pre></div><p>The horizon affects more than the output layer. It changes target construction, prediction metrics, signal aggregation, and execution timing. A one-step model should not be called a three-day model merely because its output is used repeatedly. The selected horizon should be stored in experiment metadata.</p><p><code>target_type</code> is similarly uncertain. The paper does not clearly say whether it predicts raw prices, scaled prices, returns, or percentage changes. The generated default, <code>"scaled_price"</code>, is a practical implementation choice and not a verified transcription.</p><h3>4.3 Position-sizing configuration: metadata versus enforcement</h3><p>The position-management equations are among the least reproducible parts of the paper. Equations (5) and (6) appear to describe recursive additions intended to compensate for declines and fees, but their notation is corrupted. Equation (7) does not clearly identify the independent variable of the exponential, and Equation (8) appears to normalize position scores into allocations.</p><p><code>SizingConfig</code> records a possible reconstruction:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;21be9b28-94ed-41c3-ab9b-43d9444851ce&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass
class SizingConfig:
    maximum_additions: int = 5
    initial_level: int = 1
    growth_rate: float = 0.50
    score_scale: float = 1.0
    allocation_fraction: float = 1.0
    initial_position_fraction: float = 0.20
    maximum_exposure_fraction: float = 1.0
    allow_reuse: bool = False
    calibrate_from_percentiles: bool = False
    recovery_percentile: float = 0.50
    decline_percentile: float = 0.10
    fee_buffer: float = 0.0</code></pre></div><p>The intended normalized reconstruction is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;8061810d-bc3a-4247-b3df-38ba8da93314&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">score_i      = score_scale * exp(growth_rate * i)
weight_i     = score_i / sum(score_j)
allocation_i = budget * weight_i</code></pre></div><p>However, these configuration fields should not be confused with runtime guarantees. In the generated project, <code>sizing.py</code> builds schedules through <code>build_exponential_schedule()</code> and enforces schedule-level properties through <code>SizingSchedule</code>, including finite allocations, normalized weights, a total budget, and one-time level consumption. <code>validate_config()</code> validates <code>maximum_additions</code> and related fractions, but it does not construct a schedule or verify its normalized weights.</p><p>Likewise, <code>allocation_fraction</code>, <code>maximum_exposure_fraction</code>, <code>allow_reuse</code>, and <code>fee_buffer</code> are currently configuration or documentation choices rather than completely wired controls in the generated sizing path. The schedule builder receives its own <code>budget</code> and growth arguments. It does not automatically read every corresponding field from <code>SizingConfig</code>, and <code>fee_buffer</code> is not applied to an allocation. A caller must pass the intended values into the sizing module explicitly or add the missing integration before treating them as active controls.</p><p>The safe interpretation is therefore:</p><ul><li><p><code>maximum_additions</code> documents the intended finite schedule size;</p></li><li><p>schedule normalization and budget checks belong to <code>SizingSchedule</code> and <code>sizing.py</code>;</p></li><li><p>reuse prevention is enforced when a constructed schedule marks a level as used;</p></li><li><p>exposure and cash caps must be applied by the sizer, strategy, or simulator that receives them;</p></li><li><p>the incomplete paper equations remain a reconstruction, not an exact implementation.</p></li></ul><h3>4.4 VaR configuration is an operational reconstruction</h3><p>The paper names VaR but does not define its confidence level, horizon, estimation window, return distribution, or action threshold. A usable risk filter must choose all of these values, so <code>VaRConfig</code> makes them explicit:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;fa36ea69-bb5c-44f9-902b-0a3c2ed4edbd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass
class VaRConfig:
    enabled: bool = True
    confidence: float = 0.95
    horizon: int = 1
    lookback: int = 100
    method: str = "historical"
    action: str = "scale"
    max_var_fraction: float = 0.05
    insufficient_history: str = "allow"
    min_observations: int = 30</code></pre></div><p>These defaults describe a rolling historical-VaR interpretation, not the paper's verified settings:</p><ul><li><p><code>confidence=0.95</code> selects a 95% lower-tail confidence convention;</p></li><li><p><code>horizon=1</code> uses one period by default;</p></li><li><p><code>lookback=100</code> limits the estimation history;</p></li><li><p><code>method="historical"</code> avoids assuming normally distributed returns;</p></li><li><p><code>action="scale"</code> permits reducing an order when projected risk is too high;</p></li><li><p><code>max_var_fraction=0.05</code> expresses the risk limit as a fraction of portfolio value;</p></li><li><p><code>insufficient_history="allow"</code> defines behavior before enough observations exist.</p></li></ul><p>The generated <code>risk.py</code> contains the actual historical estimator and filter. It must receive returns that were available before the decision timestamp. Configuration alone cannot enforce timestamp ordering. Also, because the paper provides no numerical VaR settings, every backtest should report these values as reconstruction metadata.</p><p>A stricter configuration can block trades until sufficient history exists:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a4a0a4df-0be5-4af6-9eb1-629a2f965166&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">risk = VaRConfig(
    confidence=0.99,
    lookback=252,
    min_observations=100,
    action="/__u/onepagecode.substack.com/block",
    insufficient_history="block",
)</code></pre></div><h3>4.5 Execution and liquidation: record the intended contract carefully</h3><p>The paper does not specify whether a signal uses the same closing price, the next opening price, or another execution price. It also does not say clearly whether a fee applies to purchases, sales, or both. The intended execution configuration is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5e0a9ed7-9dea-46e3-add2-2fe0ce9520f9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass
class ExecutionConfig:
    fee_rate: float = 0.002
    fee_on_buy: bool = True
    fee_on_sell: bool = True
    slippage_rate: float = 0.0
    allow_fractional_units: bool = True
    allow_short: bool = False
    max_leverage: float = 1.0
    execute_next_bar: bool = True
    reject_insufficient_cash: bool = True</code></pre></div><p>These fields describe the desired research contract, but the generated <code>backtest.py</code> does not fully honor all of them as written. In particular:</p><ul><li><p>the simulator applies <code>fee_rate</code> to both buys and sells unconditionally; it does not consult <code>fee_on_buy</code> or <code>fee_on_sell</code>;</p></li><li><p>unaffordable buys are clipped to the affordable quantity rather than being controlled by <code>reject_insufficient_cash</code>;</p></li><li><p>the simulator uses its own <code>allow_leverage</code> behavior and does not enforce <code>max_leverage</code> from <code>ExecutionConfig</code>;</p></li><li><p>the default simulator is nevertheless long-only and non-leveraged through its separate defaults, so cash and holdings are intended to remain nonnegative;</p></li><li><p>next-bar execution is enforced by requiring the fill timestamp to be strictly later than the signal timestamp.</p></li></ul><p>This distinction is important. A configuration field is not proof that the current runtime path consumes it. Until the wiring is corrected, results should state the simulator's actual behavior: two-sided fees, optional slippage, clipping of unaffordable buys, no shorting by default, and strict later-bar execution.</p><p><code>BacktestConfig</code> contains additional lifecycle choices:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;67da6719-bcbb-46d1-84c0-3b25e9aec78d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass
class BacktestConfig:
    initial_cash: float = 500.0
    initial_quantity: float = 0.0
    liquidate_at_end: bool = True
    risk_free_rate: float = 0.0
    periods_per_year: int = 252
    currency: str = "USD"</code></pre></div><p>The <code>$500</code> default resembles the paper's stated per-asset starting allocation, but it does not reproduce the paper's final values. Dates, trades, fees, holdings, and liquidation accounting remain unspecified. <code>liquidate_at_end=True</code> is an explicit implementation choice: it may sell remaining holdings and charge the simulator's sell-side fee. A mark-to-market result without liquidation would be different.</p><h3>4.6 Fee grids use decimal conventions</h3><p>The paper lists percentages, while the configuration stores decimal rates:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Dp71!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2078457-e10d-41be-833d-f48dde76b542_1096x214.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Dp71!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2078457-e10d-41be-833d-f48dde76b542_1096x214.png 424w, /__u/substackcdn.com/image/fetch/$s_!Dp71!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2078457-e10d-41be-833d-f48dde76b542_1096x214.png 848w, /__u/substackcdn.com/image/fetch/$s_!Dp71!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2078457-e10d-41be-833d-f48dde76b542_1096x214.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Dp71!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2078457-e10d-41be-833d-f48dde76b542_1096x214.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Dp71!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2078457-e10d-41be-833d-f48dde76b542_1096x214.png" width="1096" height="214" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e2078457-e10d-41be-833d-f48dde76b542_1096x214.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:214,&quot;width&quot;:1096,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:37329,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2078457-e10d-41be-833d-f48dde76b542_1096x214.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Dp71!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2078457-e10d-41be-833d-f48dde76b542_1096x214.png 424w, /__u/substackcdn.com/image/fetch/$s_!Dp71!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2078457-e10d-41be-833d-f48dde76b542_1096x214.png 848w, /__u/substackcdn.com/image/fetch/$s_!Dp71!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2078457-e10d-41be-833d-f48dde76b542_1096x214.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Dp71!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2078457-e10d-41be-833d-f48dde76b542_1096x214.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>The conversion is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;1c53232b-97e6-4de2-a8a1-141a2c59f2e4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">0.2% = 0.2 / 100 = 0.002
1.5% = 1.5 / 100 = 0.015</code></pre></div><p>The corresponding defaults are:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;05e10b5f-f1cc-4c5b-a458-03d8afc8f60a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass
class FeeSensitivityConfig:
    gold_rates: tuple[float, ...] = (
        0.002, 0.004, 0.006, 0.008, 0.010, 0.012
    )
    bitcoin_rates: tuple[float, ...] = (
        0.015, 0.020, 0.025, 0.030, 0.035
    )
    apply_to_buys: bool = True
    apply_to_sells: bool = True
    hold_signals_fixed: bool = True</code></pre></div><p>Using <code>0.2</code> for a 0.2% fee would represent a 20% rate. The validator checks that rates are finite and lie in the interval <code>[0, 1]</code>, but it cannot determine whether a user has confused percentage points with decimal rates. In other words, <code>0.2</code> is accepted as a numerically valid rate even though it is probably a unit mistake for this experiment. The decimal convention must therefore be documented and reviewed by the caller.</p><p>The fee-side flags are also not fully honored by the generated simulator, which currently charges both sides. <code>hold_signals_fixed</code> describes the sensitivity protocol rather than changing the backtest automatically. If fees affect feasibility or greedy ranking, actions may need to be regenerated for each fee scenario.</p><h3>4.7 What <code>validate_config()</code> checks</h3><p>The public validator accepts a complete aggregate or individual components:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c8cbd5f2-9df2-4391-aaca-367b222e51a6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from quantitative_trading_model.config import (
    ResearchConfig,
    ForecastConfig,
    validate_config,
)

config = ResearchConfig(
    forecast=ForecastConfig(
        model_name="att_bilstm",
        horizon=3,
        dropout=0.20,
        loss="mae",
        optimizer="rmsprop",
    )
)

validate_config(config)</code></pre></div><p>Component validators check basic constraints such as:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f042b3e2-1a22-42b6-bba1-77f20b857783&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">_require(config.window_length &gt; 0, "data.window_length must be positive")
_require(0 &lt;= config.dropout &lt; 1, "forecast.dropout must be in [0, 1)")
_require(0 &lt; config.confidence &lt; 1, "var.confidence must be in (0, 1)")
_require(config.fee_rate &gt;= 0, "execution.fee_rate cannot be negative")
_require(config.maximum_additions &gt; 0, "sizing.maximum_additions must be positive")</code></pre></div><p>This validation rejects invalid configuration values such as nonpositive horizons, dropout outside <code>[0, 1)</code>, unsupported optimizer or loss names, split fractions that do not sum to one, negative fees or slippage, invalid VaR actions, and empty fee grids.</p><p>It does <strong>not</strong> perform every runtime or semantic check. Specifically:</p><ul><li><p>it does not construct a sizing schedule or check that its weights sum to one;</p></li><li><p>it does not validate the contents of a price file;</p></li><li><p>it does not prove that every configuration field is consumed by the backtest;</p></li><li><p>it does not detect a percentage-unit mistake such as entering <code>0.2</code> for <code>0.2%</code>;</p></li><li><p>it does not establish that VaR inputs are timestamp-safe;</p></li><li><p>it does not verify that generated experiments reproduce the paper.</p></li></ul><p>Schedule-level invariants are handled in <code>sizing.py</code>, data invariants in <code>data.py</code>, and accounting invariants in <code>backtest.py</code>. Those layers must be reviewed together with configuration validation.</p><p>The validator also rejects supplying both a complete <code>ResearchConfig</code> and individual component overrides. This prevents a caller from validating one configuration while accidentally running another.</p><h3>4.8 Record a complete experiment configuration</h3><p>A complete experiment can combine paper-inspired and reconstructed settings:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;686519f4-a104-4ae9-a136-45a0f4da8eb8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from quantitative_trading_model.config import (
    ResearchConfig,
    ForecastConfig,
    VaRConfig,
    validate_config,
)

config = ResearchConfig(
    forecast=ForecastConfig(
        model_name="att_bilstm",
        horizon=3,
        hidden_size=32,
        dropout=0.20,
        loss="mae",
        optimizer="rmsprop",
    ),
    var=VaRConfig(
        confidence=0.95,
        horizon=1,
        lookback=100,
        action="/__u/onepagecode.substack.com/scale",
    ),
)

validate_config(config)

metadata = {
    "window_length": config.data.window_length,
    "stride": config.data.stride,
    "forecast_horizon": config.forecast.horizon,
    "dropout": config.forecast.dropout,
    "loss": config.forecast.loss,
    "optimizer": config.forecast.optimizer,
    "var_confidence": config.var.confidence,
    "var_lookback": config.var.lookback,
    "fee_rate": config.execution.fee_rate,
    "liquidate_at_end": config.backtest.liquidate_at_end,
}</code></pre></div><p>This record combines direct paper-inspired values with reconstruction choices. It should be saved alongside forecast metrics, portfolio metrics, data provenance, and software-version information.</p><p><code>ResearchConfig</code> is available from <code>quantitative_trading_model.config</code>. It is not currently included in the package-level exports in <code>quantitative_trading_model/__init__.py</code>, so importing it from the configuration module is the reliable documented path.</p><h3>4.9 Configuration is necessary but not sufficient</h3><p>The configuration contract solves an engineering problem: it makes assumptions visible, validates basic values, and gives experiments a stable record. It does not solve the paper's missing-information problem. It cannot recover:</p><ul><li><p>the unidentified gold and Bitcoin datasets;</p></li><li><p>the original dates, frequency, or test boundary;</p></li><li><p>the exact attention query and network architecture;</p></li><li><p>the corrupted recursive position-sizing equations;</p></li><li><p>the paper's VaR settings or action policy;</p></li><li><p>the modified greedy algorithm's objective;</p></li><li><p>the original execution, fee, and liquidation protocol;</p></li><li><p>the numeric values behind the fee-sensitivity figures.</p></li></ul><p>The approximately <code>$646</code> gold result, approximately <code>$215,487</code> Bitcoin result, combined approximately <code>$216,133</code> result, and reported benchmark metrics therefore remain unverified paper claims. Matching visible defaults does not reproduce those results.</p><p>The generated project also has known API inconsistencies between some configuration fields and downstream modules, and semantic code verification was skipped. This section consequently describes the intended configuration contract and the actual validation boundaries; it does not claim that the generated code was executed or that every setting passed through the complete runtime pipeline.</p><p>The safest interpretation is that configuration turns an incomplete paper into an auditable research specification. It identifies what comes from the source, what is reconstructed, what is merely illustrative, and which safety controls belong to this educational implementation rather than to the original paper.</p><h2>5. Load Local Prices Without Assuming the Paper&#8217;s Missing Dataset</h2><p>The paper says that it uses historical gold and Bitcoin prices, but it does not identify the data source, instrument, currency, frequency, date range, or exact columns. That omission is important: even a correct implementation can produce different results if it uses a different gold contract, Bitcoin exchange, timezone, sampling interval, or adjusted-price convention.</p><p>The generated project therefore does not silently select a market-data provider. Instead, it accepts a local CSV file or a caller-supplied pandas object. This keeps the experiment offline and makes data provenance the reader&#8217;s responsibility. Before treating a result as empirical evidence, record where the file came from, which instrument it represents, and how it was prepared.</p><h3>5.1 The minimum data contract</h3><p>The implementation reduces each asset to two essential fields:</p><ul><li><p><code>timestamp</code>: when the price observation was available;</p></li><li><p><code>price</code>: a finite, strictly positive price.</p></li></ul><p>A minimal CSV looks like this:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;csv&quot;,&quot;nodeId&quot;:&quot;4bf7032e-8866-4a9d-b8f9-ef87c1998053&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-csv">timestamp,price
2020-01-01,1518.2
2020-01-02,1523.7
2020-01-03,1511.4</code></pre></div><p>The <code>PriceData</code> class in <code>src/quantitative_trading_model/data.py</code> stores the cleaned result and an asset label:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4dd73e3e-5153-479f-92d8-59ab84f2bf42&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class PriceData:
    frame: pd.DataFrame
    asset: str = "asset"

    def __post_init__(self) -&gt; None:
        required = {"timestamp", "price"}
        missing = required.difference(self.frame.columns)
        if missing:
            raise ValueError(f"PriceData is missing required columns: {sorted(missing)}")

        timestamps = pd.to_datetime(self.frame["timestamp"], errors="raise")
        prices = pd.to_numeric(self.frame["price"], errors="raise").to_numpy(dtype=float)

        if timestamps.duplicated().any():
            raise ValueError("PriceData timestamps must be unique")
        if not timestamps.is_monotonic_increasing:
            raise ValueError("PriceData timestamps must be strictly increasing")
        if not np.isfinite(prices).all() or (prices &lt;= 0).any():
            raise ValueError("PriceData prices must be finite and strictly positive")</code></pre></div><p><code>PriceData.__post_init__</code> is a validation boundary. It does not decide whether a price is economically meaningful for every possible instrument, but it enforces the assumptions required by the downstream time-series code. An empty dataset, duplicate timestamp, unsorted timestamp sequence, missing value, infinity, zero, or negative price cannot silently proceed into window construction.</p><p>The <code>timestamps</code> and <code>prices</code> properties provide convenient, typed access to the two series:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;17beb2fd-68b7-4a1a-a4ed-b0781b084374&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@property
def timestamps(self) -&gt; pd.DatetimeIndex:
    return pd.DatetimeIndex(self.frame["timestamp"])

@property
def prices(self) -&gt; np.ndarray:
    return self.frame["price"].to_numpy(dtype=float, copy=True)</code></pre></div><p>Returning a copy from <code>prices</code> reduces the chance that a caller accidentally mutates the validated internal frame without re-running validation.</p><h3>5.2 Standardizing different local column names</h3><p>Real CSV files often use names such as <code>Date</code>, <code>datetime</code>, <code>Close</code>, or <code>Adj Close</code>. The <code>standardize_prices</code> function accepts explicit column names when necessary, but can also recognize common alternatives:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5e8daf45-e6f0-47cd-bc79-b561e6a75aaf&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def standardize_prices(
    data: Union[pd.DataFrame, PriceData],
    *,
    timestamp_column: Optional[str] = None,
    price_column: Optional[str] = None,
    asset: str = "asset",
    drop_invalid: bool = True,
) -&gt; PriceData:</code></pre></div><p>The function first copies the input, so cleaning does not mutate the caller&#8217;s DataFrame. It then resolves columns through <code>_resolve_column</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4b581dca-a480-46a8-b23a-f84f15b0dada&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">timestamp_column = _resolve_column(
    frame, timestamp_column,
    ("timestamp", "datetime", "date", "time"),
    "timestamp",
)
price_column = _resolve_column(
    frame, price_column,
    ("price", "close", "adj_close", "adjusted_close"),
    "price",
)</code></pre></div><p>Explicit names are safer when a file contains several possible price fields. For example, choosing <code>close</code> versus <code>adjusted_close</code> can materially change a historical experiment. Automatic inference is a convenience, not a substitute for documenting the choice.</p><p>The selected columns are renamed to the package&#8217;s stable internal names:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;64efd00c-fc7a-4ae9-b129-58503b4274bf&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">normalized = frame[[timestamp_column, price_column]].rename(
    columns={timestamp_column: "timestamp", price_column: "price"}
)
normalized["timestamp"] = pd.to_datetime(
    normalized["timestamp"], errors="coerce", utc=False
)
normalized["price"] = pd.to_numeric(normalized["price"], errors="coerce")</code></pre></div><p>Values that cannot be parsed become missing. The function then constructs a validity mask requiring a real timestamp and a finite, positive price:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;18058e46-0107-4bae-8ce1-291e05db7baf&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">valid = normalized["timestamp"].notna() &amp; normalized["price"].notna()
valid &amp;= np.isfinite(normalized["price"].to_numpy(dtype=float))
valid &amp;= normalized["price"] &gt; 0</code></pre></div><p>By default, invalid rows are removed. If <code>drop_invalid=False</code>, the function raises an error instead. The choice depends on the research context: dropping a corrupted row may be reasonable for a demonstration, but a production-quality study should usually stop and investigate why the row is invalid rather than silently deleting it.</p><h3>5.3 Sorting and duplicate timestamps</h3><p>After parsing, the function sorts by timestamp using a stable sort and removes duplicate timestamps by retaining the last source row:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e6166f47-3cc2-4e39-90d9-d4c18975aead&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">normalized = (
    normalized.sort_values("timestamp", kind="mergesort")
    .drop_duplicates("timestamp", keep="last")
    .reset_index(drop=True)
)</code></pre></div><p>This gives the downstream model one price per timestamp. It does not prove that the retained row is the correct observation. If duplicate rows represent separate trades, multiple venues, or different assets, they should be aggregated or separated before calling <code>standardize_prices</code>.</p><p>The duplicate policy is therefore an implementation convenience, not a fact recovered from the paper. The paper does not explain how duplicate timestamps or multiple records within one sampling interval were handled.</p><h3>5.4 Loading a local CSV</h3><p><code>load_price_csv</code> is a thin file-system wrapper around <code>standardize_prices</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;135de609-cc58-4011-ad70-afdb4ab63170&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def load_price_csv(
    path: Union[str, Path],
    *,
    timestamp_column: Optional[str] = None,
    price_column: Optional[str] = None,
    asset: str = "asset",
    drop_invalid: bool = True,
) -&gt; PriceData:
    source = Path(path)
    if not source.exists() or not source.is_file():
        raise FileNotFoundError(f"Price CSV does not exist: {source}")
    if source.suffix.lower() != ".csv":
        raise ValueError("load_price_csv accepts a local .csv file")
    return standardize_prices(
        pd.read_csv(source),
        timestamp_column=timestamp_column,
        price_column=price_column,
        asset=asset,
        drop_invalid=drop_invalid,
    )</code></pre></div><p>The function deliberately accepts only a local <code>.csv</code> path. There are no HTTP requests, exchange clients, credentials, or hidden downloads. This is consistent with the project&#8217;s educational and offline scope.</p><p>The CLI exposes the same idea through the <code>validate-data</code> subcommand. Its local-path helper rejects URL-like strings before loading them:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;78828c5e-16ff-469f-b6b1-ea85f35c5eab&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def _safe_input_path(value: str) -&gt; Path:
    if "://" in value:
        raise CLIError(
            "Only local files are supported; URLs and network paths are not accepted"
        )
    path = Path(value).expanduser()
    if not path.exists():
        raise CLIError(f"Input file does not exist: {path}")
    if not path.is_file():
        raise CLIError(f"Input path is not a regular file: {path}")
    return path.resolve()</code></pre></div><p>A local validation command can therefore be written as:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;8e9d0da1-1876-403b-aaa1-de3c804502d4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">quant-trading-model validate-data prices.csv \
  --timestamp-column timestamp \
  --price-column close \
  --asset gold</code></pre></div><p>The command prints or writes a compact summary rather than training a model. This makes data inspection a separate step from forecasting and backtesting.</p><h3>5.5 Keeping gold and Bitcoin separate</h3><p>The <code>PriceData</code> representation is intentionally single-asset. A file may contain multiple assets, but the research pipeline should filter it into one chronologically ordered series per asset before creating windows. Gold and Bitcoin have different price scales, trading calendars, liquidity characteristics, and likely data sources. Combining them accidentally into one price sequence would make the next observation after a gold row appear to be a Bitcoin movement.</p><p>A multi-asset table can be filtered explicitly:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;fc2b9cd3-89aa-4286-8371-e4bd101b9f2b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">raw = pd.read_csv("historical_prices.csv")

gold = standardize_prices(
    raw.loc[raw["asset"].eq("gold")],
    timestamp_column="timestamp",
    price_column="price",
    asset="gold",
)

bitcoin = standardize_prices(
    raw.loc[raw["asset"].eq("bitcoin")],
    timestamp_column="timestamp",
    price_column="price",
    asset="bitcoin",
)</code></pre></div><p>Each result can then be passed independently to the window builder and forecasting experiment. This also makes it possible to record asset-specific provenance and fee assumptions. The paper&#8217;s headline results allocate $500 to each asset, but the extracted text does not identify whether the assets were sampled on matching dates or evaluated with identical calendars.</p><h3>5.6 Synthetic data are for interfaces, not validation</h3><p>Because the paper&#8217;s original datasets are unavailable, the project includes <code>generate_synthetic_prices</code>. It creates deterministic, positive price paths using a seeded random generator:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;bd45fa12-0606-4623-9297-d96654589830&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def generate_synthetic_prices(
    *,
    asset: str = "synthetic",
    periods: int = 500,
    start: Union[str, pd.Timestamp] = "2020-01-01",
    seed: int = 7,
    initial_price: Optional[float] = None,
) -&gt; PriceData:</code></pre></div><p>The function chooses different illustrative drift and volatility defaults for labels containing <code>gold</code> or <code>bitcoin</code>, then constructs a positive path from exponentiated random shocks. The seed makes the example repeatable for teaching and invariant checks.</p><p>For example:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b969007f-9175-4c04-bf1b-ed1ec1a730ac&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from quantitative_trading_model.data import generate_synthetic_prices

prices = generate_synthetic_prices(
    asset="gold",
    periods=500,
    seed=7,
)
print(prices.asset)
print(prices.frame.head())</code></pre></div><p>The label-specific behavior is only a teaching convenience. A &#8220;Bitcoin-like&#8221; synthetic path is not Bitcoin data, and a &#8220;gold-like&#8221; path is not a gold instrument. Synthetic data can demonstrate that timestamps, windows, attention tensors, sizing budgets, and portfolio accounting have compatible interfaces. It cannot validate the paper&#8217;s forecast metrics, fee curves, or reported final values.</p><h3>5.7 Data checks before modeling</h3><p>Before moving to sliding windows, record at least the following:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!xMDU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cded6df-518e-4312-9db2-4eece7eca0ec_1282x588.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!xMDU!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cded6df-518e-4312-9db2-4eece7eca0ec_1282x588.png 424w, /__u/substackcdn.com/image/fetch/$s_!xMDU!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cded6df-518e-4312-9db2-4eece7eca0ec_1282x588.png 848w, /__u/substackcdn.com/image/fetch/$s_!xMDU!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cded6df-518e-4312-9db2-4eece7eca0ec_1282x588.png 1272w, /__u/substackcdn.com/image/fetch/$s_!xMDU!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cded6df-518e-4312-9db2-4eece7eca0ec_1282x588.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!xMDU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cded6df-518e-4312-9db2-4eece7eca0ec_1282x588.png" width="1282" height="588" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6cded6df-518e-4312-9db2-4eece7eca0ec_1282x588.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:588,&quot;width&quot;:1282,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:122560,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cded6df-518e-4312-9db2-4eece7eca0ec_1282x588.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!xMDU!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cded6df-518e-4312-9db2-4eece7eca0ec_1282x588.png 424w, /__u/substackcdn.com/image/fetch/$s_!xMDU!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cded6df-518e-4312-9db2-4eece7eca0ec_1282x588.png 848w, /__u/substackcdn.com/image/fetch/$s_!xMDU!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cded6df-518e-4312-9db2-4eece7eca0ec_1282x588.png 1272w, /__u/substackcdn.com/image/fetch/$s_!xMDU!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cded6df-518e-4312-9db2-4eece7eca0ec_1282x588.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The generated tests in <code>tests/test_data_and_streaks.py</code> are intended to check these kinds of invariants with small hand-built fixtures. They are static artifacts in this project review; no claim is made here that the generated test suite was executed successfully.</p><p>The essential principle is simple: clean only what the data contract justifies, preserve timestamps, separate assets, and document every assumption before training. Once the local series is trustworthy and its provenance is recorded, the next step is to construct the paper&#8217;s 100-observation windows without allowing scaling or temporal partitioning to leak future information.</p><h2>6. Build Leakage-Safe Sliding Windows</h2><p>The paper describes a sliding-window forecasting setup with a history of 100 observations and a stride of one. The model reads the latest 100 observations, predicts a later value, advances one observation, and creates the next overlapping example.</p><p>The important issue is not only the window size. Every forecast must use information available at its decision timestamp. The target must occur after the input window, scaling parameters must be fitted without future observations, and evaluation must preserve chronological order.</p><h3>6.1 The one-step relationship</h3><p>Let <code>z_t</code> denote a transformed price or feature value at time <code>t</code>. With a window length of 100, a one-step sample is:</p><p>\[ X<em>t = [z</em>{t-99}, z<em>{t-98}, \ldots, z</em>{t-1}, z<em>t], \qquad y</em>t = z_{t+1}. \]</p><p>The final value in <code>X_t</code> is the latest information available when the forecast is generated. The target belongs to the next observation.</p><p>The paper-inspired window settings are represented by <code>DataConfig</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;538dee89-7b54-4de4-84d9-0010889e25c7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass
class DataConfig:
    timestamp_column: str = "timestamp"
    price_column: str = "price"
    window_length: int = 100
    stride: int = 1</code></pre></div><p><code>window_length=100</code> and <code>stride=1</code> come from the paper. The paper does not clearly state whether the model uses raw prices, normalized prices, returns, or additional features, so the representation remains an explicit implementation choice.</p><h3>6.2 How <code>make_windows</code> constructs samples</h3><p>The central function is <code>make_windows</code> in <code>src/quantitative_trading_model/data.py</code>. It accepts a cleaned <code>PriceData</code> object, a DataFrame, or a numeric array. For timestamp-aware inputs, it stores the start, end, and target timestamps alongside the arrays.</p><p>The core operation is equivalent to:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;fc082e4f-8ffb-4ae0-8ce2-51b04bfdd636&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">for start in range(0, last_start + 1, stride):
    end = start + window_length - 1
    X.append(transformed[start : end + 1])
    targets.append(transformed[end + 1 : end + 1 + horizon, 0])
    target_times.append(timestamps[end + horizon])</code></pre></div><p>With the defaults, the first sample contains observations <code>0</code> through <code>99</code> and targets observation <code>100</code>. The next sample contains observations <code>1</code> through <code>100</code> and targets observation <code>101</code>.</p><p>The returned <code>WindowedDataset</code> includes temporal metadata:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e525b16f-f8a2-4a41-8c04-ab9e5c0c3fa7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class WindowedDataset:
    X: np.ndarray
    y: np.ndarray
    window_start: pd.DatetimeIndex
    window_end: pd.DatetimeIndex
    y_timestamp: pd.DatetimeIndex
    feature_names: Tuple[str, ...] = ("price",)</code></pre></div><p>Its invariant check requires every target timestamp to follow the corresponding input window:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f84ec74a-311c-4aec-9cc7-a200718f5719&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">if n and not (self.y_timestamp &gt; self.window_end).all():
    raise ValueError("Every target timestamp must be after its input window")</code></pre></div><p>This timestamp guarantee applies when the input carries timestamps, such as <code>PriceData</code> or a cleaned DataFrame. For a raw numeric array, the current implementation has no timestamp argument in that code path and creates placeholder daily timestamps beginning at <code>1970-01-01</code>. Those generated timestamps preserve array alignment for shape-based demonstrations, but they are not source-market timestamps. A trading experiment should therefore use <code>PriceData</code> or a timestamp-aware DataFrame.</p><h3>6.3 One-day versus three-day targets</h3><p>The paper is ambiguous about the forecast horizon. Its prediction discussion is compatible with one-step forecasting, while its conclusion refers to forecasts for the next three days. The implementation exposes <code>horizon</code> rather than silently choosing one interpretation.</p><p>For <code>horizon=1</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;3333df75-46a0-40bc-8412-4a694af53e49&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">X.shape = (samples, 100, features)
y.shape = (samples,)</code></pre></div><p>For <code>horizon=3</code>:</p><p>\[ y<em>t = [z</em>{t+1}, z<em>{t+2}, z</em>{t+3}]. \]</p><p>The shape is then:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;977ea76f-26cb-4ac7-a03e-8d4b1a6d0e92&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">X.shape = (samples, 100, features)
y.shape = (samples, 3)</code></pre></div><p>The final target timestamp is the timestamp of <code>z_{t+3}</code>. Converting a three-value forecast into a buy, sell, or hold decision belongs to the strategy layer. Possible rules include using the terminal forecast, the mean forecast, or a separately specified multi-day signal.</p><p>For a series of <code>N</code> observations, the number of stride-one windows is:</p><p>N&#8722;100&#8722;h+1,</p><p>where <code>h</code> is the forecast horizon. The final window must leave enough observations for every target value.</p><h3>6.4 Overlapping windows and temporal partitions</h3><p>Stride one creates heavily overlapping samples:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;2b474807-d037-4f57-bdfe-30ccfb9ef908&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">window 1: [0, 1, 2, ..., 99] -&gt; target 100
window 2: [1, 2, 3, ..., 100] -&gt; target 101</code></pre></div><p>This overlap is normal for sliding-window forecasting. The risk comes from treating neighboring samples as independent random observations. If windows are randomly shuffled before splitting, almost identical histories can appear in both partitions, producing optimistic estimates.</p><p>The generated recurrent training code therefore uses non-shuffled batches:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a0b32f31-83f2-4f2d-b05b-1a42b310033e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">train_loader = DataLoader(
    TensorDataset(torch.from_numpy(x_train), torch.from_numpy(y_train)),
    batch_size=batch_size,
    shuffle=False,
)</code></pre></div><p><code>chronological_split</code> also preserves sample order. This prevents random cross-partition mixing, but it does <strong>not</strong> guarantee that raw input histories are disjoint at the boundary. The final training window and first test window may still share observations because the windows overlap by design. That dependence should be acknowledged when interpreting metrics; walk-forward evaluation and a documented gap can provide a stricter protocol.</p><h3>6.5 Fit scaling on training observations only</h3><p>Scaling can make neural-network optimization easier, but fitting a scaler on the complete series leaks future distribution information. For example, a full-series min-max scaler exposes the training process to the future minimum and maximum.</p><p>The repository provides <code>ChronologicalScaler</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d801e448-278e-468e-9b20-ed5a007773b1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">scaler = ChronologicalScaler()
scaler.fit(training_values)                 # training rows only
train_scaled = scaler.transform(train_values)
test_scaled = scaler.transform(test_values)</code></pre></div><p>The safe sequence is:</p><p>Clean and sort the complete local file.</p><p>Establish the chronological raw-row training boundary.</p><p>Fit the scaler only on rows before that boundary.</p><p>Transform later rows with the fitted scaler.</p><p>Build windows and retain their timestamps.</p><p>Inverse-transform predictions before reporting price-level metrics when appropriate.</p><p><code>make_windows</code> accepts an already fitted scaler. Its <code>fit_scaler</code> option is intended for controlled demonstrations; setting <code>fit_scaler=True</code> on the complete dataset before an out-of-sample experiment would leak future information. The window builder does not know whether the caller selected a valid training boundary and cannot enforce this policy automatically.</p><p>A safer pattern is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6b78a406-21ff-4282-a6fe-c6e7da5e8d09&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">train_rows = price_data.frame.iloc[:training_row_count]
scaler = ChronologicalScaler().fit(
    train_rows[["price"]].to_numpy(dtype=float)
)

windows = make_windows(
    price_data,
    window_length=100,
    horizon=1,
    stride=1,
    scaler=scaler,
    fit_scaler=False,
)</code></pre></div><p>Here, <code>training_row_count</code> must be determined from a documented chronological rule, such as a date boundary or a fixed initial training period. It must not be chosen after inspecting test performance. If windows are built from the entire transformed series and then split, the scaler remains training-fitted, but boundary overlap still exists; a stricter workflow can build or select partitions around an explicit raw-row boundary.</p><h3>6.6 Chronological 70:30 splitting</h3><p>The paper reports a 70:30 train/test split. The implementation provides <code>chronological_split</code> for a paper-style comparison:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;381b4752-846f-4402-b0e1-8dcb35f61b1c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">train, test = chronological_split(
    dataset,
    train_fraction=0.70,
    validation_fraction=0.0,
)</code></pre></div><p>The function slices ordered samples rather than randomly selecting indices. With validation data, the order is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;6d8b5105-3b40-497e-8039-530c2f39796c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">training windows -&gt; validation windows -&gt; test windows</code></pre></div><p>The metadata-preserving helper is conceptually:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;3cc9069c-69d9-4b20-a3c7-7e6ca964d4d6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def dataset_slice(dataset, start, stop):
    return WindowedDataset(
        X=dataset.X[start:stop],
        y=dataset.y[start:stop],
        window_start=dataset.window_start[start:stop],
        window_end=dataset.window_end[start:stop],
        y_timestamp=dataset.y_timestamp[start:stop],
        feature_names=dataset.feature_names,
    )</code></pre></div><p>A useful ordering check is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d45ea1d6-a032-43dd-99b8-4ddc8343d8d0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">train_times = train.y_timestamp
test_times = test.y_timestamp
assert train_times.max() &lt; test_times.min()</code></pre></div><p>This confirms that target timestamps are ordered across the sample split. It does not prove that the corresponding input histories are disjoint, because stride-one windows can overlap at the boundary.</p><p>The exact boundary remains a reconstruction. The paper does not specify the split date or whether it splits raw observations before window construction or partitions already-created windows. Those choices can produce different edge behavior and should be recorded in experiment metadata.</p><h3>6.7 Walk-forward evaluation for trading conclusions</h3><p>A single 70:30 split answers only how one fixed training period performs on one later period. It does not show how results change as new observations arrive.</p><p>An expanding walk-forward procedure repeatedly preserves the information boundary:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;cda0d5fb-7238-489f-8791-295426e33852&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">train on [0, ..., t]
forecast the next block
advance the boundary
expand the training set
forecast the next block
repeat</code></pre></div><p>The <code>walk_forward_splits</code> generator is used like this:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ca739ab0-291d-44e9-b61e-665771831f41&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">for train, validation, test in walk_forward_splits(
    dataset,
    initial_train_size=500,
    validation_size=0,
    test_size=20,
    step=20,
):
    # Fit only on train, optionally tune on validation,
    # then evaluate on the later test block.
    pass</code></pre></div><p>Each test block follows its training block. The training set expands over time; a rolling fixed-length training window would be another explicit choice.</p><p>The paper does not report walk-forward results, so such results are not reproductions of its tables. They are a more appropriate protocol for trading conclusions because they repeatedly test the model on observations that follow the fitting period.</p><h3>6.8 What the generated test artifact actually demonstrates</h3><p>The generated <code>tests/test_data_and_streaks.py</code> uses a compatibility helper with a numeric price array and a separate timestamp object. In the current <code>make_windows</code> implementation, the numeric-array path does not accept or preserve that separate timestamp argument; it creates placeholder daily timestamps. Therefore, the test's meaningful checks are the numeric window boundaries and target values, while its timestamp behavior should not be interpreted as validation of real source timestamps.</p><p>A timestamp-aware illustrative example, using the public data contract directly, is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a403446a-5733-4edb-bd87-1bc2a1048516&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">prices = np.arange(10.0, 17.0)
timestamps = pd.date_range(
    "2024-01-01", periods=len(prices), freq="D"
)
price_data = PriceData(
    pd.DataFrame({"timestamp": timestamps, "price": prices})
)

dataset = make_windows(
    price_data,
    window_length=3,
    horizon=1,
    stride=1,
)

np.testing.assert_allclose(dataset.y, prices[3:])
assert (dataset.y_timestamp &gt; dataset.window_end).all()</code></pre></div><p>This is an illustrative contract for timestamp-aware input, not a claim that the generated test suite was executed. It also makes the distinction clear: use <code>PriceData</code> or a timestamp-aware DataFrame when timestamps matter; use raw arrays only when placeholder metadata is acceptable or when a higher-level API supplies the timestamp mapping separately.</p><p>The project was reviewed primarily through static checks, and semantic code verification was skipped. The intended contracts are described here, but this section does not claim that the complete test suite passed or that runtime behavior was independently confirmed.</p><h3>6.9 Practical preparation workflow</h3><p>For a local asset file, a disciplined workflow is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d9f3631f-7e18-4fa4-86f0-622c77dfa16e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">price_data = load_price_csv(
    "data/gold.csv",
    timestamp_column="timestamp",
    price_column="price",
    asset="gold",
)

# Establish this from a documented date or raw-row boundary.
training_row_count = 500
train_rows = price_data.frame.iloc[:training_row_count]
scaler = ChronologicalScaler().fit(
    train_rows[["price"]].to_numpy(dtype=float)
)

windows = make_windows(
    price_data,
    window_length=100,
    horizon=1,       # use 3 only when a three-step target is intended
    stride=1,
    scaler=scaler,
    fit_scaler=False,
)
train, test = chronological_split(windows, train_fraction=0.70)</code></pre></div><p>The row boundary must be established before fitting the scaler and must reflect the intended training period. The split shown above is a paper-style sample split applied after window construction. For a stricter out-of-time study, define the raw-row boundary and any gap explicitly, then use walk-forward partitions for repeated evaluation.</p><h3>6.10 Preparation limits</h3><p>This data layer provides:</p><ul><li><p>the paper's 100-observation, stride-one setup;</p></li><li><p>one-step and three-step target shapes;</p></li><li><p>timestamp metadata for <code>PriceData</code> and timestamp-aware DataFrame inputs;</p></li><li><p>chronological ordering and sample splits;</p></li><li><p>training-only scaling primitives;</p></li><li><p>expanding walk-forward partitions.</p></li></ul><p>It does not resolve the paper's missing data source, instrument definition, feature representation, exact split date, forecast horizon, or execution protocol. It also cannot automatically determine whether a caller fitted a scaler on an appropriate subset or whether overlapping input histories are acceptable for a particular evaluation design.</p><p>The essential causal contract is therefore: the historical window ends first, the target occurs later, scaling uses only the declared training information, and execution is aligned to a timestamp after the information endpoint. Later forecasting and backtesting layers should preserve that contract explicitly.</p><h2>7. Keep Forecast Accuracy Separate from Trading Performance</h2><p>The paper compares forecasting models with RMSE, MAE, MAPE, and R-squared. These metrics answer one question: <strong>how close were the predictions to their targets?</strong> They do not establish that a trading strategy is profitable after signal rules, position sizing, VaR, execution timing, fees, slippage, and liquidity are included.</p><p>The implementation keeps these concerns separate in <code>src/quantitative_trading_model/metrics.py</code>:</p><ul><li><p>forecast metrics operate on target and prediction arrays;</p></li><li><p>portfolio metrics operate on a marked-to-market value path and optional execution records;</p></li><li><p><code>backtest.py</code> supplies the fills, traded notionals, and fees used by portfolio reporting.</p></li></ul><p>This distinction is also documented in <code>README.md</code>, which warns that forecast accuracy does not establish profitability and recommends reporting drawdown, volatility, turnover, fees, and a buy-and-hold comparison.</p><h3>7.1 Forecast error metrics</h3><p>Let <code>y_k</code> be the observed target and <code>y_hat_k</code> the corresponding prediction for evaluation item <code>k</code>.</p><h4>Mean absolute error</h4><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot; \\operatorname{MAE} = \\frac{1}{N}\\sum{k=1}^{N}|yk-\\hat y_k|. &quot;,&quot;id&quot;:&quot;YIRKOIZOKT&quot;}" data-component-name="LatexBlockToDOM"></div><p>MAE is measured in the target's units. For raw dollar prices, it is measured in dollars. The paper-inspired recurrent models also use MAE as their training loss.</p><h4>Root mean squared error</h4><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\operatorname{RMSE} = \\sqrt{\\frac{1}{N}\\sum{k=1}^{N}(yk-\\hat y_k)^2}. &quot;,&quot;id&quot;:&quot;WNMEYLGIUU&quot;}" data-component-name="LatexBlockToDOM"></div><p>Squaring the residuals gives larger errors more influence. RMSE is scale-dependent, so raw RMSE values for gold and Bitcoin should not be compared without considering their different price magnitudes and target representations.</p><h4>Mean absolute percentage error</h4><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot; \\operatorname{MAPE} = \\frac{100}{N}\\sum{k=1}^{N}\\left|\\frac{yk-\\hat yk}{yk}\\right|. &quot;,&quot;id&quot;:&quot;IMCCRHRKBQ&quot;}" data-component-name="LatexBlockToDOM"></div><p>The implementation reports MAPE in percentage points: <code>2.5</code> means 2.5%, not 0.025. MAPE is undefined or unstable for zero and near-zero targets. <code>forecast_metrics</code> therefore exposes two controls:</p><ul><li><p><code>mape_epsilon</code> defines the near-zero threshold;</p></li><li><p><code>mape_zero_policy="exclude"</code> omits those targets, while <code>"raise"</code> rejects them.</p></li></ul><p>The number of observations included in MAPE is stored separately from the total sample count. Two MAPE values should not be compared without checking that policy and count.</p><h4>R-squared</h4><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;R^2 = 1 - \\frac{\\sumk(yk-\\hat yk)^2}{\\sumk(y_k-\\bar y)^2}. &quot;,&quot;id&quot;:&quot;LPSWKPKLXS&quot;}" data-component-name="LatexBlockToDOM"></div><p>R-squared compares residual variation with the variation of the targets around their mean. It can be negative when predictions are worse than the constant mean-target baseline. If every target is identical, the denominator is zero; the generated implementation returns <code>1.0</code> for an exact prediction and <code>0.0</code> otherwise.</p><h3>7.2 The forecast-metrics interface</h3><p>The public function is <code>forecast_metrics(actual, predicted, ...)</code>. It validates finite inputs, flattens supported array-like values, checks that targets and predictions have equal lengths, and returns a <code>ForecastMetrics</code> dataclass.</p><p>The following is a <strong>shortened tutorial excerpt</strong>, not the complete generated source. The actual module also contains helper functions for input validation, rounding, and serialization.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;1d1efa10-02e4-4c25-8eda-8068ef16da82&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">metrics = forecast_metrics(
    actual,
    predicted,
    mape_epsilon=1e-12,
    mape_zero_policy="exclude",
)

print(metrics.rmse)
print(metrics.mae)
print(metrics.mape_percent)
print(metrics.r_squared)
print(metrics.as_dict())</code></pre></div><p><code>ForecastMetrics.as_dict()</code> is the generated dataclass's serialization method. It includes the four metrics, sample counts, the MAPE threshold, and the selected MAPE policy. The full implementation should be treated as authoritative rather than this abbreviated excerpt.</p><p>For example:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f08a0bb2-26e7-4cfe-84fa-b4d9132676fa&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import numpy as np
from quantitative_trading_model.metrics import forecast_metrics

actual = np.array([100.0, 110.0, 90.0])
predicted = np.array([102.0, 107.0, 93.0])

result = forecast_metrics(actual, predicted)</code></pre></div><p>The residuals are <code>[2, -3, 3]</code>. Thus MAE is <code>8/3</code>, RMSE is <code>sqrt(22/3)</code>, and MAPE is the mean of the three absolute relative errors multiplied by 100. These values describe only the prediction sample. They do not determine whether a strategy would buy, sell, or hold.</p><p>For multi-step predictions represented as <code>(samples, horizon)</code>, the generated metric helper flattens the arrays. The caller must ensure that the target and prediction arrays have identical ordering. A production experiment should also retain target timestamps so that metrics are calculated on aligned observations.</p><h3>7.3 Why forecast accuracy is not trading performance</h3><p>A model can have favorable forecast errors and still produce a poor trading result:</p><p>A predicted move may be too small to cover entry and exit costs.</p><p>Price-level accuracy does not guarantee correct direction relative to the current price.</p><p>A forecast evaluated at a closing timestamp may not describe the price available at execution.</p><p>Position sizing can amplify a small forecast error.</p><p>A shuffled or otherwise optimistic time-series split may not represent deployment conditions.</p><p>Fees, spreads, slippage, partial fills, and liquidity are portfolio effects, not components of RMSE or MAE.</p><p>The paper reports strong Att-BiLSTM forecasting metrics, but those reported values remain unverified claims because the source data, preprocessing, target definition, and test protocol are unavailable. Even reproducing the metric ranking would not prove that the resulting trading strategy is profitable.</p><h3>7.4 Portfolio metrics use a different vocabulary</h3><p>Portfolio evaluation begins with marked-to-market values:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot; Vt = \\operatorname{cash}t + \\operatorname{quantity}t pt. &quot;,&quot;id&quot;:&quot;IHELZKMAKW&quot;}" data-component-name="LatexBlockToDOM"></div><p>The generated implementation reports or supports the following measures.</p><h4>Cumulative return</h4><p>\[ R<em>{0:T} = \frac{V</em>T}{V_0}-1. \]</p><p>This reflects trades, fees, remaining holdings, and the configured liquidation policy.</p><h4>Drawdown and maximum drawdown</h4><p>\[ D<em>t = \frac{V</em>t}{\max<em>{s\le t}V</em>s}-1. \]</p><p>Drawdown is non-positive: <code>-0.25</code> means the portfolio is 25% below its previous running peak. Maximum drawdown is the minimum value in the drawdown series.</p><h4>Volatility</h4><p>For periodic returns <code>r_t</code>, annualized volatility is approximately:</p><p>\[ \sigma<em>{annual}=\operatorname{std}(r</em>t)\sqrt{P}, \]</p><p>where <code>P</code> is the configured number of periods per year. The default of 252 is an implementation assumption for daily observations, not a paper-specified value.</p><h4>Sharpe ratio</h4><p>\[ \operatorname{Sharpe} = \frac{\operatorname{mean}(r<em>t-r</em>{f,t})} {\operatorname{std}(r<em>t-r</em>{f,t})}\sqrt{P}. \]</p><p>The implementation converts an annual risk-free rate into a periodic compound-equivalent rate. Its default risk-free rate is zero, another explicit evaluation choice rather than a paper result.</p><h4>Turnover and fees</h4><p>The generated metrics layer defines turnover as traded notional divided by mean portfolio value:</p><p>\[ \operatorname{turnover} = \frac{\sum<em>i \operatorname{traded\ notional}</em>i}{\operatorname{mean}(V_t)}. \]</p><p>Fees are supplied by the execution layer and summed separately:</p><p>\[ \operatorname{total\ fees}=\sum<em>i\operatorname{fee}</em>i. \]</p><p>This preserves the distinction between portfolio valuation and fill-level accounting.</p><h3>7.5 The portfolio-metrics interface</h3><p><code>portfolio_metrics</code> accepts a value path and optional periodic returns, traded notionals, fees, annualization frequency, and annual risk-free rate. The following is a <strong>shortened excerpt</strong>, not a complete copy of <code>metrics.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4739f1f5-c6f2-482c-85c7-fe81869336ed&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">portfolio = portfolio_metrics(
    portfolio_values,
    returns=None,              # derive returns from consecutive values
    traded_notionals=notionals,
    fees=fees,
    periods_per_year=252,
    risk_free_rate_annual=0.0,
)

print(portfolio.cumulative_return)
print(portfolio.maximum_drawdown)
print(portfolio.annualized_volatility)
print(portfolio.sharpe_ratio)
print(portfolio.turnover)
print(portfolio.total_fees)</code></pre></div><p>The full generated function additionally validates positive portfolio values, finite inputs, return-length conventions, nonnegative notionals, scalar-versus-sequence fee inputs, and risk-free-rate settings. It returns a <code>PortfolioMetrics</code> dataclass with an <code>as_dict()</code> method.</p><h3>7.6 A report should retain both layers</h3><p>A useful experiment table keeps forecast and portfolio results in separate columns:</p><p>| Forecast evaluation | Portfolio evaluation | | --- | --- | | RMSE | Final value | | MAE | Cumulative return | | MAPE and MAPE policy | Maximum drawdown | | R-squared | Annualized volatility | | Target horizon | Sharpe ratio | | Test timestamps | Turnover | | Scaling metadata | Total fees |</p><p>This prevents a low prediction error from being mistaken for a trading objective. A buy-and-hold benchmark should also use the same asset, period, initial capital, fee convention, and valuation policy.</p><h3>7.7 Verification scope</h3><p>The supplied project currently does <strong>not</strong> include a dedicated metrics test module. <code>tests/test_data_and_streaks.py</code> tests data preparation and streak behavior; it does not directly test MAPE policies, constant-target R-squared behavior, drawdown calculations, or fee aggregation. Those behaviors are documented and statically inspectable invariants, not claims of executed test coverage.</p><p>Static review can still check that:</p><ul><li><p>forecast arrays are finite and aligned;</p></li><li><p>MAPE has an explicit near-zero policy;</p></li><li><p>constant-target R-squared behavior is deterministic;</p></li><li><p>portfolio values are positive and drawdown is non-positive;</p></li><li><p>turnover and fees are supplied separately from forecast errors.</p></li></ul><p>However, semantic code verification was skipped, and the generated project was not claimed to have been executed. These checks therefore do not establish runtime correctness, profitability, or reproduction of the paper's metrics.</p><p>The practical rule is: <strong>use forecast metrics to evaluate the predictor and portfolio metrics to evaluate the trading system.</strong> Both are necessary, but neither is a substitute for the other.</p><h2>8. Implement the Attention-Enhanced BiLSTM</h2><p>The paper identifies an attention-enhanced bidirectional LSTM, or <strong>Att-BiLSTM</strong>, as its primary forecasting model. Its intended flow is:</p><p>Read a historical window, such as the latest 100 observations.</p><p>Process that window in both temporal directions with a BiLSTM.</p><p>Produce one hidden representation for every time step.</p><p>Use attention to assign larger weights to more relevant historical steps.</p><p>Combine the weighted representations into one context vector.</p><p>Map that context vector to a one-day or multi-day forecast.</p><p>The generated implementation follows this structure in <code>src/quantitative_trading_model/models/forecasters.py</code>. The attention mathematics is isolated in <code>src/quantitative_trading_model/models/attention.py</code>, while <code>tests/test_model_shapes.py</code> checks the intended tensor contracts when the optional PyTorch dependency is available.</p><p><strong>Reconstruction boundary:</strong> The paper specifies the broad Att-BiLSTM idea, dot-product attention, 20% dropout, MAE, and RMSprop. It does not specify the number of recurrent layers, hidden dimensions, attention-query construction, learning rate, training duration, or exact target representation. Those details are configurable implementation choices, not verified transcriptions of the original model.</p><h3>8.1 What the BiLSTM contributes</h3><p>An ordinary LSTM processes a sequence in one direction. Given an input window ordered from older to newer observations, its hidden state at time step <code>i</code> summarizes the observations up to that point in the forward direction.</p><p>A BiLSTM adds a second LSTM that processes the same historical window in reverse. At each time step, the model combines the forward and backward states. If the one-direction hidden size is <code>h</code>, the combined representation commonly has size <code>2h</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;d16da738-9122-4bb6-be04-2a09ce751cdb&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">input sequence:       (batch, time, features)
forward LSTM states:  (batch, time, h)
backward LSTM states: (batch, time, h)
combined states:      (batch, time, 2h)</code></pre></div><p>The important timing condition is that the complete window must end at the information timestamp. If the signal is generated after observing time <code>t</code>, the input may contain observations through <code>t</code>, but not <code>t+1</code> or any later observation. The backward LSTM is not automatically leakage because it looks backward <strong>inside the already available window</strong>. It becomes leakage only if the window itself contains future observations.</p><p>The model therefore relies on the contract established by the data layer:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;4ff48439-377d-4fee-b857-6b5c9658e79a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">window ends at t
forecast is generated using data through t
execution occurs at a later bar</code></pre></div><p>This is why the window metadata and timestamp checks discussed in the previous section matter to the neural model as well as to the backtest.</p><h3>8.2 Dot-product temporal attention</h3><p>The BiLSTM produces a sequence of hidden representations rather than one final vector. The paper's attention mechanism decides how to combine those representations.</p><p>Let <code>x_i</code> be the BiLSTM representation at time step <code>i</code>, and let <code>q</code> be a query vector. The paper gives the following operations.</p><h4>Compatibility score</h4><p>\[ S(x<em>i,q) = x</em>i^Tq. \]</p><p>This dot product produces one scalar score for every time step. A larger score indicates stronger alignment between <code>x_i</code> and the query <code>q</code>.</p><h4>Attention weights</h4><p>The scores are normalized with a softmax:</p><p>\[ a<em>i = \operatorname{softmax}</em>i(S(x<em>i,q)) = \frac{\exp(S(x</em>i,q))} {\sum<em>j \exp(S(x</em>j,q))}. \]</p><p>The subscript on softmax is important. Normalization must happen <strong>over the time dimension</strong>, so that the weights for one sequence sum to one across its historical observations. Normalizing across features or across the batch would produce a different mechanism.</p><h4>Context vector</h4><p>Finally, the representations are combined with their attention weights:</p><p>\[ \operatorname{context}(q,x)=\sum<em>i a</em>i x_i. \]</p><p>For a batch of sequences, the expected tensor shapes are:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;5bf7d72f-87cb-4780-80d4-254512018b13&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">sequence:  (batch, time, feature_dim)
scores:    (batch, time)
weights:   (batch, time)
context:   (batch, feature_dim)
forecast:  (batch, horizon)</code></pre></div><p>The implementation uses <code>DotProductTemporalAttention</code> for these operations:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2d4d5820-c3d9-4420-bbab-659e98c05343&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># src/quantitative_trading_model/models/attention.py

class DotProductTemporalAttention(nn.Module):
    def forward(
        self,
        sequence: Tensor,
        query: Optional[Tensor] = None,
        *,
        return_weights: bool = False,
    ) -&gt; Tensor | Tuple[Tensor, Tensor]:
        self._validate_sequence(sequence)
        prepared_query = self._prepare_query(sequence, query)

        # Equation (1): one dot product per time step.
        scores = torch.einsum("btd,bd-&gt;bt", sequence, prepared_query)

        # Equation (2): normalize across time, not features or batches.
        weights = torch.softmax(scores, dim=1)

        # Equation (3): weighted sum of sequence representations.
        context = torch.sum(weights.unsqueeze(-1) * sequence, dim=1)

        self._last_attention_weights = weights
        if return_weights:
            return context, weights
        return context</code></pre></div><p>The Einstein summation expression <code>"btd,bd-&gt;bt"</code> means:</p><ul><li><p><code>b</code>: batch item;</p></li><li><p><code>t</code>: time step;</p></li><li><p><code>d</code>: feature dimension;</p></li><li><p>the feature dimension is multiplied and summed away;</p></li><li><p>the result retains batch and time, producing one score per time step.</p></li></ul><p>The next line, <code>torch.softmax(scores, dim=1)</code>, is the key implementation detail. For a tensor shaped <code>(batch, time)</code>, dimension <code>1</code> is time. The result has one probability-like weight for every historical position.</p><h3>8.3 The unspecified query vector</h3><p>The paper uses a query vector <code>q</code> in its equations but does not explain where that vector comes from. This missing detail affects the architecture, so the implementation makes the choice explicit rather than hiding it.</p><p><code>DotProductTemporalAttention</code> supports two modes:</p><ul><li><p>a <strong>learned query</strong>, stored as a trainable parameter and used by default;</p></li><li><p>a <strong>caller-provided query</strong>, useful for experiments with another interpretation.</p></li></ul><p>The learned-query choice is visible in the constructor:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;9485cce2-5aa1-4bf1-879c-1f72f1adc238&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># src/quantitative_trading_model/models/attention.py

if learned_query:
    query = torch.empty(self.feature_dim)
    nn.init.normal_(query, mean=0.0, std=float(query_init_scale))
    self.query = nn.Parameter(query)
else:
    self.register_parameter("query", None)</code></pre></div><p>When no explicit query is supplied, <code>_prepare_query</code> retrieves the learned vector and broadcasts it across the batch. When a query is supplied, the method checks its dimensionality, device, data type, and finiteness before using it.</p><p>This is a sound interface decision for an educational implementation, but it must be labeled correctly:</p><p>The learned query is a reconstruction of an unspecified paper detail. It is not evidence that the original authors used a learned query rather than a final hidden state, a separate projection, or another query-generation mechanism.</p><h3>8.4 From sequence states to a forecast</h3><p>The primary model in <code>forecasters.py</code> first creates a bidirectional LSTM whose <code>batch_first=True</code> input convention is <code>(batch, time, features)</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f38a226f-ea72-4da6-b387-60ceec27483f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># src/quantitative_trading_model/models/forecasters.py

self.encoder = nn.LSTM(
    input_size=input_size,
    hidden_size=hidden_size,
    num_layers=num_layers,
    batch_first=True,
    dropout=recurrent_dropout,
    bidirectional=True,
)

self.dropout = nn.Dropout(dropout)
self.attention = DotProductTemporalAttention(self.output_size)
self.head = nn.Linear(self.output_size, horizon)</code></pre></div><p><code>self.output_size</code> is <code>hidden_size * 2</code>, because the forward and backward states are concatenated. The dense <code>head</code> converts the attention context into the requested forecast horizon. For example:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;32654e47-7f1f-446c-8dc1-43eaf3a93e40&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">horizon = 1  -&gt; output shape (batch, 1)
horizon = 3  -&gt; output shape (batch, 3)</code></pre></div><p>The forward pass is intentionally short because each responsibility is delegated to a clear layer:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c6a9d71b-2850-4844-add3-a28ad819e5ea&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class AttBiLSTMForecaster(nn.Module):
    def forward(self, inputs: Tensor) -&gt; Tensor:
        _validate_input_tensor(inputs)
        sequence, _ = self.encoder(inputs)
        sequence = self.dropout(sequence)

        context = self.attention(sequence)
        return self.head(self.dropout(context))</code></pre></div><p>Conceptually, this is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;ebe26cd4-9fbe-47d9-99c1-e0c69ebe756e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">historical window
    -&gt; bidirectional LSTM sequence states
    -&gt; dropout on sequence states
    -&gt; temporal dot-product attention
    -&gt; context vector
    -&gt; dropout on context
    -&gt; linear forecast head</code></pre></div><p>The exact generated implementation also retains the latest attention weights when the attention layer exposes them. That is useful for inspection and debugging, but attention weights should not automatically be interpreted as a complete causal explanation of the prediction.</p><h3>8.5 Dropout, MAE, and RMSprop</h3><p>The paper reports a dropout rate of 20%, represented by the default <code>dropout=0.20</code>. Dropout randomly suppresses a fraction of activations during training, which can reduce reliance on a narrow set of internal features. The paper's wording does not establish whether its 20% applies to ordinary feed-forward dropout, recurrent dropout, or both, so the generated code uses explicit <code>nn.Dropout</code> layers and documents the recurrent-layer behavior.</p><p>The training function uses MAE loss:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6f75ba1a-a232-4edf-ae7a-92dd1c5b7ede&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">criterion = nn.L1Loss()</code></pre></div><p>This corresponds to:</p><p>\[ \operatorname{MAE} = \frac{1}{N}\sum<em>k |y</em>k-\hat y_k|. \]</p><p>The optimizer is RMSprop:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;57503a42-d898-4b8e-8328-88b1d51a50e8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">optimizer = torch.optim.RMSprop(
    model.parameters(),
    lr=learning_rate,
)</code></pre></div><p>This corrects an important terminology issue in the paper interpretation. <strong>RMSprop is an optimizer, not an activation function.</strong> It controls how model parameters are updated after gradients are calculated. The activation behavior of the recurrent cells and the linear output head is a separate architectural matter.</p><p>The implementation keeps learning rate, hidden size, number of layers, batch size, epochs, and patience configurable because the paper does not provide enough information to hard-code them as original values.</p><h3>8.6 Why the model trains on chronological batches</h3><p><code>train_forecaster</code> constructs a <code>DataLoader</code> with <code>shuffle=False</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d5fe2069-eb3c-4168-bb39-8aee043257b8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">train_loader = DataLoader(
    TensorDataset(
        torch.from_numpy(x_train),
        torch.from_numpy(y_train),
    ),
    batch_size=batch_size,
    shuffle=False,
)</code></pre></div><p>Adjacent sliding windows are highly dependent because they share most of their observations. Turning off shuffling preserves the temporal order and makes the training protocol easier to audit. This does not remove all dependence between windows, but it avoids adding a separate randomization step that could make the time-series experiment harder to interpret.</p><p>This is a conservative engineering choice rather than a claim that the paper used the same batching policy. The paper mentions shuffled or randomly arranged training behavior in places, but does not provide enough detail to establish a reproducible protocol. For trading conclusions, chronological and walk-forward evaluation remain more important than reproducing an unclear shuffling choice.</p><h3>8.7 Model-shape checks</h3><p>The generated tests focus on architecture contracts rather than empirical performance. For example, <code>tests/test_model_shapes.py</code> checks that a configured attention layer returns the expected context shape and that its weights normalize across time:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ba9492b8-8f98-45ec-84c0-8c3f4db8d7f2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@pytest.mark.skipif(not TORCH_AVAILABLE, reason="PyTorch is an optional dependency")
class TestAttentionShapes:
    def test_attention_weights_are_normalized_across_time(self) -&gt; None:
        attention = _construct_attention()
        sequence = torch.randn(3, 7, 8)
        output = attention(sequence)
        _, weights = _unpack_attention_output(output)

        if weights is None:
            weights = getattr(attention, "last_attention_weights", None)

        assert tuple(weights.shape) == (3, 7)
        assert torch.isfinite(weights).all()
        assert torch.all(weights &gt;= 0)
        assert torch.allclose(
            weights.sum(dim=1),
            torch.ones(3),
            atol=1e-5,
        )</code></pre></div><p>The test expresses the mathematical invariant directly:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;6b5e5919-28d3-42cb-bbff-52ac8ba63886&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">weights.shape == (batch, time)
weights &gt;= 0
sum(weights over time) == 1</code></pre></div><p>Additional shape tests check one-step and three-step outputs:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b41c9ec4-72af-4365-824f-f5e2c7cc5f63&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@pytest.mark.parametrize("horizon", [1, 3])
def test_attention_bilstm_output_horizon(self, horizon: int) -&gt; None:
    model = _construct_forecaster("AttBiLSTMForecaster", horizon=horizon)
    inputs = torch.randn(2, 100, 3)
    predictions = _model_output(model, inputs)

    assert predictions.shape[0] == 2
    assert predictions.shape[-1] == horizon</code></pre></div><p>These checks do not demonstrate that the model matches the paper's forecast metrics. They only verify the intended tensor interface and attention normalization when the optional dependency is installed. The local verification record should therefore be read as static and structural evidence, not as a claim that the model was executed successfully or that the paper's numerical results were reproduced.</p><h3>8.8 What is direct and what is reconstructed here?</h3><p>The implementation status of this model can be summarized as follows:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!H9Oq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79a4afda-86f8-4b46-9233-5d0855b5eaca_1408x976.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!H9Oq!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79a4afda-86f8-4b46-9233-5d0855b5eaca_1408x976.png 424w, /__u/substackcdn.com/image/fetch/$s_!H9Oq!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79a4afda-86f8-4b46-9233-5d0855b5eaca_1408x976.png 848w, /__u/substackcdn.com/image/fetch/$s_!H9Oq!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79a4afda-86f8-4b46-9233-5d0855b5eaca_1408x976.png 1272w, /__u/substackcdn.com/image/fetch/$s_!H9Oq!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79a4afda-86f8-4b46-9233-5d0855b5eaca_1408x976.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!H9Oq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79a4afda-86f8-4b46-9233-5d0855b5eaca_1408x976.png" width="1408" height="976" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/79a4afda-86f8-4b46-9233-5d0855b5eaca_1408x976.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:976,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:197770,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79a4afda-86f8-4b46-9233-5d0855b5eaca_1408x976.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!H9Oq!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79a4afda-86f8-4b46-9233-5d0855b5eaca_1408x976.png 424w, /__u/substackcdn.com/image/fetch/$s_!H9Oq!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79a4afda-86f8-4b46-9233-5d0855b5eaca_1408x976.png 848w, /__u/substackcdn.com/image/fetch/$s_!H9Oq!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79a4afda-86f8-4b46-9233-5d0855b5eaca_1408x976.png 1272w, /__u/substackcdn.com/image/fetch/$s_!H9Oq!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79a4afda-86f8-4b46-9233-5d0855b5eaca_1408x976.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The model is therefore a useful translation of the paper's stated architecture, but not a verified reproduction of its original neural network. Exact comparison would require the original data, preprocessing, target definition, full architecture, training hyperparameters, and untouched evaluation protocol.</p><h2>9. Walk Through Attention, Training, and Prediction Functions</h2><p>This section follows the generated Python implementation from tensor validation through attention, recurrent encoding, training, and prediction. The primary model is an attention-enhanced bidirectional LSTM, or Att-BiLSTM. The paper specifies the broad architecture, 20% dropout, MAE loss, and RMSprop optimization, but it does not specify several architectural and training details. Those details remain explicit reconstruction choices.</p><p>The examples below describe the generated interfaces and intended invariants. They do not claim that the code was executed, that the tests passed, or that the paper's empirical results were reproduced.</p><h3>9.1 Optional PyTorch imports</h3><p>PyTorch is an optional dependency. The data, streak, sizing, risk, and portfolio components are intended to remain usable without installing the deep-learning stack. The model modules therefore import PyTorch conditionally:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a0c771a3-c249-4014-a428-a71b3bafc1cd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">try:
    import torch
    from torch import Tensor, nn
except ImportError:
    torch = None
    Tensor = Any
    nn = type("nn", (), {"Module": _ModuleFallback})</code></pre></div><p>The fallback permits the module to be imported in a minimal installation. Constructing <code>DotProductTemporalAttention</code> or a recurrent forecaster still requires PyTorch and raises a clear <code>ImportError</code> if it is unavailable. This is preferable to silently substituting a different model.</p><p>The package initializer is also side-effect free. Importing <code>quantitative_trading_model</code> does not load market data, train a model, access a network, or create a trading connection.</p><h3>9.2 The attention layer's tensor contract</h3><p><code>DotProductTemporalAttention</code> expects a sequence tensor with shape:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;9cb4751e-cd89-4dea-a05b-0cf373d88b72&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">(batch, time, feature_dim)</code></pre></div><p>For the paper-inspired data setup, <code>time</code> is normally 100. The feature dimension is the size of each recurrent representation. A bidirectional LSTM commonly produces a representation whose width is twice its hidden size because its forward and backward states are concatenated.</p><p>The attention layer returns a context tensor with shape:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;3ecbf9f9-8875-484e-81fb-bfec024c77cd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">(batch, feature_dim)</code></pre></div><p>Attention weights have shape:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;f94ccadf-6c4a-4149-a7a9-fb4c9219f081&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">(batch, time)</code></pre></div><p>They can be returned explicitly or inspected after a forward pass. The weights are useful for diagnostics and shape tests, but they should not automatically be treated as a complete causal explanation of the model.</p><h3>9.3 Constructing attention and representing the missing query</h3><p>The paper gives the score equation</p><p>\[ S(x<em>i,q)=x</em>i^Tq, \]</p><p>but does not explain how the query vector <code>q</code> is created. The constructor exposes this ambiguity:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5ba9f8fd-31f5-4753-a142-cb17710c8266&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class DotProductTemporalAttention(nn.Module):
    def __init__(
        self,
        feature_dim: Optional[int] = None,
        *,
        input_dim: Optional[int] = None,
        learned_query: bool = True,
        query_init_scale: float = 0.02,
    ) -&gt; None:
        ...</code></pre></div><p><code>feature_dim</code> and <code>input_dim</code> are aliases. If both are supplied, they must agree. The dimension must be a positive integer because the unprojected dot product requires the query and sequence representations to have the same width.</p><p>With <code>learned_query=True</code>, the layer creates a trainable query parameter:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2754d81f-be66-458e-a4a3-bd2976eb20a7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">if learned_query:
    query = torch.empty(self.feature_dim)
    nn.init.normal_(query, mean=0.0, std=float(query_init_scale))
    self.query = nn.Parameter(query)
else:
    self.register_parameter("query", None)</code></pre></div><p>The learned query is a reconstruction choice, not a verified detail from the paper. With <code>learned_query=False</code>, the caller must provide a query to every forward call.</p><h3>9.4 Validating the sequence tensor</h3><p>Before computing attention scores, <code>_validate_sequence</code> checks the input shape and values:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f6461671-44da-4552-bafc-0842cd4911c4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def _validate_sequence(self, sequence: Tensor) -&gt; Tuple[int, int, int]:
    if sequence.ndim != 3:
        raise ValueError("sequence must have shape (batch, time, feature_dim)")

    batch, time, feature = (int(value) for value in sequence.shape)
    if batch &lt;= 0 or time &lt;= 0:
        raise ValueError("batch and time dimensions must be positive")
    if feature != self.feature_dim:
        raise ValueError("sequence feature dimension does not match configuration")
    if not sequence.is_floating_point() and not sequence.is_complex():
        raise TypeError("sequence must use a floating-point or complex dtype")
    if not torch.isfinite(sequence).all():
        raise ValueError("sequence contains NaN or infinite values")
    return batch, time, feature</code></pre></div><p>The checks protect separate assumptions:</p><ul><li><p><strong>Rank:</strong> the layer needs batch, time, and feature axes.</p></li><li><p><strong>Nonempty dimensions:</strong> an empty sequence cannot produce a context vector.</p></li><li><p><strong>Feature agreement:</strong> the dot product requires matching dimensions.</p></li><li><p><strong>Numeric type:</strong> the implementation technically accepts floating-point and complex PyTorch tensors.</p></li><li><p><strong>Finite values:</strong> NaNs and infinities could propagate through scores and softmax.</p></li></ul><p>In normal financial-model usage, inputs are real-valued floating-point tensors. Complex-tensor acceptance is a property of the generated validation condition, not a paper requirement or a recommended financial-data representation.</p><p>These checks do not verify temporal causality. A tensor can have the correct shape while containing future observations. Timestamp alignment and target-after-window rules remain responsibilities of the data and experiment layers.</p><h3>9.5 Preparing and broadcasting the query</h3><p>The query may be a single vector shared across the batch or a separate vector for every batch item:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;900eeb26-07ce-4c3c-a0e3-92670b54ba08&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def _prepare_query(self, sequence: Tensor, query: Optional[Tensor]) -&gt; Tensor:
    batch = sequence.shape[0]
    candidate = self.query if query is None else query

    if candidate is None:
        raise ValueError("query is required when learned_query=False")

    if candidate.ndim == 1:
        if candidate.shape[0] != self.feature_dim:
            raise ValueError("query dimension does not match feature_dim")
        candidate = candidate.unsqueeze(0).expand(batch, -1)
    elif candidate.ndim == 2:
        if candidate.shape != (batch, self.feature_dim):
            raise ValueError("batched query has the wrong shape")
    else:
        raise ValueError("query must have shape (feature_dim,) or (batch, feature_dim)")

    if candidate.device != sequence.device:
        raise ValueError("sequence and query must be on the same device")
    if candidate.dtype != sequence.dtype:
        raise ValueError("sequence and query must use the same dtype")
    if not torch.isfinite(candidate).all():
        raise ValueError("query contains NaN or infinite values")
    return candidate</code></pre></div><p>A one-dimensional query is expanded across the batch. A two-dimensional query must have one row per batch item. Device and dtype checks prevent errors such as combining tensors on different devices or with incompatible precision.</p><p>The method does not project either operand into another space. That follows the displayed paper equation, although the paper does not establish that its actual implementation avoided projections.</p><h3>9.6 Forward attention: scores, weights, and context</h3><p>The core <code>forward</code> method is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;04a88858-2ef0-46e9-aec8-00faec8e8289&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def forward(
    self,
    sequence: Tensor,
    query: Optional[Tensor] = None,
    *,
    return_weights: bool = False,
) -&gt; Tensor | Tuple[Tensor, Tensor]:
    self._validate_sequence(sequence)
    prepared_query = self._prepare_query(sequence, query)

    scores = torch.einsum("btd,bd-&gt;bt", sequence, prepared_query)
    weights = torch.softmax(scores, dim=1)
    context = torch.sum(weights.unsqueeze(-1) * sequence, dim=1)

    self._last_attention_weights = weights
    if return_weights:
        return context, weights
    return context</code></pre></div><p>The three central operations correspond to the paper's equations.</p><h4>Dot-product scores</h4><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;897b8c8c-8679-43f3-aef4-57d4af30bf19&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">scores = torch.einsum("btd,bd-&gt;bt", sequence, prepared_query)</code></pre></div><p>Here <code>b</code> is the batch index, <code>t</code> is the time index, and <code>d</code> is the feature index. The result has shape <code>(batch, time)</code>, with one scalar score for each historical time step.</p><h4>Softmax across time</h4><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a94cca76-9cfe-46d2-8bc8-4bc322c1f89a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">weights = torch.softmax(scores, dim=1)</code></pre></div><p><code>dim=1</code> is essential because the time axis is the second dimension. For each batch item:</p><p>\[ \sum<em>i a</em>i=1, \qquad a_i\geq 0. \]</p><p>Applying softmax across features would produce a different operation and would not assign relative weights to historical time steps.</p><h4>Weighted context</h4><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;58c1e6fa-fe20-4918-a364-4d0248ef3a76&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">context = torch.sum(weights.unsqueeze(-1) * sequence, dim=1)</code></pre></div><p><code>weights.unsqueeze(-1)</code> changes the weight shape from <code>(batch, time)</code> to <code>(batch, time, 1)</code>. Each representation is multiplied by its scalar attention weight, and summing over time produces one context vector per batch item.</p><p>The generated implementation stores the latest weights and checks that both weights and context are finite. Callers that retain weights across many batches should detach them first if they do not need gradients, otherwise the retained tensors may keep autograd graphs alive.</p><h3>9.7 Retrieving attention weights correctly</h3><p>The default call returns only the context:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;053ff704-4398-497f-b12f-a6d8d897ff51&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">context = attention(sequence)</code></pre></div><p>To receive both context and weights in one call, pass <code>return_weights=True</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;088130cc-b841-461b-a6f0-c38f44d3e09a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">context, weights = attention(sequence, return_weights=True)
assert context.shape == (sequence.shape[0], sequence.shape[2])
assert weights.shape == (sequence.shape[0], sequence.shape[1])</code></pre></div><p>Alternatively, the caller can use the inspection method:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;10c04f10-6126-440e-b67d-2cfcdced8611&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">context = attention(sequence)
weights = attention.last_attention_weights</code></pre></div><p>or recompute them explicitly:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8af9e4c8-2941-4d10-81ff-62c77423516b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">weights = attention.attention_weights(sequence)</code></pre></div><p>This distinction matters because <code>attention(sequence)</code> does <strong>not</strong> return a tuple by default. The generated model uses the default context-only call and reads the most recent weights from the attention layer when they are available.</p><h3>9.8 The shared recurrent baseline</h3><p><code>forecasters.py</code> defines <code>_RecurrentForecaster</code>, which supplies common behavior for LSTM, GRU, and non-attention BiLSTM baselines. The recurrent class is selected from the requested kind:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8f93fe83-66c7-468d-8aa9-6961b865024b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">recurrent_class = nn.LSTM if recurrent_kind == "lstm" else nn.GRU
self.encoder = recurrent_class(
    input_size=input_size,
    hidden_size=hidden_size,
    num_layers=num_layers,
    batch_first=True,
    dropout=recurrent_dropout,
    bidirectional=bidirectional,
)
self.dropout = nn.Dropout(dropout)
self.head = nn.Linear(self.output_size, horizon)</code></pre></div><p>The representation width is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d5d9c587-6844-43b5-8731-43efc05ac388&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">self.output_size = hidden_size * (2 if bidirectional else 1)</code></pre></div><p>The factor of two accounts for the forward and backward hidden streams in a bidirectional layer. The linear head maps the representation to the configured forecast horizon.</p><p>PyTorch recurrent dropout is applied between recurrent layers. Therefore the generated code uses:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;88704a97-4d4b-481f-b87f-b376b36824b1&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">recurrent_dropout = dropout if num_layers &gt; 1 else 0.0</code></pre></div><p>It also includes a separate <code>nn.Dropout(dropout)</code> stage. Whether the paper used recurrent dropout, feed-forward dropout, or both is unspecified, so this remains a reconstruction choice.</p><h3>9.9 Encoding and baseline prediction</h3><p>For recurrent baselines, the encoder returns the final time-step representation:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a8a79a16-704e-4b8b-90b4-73fbac1aa5f2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def encode(self, inputs: Tensor) -&gt; Tensor:
    _validate_input_tensor(inputs)
    sequence, _ = self.encoder(inputs)
    return self.dropout(sequence[:, -1, :])

def forward(self, inputs: Tensor) -&gt; Tensor:
    return self.head(self.encode(inputs))</code></pre></div><p>The input contract is <code>(batch, time, features)</code>. The data layer normally supplies <code>time=100</code>, but the model does not hard-code that number. The configured window builder determines the time length.</p><h3>9.10 Constructing the Att-BiLSTM</h3><p>The primary model returns a representation for every time step so attention can operate over the full historical window:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a9452f4d-34e2-4874-a787-c248c1fe3cf9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">self.encoder = nn.LSTM(
    input_size=input_size,
    hidden_size=hidden_size,
    num_layers=num_layers,
    batch_first=True,
    dropout=recurrent_dropout,
    bidirectional=True,
)
self.dropout = nn.Dropout(dropout)
self.attention = DotProductTemporalAttention(self.output_size)
self.head = nn.Linear(self.output_size, horizon)</code></pre></div><p>The tensor flow is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;df40d205-4e31-460a-8aac-ce5feae9cfbc&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">inputs                 (batch, time, input_size)
BiLSTM sequence        (batch, time, 2 * hidden_size)
dropout sequence       (batch, time, 2 * hidden_size)
attention context      (batch, 2 * hidden_size)
linear output head     (batch, horizon)</code></pre></div><p>The paper does not specify hidden size, recurrent depth, query construction, dense-layer structure, or output representation. These values must therefore be recorded as configuration metadata rather than presented as recovered paper details.</p><h3>9.11 Connecting the Att-BiLSTM to attention</h3><p>The primary model's forward path is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;da4c86aa-fddf-4608-90d8-601a0e12c131&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def forward(self, inputs: Tensor) -&gt; Tensor:
    _validate_input_tensor(inputs)
    sequence, _ = self.encoder(inputs)
    sequence = self.dropout(sequence)

    attended = self.attention(sequence)
    if isinstance(attended, tuple):
        context, weights = attended
        self._last_attention_weights = weights
    else:
        context = attended
        weights = getattr(self.attention, "last_attention_weights", None)
        if isinstance(weights, Tensor):
            self._last_attention_weights = weights

    return self.head(self.dropout(context))</code></pre></div><p>The intended sequence is:</p><p>Validate the input rank and values.</p><p>Process the historical window with the bidirectional LSTM.</p><p>Apply dropout to the sequence of hidden representations.</p><p>Use temporal attention to form one context vector.</p><p>Apply dropout to that context.</p><p>Map the context to one or more forecast values.</p><p>The tuple branch is defensive compatibility logic. The generated attention layer returns only the context unless <code>return_weights=True</code>, but the model can also accommodate an attention implementation that returns <code>(context, weights)</code>.</p><h3>9.12 Prediction methods and evaluation mode</h3><p>The model classes provide a prediction method that disables training behavior and gradient tracking:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;20c98dbb-a375-4588-abf2-5bc296e75db8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def predict(self, inputs: np.ndarray | Tensor) -&gt; np.ndarray:
    self.eval()
    with torch.no_grad():
        tensor = _as_input_tensor(inputs)
        return self(tensor).detach().cpu().numpy()</code></pre></div><p>The three operations have distinct purposes:</p><ul><li><p><code>self.eval()</code> disables training-specific behavior such as dropout.</p></li><li><p><code>torch.no_grad()</code> avoids constructing gradient graphs during inference.</p></li><li><p><code>.detach().cpu().numpy()</code> produces a NumPy result for metrics and signal generation.</p></li></ul><p><code>predict_forecaster</code> additionally extracts <code>X</code> from a windowed dataset and can apply an inverse transformation. The scaler must have been fitted using training data only. The model does not determine whether its target is a raw price, scaled price, return, or percentage change; that convention belongs to the configuration and experiment protocol.</p><h3>9.13 Model construction with <code>build_forecaster</code></h3><p>The factory normalizes common names and maps them to model classes:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f1aafa04-b8b5-4172-a1a1-66e27466ba47&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">aliases = {
    "att-bilstm": AttBiLSTMForecaster,
    "attbilstm": AttBiLSTMForecaster,
    "lstm": LSTMForecaster,
    "gru": GRUForecaster,
    "bilstm": BiLSTMForecaster,
}</code></pre></div><p>Unspecified settings are read from configuration:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2b212502-7280-44b6-9ae9-c9edd3343120&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">kwargs = {
    "hidden_size": hidden_size if hidden_size is not None else _config_value(config, "hidden_size", 64),
    "num_layers": num_layers if num_layers is not None else _config_value(config, "num_layers", 1),
    "dropout": dropout if dropout is not None else _config_value(config, "dropout", 0.20),
    "horizon": horizon if horizon is not None else _config_value(config, "horizon", 1),
}</code></pre></div><p>The 20% dropout default is paper-inspired. The default hidden size and layer count are engineering defaults, not findings recovered from the paper.</p><h3>9.14 Deterministic seeding</h3><p><code>train_forecaster</code> calls <code>_set_seed</code> before training:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c92ffd50-f64d-455c-a0d2-39fa2847d6ec&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def _set_seed(seed: int) -&gt; None:
    random.seed(seed)
    np.random.seed(seed)
    torch.manual_seed(seed)
    if torch.cuda.is_available():
        torch.cuda.manual_seed_all(seed)</code></pre></div><p>This aligns the main Python, NumPy, and PyTorch random generators. A seed improves repeatability but does not guarantee bit-for-bit equality across hardware, PyTorch versions, CUDA kernels, or data pipelines. It should be recorded as experiment metadata, not treated as proof of reproducibility.</p><h3>9.15 Extracting training data and normalizing target shape</h3><p>The <code>_extract_xy</code> helper accepts a <code>WindowedDataset</code>-like object exposing <code>X</code> and <code>y</code>, or a two-item <code>(X, y)</code> pair:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5c0f125f-ee82-4a9e-b996-223f64c01fda&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def _extract_xy(data: Any) -&gt; tuple[np.ndarray, np.ndarray]:
    if hasattr(data, "X") and hasattr(data, "y"):
        x, y = data.X, data.y
    elif isinstance(data, (tuple, list)) and len(data) == 2:
        x, y = data
    else:
        raise TypeError("data must expose X and y or be a two-item (X, y) pair")

    x_array = np.asarray(x, dtype=np.float32)
    y_array = np.asarray(y, dtype=np.float32)
    if x_array.ndim != 3:
        raise ValueError("X must have shape (samples, time, features)")
    if y_array.ndim == 1:
        y_array = y_array[:, None]
    if y_array.ndim != 2:
        raise ValueError("y must have shape (samples, horizon)")
    if len(x_array) != len(y_array):
        raise ValueError("X and y must contain the same number of samples")
    return x_array, y_array</code></pre></div><p>Converting a one-step target from <code>(samples,)</code> to <code>(samples, 1)</code> gives both one-step and three-step models the same output contract. The helper also checks for empty datasets and non-finite values. It does not accept a final test set implicitly; callers must pass training and validation partitions explicitly.</p><h3>9.16 Training with MAE and RMSprop</h3><p><code>train_forecaster</code> first checks that the target horizon agrees with the configuration and model:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;21e610da-6fa1-426b-a61c-2fc267c22446&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">horizon = int(_config_value(config, "horizon", y_train.shape[1]))
if y_train.shape[1] != horizon:
    raise ValueError("training target horizon does not match configured horizon")
if getattr(model, "horizon", horizon) != horizon:
    raise ValueError("model horizon and target horizon do not match")</code></pre></div><p>It then creates a chronological data loader:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;bf5212a0-60ec-4e48-a1f2-9945cd299ae2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">train_loader = DataLoader(
    TensorDataset(torch.from_numpy(x_train), torch.from_numpy(y_train)),
    batch_size=batch_size,
    shuffle=False,
)</code></pre></div><p><code>shuffle=False</code> is a deliberate time-series choice. Adjacent stride-one windows share most of their observations. Preserving order does not remove dependence, but it avoids adding random mixing and keeps the temporal protocol visible.</p><p>The paper-inspired loss and optimizer are configured as follows:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7271b12a-91c8-4c0d-80ec-094ff5147368&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">criterion = nn.L1Loss()
optimizer = torch.optim.RMSprop(model.parameters(), lr=learning_rate)</code></pre></div><p><code>nn.L1Loss()</code> implements mean absolute error. RMSprop updates model parameters; it is not an activation function. The learning rate and other optimizer details remain configurable because the paper does not provide them.</p><p>The core update loop is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;84cd3239-c6e7-4402-b191-756c4d5aa019&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">for batch_x, batch_y in train_loader:
    batch_x = batch_x.to(selected_device)
    batch_y = batch_y.to(selected_device)
    optimizer.zero_grad(set_to_none=True)
    predictions = model(batch_x)
    loss = criterion(predictions, batch_y)
    loss.backward()
    optimizer.step()</code></pre></div><p>The sequence is standard gradient-based training: move the batch, clear gradients, compute predictions, calculate MAE, backpropagate, and update with RMSprop. The generated implementation also checks that prediction and target shapes match before the update, preventing accidental broadcasting between one-step and multi-step targets.</p><h3>9.17 Validation loss, early stopping, and best-state restoration</h3><p>A validation dataset is optional. When supplied, it is evaluated without gradient calculation:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;334658b7-023e-4d98-9f7c-08977ee73153&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">model.eval()
with torch.no_grad():
    validation_prediction = model(validation_tensors[0])
    validation_loss = criterion(validation_prediction, validation_tensors[1])</code></pre></div><p>The final test partition should not be passed as validation data. Validation can guide early stopping, while the untouched test period remains reserved for final evaluation.</p><p>The function saves a copy of the best model state:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c7f79010-962d-4cb5-a01b-975304fc8df9&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">if monitored_loss &lt; best_loss - 1e-12:
    best_loss = monitored_loss
    history.best_epoch = epoch
    best_state = copy.deepcopy(model.state_dict())
    stale_epochs = 0
else:
    stale_epochs += 1
    if validation_tensors is not None and stale_epochs &gt; patience:
        history.stopped_early = True
        break</code></pre></div><p>After training, the best state is restored:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;99e5307e-9b9a-4024-b96b-ee0864030840&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">if best_state is not None:
    model.load_state_dict(best_state)</code></pre></div><p>This prevents the returned model from being an arbitrary final epoch when an earlier epoch had lower monitored loss. <code>TrainingHistory</code> records training loss, validation loss, the best epoch, and whether early stopping occurred.</p><p>When no validation set is supplied, training loss is monitored and validation-based early stopping is not used. This is a weaker selection protocol, so a research experiment should normally provide a separate validation period or use a documented walk-forward retraining procedure.</p><h3>9.18 Shape tests and the verification boundary</h3><p>The model-shape tests are intended to verify tensor contracts rather than financial performance. An API-accurate attention test is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;85edce4a-7634-48ee-bf3b-3b1051e81809&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">sequence = torch.randn(3, 7, 8)
context, weights = attention(sequence, return_weights=True)

assert tuple(context.shape) == (3, 8)
assert tuple(weights.shape) == (3, 7)
assert torch.all(weights &gt;= 0)
assert torch.allclose(weights.sum(dim=1), torch.ones(3), atol=1e-5)</code></pre></div><p>The equivalent context-only call is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;728483a7-e76b-4d70-aa14-06c8db6e6a28&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">context = attention(sequence)
weights = attention.last_attention_weights</code></pre></div><p>Recurrent tests use the paper-inspired time length and check one-step and three-step output contracts:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;948d45ec-6853-4658-b2e9-bfb2e291396e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">inputs = torch.randn(2, 100, 3)
predictions = model(inputs)

assert predictions.shape[0] == 2
assert predictions.shape[-1] == horizon</code></pre></div><p>These checks can detect wrong attention axes, shape regressions, missing dropout, and horizon mismatches. They cannot establish that the model learned useful financial structure, matched the paper's metrics, or produced a profitable strategy.</p><p>The available verification record must also be interpreted cautiously. Semantic code verification was skipped. Static checks reported findings involving placeholder detection in <code>experiments.py</code> and <code>cli.py</code>, and Markdown-fence checks involving the generated <code>README.md</code> and <code>docs/tutorial.md</code> code fields. In addition, some generated tests and orchestration calls have API inconsistencies with the generated modules. Consequently, this section documents intended interfaces and invariants; it does not claim a validated runtime contract, successful test execution, or empirical correctness.</p><h3>9.19 Practical model-use sequence</h3><p>A complete forecasting workflow should follow this order:</p><p>Clean and standardize local prices.</p><p>Establish a chronological training boundary.</p><p>Fit scaling parameters on training observations only.</p><p>Build 100-observation windows and future targets.</p><p>Split training and validation data without using final test targets.</p><p>Build <code>AttBiLSTMForecaster</code> with an explicit configuration.</p><p>Train with MAE and RMSprop.</p><p>Restore the best validation state through the training result.</p><p>Predict on untouched test windows.</p><p>Inverse-transform predictions if the target was scaled.</p><p>Compute forecast metrics separately from any portfolio backtest.</p><p>Pass only timestamp-aligned forecasts to the signal layer.</p><p>The model produces forecasts. It does not choose an order, determine a position size, apply VaR, or execute a trade. Those responsibilities belong to later modules.</p><h3>9.20 Reconstruction boundary</h3><p>The generated functions implement mechanics that can be stated precisely: recurrent sequence processing, dot-product attention, temporal softmax normalization, weighted context formation, dropout, MAE, RMSprop, and tensor-shape validation.</p><p>The following remain assumptions rather than verified paper details:</p><ul><li><p>hidden size and recurrent depth;</p></li><li><p>learned-query construction;</p></li><li><p>input feature set and target representation;</p></li><li><p>learning rate, batch size, epoch count, and stopping policy;</p></li><li><p>exact scaling procedure;</p></li><li><p>one-day versus three-day operational target;</p></li><li><p>hardware, random seed, and software-version details.</p></li></ul><p>Keeping these boundaries visible is part of the implementation. It prevents a clean model class from being mistaken for an exact reproduction of an incompletely specified experiment.</p><h2>10. Compare Recurrent, Statistical, HMM, and Boosted-Tree Baselines Carefully</h2><p>The paper compares its attention-enhanced BiLSTM with several other forecasting families: LSTM, GRU, BiLSTM, Holt-Winters, ARIMA, HMM, and XGBoost. This comparison is useful because it asks whether the attention-enhanced recurrent model adds value beyond simpler neural, statistical, state-based, and tree-based approaches.</p><p>However, a benchmark table is only meaningful when the models receive comparable data and are evaluated under the same target definition and timestamps. The paper does not provide enough configuration detail to reproduce its table exactly. The generated project therefore implements benchmark <strong>interfaces and adapters</strong>, while labeling locally computed metrics separately from the paper's reported claims.</p><p><strong>Reproduction boundary:</strong> The code below shows how the benchmark layer is intended to work. It does not establish that the generated code was executed, that every optional dependency is installed, or that the paper's numerical metrics were reproduced.</p><h3>10.1 A common result contract</h3><p>Different forecasting libraries return different model objects and prediction types. A benchmark experiment needs a common record so that results can be compared without losing important status information. The generated <code>BenchmarkResult</code> class in <code>src/quantitative_trading_model/models/benchmarks.py</code> provides that record:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f21a3649-9a18-4854-ae8f-eab6b988d478&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass
class BenchmarkResult:
    model_name: str
    predictions: np.ndarray | None = None
    metrics: dict[str, float] = field(default_factory=dict)
    status: str = "ok"
    message: str = ""
    metadata: dict[str, Any] = field(default_factory=dict)

    def to_dict(self, include_predictions: bool = False) -&gt; dict[str, Any]:
        output = {
            "model_name": self.model_name,
            "metrics": dict(self.metrics),
            "status": self.status,
            "message": self.message,
            "metadata": dict(self.metadata),
        }
        if include_predictions:
            output["predictions"] = (
                None if self.predictions is None
                else self.predictions.tolist()
            )
        return output</code></pre></div><p>The fields have distinct purposes:</p><ul><li><p><code>model_name</code> identifies the adapter or experiment label.</p></li><li><p><code>predictions</code> contains the aligned test forecasts when a model fitted successfully.</p></li><li><p><code>metrics</code> contains values such as RMSE, MAE, MAPE, and R-squared.</p></li><li><p><code>status</code> distinguishes a successful run from an unavailable optional package or a model failure.</p></li><li><p><code>message</code> preserves an actionable explanation instead of silently substituting another model.</p></li><li><p><code>metadata</code> records choices such as the backend, horizon, feature construction, or HMM interpretation.</p></li></ul><p>This status handling matters in a research environment. If <code>statsmodels</code> is not installed, the result should say that Holt-Winters or ARIMA is unavailable. It should not quietly replace the missing model with a naive forecast and then present the resulting number as an apples-to-apples comparison.</p><p>The <code>to_dict</code> method also provides a local serialization boundary. Predictions can be omitted from a summary table while retaining metrics, configuration metadata, and failure explanations.</p><h3>10.2 Holt-Winters: a statistical smoothing baseline</h3><p>Holt-Winters, also called exponential smoothing, models a series through components such as level, trend, and seasonality. For a financial price series, the generated adapter defaults to an additive trend and no seasonal component:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7cf95fdd-fee9-4ec3-b733-4e84056b2792&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class HoltWintersForecaster:
    name = "Holt-Winters"

    def __init__(
        self,
        trend: str | None = "add",
        seasonal: str | None = None,
        seasonal_periods: int | None = None,
        damped_trend: bool = False,
    ) -&gt; None:
        self.trend = trend
        self.seasonal = seasonal
        self.seasonal_periods = seasonal_periods
        self.damped_trend = damped_trend
        self._fit_result: Any = None

    def fit(self, values: ArrayLike) -&gt; "HoltWintersForecaster":
        series = _as_float_array(values, "training values")
        ...
        model = ExponentialSmoothing(
            series,
            trend=self.trend,
            damped_trend=(
                self.damped_trend if self.trend is not None else False
            ),
            seasonal=self.seasonal,
            seasonal_periods=self.seasonal_periods,
            initialization_method="estimated",
        )
        self._fit_result = model.fit(optimized=True, use_brute=True)
        return self

    def predict(self, horizon: int) -&gt; np.ndarray:
        if self._fit_result is None:
            raise RuntimeError("fit must be called before predict")
        return np.asarray(self._fit_result.forecast(horizon), dtype=float)</code></pre></div><p>The <code>fit</code> method validates the input, imports <code>statsmodels</code> only when needed, constructs the smoothing model, and stores the fitted result. The <code>predict</code> method requires a prior fit and asks the backend for a specified number of future observations.</p><p>Several choices here are not recoverable from the paper:</p><ul><li><p>whether the series has a meaningful seasonal period;</p></li><li><p>whether trend should be additive, multiplicative, damped, or absent;</p></li><li><p>how initialization should be performed;</p></li><li><p>whether hyperparameters were tuned on a validation set;</p></li><li><p>whether the model forecasted raw prices, returns, or transformed values.</p></li></ul><p>Consequently, a locally fitted Holt-Winters model is a documented benchmark configuration, not proof that it matches the paper's implementation.</p><h3>10.3 ARIMA and auto-ARIMA</h3><p>ARIMA models temporal dependence through autoregressive terms, differencing, and moving-average terms. The generated <code>ARIMAForecaster</code> supports either an explicit order or an optional <code>pmdarima.auto_arima</code> path:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;31c9554a-b774-498e-b9db-78594f5a5521&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class ARIMAForecaster:
    name = "ARIMA"

    def __init__(
        self,
        order: tuple[int, int, int] = (1, 1, 1),
        use_auto: bool = False,
        seasonal: bool = False,
        suppress_warnings: bool = True,
    ) -&gt; None:
        self.order = tuple(int(part) for part in order)
        self.use_auto = use_auto
        self.seasonal = seasonal
        self.suppress_warnings = suppress_warnings
        self._fit_result: Any = None
        self._backend = ""</code></pre></div><p>When <code>use_auto=False</code>, the adapter uses the configured <code>(p, d, q)</code> order with <code>statsmodels</code>. When <code>use_auto=True</code>, it requires <code>pmdarima</code> and delegates model-order selection to <code>auto_arima</code>. These are materially different experiments: a fixed ARIMA order and an automatically selected order should not be reported under one undifferentiated label.</p><p>The paper mentions <code>auto_arima</code>, but it does not state the search bounds, seasonal settings, differencing policy, information criterion, or whether the test period influenced model selection. Those omissions affect both accuracy and fairness. A benchmark must select its order using training data or a separate validation procedure, never by inspecting final test targets.</p><h3>10.4 HMM: state decoding is not automatically value forecasting</h3><p>A hidden Markov model represents observations as being generated by an unobserved state sequence. The Viterbi algorithm finds the most likely sequence of hidden states given the observations. That is a <strong>state-decoding</strong> task; it is not automatically a forecast of the next numerical price.</p><p>This distinction corrects an ambiguity in the paper. Saying that an HMM uses Viterbi does not, by itself, specify how the next gold or Bitcoin price is calculated. A value forecast needs an additional observation model or transition-based prediction rule.</p><p>The generated adapter makes its reconstruction explicit. It fits a Gaussian HMM to one-step price changes, estimates the mean change associated with each latent state, and uses transition probabilities to estimate a future change. Its metadata records that interpretation:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e59161e8-e52c-4325-968b-f42815c4a0a0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class HMMForecaster:
    name = "HMM"

    def fit(self, values: ArrayLike) -&gt; "HMMForecaster":
        series = _as_float_array(values, "training values")
        changes = np.diff(series).reshape(-1, 1)
        model = GaussianHMM(
            n_components=self.n_states,
            covariance_type="diag",
            n_iter=self.n_iter,
            random_state=self.random_state,
        )
        model.fit(changes)
        states = model.predict(changes)
        ...
        self._model = model
        self._state_means = means
        self._last_value = float(series[-1])
        return self</code></pre></div><p>The important design choice is not the exact HMM implementation; it is the separation between:</p><p>fitting latent states;</p><p>decoding or identifying a current state;</p><p>converting state transitions into an expected future change;</p><p>adding that change to a price level.</p><p>The paper does not specify the number of states, emissions, covariance structure, observation variable, or prediction equation. Therefore, the HMM result must be labeled <strong>reconstructed</strong>, not presented as a verified implementation of the paper's HMM procedure.</p><h3>10.5 XGBoost is gradient boosting, not a random forest</h3><p>The paper's description of XGBoost as reducing computation through random forest is technically inaccurate. XGBoost is an implementation of gradient-boosted decision trees. A random forest builds many independently randomized trees and averages them; gradient boosting builds trees sequentially, with later trees attempting to correct earlier errors.</p><p>The generated <code>XGBoostForecaster</code> uses lagged or flattened window features:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f8999c48-51c6-4f47-9b1a-1193e03e6044&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class XGBoostForecaster:
    name = "XGBoost"

    @staticmethod
    def _features(values: Any) -&gt; np.ndarray:
        array = np.asarray(values, dtype=float)
        if array.ndim == 1:
            array = array.reshape(-1, 1)
        elif array.ndim == 3:
            array = array.reshape(array.shape[0], -1)
        elif array.ndim != 2:
            raise ValueError(
                "XGBoost features must be one-, two-, or three-dimensional"
            )
        return array</code></pre></div><p>A recurrent network consumes the temporal dimension as a sequence. A tree model generally needs a two-dimensional table, so a window such as <code>(samples, 100, features)</code> is flattened into <code>(samples, 100 * features)</code>. That transformation is a modeling decision and should be recorded because it affects comparability.</p><p>The adapter also supports recursive prediction when no future feature rows are supplied. That behavior is only a practical reconstruction: the paper does not say whether XGBoost used lag features, technical indicators, rolling statistics, or another feature representation. A rigorous comparison should define the feature matrix before looking at test targets and should ensure that every lag uses information available at the corresponding timestamp.</p><h3>10.6 Optional dependencies are part of the result</h3><p>The benchmark module imports optional libraries inside the methods that need them. This keeps the core package usable when only NumPy and pandas are installed. The adapter catches missing dependencies and creates an explicit unavailable result:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;cbbc466c-aee8-4d07-a7c5-c164eba54d9d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">try:
    prediction = np.asarray(
        _fit_and_predict(model, train_values, test_values, horizon),
        dtype=float,
    ).reshape(-1)
    metrics = _metric_dict(actual, prediction)
    results[model_name] = BenchmarkResult(
        model_name=model_name,
        predictions=prediction,
        metrics=metrics,
        metadata=model_metadata,
    )
except _OptionalDependencyError as exc:
    results[model_name] = BenchmarkResult(
        model_name=model_name,
        status="unavailable",
        message=str(exc),
        metadata=common_metadata,
    )
except Exception as exc:
    results[model_name] = BenchmarkResult(
        model_name=model_name,
        status="failed",
        message=f"{type(exc).__name__}: {exc}",
        metadata=common_metadata,
    )</code></pre></div><p>This structure distinguishes at least three outcomes:</p><ul><li><p><strong>`ok`</strong>: the adapter returned predictions and metrics;</p></li><li><p><strong>`unavailable`</strong>: an optional package was not installed;</p></li><li><p><strong>`failed`</strong>: the package was available, but fitting or prediction raised an error.</p></li></ul><p>That distinction should remain visible in reports. A table containing only numeric columns can make a missing model look like a successful zero-result experiment.</p><p>The project's <code>pyproject.toml</code> separates deep-learning dependencies from classical benchmark dependencies. This supports lightweight use of data preparation, streak analysis, sizing, risk, and accounting without requiring PyTorch, <code>statsmodels</code>, <code>pmdarima</code>, <code>hmmlearn</code>, or XGBoost.</p><h3>10.7 Shared metric calculation</h3><p>The adapters pass predictions through the common metrics layer instead of implementing a different error formula for each library:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;fa712bbd-62a6-4f25-b04b-becb9cfaf51d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def _metric_dict(actual: ArrayLike, predicted: ArrayLike) -&gt; dict[str, float]:
    result: ForecastMetrics = forecast_metrics(actual, predicted)
    return {
        "rmse": float(result.rmse),
        "mae": float(result.mae),
        "mape_percent": float(result.mape_percent),
        "r_squared": float(result.r_squared),
    }</code></pre></div><p>The exact generated helper also preserves the metric object's serializable fields. The important contract is that every model is evaluated with the same definitions from <code>src/quantitative_trading_model/metrics.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e600fba8-8aa4-4d86-b50d-18a694612c56&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class ForecastMetrics:
    rmse: float
    mae: float
    mape_percent: float
    r_squared: float
    sample_count: int
    mape_sample_count: int
    mape_epsilon: float = 1e-12
    mape_zero_policy: str = "exclude"</code></pre></div><p>The metrics layer handles several edge cases explicitly:</p><ul><li><p>RMSE and MAE require equal-length, finite target and prediction arrays.</p></li><li><p>MAPE excludes targets at or below a configured near-zero threshold by default.</p></li><li><p>A caller can instead request an error when MAPE encounters a near-zero target.</p></li><li><p>R-squared has defined behavior for a constant target series: exact predictions receive <code>1.0</code>, while non-exact predictions receive <code>0.0</code>.</p></li><li><p>Metrics are computed on the same target scale used for comparison, normally after inverse transformation from a training-fitted scaler.</p></li></ul><p>MAPE units also need attention. The implementation reports percentage points, so <code>2.5</code> means 2.5%, not <code>0.025</code>. The paper's reported MAPE values should not be compared blindly until the target scale and unit convention are confirmed.</p><h3>10.8 What makes a benchmark comparison fair?</h3><p>A fair comparison requires more than calling eight <code>fit</code> methods. At minimum, document and align the following:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!nWRf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e62f8e8-0ee2-4b65-8029-0bfa27e689ab_1216x744.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!nWRf!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e62f8e8-0ee2-4b65-8029-0bfa27e689ab_1216x744.png 424w, /__u/substackcdn.com/image/fetch/$s_!nWRf!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e62f8e8-0ee2-4b65-8029-0bfa27e689ab_1216x744.png 848w, /__u/substackcdn.com/image/fetch/$s_!nWRf!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e62f8e8-0ee2-4b65-8029-0bfa27e689ab_1216x744.png 1272w, /__u/substackcdn.com/image/fetch/$s_!nWRf!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e62f8e8-0ee2-4b65-8029-0bfa27e689ab_1216x744.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!nWRf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e62f8e8-0ee2-4b65-8029-0bfa27e689ab_1216x744.png" width="1216" height="744" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7e62f8e8-0ee2-4b65-8029-0bfa27e689ab_1216x744.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:744,&quot;width&quot;:1216,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:138090,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e62f8e8-0ee2-4b65-8029-0bfa27e689ab_1216x744.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!nWRf!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e62f8e8-0ee2-4b65-8029-0bfa27e689ab_1216x744.png 424w, /__u/substackcdn.com/image/fetch/$s_!nWRf!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e62f8e8-0ee2-4b65-8029-0bfa27e689ab_1216x744.png 848w, /__u/substackcdn.com/image/fetch/$s_!nWRf!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e62f8e8-0ee2-4b65-8029-0bfa27e689ab_1216x744.png 1272w, /__u/substackcdn.com/image/fetch/$s_!nWRf!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e62f8e8-0ee2-4b65-8029-0bfa27e689ab_1216x744.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The generated <code>run_benchmark_suite</code> function attempts to provide one uniform workflow: extract training and test targets, fit each requested model on training data, predict the test interval, verify prediction length, calculate common metrics, and store implementation metadata. <code>compare_forecasts</code> then creates a stable pandas table and retains failed or unavailable models with status fields.</p><p>The common interface improves auditability, but it does not erase model-specific differences. An ARIMA model trained on a continuous series, an XGBoost model trained on flattened lag windows, and an HMM trained on price changes do not necessarily solve exactly the same statistical problem. Their input construction and forecast semantics must be described alongside their metric values.</p><h3>10.9 Interpreting the paper's benchmark table</h3><p>The paper reports Att-BiLSTM as the best listed predictor for both gold and Bitcoin. The extracted benchmark claims include values such as:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!IU4Y!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8a8295b-fe02-42f3-89cc-6df1cc49d3dd_1010x224.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!IU4Y!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8a8295b-fe02-42f3-89cc-6df1cc49d3dd_1010x224.png 424w, /__u/substackcdn.com/image/fetch/$s_!IU4Y!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8a8295b-fe02-42f3-89cc-6df1cc49d3dd_1010x224.png 848w, /__u/substackcdn.com/image/fetch/$s_!IU4Y!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8a8295b-fe02-42f3-89cc-6df1cc49d3dd_1010x224.png 1272w, /__u/substackcdn.com/image/fetch/$s_!IU4Y!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8a8295b-fe02-42f3-89cc-6df1cc49d3dd_1010x224.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!IU4Y!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8a8295b-fe02-42f3-89cc-6df1cc49d3dd_1010x224.png" width="1010" height="224" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b8a8295b-fe02-42f3-89cc-6df1cc49d3dd_1010x224.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:224,&quot;width&quot;:1010,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:40932,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8a8295b-fe02-42f3-89cc-6df1cc49d3dd_1010x224.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!IU4Y!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8a8295b-fe02-42f3-89cc-6df1cc49d3dd_1010x224.png 424w, /__u/substackcdn.com/image/fetch/$s_!IU4Y!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8a8295b-fe02-42f3-89cc-6df1cc49d3dd_1010x224.png 848w, /__u/substackcdn.com/image/fetch/$s_!IU4Y!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8a8295b-fe02-42f3-89cc-6df1cc49d3dd_1010x224.png 1272w, /__u/substackcdn.com/image/fetch/$s_!IU4Y!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8a8295b-fe02-42f3-89cc-6df1cc49d3dd_1010x224.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>These numbers are preserved in the project documentation and reproduction matrix as <strong>unverified claims</strong>. They are not hard-coded acceptance targets for the benchmark adapters. The source data, preprocessing, architecture, hyperparameters, target scale, and test protocol are unavailable, so a different local experiment can legitimately produce different values without indicating a coding error.</p><p>The reproduction matrix makes this distinction explicit:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;markdown&quot;,&quot;nodeId&quot;:&quot;4987ded2-1ccc-4078-a5f5-e6ddc912c3a6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-markdown">| Reported Att-BiLSTM gold metrics | Reported claim / Unavailable |
|---|---|
| RMSE 19.01, MAE 13.98, MAPE 0.7626, R-squared 0.9392 | Source data, preprocessing, target definition, and test protocol are missing. |</code></pre></div><p>The same rule applies to the paper's LSTM, GRU, Holt-Winters, ARIMA, HMM, and XGBoost figures. A local comparison table can show what the supplied implementation produced under a documented protocol, but it cannot be called a reproduction of the paper's table until the missing inputs and decisions are recovered.</p><h3>10.10 Forecast benchmarks are not trading benchmarks</h3><p>Even a successful forecasting comparison is only one layer of the larger pipeline. The model with the lowest RMSE might not produce the best portfolio because:</p><ul><li><p>small forecast errors may not overcome transaction fees;</p></li><li><p>the best numerical forecast may have poor directional timing;</p></li><li><p>a position-sizing rule can amplify a modest error;</p></li><li><p>different forecast horizons can produce different trading turnover;</p></li><li><p>execution timing and slippage can reverse the order of model rankings;</p></li><li><p>a high R-squared on a trending price level does not guarantee useful return forecasts.</p></li></ul><p>For this reason, benchmark reporting should have two separate outputs:</p><p><strong>Forecast comparison:</strong> target-level metrics using aligned test observations.</p><p><strong>Trading comparison:</strong> portfolios generated from each model using the same signal, sizing, risk, fee, execution, and benchmark protocol.</p><p>The paper's reported Att-BiLSTM superiority addresses the first category as a claim. It does not, by itself, validate the second category.</p><h3>10.11 Practical benchmark checklist</h3><p>Before comparing local benchmark results, verify:</p><ul><li><p>Every model is trained only on its permitted training data.</p></li><li><p>The final test targets are not used for model selection or hyperparameter tuning.</p></li><li><p>All models forecast the same target representation and horizon, where comparison is intended.</p></li><li><p>Forecast timestamps are aligned before calculating metrics.</p></li><li><p>Price scaling is fitted on training data only and predictions are inverse-transformed consistently.</p></li><li><p>MAPE near-zero behavior is documented.</p></li><li><p>HMM state decoding is not mislabeled as a price forecast.</p></li><li><p>XGBoost is described as gradient boosting, not random forest.</p></li><li><p>Optional dependency failures remain visible in the results.</p></li><li><p>Model-specific features and hyperparameters are recorded.</p></li><li><p>Forecast metrics are reported separately from portfolio outcomes.</p></li><li><p>Paper numbers are labeled as reported claims unless independently reproduced.</p></li></ul><p>The generated benchmark adapters provide a useful scaffold for this process. Their main contribution is not a promise of identical numbers; it is a uniform, inspectable place to document how each alternative model was trained, what it predicted, whether it was available, and how its metrics were calculated.</p><h2>11. Measure Returns and Consecutive Rise/Decline Streaks</h2><p>The paper uses consecutive price increases and decreases to inform its position-management model. In practical terms, the analysis asks how often positive or negative price movements persist and how large representative gains and declines are.</p><p>The paper calls this procedure <strong>Apriori</strong>, but the extracted description does not define itemsets, support, confidence, or association rules. Standard Apriori is a frequent-itemset-mining algorithm, whereas the described procedure counts consecutive movements. The implementation therefore treats this component as <strong>streak analysis</strong> or <strong>sequential-pattern analysis</strong>. This is a terminology and implementation correction, not a claim that the paper's original authors intended a different method.</p><h3>11.1 Define the return convention first</h3><p>For adjacent prices, the standard simple return is:</p><p>\[ r<em>t = \frac{p</em>t-p<em>{t-1}}{p</em>{t-1}}. \]</p><p>The return is aligned with the later timestamp: it describes the move from the previous observation into <code>p_t</code>. For <code>[100, 110, 99]</code>, the returns are:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;839e4d1c-36b4-49fd-9e88-096eab75e89a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">100 -&gt; 110: (110 - 100) / 100  =  0.10
110 -&gt;  99: ( 99 - 110) / 110  = -0.10</code></pre></div><p>The extracted paper equation appears to use the current price as the denominator:</p><p>\[ \tilde r<em>t = \frac{p</em>t-p<em>{t-1}}{p</em>t}. \]</p><p>Because the equation layout is ambiguous, <code>compute_returns</code> uses the conventional previous-price denominator by default and exposes the current-price form through an explicit option. The choice must be recorded and used consistently for classification, percentile estimation, sizing calibration, and any backtest signal.</p><p>The relevant implementation is in <code>src/quantitative_trading_model/streaks.py</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;56e4cffa-cbee-4fe5-aa71-493fffc16f18&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def compute_returns(
    prices: Iterable[float],
    denominator: ReturnDenominator = "previous",
) -&gt; np.ndarray:
    values = _as_float_array(prices, name="prices")
    if np.any(values &lt;= 0.0):
        raise ValueError("prices must be strictly positive")

    if values.size &lt; 2:
        return np.empty(0, dtype=float)

    normalized = str(denominator).lower()
    if normalized in {"previous", "previous_price"}:
        base = values[:-1]
    elif normalized in {"current", "current_price", "paper"}:
        base = values[1:]
    else:
        raise ValueError(
            "denominator must be 'previous', 'current', or 'paper'"
        )

    return (values[1:] - values[:-1]) / base</code></pre></div><p>The function checks that prices are one-dimensional, finite, and strictly positive. A sequence with fewer than two prices has no adjacent return. Rejecting nonpositive prices is important because the return denominator would otherwise be invalid.</p><h3>11.2 Classify rises, declines, and zero changes</h3><p>Returns are classified numerically:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!9-gp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ab4e5eb-8279-4cba-aea3-74894d7bfa24_450x266.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!9-gp!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ab4e5eb-8279-4cba-aea3-74894d7bfa24_450x266.png 424w, /__u/substackcdn.com/image/fetch/$s_!9-gp!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ab4e5eb-8279-4cba-aea3-74894d7bfa24_450x266.png 848w, /__u/substackcdn.com/image/fetch/$s_!9-gp!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ab4e5eb-8279-4cba-aea3-74894d7bfa24_450x266.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9-gp!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ab4e5eb-8279-4cba-aea3-74894d7bfa24_450x266.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!9-gp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ab4e5eb-8279-4cba-aea3-74894d7bfa24_450x266.png" width="450" height="266" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5ab4e5eb-8279-4cba-aea3-74894d7bfa24_450x266.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:266,&quot;width&quot;:450,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:23889,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ab4e5eb-8279-4cba-aea3-74894d7bfa24_450x266.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!9-gp!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ab4e5eb-8279-4cba-aea3-74894d7bfa24_450x266.png 424w, /__u/substackcdn.com/image/fetch/$s_!9-gp!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ab4e5eb-8279-4cba-aea3-74894d7bfa24_450x266.png 848w, /__u/substackcdn.com/image/fetch/$s_!9-gp!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ab4e5eb-8279-4cba-aea3-74894d7bfa24_450x266.png 1272w, /__u/substackcdn.com/image/fetch/$s_!9-gp!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ab4e5eb-8279-4cba-aea3-74894d7bfa24_450x266.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><code>classify_returns</code> returns integer labels, not strings such as <code>"positive"</code> or <code>"neutral"</code>. Keeping the labels numeric makes run detection straightforward and avoids ambiguity in comparisons.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;0c84c01f-1a0e-4537-be34-6aad6f33362a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def classify_returns(
    returns: Iterable[float],
    *,
    zero_policy: ZeroPolicy = "neutral",
    tolerance: float = 0.0,
) -&gt; np.ndarray:
    values = _as_float_array(returns, name="returns")
    labels = np.zeros(values.size, dtype=np.int8)
    labels[values &gt; tolerance] = 1
    labels[values &lt; -tolerance] = -1

    if zero_policy == "raise" and np.any(labels == 0):
        raise ValueError("zero or near-zero returns are not allowed")
    if zero_policy == "ignore":
        labels = labels[labels != 0]
    return labels</code></pre></div><p>The zero policies are:</p><ul><li><p><code>neutral</code>: retain zero labels, so they break maximal runs;</p></li><li><p><code>terminate</code>: treat zero changes as streak boundaries;</p></li><li><p><code>ignore</code>: remove zero labels before counting, making surrounding nonzero observations adjacent;</p></li><li><p><code>raise</code>: reject any zero or near-zero return.</p></li></ul><p>The default is <code>neutral</code>. A nonzero <code>tolerance</code> can classify very small returns as zero; that setting should be included in experiment metadata because it changes the streak counts.</p><p>For example:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;7d9e7962-db72-4a98-aa28-5c5d24290b8e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">returns:  +   +   -   -   -   +   0   +
labels:    1   1  -1  -1  -1   1   0   1</code></pre></div><p>Under the default policy, the zero separates the final positive movements. The positive runs have lengths two, one, and one, while the negative run has length three.</p><h3>11.3 Maximal runs and overlapping subsequences</h3><p>&#8220;Two consecutive rises&#8221; can refer to different statistics. The implementation supports both interpretations.</p><p>A <strong>maximal run</strong> is a complete uninterrupted sequence of one sign. For:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;0536cf61-b53e-496c-ae67-03965c7e06d6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">1, 1, -1, -1, -1, 1</code></pre></div><p>the maximal runs are:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;ebe958e4-eb37-4346-b12a-31e9fbee0a32&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">rise of length 2
decline of length 3
rise of length 1</code></pre></div><p>The corresponding summaries are:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;24c3786e-d9bb-4b46-9ddb-324a4659e4e5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">rise_streaks = {1: 1, 2: 1}
decline_streaks = {3: 1}</code></pre></div><p>This answers: &#8220;How many complete runs of each exact length occurred?&#8221;</p><p>An <strong>overlapping subsequence</strong> count includes shorter patterns inside longer runs. A run of four rises contains three two-rise subsequences:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;ec5ad74d-a8a7-41c9-9a06-2379ffc8639d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">[1, 1] at positions 1-2
[1, 1] at positions 2-3
[1, 1] at positions 3-4</code></pre></div><p>It also contains two three-rise subsequences and one four-rise subsequence. This answers: &#8220;How many contiguous patterns of each length occurred?&#8221;</p><p>The public function makes the distinction explicit:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;0ba16f06-dc01-4e4d-bf6d-615eb4b73e1d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def count_streaks(
    labels: Iterable[int],
    *,
    value: int,
    mode: StreakMode = "maximal",
    minimum_length: int = 1,
) -&gt; dict[int, int]:
    result: dict[int, int] = {}
    run_length = 0

    for label in labels:
        if int(label) == value:
            run_length += 1
            continue

        if run_length:
            _add_run(
                result,
                run_length,
                mode=mode,
                minimum_length=minimum_length,
            )
            run_length = 0

    if run_length:
        _add_run(
            result,
            run_length,
            mode=mode,
            minimum_length=minimum_length,
        )
    return dict(sorted(result.items()))</code></pre></div><p>The required <code>value</code> argument specifies whether to count rises (<code>value=1</code>) or declines (<code>value=-1</code>). In maximal mode, a run contributes one count for its exact length. In overlapping mode, a run of length <code>r</code> contributes <code>r-j+1</code> occurrences for each subsequence length <code>j</code>.</p><p>The paper does not specify which counting convention it uses. Both are therefore implementation interpretations, not verified transcriptions. A research result should record the selected mode and minimum streak length.</p><h3>11.4 Gain and decline percentiles</h3><p>The paper connects streak analysis with percentile statistics used by its position-management method. The implementation keeps these quantities separate:</p><ul><li><p>the 90th percentile of positive returns;</p></li><li><p>the 10th percentile of signed negative returns;</p></li><li><p>the median positive return;</p></li><li><p>the 10th percentile of positive decline magnitudes;</p></li><li><p>the median decline magnitude.</p></li></ul><p>For a negative return of <code>-0.04</code>, the signed decline is <code>-0.04</code>, while its magnitude is <code>0.04</code>. Both forms are useful: the signed value preserves the return convention, and the magnitude is convenient for a sizing calculation that expects a nonnegative loss amount.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e8108884-4793-48c1-9f50-2ae461192fa2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def estimate_gain_decline_percentiles(
    returns: Iterable[float],
    *,
    gain_percentile: float = 90.0,
    decline_percentile: float = 10.0,
) -&gt; dict[str, float]:
    values = _as_float_array(returns, name="returns")
    gains = values[values &gt; 0.0]
    declines = values[values &lt; 0.0]
    decline_magnitudes = -declines

    return {
        "gain_percentile": _percentile(gains, gain_percentile),
        "decline_percentile": _percentile(
            declines, decline_percentile
        ),
        "gain_median": _percentile(gains, 50.0),
        "decline_magnitude_percentile": _percentile(
            decline_magnitudes, decline_percentile
        ),
        "decline_magnitude_median": _percentile(
            decline_magnitudes, 50.0
        ),
    }</code></pre></div><p>For the return history:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;b63b8449-69a5-4a2f-a093-6d5cf11e0e01&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">[0.01, 0.02, 0.04, -0.01, -0.03, -0.06]</code></pre></div><p>the gain sample is <code>[0.01, 0.02, 0.04]</code>, and the signed decline sample is <code>[-0.01, -0.03, -0.06]</code>. The magnitude sample is <code>[0.01, 0.03, 0.06]</code>. Empty gain or decline samples produce <code>NaN</code>, rather than an invented zero.</p><p>These statistics are descriptive inputs to the later normalized exponential sizing reconstruction. They are not forecasts and do not imply that the next movement will match a historical percentile.</p><h3>11.5 The combined summary record</h3><p><code>compute_streak_statistics</code> combines return calculation, classification, run counting, and percentile estimation. It returns a frozen <code>StreakStatistics</code> record containing the arrays, mappings, scalar statistics, denominator, zero policy, and streak mode.</p><p>&#8220;Frozen&#8221; means that the dataclass fields cannot be reassigned through normal dataclass operations. It does <strong>not</strong> make the stored NumPy arrays deeply immutable: a caller can still mutate an array in place unless it copies the arrays or marks them read-only. The record should therefore be treated as an audit-oriented summary object, not as a complete deep-immutability guarantee.</p><p>The returned record includes fields such as:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6228a63f-9fc8-42c6-abb2-ae995fd144fa&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">statistics = compute_streak_statistics(
    prices,
    denominator="previous",
    zero_policy="neutral",
    streak_mode="maximal",
)

print(statistics.rise_streaks)
print(statistics.decline_streaks)
print(statistics.gain_percentile_90)
print(statistics.decline_magnitude_percentile_10)</code></pre></div><p>Its <code>to_dict</code> method is intended for local result files and preserves the selected conventions and count summaries.</p><h3>11.6 Use only pre-decision information online</h3><p>A full-history summary is useful for exploration but is not automatically valid for an out-of-sample strategy. If a signal is generated at time <code>t</code>, the statistics used to choose its position size must be estimated from returns known no later than <code>t</code>.</p><p>This is unsafe:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;54635223-2ae3-4b29-ab4a-ea2c67db2785&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">compute returns over the complete dataset
compute percentiles over the complete dataset
use those percentiles to size trades in the earlier test period</code></pre></div><p>Future observations have influenced the historical calibration. The safer walk-forward pattern is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;bd87d622-f574-40a0-b880-b3746f43f98f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">for each decision timestamp t:
    use returns observed before t
    calculate or update streak statistics
    choose the next sizing level
    generate the order
    execute at a later permitted bar</code></pre></div><p>The same rule applies when passing calibration statistics to <code>build_exponential_schedule</code> in <code>sizing.py</code>. A statistic calculated once from the entire dataset should not be used as though it had been available at every earlier decision time.</p><h3>11.7 What the generated tests currently represent</h3><p><code>tests/test_data_and_streaks.py</code> is intended to document small, hand-calculable invariants such as the previous-price return convention and the difference between maximal and overlapping counts. However, the generated test file currently contains known API mismatches with the generated <code>streaks.py</code> implementation:</p><p><code>classify_returns</code> returns integer labels <code>1</code>, <code>-1</code>, and <code>0</code>, while some test assertions expect textual labels such as <code>"positive"</code>, <code>"negative"</code>, and <code>"neutral"</code>.</p><p><code>count_streaks</code> requires the keyword-only <code>value</code> argument, but the generated streak-count test calls it without <code>value=1</code> or <code>value=-1</code>.</p><p>These are <strong>planned/static test artifacts with known mismatches</strong>, not evidence of executable verification. They should be repaired before being treated as runnable tests: assertions should use numeric labels, and rise and decline counts should call <code>count_streaks(labels, value=1, ...)</code> and <code>count_streaks(labels, value=-1, ...)</code> separately.</p><p>Semantic code verification was skipped, and this tutorial does not claim that the tests were executed, that they pass, or that the paper's empirical results were reproduced. The useful review targets remain the return convention, zero policy, run-counting mode, percentile population, and pre-decision information boundary.</p><h3>11.8 Summary</h3><p>The paper's consecutive-movement component can be translated into an auditable pipeline:</p><p>Compute adjacent returns with an explicitly recorded denominator.</p><p>Classify each return as a rise, decline, or zero.</p><p>Decide how zero returns affect streak boundaries.</p><p>Count maximal runs or overlapping subsequences.</p><p>Estimate gain, signed-decline, and decline-magnitude percentiles.</p><p>Pass statistics to the later sizing reconstruction only when they could have been known at the decision timestamp.</p><p>The implementation in <code>streaks.py</code> provides these mechanics, but it does not claim to recover a standard Apriori algorithm or the paper's exact counting procedure. The appropriate description is <strong>sequential streak analysis feeding a documented, finite position-sizing reconstruction</strong>.</p><h2>12. Reconstruct the Exponential Position-Sizing Rule Safely</h2><p>The paper's position-management component addresses a different question from forecasting: <strong>after a strategy decides to trade, how much capital should it commit?</strong> The paper appears to increase later additions after successive declines, using empirical gains, declines, transaction costs, and an exponential sizing curve.</p><p>This section is not an exact reproduction of that method. The extracted forms of equations (5) through (8) contain damaged notation and do not define several important variables. The implementation in <code>src/quantitative_trading_model/sizing.py</code> therefore provides a finite, normalized-exponential alternative with explicit budget and exposure controls.</p><p><strong>Reconstruction boundary:</strong> The schedule below is an auditable interpretation of the paper's sizing idea. It is not a faithful transcription of equations (5)&#8211;(8), and it does not establish that averaging down is safe or profitable.</p><h3>12.1 Why the paper's recursive equations cannot be transcribed faithfully</h3><p>The paper appears to derive later capital additions from a recovery or break-even condition: after one or more declines, a future gain should compensate for earlier losses and transaction costs. It then represents the resulting additions with an exponential curve.</p><p>The extracted equations are insufficient to reproduce that calculation reliably:</p><ul><li><p><code>p_i</code> is not clearly defined as cash, asset quantity, or notional exposure.</p></li><li><p>Signs and powers in the recursive expressions are corrupted.</p></li><li><p>The fee convention is incomplete.</p></li><li><p>The independent variable and exponent coefficient in the printed exponential expression are unclear.</p></li><li><p>The maximum number of additions is not reliably specified.</p></li><li><p>It is unclear whether an addition follows every observed decline, a forecast-confirmed decline, or a separate decision policy.</p></li></ul><p>Those choices affect both risk and accounting. Filling them in silently would make the implementation appear more faithful than it is. The replacement below makes the missing decisions explicit.</p><h3>12.2 Normalized exponential reconstruction</h3><p>The practical reconstruction assigns a positive score to each finite addition level:</p><p>si=Aexp&#8289;(Bi),i=0,1,&#8230;,n&#8722;1.</p><p>Here <code>A</code> is a scale factor, <code>B</code> is the growth rate, and <code>n</code> is the configured maximum number of additions. The scores are converted to normalized weights:</p><p>\[ w<em>i=\frac{s</em>i}{\sum<em>{j=0}^{n-1}s</em>j}. \]</p><p>For a sizing budget <code>U</code>, the planned allocation is:</p><p>\[ a<em>i=Uw</em>i. \]</p><p>Thus, subject to floating-point rounding,</p><p>\[ \sum<em>i a</em>i=U. \]</p><p>This is a plausible interpretation of the paper's normalized allocation expression, but it is not a verified transcription. In the generated implementation, <code>build_exponential_schedule</code> receives the number of additions and budget explicitly:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c82c93f1-689c-4b02-bfd5-b8485254d05e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def build_exponential_schedule(
    max_additions: int,
    budget: float,
    *,
    growth: float = 0.5,
    amplitude: float = 1.0,
    gain_percentile: Optional[float] = None,
    decline_percentile: Optional[float] = None,
    fee_rate: float = 0.0,
    calibration_source: Any = None,
) -&gt; SizingSchedule:
    """Build normalized ``A * exp(B*i)`` addition amounts."""</code></pre></div><p>The absolute budget is therefore supplied to the sizing builder or sizer; it is not hidden inside the paper-inspired <code>SizingConfig</code> dataclass.</p><h3>12.3 Constructing the schedule safely</h3><p>The implementation computes the exponential scores and subtracts the largest exponent before calling <code>exp</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;788a2002-3dfb-4376-bc2c-4ad5865cd886&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">indices = np.arange(max_additions, dtype=float)
exponents = growth * indices
exponents -= float(np.max(exponents))
scores_array = float(amplitude) * np.exp(exponents)

score_sum = float(scores_array.sum())
weights_array = scores_array / score_sum
allocations_array = weights_array * budget</code></pre></div><p>Subtracting the maximum exponent does not change the normalized weights. It multiplies every score by the same positive constant, so each ratio <code>s_i / sum(s_j)</code> remains unchanged while reducing overflow risk.</p><p><code>SizingSchedule</code> stores:</p><ul><li><p><code>scores</code>: unnormalized exponential scores;</p></li><li><p><code>weights</code>: normalized proportions that sum to one; and</p></li><li><p><code>allocations</code>: planned currency amounts obtained from the budget.</p></li></ul><p>It also records metadata identifying the formula as a reconstruction. This makes the sizing assumptions available when a later experiment is inspected.</p><h3>12.4 Worked allocation example</h3><p>Suppose the total planned budget is <code>$1,000</code> and three levels have scores:</p><p>\[ s<em>0=1, \qquad s</em>1=2, \qquad s_2=4. \]</p><p>The score sum is <code>7</code>:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Adof!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba10dd43-5ed8-441f-ac72-7e69b8c0ad6e_800x356.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Adof!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba10dd43-5ed8-441f-ac72-7e69b8c0ad6e_800x356.png 424w, /__u/substackcdn.com/image/fetch/$s_!Adof!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba10dd43-5ed8-441f-ac72-7e69b8c0ad6e_800x356.png 848w, /__u/substackcdn.com/image/fetch/$s_!Adof!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba10dd43-5ed8-441f-ac72-7e69b8c0ad6e_800x356.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Adof!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba10dd43-5ed8-441f-ac72-7e69b8c0ad6e_800x356.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Adof!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba10dd43-5ed8-441f-ac72-7e69b8c0ad6e_800x356.png" width="800" height="356" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ba10dd43-5ed8-441f-ac72-7e69b8c0ad6e_800x356.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:356,&quot;width&quot;:800,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:43895,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba10dd43-5ed8-441f-ac72-7e69b8c0ad6e_800x356.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Adof!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba10dd43-5ed8-441f-ac72-7e69b8c0ad6e_800x356.png 424w, /__u/substackcdn.com/image/fetch/$s_!Adof!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba10dd43-5ed8-441f-ac72-7e69b8c0ad6e_800x356.png 848w, /__u/substackcdn.com/image/fetch/$s_!Adof!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba10dd43-5ed8-441f-ac72-7e69b8c0ad6e_800x356.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Adof!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba10dd43-5ed8-441f-ac72-7e69b8c0ad6e_800x356.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>These are planned amounts, not automatic orders. If only <code>$400</code> remains available when the final level is requested, the order must be reduced or rejected according to the current cash and exposure constraints.</p><p>The same relative scores arise from <code>A=1</code> and <code>B=log(2)</code>, because <code>exp(Bi)</code> then gives approximately <code>1</code>, <code>2</code>, and <code>4</code> for levels zero through two.</p><h3>12.5 Optional percentile and fee calibration</h3><p>The paper connects position sizing to empirical gain and decline statistics. The implementation can read values such as the 90th-percentile gain, a signed decline percentile, a decline-magnitude percentile, a median gain, and a fee rate.</p><p>When percentile values or a <code>calibration_source</code> are supplied to <code>build_exponential_schedule</code>, the helper <code>_calibrated_growth</code> derives a conservative growth coefficient:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;96372b2f-12c2-4a21-b01a-748279bb9555&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def _calibrated_growth(
    gain_percentile: Optional[float],
    decline_percentile: Optional[float],
    fee_rate: float,
) -&gt; tuple[float, dict[str, float]]:
    gain = max(float(gain_percentile or 0.0), 0.0)
    decline = abs(float(decline_percentile or 0.0))
    fee = _finite_nonnegative(fee_rate, "fee_rate")

    recovery = gain + _EPSILON
    burden = decline + 2.0 * fee
    growth = 0.05 + min(1.5, burden / recovery)
    return growth, {
        "gain_percentile": gain,
        "decline_magnitude": decline,
        "fee_rate": fee,
        "growth_coefficient": growth,
    }</code></pre></div><p>This is an engineering interpretation, not the paper's recovered break-even equation. The fee appears twice in the burden term as a conservative allowance for a purchase and later sale; actual fees are still charged by the execution simulator.</p><p><code>SizingConfig</code> contains fields such as <code>calibrate_from_percentiles</code>, <code>recovery_percentile</code>, <code>decline_percentile</code>, and <code>fee_buffer</code>. These fields document intended experiment settings, but the current generated schedule builder does <strong>not</strong> accept a <code>SizingConfig</code> directly and does not automatically consult <code>calibrate_from_percentiles</code>. Calibration is activated in the current implementation by supplying percentile arguments or a <code>calibration_source</code> to <code>build_exponential_schedule</code>. An experiment layer must explicitly connect configuration fields to those arguments if automatic configuration-driven calibration is desired.</p><h3>12.6 Finite levels and budget invariants</h3><p>The schedule validates finite values, nonnegative scores and allocations, normalized weights, and the relationship between planned allocations and its supplied budget. It also tracks consumed levels:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f26d95d3-c50f-4245-9bf6-27f351259f44&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">used_levels: set[int] = field(default_factory=set, repr=False)</code></pre></div><p><code>next_unused_level()</code> returns the first unconsumed level, and <code>mark_used(level)</code> records a level while rejecting reuse. This prevents a strategy loop from accidentally placing the same planned addition repeatedly.</p><p>The schedule's budget is an absolute amount supplied when the schedule is built. It is distinct from the fraction fields in <code>SizingConfig</code>. The configuration validates settings such as:</p><ul><li><p>a positive <code>maximum_additions</code> value;</p></li><li><p>an <code>initial_level</code> inside the configured level range;</p></li><li><p>a positive <code>score_scale</code>;</p></li><li><p>allocation fractions between zero and one; and</p></li><li><p>a maximum exposure fraction in <code>(0, 1]</code>.</p></li></ul><p>These checks validate fraction and level ranges. They do not compare an exposure fraction with an absolute budget, because <code>SizingConfig</code> does not contain an absolute budget field. Absolute budget and exposure constraints are enforced later by <code>build_exponential_schedule</code> and <code>allocate_next_addition</code>.</p><h3>12.7 Cash and exposure caps at allocation time</h3><p>The schedule is only a plan. At order time, <code>allocate_next_addition</code> applies the current portfolio state:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;85bfa6b7-8308-48d2-be7b-483fe408b774&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def allocate_next_addition(
    schedule: SizingSchedule,
    available_cash: float,
    *,
    current_exposure: float = 0.0,
    max_exposure: Optional[float] = None,
    level: Optional[int] = None,
    allow_partial: bool = True,
) -&gt; AllocationDecision:</code></pre></div><p>The effective capacity is calculated as:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;9b60a2d1-712e-41d3-94d3-8500d188e762&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">cap = schedule.budget if max_exposure is None else max_exposure
capacity = min(cash, max(0.0, cap - exposure))
approved = min(requested, capacity)</code></pre></div><p>This provides three separate protections:</p><p>A request cannot exceed current available cash.</p><p>Available cash does not override a maximum exposure limit.</p><p>Partial approval is explicit. With <code>allow_partial=True</code>, an amount may be reduced to capacity; with <code>allow_partial=False</code>, an undersized request is rejected.</p><p>The result is an <code>AllocationDecision</code> containing the requested and approved amounts, selected level, reason, remaining cash, and remaining exposure capacity.</p><p>The approved amount is a <strong>pre-fee notional</strong>. The sizing module does not deduct execution fees from that amount. The backtest applies fees separately, so the final cash cost of a buy can be higher than the approved notional. Consequently, the execution layer may clip the order again when fees are applied, even when the sizing layer approved it under a pre-fee cash calculation. Keeping these stages separate makes the fee convention auditable, but callers must not treat a pre-fee approval as a guarantee that the full order will fill.</p><h3>12.8 Why this is not unlimited averaging down</h3><p>Increasing allocations after declines can become an uncontrolled averaging-down strategy. The generated design limits that behavior by using:</p><ul><li><p>a finite <code>maximum_additions</code> value;</p></li><li><p>one-time consumption of each addition level;</p></li><li><p>a fixed schedule budget;</p></li><li><p>current-cash checks;</p></li><li><p>optional maximum exposure;</p></li><li><p>no leverage by default in the execution layer; and</p></li><li><p>optional risk filtering by later strategy components.</p></li></ul><p>These are safety and reproducibility choices, not recovered paper details. They may differ from the unknown original behavior, but they prevent the implementation from silently authorizing unlimited purchases.</p><p>The sizing module also does not decide when to add. It supplies an amount after another layer requests a level. A signal, forecast, decline condition, VaR decision, or greedy policy must determine whether that level is activated.</p><h3>12.9 Correct API example</h3><p>The generated sizing implementation uses <code>growth</code>, not <code>growth_rate</code>, when constructing a schedule:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;9aac6f35-b01a-4058-8bec-4a51afc7e9ad&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">schedule = build_exponential_schedule(
    budget=1_000.0,
    max_additions=4,
    growth=0.5,
)

assert np.isclose(sum(schedule.weights), 1.0)
assert np.isclose(sum(schedule.allocations), 1_000.0)</code></pre></div><p>The supplied generated test file uses <code>growth_rate</code> in one example and constructs <code>ExponentialPositionSizer</code> with a configuration object in a way that does not match the shown implementation. It also contains related naming differences in the risk and strategy tests. These are static test artifacts describing intended invariants, not evidence that the tests were executed successfully. The corrected example above follows the actual generated <code>build_exponential_schedule</code> signature.</p><h3>12.10 What the sizing tests are intended to verify</h3><p>The tests are intended to check properties of the reconstruction rather than the paper's unverified returns:</p><ul><li><p>scores, weights, and allocations are finite;</p></li><li><p>weights are nonnegative and sum to one;</p></li><li><p>planned allocations do not exceed the supplied budget;</p></li><li><p>a request cannot exceed available cash or exposure capacity;</p></li><li><p>an addition level cannot be reused; and</p></li><li><p>future realized prices do not influence the sizing decision.</p></li></ul><p>These are static and semantic expectations. They do not prove that equations (5)&#8211;(8) were recovered or that the resulting strategy is economically sound.</p><h3>12.11 Practical interpretation</h3><p>The reconstructed data flow is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;54d977d7-f6fe-430f-8960-78d0e2eb5b4e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">historical gain/decline statistics
        |
        v
optional explicit calibration arguments
        |
        v
finite exponential scores
        |
        v
normalized budget allocations
        |
        v
cash and exposure checks
        |
        v
pre-fee approved notional
        |
        v
fee-aware execution and final fill</code></pre></div><p>A larger allocation after a decline increases potential exposure if prices recover, but it also increases losses if the decline continues. Fees, slippage, liquidity, and forecast error can invalidate the recovery assumptions. Evaluation should therefore include drawdown, turnover, total fees, and risk measures rather than final value alone.</p><p>The narrow claim supported by this implementation is that it provides a finite, normalized, auditable reconstruction of the paper's exponential position-sizing concept. It does not reproduce equations (5)&#8211;(8), validate the paper's reported <code>$646</code> gold or <code>$215,487</code> Bitcoin outcomes, or establish profitability.</p><h2>13. Turn Forecasts into Signals and Reconstruct the Greedy Policy</h2><p>A forecast is not yet an order. If a model predicts Bitcoin will reach <code>$31,000</code>, the strategy must still compare that prediction with the current price, decide whether the expected move is large enough to trade, choose a notional amount, check cash and exposure limits, apply risk controls, and select an action.</p><p>The paper describes this stage using forecasting, position management, risk control, and a modified greedy algorithm. It does not define the greedy algorithm operationally: the candidate actions, objective, constraints, and tie-breaking rules are missing. The implementation therefore provides a transparent <strong>reconstructed policy</strong>, not a verified reproduction.</p><p>The intended decision flow is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;2a671660-d5a6-4723-b892-ed857020f807&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">forecast values
    -&gt; aggregate the forecast horizon
    -&gt; compare the forecast with the current price
    -&gt; create a buy, sell, add, or hold signal
    -&gt; construct and constrain candidate actions
    -&gt; estimate fees and expected benefit
    -&gt; apply a risk filter
    -&gt; rank feasible candidates deterministically
    -&gt; execute the selected action on a later bar</code></pre></div><p>The snippets in this section are abridged explanatory excerpts from <code>src/quantitative_trading_model/strategy.py</code>; they are not complete replacement functions. The generated strategy and risk modules currently have an incompatible runtime interface, so this section describes the intended contract and the required repair rather than claiming that the end-to-end policy runs successfully.</p><h3>13.1 Direction and order size are separate decisions</h3><p>The strategy module defines four directional states:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7264f1e7-4625-4c21-b54f-5ca519cf915d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class Signal(str, Enum):
    BUY = "buy"
    SELL = "sell"
    HOLD = "hold"
    ADD = "add"</code></pre></div><p>These values describe intent, not quantity:</p><ul><li><p><code>BUY</code> establishes or increases exposure when the forecast exceeds a threshold.</p></li><li><p><code>SELL</code> reduces existing long exposure when the forecast is sufficiently negative.</p></li><li><p><code>ADD</code> requests the next finite position-sizing level.</p></li><li><p><code>HOLD</code> submits no trade.</p></li></ul><p>Separating direction from size lets the forecasting model answer &#8220;should an opportunity be considered?&#8221; while the sizing module answers &#8220;how much may be allocated?&#8221; Cash limits, exposure caps, transaction fees, and risk controls can then reduce or reject the proposed amount.</p><p>The paper does not explain how an <code>ADD</code> signal is triggered. In particular, it does not state whether additions follow every decline, a forecast-confirmed decline, or a separate greedy rule. The generated <code>strategy.py</code> does not derive <code>ADD</code> directly from the streak analysis; it creates an addition candidate only when the supplied decision is <code>BUY</code> or <code>ADD</code> and a finite addition level is available. A production revision must define that trigger explicitly.</p><h3>13.2 Aggregate one-day or multi-day forecasts explicitly</h3><p>The paper is inconsistent about its forecast horizon. Its prediction discussion supports one-step forecasting, while its conclusion refers to the next three days. <code>SignalGenerator</code> makes the interpretation configurable.</p><p>The generator accepts predicted prices and supports three aggregation choices:</p><ul><li><p><code>terminal</code>: use the final forecast, such as day three;</p></li><li><p><code>mean</code>: use the average forecast over the horizon;</p></li><li><p><code>max</code>: use the largest forecast.</p></li></ul><p>The following is an abridged excerpt; validation and error handling remain in the generated class:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4ad2837e-d18d-4367-800d-d9cba536cad5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">class SignalGenerator:
    def __init__(self, buy_threshold=0.0, sell_threshold=0.0,
                 aggregation="terminal", minimum_forecast_points=1):
        self.buy_threshold = buy_threshold
        self.sell_threshold = sell_threshold
        self.aggregation = aggregation
        self.minimum_forecast_points = minimum_forecast_points

    def aggregate(self, forecasts):
        values = [float(value) for value in forecasts]
        if self.aggregation == "terminal":
            return values[-1]
        if self.aggregation == "mean":
            return sum(values) / len(values)
        return max(values)</code></pre></div><p>For a current price of <code>$100</code> and forecasts <code>[101, 103, 104]</code>, terminal and maximum aggregation produce <code>$104</code>, while mean aggregation produces approximately <code>$102.67</code>. These are reconstruction choices, not rules confirmed by the paper.</p><h3>13.3 Convert the selected forecast into a signal</h3><p>After aggregation, the strategy computes the expected price return:</p><p>r^=p^&#8722;pp,</p><p>where <code>p</code> is the current price and <code>p_hat</code> is the selected forecast.</p><p>An abridged version of <code>SignalGenerator.decide</code> is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;1a1c14bf-6736-49ef-8f4c-2ad470480b98&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def decide(self, current_price, forecasts):
    forecast_value = self.aggregate(forecasts)
    expected_return = forecast_value / current_price - 1.0

    if expected_return &gt; self.buy_threshold:
        signal = Signal.BUY
    elif expected_return &lt; -self.sell_threshold:
        signal = Signal.SELL
    else:
        signal = Signal.HOLD

    return SignalDecision(
        signal=signal,
        expected_return=expected_return,
        forecast_value=forecast_value,
        horizon=len(forecasts),
        rationale="threshold-based forecast decision",
    )</code></pre></div><p>With one-percent buy and sell thresholds:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!DZ4f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d66b2d-c298-46ae-b93a-89c851052492_812x292.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!DZ4f!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d66b2d-c298-46ae-b93a-89c851052492_812x292.png 424w, /__u/substackcdn.com/image/fetch/$s_!DZ4f!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d66b2d-c298-46ae-b93a-89c851052492_812x292.png 848w, /__u/substackcdn.com/image/fetch/$s_!DZ4f!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d66b2d-c298-46ae-b93a-89c851052492_812x292.png 1272w, /__u/substackcdn.com/image/fetch/$s_!DZ4f!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d66b2d-c298-46ae-b93a-89c851052492_812x292.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!DZ4f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d66b2d-c298-46ae-b93a-89c851052492_812x292.png" width="812" height="292" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/45d66b2d-c298-46ae-b93a-89c851052492_812x292.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:292,&quot;width&quot;:812,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:33912,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d66b2d-c298-46ae-b93a-89c851052492_812x292.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!DZ4f!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d66b2d-c298-46ae-b93a-89c851052492_812x292.png 424w, /__u/substackcdn.com/image/fetch/$s_!DZ4f!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d66b2d-c298-46ae-b93a-89c851052492_812x292.png 848w, /__u/substackcdn.com/image/fetch/$s_!DZ4f!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d66b2d-c298-46ae-b93a-89c851052492_812x292.png 1272w, /__u/substackcdn.com/image/fetch/$s_!DZ4f!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d66b2d-c298-46ae-b93a-89c851052492_812x292.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The thresholds are not specified by the paper and must be stored with experiment metadata. A zero threshold can produce frequent trades when forecasts fluctuate around the current price; a positive threshold creates a hold band but may delay decisions.</p><p>A signal must contain current-information fields only: direction, expected return, forecast value, horizon, and rationale. It must not contain a realized future price or realized future return. Execution belongs to the later backtest stage.</p><h3>13.4 Candidate actions contain size and audit information</h3><p>A <code>CandidateAction</code> combines a proposed notional with the information needed for ranking:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;3e90047e-a64e-49b9-bf35-4fbe8199e323&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class CandidateAction:
    signal: Signal
    notional: float              # pre-fee order value
    expected_return: float
    estimated_fee: float
    score: float
    feasible: bool = True
    reason: str = ""
    addition_level: Optional[int] = None
    risk_approved: bool = True
    risk_scale: float = 1.0</code></pre></div><p>For buys and additions, <code>notional</code> is the cash committed before fees. For sells, it is the market value sold before the sell fee. <code>reason</code> and <code>addition_level</code> make decisions easier to inspect.</p><p><code>PortfolioSnapshot</code> supplies the current state:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4efee9a7-3695-4c1f-9604-c7736c6d631a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class PortfolioSnapshot:
    cash: float
    holdings: float
    price: float
    addition_level: int = 0
    portfolio_value: Optional[float] = None
    max_exposure: Optional[float] = None

    @property
    def marked_value(self):
        return self.cash + self.holdings * self.price</code></pre></div><p>This state is intentionally limited to information available at the signal timestamp.</p><h3>13.5 Apply feasibility constraints before ranking</h3><p>The policy should never rank an unaffordable order as though it were executable. For a long-only account, the main constraints are:</p><p>A buy must fit available cash, including fees.</p><p>A buy must not exceed the maximum exposure.</p><p>A sell must not exceed current holdings unless shorting is explicitly enabled.</p><p>An addition must use an available finite sizing level.</p><p>A risk filter may subsequently block or scale the candidate.</p><p>For a fee rate <code>f</code>, the maximum pre-fee buy notional from cash is:</p><p>Nmax=cash1+f.</p><p>Thus <code>$75</code> of cash and a one-percent fee permit at most approximately <code>$74.26</code> of pre-fee buy notional.</p><p>The generated candidate builder caps buy amounts by cash and exposure. It does <strong>not</strong> consistently retain a separate rejected candidate for every failed cash or exposure check. Some rejected requests are reduced to an affordable candidate, while exhausted addition levels may be recorded as infeasible. Therefore, the accurate audit statement is that feasible capped candidates and selected/rejected addition cases are recorded, but the current implementation does not provide a uniform record of every failed constraint attempt. A stricter audit design would preserve both the original request and the resulting rejection reason.</p><p>A conceptual candidate calculation is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;bb9251f0-1f5a-4e1d-80ec-160ba44d4d30&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">available_buy = snapshot.cash / (1.0 + fee_rate)
current_exposure = snapshot.holdings * snapshot.price

if snapshot.max_exposure is not None:
    available_buy = min(
        available_buy,
        max(0.0, snapshot.max_exposure - current_exposure),
    )</code></pre></div><p>The backtest must enforce the same limits again at execution. Strategy-level feasibility is not a substitute for accounting-level validation.</p><h3>13.6 Connect finite sizing to <code>ADD</code> candidates</h3><p>The normalized exponential sizer supplies finite addition amounts. The strategy layer should not duplicate the sizing formula. It requests the next level from a previously constructed sizer or schedule, then applies current cash and exposure limits.</p><p>For example, the sizing object must be defined before calling the policy:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;27bbf554-d4b2-4b96-956b-43efcef98a26&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from quantitative_trading_model.sizing import (
    ExponentialPositionSizer,
    build_exponential_schedule,
)

schedule = build_exponential_schedule(
    max_additions=3,
    budget=500.0,
    growth=0.5,
)
sizer = ExponentialPositionSizer(schedule=schedule)</code></pre></div><p>The exact generated API may need alignment before runtime use, but the intended data flow is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;041af851-a924-4f2e-8b4b-ffc85b24742c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">BUY or ADD decision
    -&gt; request the next sizing level
    -&gt; cap it by cash and exposure
    -&gt; create an ADD candidate</code></pre></div><p>The paper does not define the trigger for <code>ADD</code>, so the example must not imply that the strategy automatically reacts to every consecutive decline. That behavior would need an explicit rule connecting <code>streaks.py</code>, the forecast, and the policy.</p><h3>13.7 Use a transparent fee-aware score</h3><p>The paper gives no objective function for its modified greedy algorithm. The reconstruction uses expected net benefit:</p><p>score(a)=notional(a)(benefit&nbsp;rate(a)&#8722;f).</p><p>For a buy or add, the benefit rate is the forecasted return. For a sell, the reconstruction treats a negative forecast as a potential avoided loss, so the benefit rate is <code>-expected_return</code>. Hold has a score of zero.</p><p>An abridged scoring function is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a5464a18-16b7-483c-a44e-bb0314f2a07f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def score_candidate_action(action, *, expected_return, fee_rate):
    if action.signal in {Signal.BUY, Signal.ADD}:
        benefit_rate = expected_return
    elif action.signal is Signal.SELL:
        benefit_rate = -expected_return
    else:
        benefit_rate = 0.0
    return action.notional * (benefit_rate - fee_rate)</code></pre></div><p>A <code>$100</code> buy with a three-percent expected return and one-percent fee scores:</p><p>100(0.03&#8722;0.01)=2.</p><p>A <code>$100</code> buy with a 0.5% expected return scores <code>-0.50</code> and should lose to the zero-score hold action. This score ignores uncertainty, spread, liquidity, and forecast calibration. It is a transparent replacement for an objective that the paper does not provide.</p><h3>13.8 Rank candidates deterministically</h3><p>The intended ranking process is:</p><p>Remove candidates that are infeasible or not risk-approved.</p><p>Choose the highest score.</p><p>Break equal-score ties with the smaller notional.</p><p>Apply a fixed signal priority for any remaining tie.</p><p>An abridged selection expression is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c58222e0-2f5d-431d-8b2e-c267b279b79c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">feasible = [
    candidate for candidate in candidates
    if candidate.feasible and candidate.risk_approved
]

selected = max(
    feasible,
    key=lambda candidate: (
        candidate.score,
        -candidate.notional,
        -priority[candidate.signal],
    ),
)</code></pre></div><p>The priority mapping is an implementation choice, not a paper rule. If no candidate is feasible, the policy should return a hold decision with a clear reason such as <code>"no feasible candidate"</code>.</p><h3>13.9 Risk integration: intended contract versus current generated code</h3><p>The intended ordering is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;0f44f093-ad74-46d0-bbac-e143682498b2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">candidate
    -&gt; risk filter
    -&gt; accept, scale, or block
    -&gt; recompute fee and score
    -&gt; rank the resulting candidate</code></pre></div><p>Risk must be applied before final ranking because scaling a candidate can change its score. However, the current generated <code>strategy.py</code> and <code>risk.py</code> do not implement one reliable shared protocol:</p><ul><li><p><code>GreedyPolicy._apply_risk</code> calls the filter using <code>proposed_notional</code>, while <code>VaRTradeFilter.decide</code> expects <code>proposed_quantity</code> and also requires <code>current_exposure</code> and <code>portfolio_value</code>.</p></li><li><p>The fallback positional call does not reliably map those required values to the risk method's parameters.</p></li><li><p><code>VaRTradeFilter</code> returns <code>RiskDecision</code> with <code>approved_quantity</code> and <code>blocked</code>, not an <code>approved</code> field. The current strategy adapter looks first for <code>approved</code> or <code>allowed</code>, so a blocked result can be misinterpreted as approved.</p></li><li><p>Consequently, the current code must be repaired before its VaR integration can be treated as a working runtime policy.</p></li></ul><p>A strict repaired protocol should pass the same named fields every time:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d443507e-ed85-419a-ba59-c5c5b610805c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">risk_decision = risk_filter.decide(
    proposed_quantity=action.notional,
    current_exposure=snapshot.holdings * snapshot.price,
    portfolio_value=snapshot.value,
    returns=pre_decision_returns,
)

if risk_decision.blocked:
    action=/__u/onepagecode.substack.com/replace(action, feasible=False, risk_approved=False, notional=0.0)
elif risk_decision.approved_quantity &lt; action.notional:
    scale = risk_decision.approved_quantity / action.notional
    action=/__u/onepagecode.substack.com/rescale_and_recompute_score(action, scale, fee_rate)</code></pre></div><p>This is a repair specification, not a claim about the current generated files. The paper itself does not define how VaR interacts with greedy selection, so the ordering remains a reconstruction even after the interfaces are aligned.</p><h3>13.10 Complete decision example with explicit assumptions</h3><p>The following is an explanatory excerpt showing the objects that must exist. It is not a complete runnable example, and the risk interface must be repaired as described above:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;34bef7d0-42d5-4ccb-917d-dbac3e43bddf&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from quantitative_trading_model.strategy import (
    GreedyPolicy, PortfolioSnapshot, SignalGenerator,
)

schedule = build_exponential_schedule(
    max_additions=3,
    budget=500.0,
    growth=0.5,
)
sizer = ExponentialPositionSizer(schedule=schedule)

policy = GreedyPolicy(fee_rate=0.01)
signal_generator = SignalGenerator(
    buy_threshold=0.01,
    sell_threshold=0.01,
    aggregation="terminal",
)

snapshot = PortfolioSnapshot(
    cash=500.0,
    holdings=0.0,
    price=100.0,
    addition_level=0,
    max_exposure=500.0,
)

# The intended call; risk_filter integration requires the strict protocol above. action=/__u/onepagecode.substack.com/policy.choose(
    snapshot,
    forecasts=[101.0, 103.0, 104.0],
    signal_generator=signal_generator,
    sizer=sizer,
    max_additions=3,
)</code></pre></div><p>Conceptually, this selects the terminal forecast of <code>$104</code>, computes a four-percent expected return, creates a buy decision, obtains finite sizing candidates, applies cash and exposure limits, accounts for fees, and selects the best feasible action. The selected action must be passed to the backtest for execution on a later bar.</p><h3>13.11 Verification status and no-look-ahead boundary</h3><p>The intended tests in <code>tests/test_sizing_risk_strategy.py</code> cover useful invariants:</p><ul><li><p>scores use forecasts and fees rather than realized future returns;</p></li><li><p>order sizes do not exceed cash or exposure limits;</p></li><li><p>risk scaling never increases a proposed order;</p></li><li><p>tie-breaking is deterministic;</p></li><li><p>signal timestamps precede execution timestamps.</p></li></ul><p>The supplied generated test file is currently API-inconsistent with <code>strategy.py</code>. It uses older names such as <code>side</code> instead of <code>signal</code> and calls methods with incompatible signatures. It therefore cannot substantiate that the current strategy implementation passes its intended tests without repair.</p><p>More broadly, semantic code verification was skipped. Static review identified issues in the generated project, and no code execution or passing test suite is being claimed. The excerpts above explain the intended contracts and invariants; they should not be read as evidence of runtime correctness.</p><h3>13.12 Reproduction boundary</h3><p>The paper's greedy policy remains unverified because it does not specify:</p><ul><li><p>the complete candidate-action set;</p></li><li><p>whether one action or a multi-step plan is selected;</p></li><li><p>the objective function;</p></li><li><p>how one-day and three-day forecasts are combined;</p></li><li><p>the treatment of fees, spreads, and slippage;</p></li><li><p>the interaction between VaR and action selection;</p></li><li><p>leverage, exposure, and addition constraints; or</p></li><li><p>tie-breaking behavior.</p></li></ul><p>The generated policy is therefore best understood as an auditable baseline. It demonstrates how forecasts, finite sizing, constraints, fees, risk decisions, and deterministic selection can be connected without using realized future prices. The realized price enters only during later execution and evaluation, preserving the causal boundary required for a meaningful backtest.</p><h2>14. Add Historical VaR as an Explicit Risk Filter</h2><p>The paper names Value-at-Risk (VaR) as part of its trading framework, but it does not define the parameters or the action taken when risk is high. It does not specify the confidence level, horizon, lookback window, return distribution, portfolio aggregation, acceptable loss, or whether VaR blocks or scales trades.</p><p>This project therefore uses a <strong>rolling historical VaR reconstruction</strong>. The convention in this section is explicit:</p><ul><li><p>return VaR is a nonnegative loss fraction, such as <code>0.04</code> for a 4% loss threshold;</p></li><li><p>projected VaR is a <strong>currency amount</strong>;</p></li><li><p>the risk limit is a currency amount equal to a configured fraction of portfolio value;</p></li><li><p>trades may be accepted, scaled, or blocked;</p></li><li><p>only returns available before the decision timestamp may be used.</p></li></ul><p>These choices are operational assumptions, not a verified implementation of the paper's VaR method.</p><h3>14.1 VaR sign convention and units</h3><p>Let <code>r</code> be a collection of historical returns and let <code>alpha</code> be the confidence level. The lower-tail return quantile is <code>q_(1-alpha)(r)</code>. A positive return-loss VaR is:</p><p>\[ \operatorname{VaR}^{\text{return}}<em>\alpha = -q</em>{1-\alpha}(r). \]</p><p>If the 5th-percentile return at 95% confidence is <code>-0.04</code>, return VaR is <code>0.04</code>.</p><p>For an exposure <code>E</code>, projected currency VaR is:</p><p>\[ \operatorname{VaR}^{\text{currency}}<em>\alpha = E\times \operatorname{VaR}^{\text{return}}</em>\alpha. \]</p><p>For example, a <code>$2,000</code> exposure and a 4% return-loss VaR produce:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;dd785192-5112-4794-b5b2-08a8d046e7d6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">projected currency VaR = 2,000 * 0.04 = $80</code></pre></div><p>The units are important. A currency VaR must be compared with a currency limit, not with a fraction such as <code>0.05</code>.</p><p>The risk limit is calculated from portfolio value <code>V</code> and the configured maximum fraction <code>m</code>:</p><p>limitcurrency=mV.</p><p>Thus, for a <code>$10,000</code> portfolio and <code>m=0.05</code>, the limit is <code>$500</code>.</p><p>VaR is a loss quantile, not a guaranteed maximum loss. A realized loss can exceed it, particularly when market conditions change or the historical window does not represent the current regime.</p><h3>14.2 Explicit configuration</h3><p><code>VaRConfig</code> in <code>src/quantitative_trading_model/config.py</code> makes the unspecified choices visible:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;890d0b0a-77dc-40f5-b56e-e0feb27bb6aa&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass
class VaRConfig:
    enabled: bool = True
    confidence: float = 0.95
    horizon: int = 1
    lookback: int = 100
    method: str = "historical"
    action: str = "scale"
    max_var_fraction: float = 0.05
    insufficient_history: str = "allow"
    min_observations: int = 30</code></pre></div><p>The defaults are implementation choices:</p><p>| Setting | Default | Meaning | | --- | ---: | --- | | <code>confidence</code> | <code>0.95</code> | Use the lower 5% return tail. | | <code>horizon</code> | <code>1</code> | Estimate a one-period loss threshold. | | <code>lookback</code> | <code>100</code> | Use at most the latest 100 available returns. | | <code>method</code> | <code>historical</code> | Use an empirical quantile. | | <code>action</code> | <code>scale</code> | Reduce an order that exceeds the risk limit. | | <code>max_var_fraction</code> | <code>0.05</code> | Permit VaR up to 5% of portfolio value. | | <code>min_observations</code> | <code>30</code> | Require this many horizon observations for an actionable estimate. |</p><p>Configuration validation rejects invalid confidence levels, nonpositive horizons and lookbacks, unsupported methods, and unknown action modes. Decimal transaction fees, forecast settings, sizing settings, and execution rules remain separate from VaR.</p><h3>14.3 Estimating a rolling historical quantile</h3><p><code>HistoricalVaR</code> in <code>src/quantitative_trading_model/risk.py</code> performs the following steps:</p><p>Convert the supplied returns to a finite one-dimensional array.</p><p>Keep only the configured lookback window.</p><p>Optionally construct overlapping compounded returns for a multi-period horizon.</p><p>Select the lower-tail empirical quantile.</p><p>Store the signed quantile and its nonnegative loss magnitude.</p><p>The central calculation is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;fc51d166-a798-4c01-8ab4-9b2ec413e812&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">values = self._finite_returns(returns)
values = values[-self.lookback:]
horizon_values = self._horizon_returns(values)

if len(horizon_values) &lt; self.min_observations:
    return VaREstimate(
        var=float("nan"),
        confidence=self.confidence,
        horizon=self.horizon,
        observations=len(horizon_values),
        return_quantile=float("nan"),
        sufficient_history=False,
    )

quantile = np.quantile(
    horizon_values,
    1.0 - self.confidence,
    method="linear",
)
return VaREstimate(
    var=max(0.0, -float(quantile)),
    confidence=self.confidence,
    horizon=self.horizon,
    observations=len(horizon_values),
    return_quantile=float(quantile),
    sufficient_history=True,
)</code></pre></div><p><code>VaREstimate.var</code> is a return fraction. It is not a currency amount until it is multiplied by an exposure. <code>return_quantile</code> retains the signed empirical value for auditing.</p><p>The <code>1-confidence</code> probability is deliberate. With <code>confidence=0.95</code>, the implementation selects the 5th percentile. The sign conversion prevents a negative return quantile from being confused with a negative risk amount.</p><h3>14.4 Optional horizon compounding</h3><p>For a one-period horizon, the historical returns can be used directly. For a longer horizon, consecutive simple returns are compounded. A two-period return is:</p><p>\[ (1+r<em>t)(1+r</em>{t+1})-1. \]</p><p>The corresponding helper is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d9898f9d-cfc4-4bf6-a3fa-60605c9f6625&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def _horizon_returns(self, values: np.ndarray) -&gt; np.ndarray:
    if self.horizon == 1:
        return values
    if values.size &lt; self.horizon:
        return np.empty(0, dtype=float)

    windows = np.lib.stride_tricks.sliding_window_view(
        values,
        self.horizon,
    )
    return np.prod(1.0 + windows, axis=1) - 1.0</code></pre></div><p>Overlapping historical blocks are a transparent reconstruction. The paper does not say whether its VaR uses one-day returns, compounded multi-day returns, a parametric distribution, or simulation.</p><h3>14.5 Pre-decision history and lookback boundaries</h3><p>VaR must obey the same information boundary as the forecast. If a decision is made after observations through time <code>t</code>, the risk history may contain returns known by that point, but not returns from <code>t+1</code> or later.</p><p><code>HistoricalVaR</code> accepts caller-supplied returns and cannot infer timestamps. The caller is responsible for constructing the correct information set:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8b1d9348-ae9c-4782-9ba7-b644c235b5e4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># known_returns must contain only returns available at the signal time
estimate = var_estimator.estimate(known_returns, include_last=True)</code></pre></div><p>If the final element represents a current-bar return that is not yet available when the signal is formed, use <code>include_last=False</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ddac7b20-872e-44ad-883d-a10d0229ff8d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">estimate = var_estimator.estimate(
    known_returns,
    include_last=False,
)</code></pre></div><p>A leakage-aware sequence is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;f3b62310-53ec-4ac6-9df3-cde646208d9a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">1. Observe prices through the information timestamp.
2. Compute only returns available at that timestamp.
3. Estimate VaR from the permitted lookback window.
4. Generate the forecast and candidate order.
5. Apply the VaR filter.
6. Execute no earlier than the next permitted bar.</code></pre></div><p>The risk module must not be passed the complete historical return series during every past decision. That would use future information.</p><h3>14.6 Insufficient history</h3><p>When fewer than <code>min_observations</code> horizon returns are available, <code>HistoricalVaR.estimate</code> returns an estimate with <code>sufficient_history=False</code> and unavailable numerical values. It does not fabricate a zero-risk estimate.</p><p><code>VaRTradeFilter</code> then applies the configured policy:</p><ul><li><p><code>insufficient_history="allow"</code> permits the order while recording that VaR was unavailable;</p></li><li><p><code>insufficient_history="block"</code> rejects positive buy exposure until enough history exists.</p></li></ul><p>This is a policy choice that belongs in experiment metadata. Missing risk information is not evidence that risk is zero.</p><h3>14.7 Projecting VaR consistently in currency units</h3><p>The corrected unit contract for the risk filter is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;9d795833-9360-4f58-b21c-655947ce1f7e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def projected_var(
    self,
    estimate: VaREstimate,
    *,
    exposure: float,
    portfolio_value: float,
) -&gt; float:
    """Return projected VaR as a currency amount."""
    if not estimate.sufficient_history:
        return float("nan")
    if portfolio_value &lt;= 0:
        raise ValueError("portfolio_value must be positive")
    if exposure &lt; 0:
        raise ValueError("exposure must be nonnegative")
    return estimate.var * exposure</code></pre></div><p>The portfolio value is needed for the limit, not for dividing the projected loss:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;3f08ab2e-9923-4373-9af6-10f94e617548&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">limit = self.max_var_fraction * portfolio_value
projected = self.projected_var(
    estimate,
    exposure=projected_exposure,
    portfolio_value=portfolio_value,
)</code></pre></div><p>For a 4% return VaR and <code>$2,000</code> exposure, <code>projected_var</code> returns <code>$80</code>. For a <code>$10,000</code> portfolio and a 5% limit, the comparison is <code>$80 &lt;= $500</code>; both sides are currency amounts.</p><p>This avoids mixing a portfolio-relative fraction with a currency limit. If a project instead chooses to return a relative fraction, it must compare that result with <code>max_var_fraction</code> directly. The implementation described here uses currency amounts throughout the decision path.</p><h3>14.8 Accept, scale, and block behavior</h3><p><code>VaRTradeFilter</code> applies a risk decision to a proposed <strong>notional quantity</strong>. Its modes are:</p><h4>Accept</h4><p>Approve the proposed quantity. This can be used when VaR is informational or when the order is already within the limit.</p><h4>Block</h4><p>Return an approved quantity of zero when projected currency VaR exceeds the currency limit.</p><h4>Scale</h4><p>Approve the largest quantity that remains within the limit. With portfolio value <code>V</code>, limit fraction <code>m</code>, and return VaR <code>v</code>, the maximum total exposure is:</p><p>Emax=mVv,</p><p>provided <code>v &gt; 0</code>. The maximum additional buy is this amount minus current exposure.</p><p>A consistent decision core is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;209a9325-8221-4b9b-9b76-176e4801676d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">limit = self.max_var_fraction * portfolio_value
buy_amount = max(proposed_quantity, 0.0)
full_exposure = current_exposure + buy_amount
full_var = self.projected_var(
    estimate,
    exposure=full_exposure,
    portfolio_value=portfolio_value,
)

if proposed_quantity &lt;= 0 or full_var &lt;= limit:
    return RiskDecision(
        "accept",
        proposed_quantity,
        proposed_quantity,
        1.0,
        full_var,
        limit,
        current_exposure,
        full_exposure,
        "within VaR limit or exposure-reducing order",
        True,
    )

if self.action == "block" or estimate.var &lt;= 0:
    return RiskDecision(
        "block",
        proposed_quantity,
        0.0,
        0.0,
        full_var,
        limit,
        current_exposure,
        current_exposure,
        "proposed exposure exceeds VaR limit",
        True,
    )

allowed_exposure = limit / estimate.var
allowed_buy = max(0.0, allowed_exposure - current_exposure)
approved = min(proposed_quantity, allowed_buy)
scale = approved / proposed_quantity if proposed_quantity else 0.0</code></pre></div><p>The critical correction is <code>allowed_exposure = limit / estimate.var</code>, not <code>limit * portfolio_value / estimate.var</code>, because <code>limit</code> is already a currency amount. Similarly, <code>projected_var</code> returns <code>estimate.var * exposure</code>, not <code>estimate.var * exposure / portfolio_value</code>.</p><h3>14.9 Connecting VaR to the strategy layer</h3><p>The strategy flow is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;6c0555fb-4b42-4dbb-b456-514d1f952f4a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">forecast
  -&gt; signal
  -&gt; candidate notional
  -&gt; cash and exposure checks
  -&gt; VaR accept / scale / block
  -&gt; greedy ranking
  -&gt; next-bar execution</code></pre></div><p>The VaR call must use the same argument names and units as <code>VaRTradeFilter.decide</code>. In particular, <code>GreedyPolicy</code> should pass the proposed order as <code>proposed_quantity</code>, current marked exposure as <code>current_exposure</code>, and portfolio value as <code>portfolio_value</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d74926f7-8801-49b0-90c8-112cf8bf245d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">risk_decision = self.risk_filter.decide(
    proposed_quantity=action.notional,
    current_exposure=snapshot.holdings * snapshot.price,
    portfolio_value=snapshot.value,
    returns=returns,
)</code></pre></div><p>The selected quantity is then taken from <code>risk_decision.approved_quantity</code>. A blocked action must be marked infeasible, and a scaled action must be rescored using its approved notional before greedy ranking. The risk filter must never increase the proposed quantity.</p><p>This explicit adapter is preferable to relying on incompatible keyword aliases such as <code>proposed_notional</code> when the risk API requires <code>proposed_quantity</code>. The strategy and risk modules should share one documented contract rather than silently catching argument errors.</p><h3>14.10 Treatment of sells</h3><p>The default filter is a long-exposure filter. A negative proposed quantity represents a sell. A sell that reduces long exposure is accepted by this filter because it does not increase the projected long exposure.</p><p>This does not constitute a complete portfolio-risk model. It does not model short positions, leverage, cross-asset correlations, liquidity, spread, market impact, or nonlinear instruments. The paper's brief mention of VaR does not justify assuming those features.</p><h3>14.11 Numerical examples</h3><p>Suppose:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;216eaadf-9c85-4010-be84-635ff7154763&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">portfolio value       = $10,000
current exposure       = $2,000
proposed buy           = $3,000
return VaR             = 4%
maximum VaR fraction   = 5%</code></pre></div><p>The complete proposed exposure is <code>$5,000</code>, so projected currency VaR is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;5338afa7-cbd6-468e-88f8-be1984d593d0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">$5,000 * 0.04 = $200</code></pre></div><p>The currency limit is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;35ff3ebf-34c6-4dc9-b6da-af41b634e9cf&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">$10,000 * 0.05 = $500</code></pre></div><p>Because <code>$200 &lt;= $500</code>, the VaR layer accepts the proposed order, subject to cash, sizing, fee, and exposure constraints elsewhere.</p><p>If return VaR is 12% instead:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;38967bd8-ea11-44ba-a569-4e2be4e7890f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">full projected VaR = $5,000 * 0.12 = $600
risk limit         = $500</code></pre></div><p>In scale mode, maximum total exposure is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;68455c85-a875-4950-98d2-65cc83abde7d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">$500 / 0.12 = approximately $4,166.67</code></pre></div><p>After subtracting the existing <code>$2,000</code> exposure, the maximum additional buy is approximately <code>$2,166.67</code>. In block mode, the <code>$3,000</code> proposed order is rejected.</p><p>These examples describe the reconstructed unit-consistent policy. They do not establish what the paper's original VaR procedure did.</p><h3>14.12 Tests and verification boundary</h3><p>Tests for the corrected API should call the actual interfaces directly. For example:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8ca4e4bd-b3d2-49d4-ae21-f1609f37b0c7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">var = HistoricalVaR(
    confidence=0.95,
    horizon=1,
    lookback=len(returns),
    min_observations=2,
)
estimate = var.estimate(returns)

risk_filter = VaRTradeFilter(
    var_estimator=var,
    max_var_fraction=0.005,
    action="/__u/onepagecode.substack.com/block",
)
decision = risk_filter.decide(
    proposed_quantity=500.0,
    current_exposure=0.0,
    portfolio_value=1_000.0,
    estimate=estimate,
)

assert decision.approved_quantity &lt;= 500.0</code></pre></div><p>The test should not pass <code>portfolio_value</code> to <code>HistoricalVaR.estimate</code>, because historical estimation uses returns and VaR configuration; portfolio value belongs to exposure projection and trade filtering. It should also use <code>max_var_fraction</code>, not an unsupported <code>max_var</code> constructor argument, and should call <code>decide</code> with <code>proposed_quantity</code>, <code>current_exposure</code>, and <code>portfolio_value</code>.</p><p>These are invariant-level checks, not empirical validation. Static and semantic review can inspect units, argument contracts, and no-look-ahead behavior, but no claim is made here that the generated code was executed. The supplied verification record states that semantic code verification was skipped, so the risk tests and integration snippets should be treated as corrected intended contracts until the source artifacts are aligned and independently checked.</p><h3>14.13 Reproducibility boundary</h3><p>This VaR layer remains <strong>reconstructed</strong>. A verified implementation of the paper would require at least:</p><p>the confidence level;</p><p>the VaR horizon;</p><p>the return definition and portfolio aggregation method;</p><p>the estimation window;</p><p>the historical or parametric estimation method;</p><p>the maximum acceptable loss or exposure threshold;</p><p>whether VaR blocks, scales, or merely reports trades;</p><p>how VaR interacts with the greedy policy and position-sizing schedule.</p><p>The accurate claim is therefore: <strong>the project defines a documented rolling historical VaR filter with consistent currency units and explicit action rules, inspired by the paper's risk-control description</strong>. It does not claim to reproduce the paper's unspecified VaR method or its reported investment outcomes.</p><h2>15. Simulate Next-Bar Execution and Portfolio Accounting</h2><p>A forecast becomes useful to a backtest only after it is converted into an order and accounted for consistently. The paper describes a combined trading framework, but it does not fully specify signal timing, execution prices, fee conventions, slippage, liquidation, or portfolio accounting. The generated implementation therefore makes these choices explicit in <code>src/quantitative_trading_model/backtest.py</code>.</p><p>These rules are reconstructions, not claims about the paper's exact backtest. The simulator is offline: it accepts caller-supplied local prices and orders and does not connect to an exchange or submit live orders.</p><h3>15.1 Signal time and execution time</h3><p>A strategy forms a signal using information available at time <code>t</code>. The order must execute at a later timestamp. The simulator enforces that ordering:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;1e72bd31-6c1b-4876-8c03-29020e8fa95a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">historical window ends at t
    -&gt; forecast and signal are generated at t
    -&gt; order is submitted for a later bar
    -&gt; fill occurs after t
    -&gt; portfolio is marked to market</code></pre></div><p>The <code>Order</code> dataclass records both the signal and intended execution timestamps:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;508a03aa-1ccf-4c96-a8e6-ea5278c2ecfa&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class Order:
    side: str
    quantity: float | None = None
    signal_timestamp: Any = None
    execution_timestamp: Any = None
    notional: float | None = None
    addition_level: int | None = None
    reason: str = ""
    metadata: Mapping[str, Any] = field(default_factory=dict)</code></pre></div><p>An order may specify asset units through <code>quantity</code> or a currency amount through <code>notional</code>. In the current implementation, a notional order is converted to units using the supplied <strong>reference price</strong> before slippage is applied. Therefore, with nonzero slippage, the executed notional can differ from the requested notional. This is an implementation detail that should be recorded in experiment metadata rather than described as conversion at the final slippage-adjusted price.</p><p><code>PortfolioSimulator.execute</code> rejects same-bar and earlier execution:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e11e2106-9916-41ce-a97a-095a3caf0a77&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">fill_time = _timestamp(timestamp)
signal_time = _timestamp(order.signal_timestamp)
if fill_time &lt;= signal_time:
    raise ValueError(
        "Orders must execute strictly after their signal timestamp"
    )</code></pre></div><p>This prevents a signal generated from a closing price from being filled at that same closing price. The simulator enforces a later timestamp, but it does not independently determine that the timestamp is exactly the next bar. A caller must supply next-bar orders or use a scheduling layer that does so.</p><p>Each completed order becomes a <code>Fill</code> record containing its timing and accounting details:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;384e4d94-f5e9-4e20-8469-67c930b75799&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class Fill:
    timestamp: Timestamp
    signal_timestamp: Timestamp
    side: str
    quantity: float
    reference_price: float
    execution_price: float
    notional: float
    fee: float
    cash_change: float
    reason: str = ""
    addition_level: int | None = None</code></pre></div><p>Thus, a fill is an auditable event rather than merely a change in portfolio value.</p><h3>15.2 Fees and executed notional</h3><p>For an executed notional <code>N</code> and fee rate <code>f</code>, the configured paper-inspired convention is:</p><p>buy cash cost=N(1+f),</p><p>sell cash proceeds=N(1&#8722;f).</p><p>The current simulator applies one fee rate to both buys and sells whenever that rate is configured. However, the bare <code>PortfolioSimulator</code> constructor defaults to a fee rate of <code>0.0</code>. The two-sided fee behavior is therefore a configured convention, not an unconditional simulator default. <code>ExecutionConfig</code> supplies a paper-inspired fee rate when it is passed to the simulator, but the current simulator does not consult the <code>fee_on_buy</code> and <code>fee_on_sell</code> boolean fields. Those flags should not be described as controlling accounting until the implementation is extended to enforce them.</p><p>The core fill calculation is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;3f7b3581-49a3-4cb5-b1a4-070fd67f647a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">notional = quantity * effective_price
fee = notional * self.fee_rate

if side == "buy":
    cash_change = -(notional + fee)
    self.state.quantity += quantity
    self.state.cash += cash_change
else:
    cash_change = notional - fee
    self.state.quantity -= quantity
    self.state.cash += cash_change

self.state.fees_paid += fee
self.state.traded_notional += notional</code></pre></div><p><code>PortfolioState</code> stores the mutable account state:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ae8dc651-4111-4a0c-a717-62db99ced9c8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass
class PortfolioState:
    cash: float
    quantity: float = 0.0
    average_price: float = 0.0
    fees_paid: float = 0.0
    traded_notional: float = 0.0
    addition_level: int = 0</code></pre></div><p>The average price is informational. Portfolio value is calculated from cash, quantity, and the current market price. Cumulative fees and traded notional support later turnover and fee-sensitivity reporting.</p><h3>15.3 Slippage</h3><p>The reference price is the observed price supplied to the simulator. Optional proportional slippage adjusts it against the trader:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;3866e612-3219-4657-8b6a-ff51306404e4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">effective_price = price * (
    1.0 + self.slippage_rate if side == "buy"
    else 1.0 - self.slippage_rate
)</code></pre></div><p>A buy therefore receives a higher effective price and a sell receives a lower one. Fees are calculated from the resulting executed notional.</p><p>This is a simple slippage model. It does not represent bid-ask spread, market impact, liquidity, partial fills, or intrabar price paths. The paper does not provide those details, so proportional slippage is an explicit reconstruction rather than a paper-faithful assumption.</p><p>Because notional-to-unit conversion currently uses the reference price before this adjustment, a request for a fixed notional should be interpreted as a reference-price notional. If exact post-slippage notional sizing is required, the conversion logic in <code>backtest.py</code> must be changed before using that interpretation.</p><h3>15.4 Units, cash limits, and leverage</h3><p>The current simulator always permits fractional quantities. Although <code>ExecutionConfig</code> contains an <code>allow_fractional_units</code> field, <code>PortfolioSimulator</code> does not currently enforce it. Fractional-unit support should therefore be described as the simulator's present behavior, not as a configurable guarantee.</p><p>When leverage is disabled, buys are clipped to available cash:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;802a0ab0-c06f-437f-8cef-f37aeaced5b8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">if side == "buy" and not self.allow_leverage:
    affordable = self.state.cash / (
        effective_price * (1.0 + self.fee_rate)
    )
    quantity = min(quantity, max(0.0, affordable))</code></pre></div><p>Likewise, a long-only account cannot sell more than it owns:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6d280058-234b-4c8a-9d88-495752fe9e34&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">if side == "sell" and not self.allow_short:
    quantity = min(quantity, max(0.0, self.state.quantity))</code></pre></div><p>These safeguards prevent negative cash and negative holdings under the simulator's default long-only policy. They are safety choices made by the implementation, not evidence that the paper used the same constraints. The paper does not state whether it allowed leverage, margin, short positions, or unlimited reinvestment.</p><p>For example, an account with <code>$1,000</code> buys <code>$400</code> of an asset at a configured fee rate of <code>0.5%</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;9574e3b0-44df-491b-89d7-2cb5573768a4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">reference notional = $400.00
fee                = $400.00 * 0.005 = $2.00
cash decrease      = $402.00
cash remaining     = $598.00</code></pre></div><p>At a reference price of <code>$100</code>, the requested notional becomes four units before any slippage adjustment.</p><h3>15.5 Mark-to-market valuation</h3><p>At each valuation timestamp, the simulator computes:</p><p>\[ V<em>t=\text{cash}</em>t+\text{quantity}<em>t\times p</em>t. \]</p><p>The corresponding method is straightforward:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;276f0e0d-b351-4f2e-b8f0-1f326e66ad7b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def portfolio_value(self, price: float) -&gt; float:
    price = _number(price)
    return self.cash + self.quantity * price</code></pre></div><p><code>mark_to_market</code> stores the result in a <code>BacktestRecord</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d53121a5-1614-4ec3-bd37-2da6ee7f37fe&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass(frozen=True)
class BacktestRecord:
    timestamp: Timestamp
    price: float
    cash: float
    quantity: float
    portfolio_value: float
    fees_paid: float
    traded_notional: float
    fill_count: int = 0</code></pre></div><p>Valuation timestamps must increase strictly. The resulting value must not be materially negative. These rules provide an accounting invariant that can be checked independently for every record.</p><p>Continuing the example, four units held at a market price of <code>$105</code> produce:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;b4c55f71-086f-4113-8f3a-dad42f89898e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">cash             = $598.00
marked holdings  = 4 * $105 = $420.00
portfolio value  = $1,018.00</code></pre></div><p>The entry fee is already reflected in cash. A later sale incurs a further fee when the configured fee rate is nonzero.</p><h3>15.6 Explicit liquidation</h3><p>The final value of an open position can mean either its marked value or the cash remaining after liquidation. These are different when selling incurs fees. The simulator therefore does not silently assume liquidation.</p><p><code>liquidate</code> creates a sell order for the remaining long quantity and places it after the last valuation timestamp:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d7050202-0d4e-4e71-8f4c-6fab04d387df&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def liquidate(
    self,
    timestamp: Any,
    price: float,
    reason: str = "end_of_backtest",
) -&gt; Fill | None:
    if self.state.quantity &lt;= 0:
        return None
    signal_time = self._last_valuation_timestamp
    return self.execute(
        Order(
            side="sell",
            quantity=self.state.quantity,
            signal_timestamp=signal_time,
            reason=reason,
        ),
        timestamp,
        price,
    )</code></pre></div><p>The liquidation fill remains in the event ledger and any configured sell fee is included in cumulative costs. Whether to liquidate depends on the research question. A cash-outcome report generally requires liquidation, while a marked portfolio report may not. The paper does not specify which convention produced its headline values, so the choice must accompany every result.</p><h3>15.7 Buy-and-hold benchmark</h3><p>A strategy should be compared with passive ownership over the same period. The <code>buy_and_hold</code> function buys at the first available bar, applies the configured purchase assumptions, holds the asset, and optionally liquidates at the end.</p><p>Its entry signal is placed just before the first bar so that it satisfies the same timestamp contract:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;fb29c74e-6cb3-4207-ac29-45589f51600b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">first = _timestamp(frame.index[0])
signal_time = first - pd.Timedelta(nanoseconds=1)

quantity = initial_cash / (
    float(frame.iloc[0, 0])
    * (1.0 + slippage_rate)
    * (1.0 + fee_rate)
)

order = Order(
    side="buy",
    quantity=quantity,
    signal_timestamp=signal_time,
    execution_timestamp=first,
    reason="buy_and_hold_entry",
)</code></pre></div><p>This benchmark uses no forecasts, streak statistics, VaR, or greedy selection. It helps distinguish strategy-specific behavior from the return generated simply by holding the asset. The paper does not report a buy-and-hold comparison, so adding one is an evaluation improvement rather than a claim about the original method.</p><h3>15.8 Accounting tests and verification limits</h3><p><code>tests/test_backtest_accounting.py</code> contains hand-calculable tests for intended contracts such as fee arithmetic, later execution, valuation identity, no-leverage behavior, liquidation, and buy-and-hold accounting. A representative invariant is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d22e9829-e66f-4653-adab-92ada50add96&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">cash = _field(state, "cash")
quantity = _field(state, "quantity", "holdings", "asset_quantity")
value = _field(record, "portfolio_value", "value", "equity")
assert value == pytest.approx(cash + quantity * 12.0)</code></pre></div><p>These files are static, intended test artifacts. They should not be presented as executed or passing. The supplied verification record reports that semantic code verification was skipped. It also notes apparent API inconsistencies between some generated tests and the generated sizing, risk, strategy, and configuration modules. Consequently, these tests may require API alignment before they can serve as executable checks.</p><p>The useful claim is narrower: the tests document the accounting invariants that a corrected implementation should satisfy. They do not establish runtime correctness, profitability, or reproduction of the paper.</p><h3>15.9 Why the paper's headline values cannot be validated</h3><p>The paper reports approximately <code>$646</code> from a <code>$500</code> gold allocation and <code>$215,487</code> from a <code>$500</code> Bitcoin allocation. The simulator does not treat these figures as expected outputs. Validation would require, at minimum:</p><ul><li><p>exact assets and historical dates;</p></li><li><p>sampling frequency and price field;</p></li><li><p>forecast target and signal timestamps;</p></li><li><p>buy, sell, hold, and add-position rules;</p></li><li><p>original position-sizing parameters;</p></li><li><p>VaR confidence, horizon, and action rule;</p></li><li><p>greedy objective and constraints;</p></li><li><p>fee, spread, and slippage conventions;</p></li><li><p>fractional-unit, leverage, and reinvestment rules;</p></li><li><p>liquidation or mark-to-market convention.</p></li></ul><p>Without these details, matching a final number would not demonstrate faithful reproduction. It could result from selecting a similar period or from offsetting errors in data preparation and accounting.</p><p>The appropriate interpretation is:</p><p>The backtest module demonstrates timestamped execution and auditable portfolio accounting under explicit assumptions. It does not verify the paper's reported profits.</p><p>This distinction keeps the implementation useful for research while avoiding unsupported claims about profitability or live-trading readiness.</p><h2>16. Orchestrate Forecast Experiments and Fee Sensitivity</h2><p>The experiment layer is intended to make research runs repeatable, but its current generated implementation has a narrower role than the complete architecture described earlier. <code>src/quantitative_trading_model/experiments.py</code> currently provides:</p><ul><li><p>chronological forecast orchestration;</p></li><li><p>expanding walk-forward forecast orchestration;</p></li><li><p>optional benchmark dispatch;</p></li><li><p>fee-sensitivity callback infrastructure; and</p></li><li><p>result metadata and local JSON/CSV persistence.</p></li></ul><p>It does <strong>not yet directly compose</strong> <code>strategy.py</code>, <code>sizing.py</code>, <code>risk.py</code>, and <code>backtest.py</code> into one end-to-end trading run. A complete integration would need an additional orchestration function that converts forecasts into signals, applies sizing and VaR, selects actions, submits next-bar orders, and collects portfolio metrics. The current fee-sensitivity function instead receives that behavior through an injected callback.</p><p>The distinction matters because an experiment runner can preserve assumptions without necessarily implementing every stage itself. The paper's data, execution protocol, position-sizing equations, VaR settings, and greedy policy are also incomplete, so the orchestration layer must report both its own implementation status and the paper-fidelity limitations.</p><h3>16.1 Current position in the research pipeline</h3><p>The intended overall workflow is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;3a1c951a-2c8c-406e-a43f-8a0a080eb1e4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">local price data
    -&gt; cleaning and sliding windows
    -&gt; chronological split or walk-forward folds
    -&gt; forecast model and optional benchmarks
    -&gt; signals, sizing, VaR, and greedy policy
    -&gt; event-driven backtest
    -&gt; forecast and portfolio metrics
    -&gt; result metadata and local files</code></pre></div><p>The current <code>experiments.py</code> implementation covers the forecast and result-recording portions directly. The trading stages can be supplied by a caller through a backtest callback, but they are not automatically wired together by <code>run_forecast_experiment</code> or <code>run_fee_sensitivity</code>.</p><p>Its common result container is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;28138bde-d030-460f-8d4b-883a462651ba&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass
class ExperimentResult:
    experiment_type: str
    asset: str | None = None
    model: str | None = None
    status: str = "completed"
    metrics: dict[str, Any] = field(default_factory=dict)
    portfolio: dict[str, Any] = field(default_factory=dict)
    metadata: dict[str, Any] = field(default_factory=dict)
    predictions: list[dict[str, Any]] = field(default_factory=list)
    warnings: list[str] = field(default_factory=list)</code></pre></div><p>The separate <code>metrics</code> and <code>portfolio</code> fields are useful even though the current forecast runner primarily fills <code>metrics</code>. Forecast errors such as RMSE and MAE belong in <code>metrics</code>; final value, drawdown, turnover, and fees belong in <code>portfolio</code> when a backtest callback supplies them. The <code>warnings</code> field should identify unavailable dependencies, unresolved interfaces, and paper claims that were not reproduced.</p><p><code>to_dict()</code>, <code>to_json()</code>, <code>summarize_results()</code>, and <code>save_local_results()</code> are serialization utilities. They preserve local research records; they do not send results to a market service or submit orders.</p><h3>16.2 Paper-style chronological forecast experiments</h3><p>The paper reports a 70:30 train/test split. The current <code>run_forecast_experiment</code> function is the generated implementation's paper-style forecast path:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c520a358-8770-4922-b2ae-f7affcae5767&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def run_forecast_experiment(
    prices: Any,
    *,
    asset: str = "asset",
    config: ResearchConfig | None = None,
    model_name: str = "att_bilstm",
    model: Any = None,
    include_benchmarks: bool = False,
    seed: int | None = 7,
) -&gt; ExperimentResult:</code></pre></div><p>Its intended sequence is:</p><p>Validate a supplied <code>ResearchConfig</code>.</p><p>Standardize local prices.</p><p>Create windows using the configured window length, stride, and horizon.</p><p>Split samples chronologically.</p><p>Fit the requested model on the training partition.</p><p>Predict the held-out test partition.</p><p>Calculate forecast metrics.</p><p>Store model, split, horizon, seed, and warning metadata.</p><p>This is forecast orchestration, not a complete portfolio experiment. It does not itself call the signal generator, position sizer, VaR filter, greedy policy, or portfolio simulator. A caller that needs trading results must connect those modules separately or provide a higher-level callback.</p><p>The intended metadata shape is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;1e1b66ef-5d54-48c3-8c69-58ae373f0ebb&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">metadata = {
    "asset": asset,
    "model": model_name,
    "split": "chronological_70_30_or_configured",
    "window_length": 100,
    "horizon": 1,
    "seed": 7,
    "data_points": len(values),
    "synthetic_data_is_not_paper_validation": True,
}</code></pre></div><p>This metadata design is useful, but it should not be treated as proof that every field is currently populated correctly. The generated code contains an unresolved interface mismatch: <code>run_forecast_experiment</code> attempts to pass <code>timestamps</code> to <code>standardize_prices</code>, while the generated <code>standardize_prices</code> signature accepts a DataFrame or <code>PriceData</code> and does not define a <code>timestamps</code> keyword. Timestamp-aware runs therefore require API alignment before they can be considered runtime-ready.</p><p>The paper's exact data source, split date, scaling boundary, and target protocol are also unavailable. A chronological split is therefore a documented comparison protocol, not an exact reproduction of the paper's experiment.</p><h3>16.3 Walk-forward forecast experiments</h3><p>For trading conclusions, expanding walk-forward evaluation is preferable to one fixed split:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;ba9870fb-f295-43fb-84b0-3f6b62c6a197&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">train through t
forecast the next test block
move the boundary forward
expand the training history
forecast again
repeat</code></pre></div><p><code>run_walk_forward_experiment</code> is intended to retrain or invoke a supplied <code>train_predict</code> callback for each fold, then aggregate the fold predictions and targets. This protocol better reflects deployment because each forecast uses only the historical information available at that point.</p><p>Walk-forward evaluation does not remove all sources of optimism. Feature construction, scaling, model selection, and hyperparameter tuning must still be restricted to the information available in each fold. Walk-forward results also should not be compared directly with the paper's 70:30 metrics without explaining the protocol difference.</p><h3>16.4 Fee grids are experiment inputs, not recovered results</h3><p>The paper describes the following transaction-fee scenarios:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!2aYh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a32d320-ffc9-4eb8-a545-61e205883377_1114x224.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!2aYh!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a32d320-ffc9-4eb8-a545-61e205883377_1114x224.png 424w, /__u/substackcdn.com/image/fetch/$s_!2aYh!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a32d320-ffc9-4eb8-a545-61e205883377_1114x224.png 848w, /__u/substackcdn.com/image/fetch/$s_!2aYh!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a32d320-ffc9-4eb8-a545-61e205883377_1114x224.png 1272w, /__u/substackcdn.com/image/fetch/$s_!2aYh!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a32d320-ffc9-4eb8-a545-61e205883377_1114x224.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!2aYh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a32d320-ffc9-4eb8-a545-61e205883377_1114x224.png" width="1114" height="224" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9a32d320-ffc9-4eb8-a545-61e205883377_1114x224.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:224,&quot;width&quot;:1114,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:38595,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a32d320-ffc9-4eb8-a545-61e205883377_1114x224.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!2aYh!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a32d320-ffc9-4eb8-a545-61e205883377_1114x224.png 424w, /__u/substackcdn.com/image/fetch/$s_!2aYh!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a32d320-ffc9-4eb8-a545-61e205883377_1114x224.png 848w, /__u/substackcdn.com/image/fetch/$s_!2aYh!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a32d320-ffc9-4eb8-a545-61e205883377_1114x224.png 1272w, /__u/substackcdn.com/image/fetch/$s_!2aYh!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a32d320-ffc9-4eb8-a545-61e205883377_1114x224.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><code>FeeSensitivityConfig</code> defines these fields:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e55be5c2-df11-410d-809f-4b9d1bdb0d84&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">@dataclass
class FeeSensitivityConfig:
    gold_rates: tuple[float, ...] = (
        0.002, 0.004, 0.006, 0.008, 0.010, 0.012
    )
    bitcoin_rates: tuple[float, ...] = (
        0.015, 0.020, 0.025, 0.030, 0.035
    )
    apply_to_buys: bool = True
    apply_to_sells: bool = True
    hold_signals_fixed: bool = True</code></pre></div><p>A decimal rate of <code>0.002</code> means 0.2%. Passing <code>0.2</code> would mean 20% and would be a different experiment.</p><p>There is an important current implementation limitation: <code>run_fee_sensitivity</code> looks for configuration attributes named <code>gold_fees</code> and <code>bitcoin_fees</code>, but <code>FeeSensitivityConfig</code> defines <code>gold_rates</code> and <code>bitcoin_rates</code>. As generated, custom configured grids are therefore not reliably consumed; the function can fall back to its module-level default grids. This naming mismatch must be repaired before claiming that arbitrary <code>FeeSensitivityConfig</code> values control the sweep.</p><h3>16.5 Fixed signals versus regenerated policy decisions</h3><p>Fee sensitivity can answer two different questions.</p><h4>Fixed-signal sensitivity</h4><p>Generate forecasts and orders once, then replay the same orders under each fee rate:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;f52193d7-4377-4b60-9b51-95d18539236b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">same data
same forecasts
same signals
same sizing schedule
fee rate changes
backtest reruns</code></pre></div><p>This isolates the direct accounting effect of fees.</p><h4>Fee-aware policy sensitivity</h4><p>Regenerate candidate actions for each fee rate when fees affect:</p><ul><li><p>whether an order is affordable;</p></li><li><p>whether its expected return exceeds its cost;</p></li><li><p>which greedy candidate has the highest score; or</p></li><li><p>whether exposure or risk constraints permit it.</p></li></ul><p>The current <code>run_fee_sensitivity</code> function accepts a <code>run_backtest</code> callback and passes <code>fee_rate</code> to that callback:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ba66d656-3438-4df4-b2e9-79eed57127b7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def run_fee_sensitivity(
    run_backtest: Callable[..., Any],
    *,
    asset: str,
    config: FeeSensitivityConfig | None = None,
    fees: Sequence[float] | None = None,
    base_kwargs: Mapping[str, Any] | None = None,
) -&gt; list[ExperimentResult]:</code></pre></div><p>This callback-based design is intentional infrastructure, not direct strategy integration. The callback must decide whether to reuse signals or regenerate them. The current function does <strong>not</strong> inspect or enforce <code>FeeSensitivityConfig.hold_signals_fixed</code>, and it does not currently record an explicit <code>signals_fixed</code> or <code>policy_regenerated</code> field. Its hardcoded metadata field <code>non_fee_assumptions_held_fixed=True</code> only states the intended sweep contract; it does not establish how the callback selected actions.</p><p>A repaired orchestration contract should record the policy mode explicitly, for example:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;dd77051f-77dd-48d3-b751-bf73630b8290&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">result.metadata.update(
    {
        "fee_rate": rate,
        "fee_percent": rate * 100.0,
        "signals_fixed": True,
        "policy_regenerated": False,
        "non_fee_assumptions_held_fixed": True,
    }
)</code></pre></div><p>For a fee-aware callback, the last two policy fields should instead indicate that decisions were regenerated. Until that metadata and configuration handling are aligned, fee-sweep results should be treated as callback outputs whose policy behavior must be inspected separately.</p><h3>16.6 Optional benchmark orchestration</h3><p><code>run_forecast_experiment(include_benchmarks=True)</code> is intended to request the optional benchmark suite. However, the generated call currently passes names such as <code>train_data</code> and <code>test_data</code>, while <code>run_benchmark_suite</code> requires <code>train_values</code> and <code>test_values</code>. The compatibility helper filters keyword arguments but cannot invent missing required parameter names. Consequently, the benchmark path may fail and be converted into a warning rather than producing benchmark results.</p><p>The required interface should be aligned explicitly, for example:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8558fe0b-f85b-47ae-ae74-037cd80542f0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">benchmark = run_benchmark_suite(
    train_values=train_data,
    test_values=test_data,
    metadata={"asset": asset, "horizon": horizon},
)</code></pre></div><p>Until that repair is made, benchmark orchestration should be described as an intended dispatch path, not as a verified working connection. Optional dependency status and model semantics remain important: HMM Viterbi decoding is not automatically a price forecast, and XGBoost is gradient-boosted trees rather than a random forest. Benchmark metrics also require aligned targets, timestamps, preprocessing, and untouched test data.</p><h3>16.7 Recording results and paper claims</h3><p>The result layer can save local artifacts:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;da28ea9c-5bab-4e0f-84fe-d312782f255c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def save_local_results(
    results: ExperimentResult | Iterable[ExperimentResult],
    path: str | Path,
    *,
    format: str | None = None,
) -&gt; Path:</code></pre></div><p>JSON preserves nested predictions, warnings, and metadata. CSV is convenient for scalar comparison tables. Neither format establishes empirical validity.</p><p>The paper's headline outcomes are stored as claims rather than expected outputs:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;707ca046-bb7c-4b90-95d4-bee9d00b0f61&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">PAPER_REPORTED_VALUES = {
    "gold_final_value_usd": {
        "value": 646.0,
        "status": "reported_unverified_claim",
    },
    "bitcoin_final_value_usd": {
        "value": 215487.0,
        "status": "reported_unverified_claim",
    },
    "combined_final_value_usd": {
        "value": 216133.0,
        "status": "reported_unverified_claim",
    },
}</code></pre></div><p>These values must not be used as unit-test expectations. The original data, dates, signals, sizing equations, VaR settings, greedy rules, fee convention, and liquidation protocol are missing. Agreement with one number would not prove that the full method was reproduced; disagreement under a different local protocol would not by itself identify a coding error.</p><h3>16.8 What a trustworthy fee-sensitivity record should contain</h3><p>Each scenario should preserve at least:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!g9K1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10015f10-c7d8-4205-95b2-99cfcd494293_1236x666.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!g9K1!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10015f10-c7d8-4205-95b2-99cfcd494293_1236x666.png 424w, /__u/substackcdn.com/image/fetch/$s_!g9K1!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10015f10-c7d8-4205-95b2-99cfcd494293_1236x666.png 848w, /__u/substackcdn.com/image/fetch/$s_!g9K1!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10015f10-c7d8-4205-95b2-99cfcd494293_1236x666.png 1272w, /__u/substackcdn.com/image/fetch/$s_!g9K1!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10015f10-c7d8-4205-95b2-99cfcd494293_1236x666.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!g9K1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10015f10-c7d8-4205-95b2-99cfcd494293_1236x666.png" width="1236" height="666" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/10015f10-c7d8-4205-95b2-99cfcd494293_1236x666.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:666,&quot;width&quot;:1236,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:116265,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10015f10-c7d8-4205-95b2-99cfcd494293_1236x666.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!g9K1!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10015f10-c7d8-4205-95b2-99cfcd494293_1236x666.png 424w, /__u/substackcdn.com/image/fetch/$s_!g9K1!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10015f10-c7d8-4205-95b2-99cfcd494293_1236x666.png 848w, /__u/substackcdn.com/image/fetch/$s_!g9K1!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10015f10-c7d8-4205-95b2-99cfcd494293_1236x666.png 1272w, /__u/substackcdn.com/image/fetch/$s_!g9K1!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10015f10-c7d8-4205-95b2-99cfcd494293_1236x666.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The current <code>summarize_results</code> function flattens scalar fields from result metrics, portfolio summaries, and metadata. <code>save_local_results</code> writes those records locally. Before relying on the files, inspect whether the callback supplied portfolio fields and whether policy metadata accurately describes the run.</p><h3>16.9 Practical checklist and current readiness</h3><p>Before preparing a forecast or fee experiment, record:</p><p>Local data path and provenance.</p><p>Asset, instrument, currency, frequency, and date range.</p><p>Window length, stride, target representation, and horizon.</p><p>Chronological or walk-forward boundaries.</p><p>Scaling method and fitting rows.</p><p>Model architecture, seed, and optional dependency versions.</p><p>Signal aggregation and thresholds.</p><p>Sizing schedule and exposure limits.</p><p>VaR confidence, horizon, lookback, and action mode.</p><p>Fee convention, slippage, execution timing, and liquidation policy.</p><p>Whether signals are fixed or regenerated for each fee scenario.</p><p>Forecast metrics and portfolio metrics in separate fields.</p><p>Any warnings caused by unavailable dependencies or API mismatches.</p><p>The generated project should currently be treated as a static, educational artifact rather than an executed experiment system. Known issues include the timestamp argument mismatch in <code>standardize_prices</code>, the benchmark argument-name mismatch, the unused <code>hold_signals_fixed</code> setting, and the fee-grid naming mismatch. Static checks also reported findings in <code>experiments.py</code>, while semantic code verification was skipped. Therefore, this section documents the intended orchestration contracts and the current limitations without claiming that the experiment paths were executed or that their outputs are correct.</p><h3>16.10 Interpreting the missing fee figures</h3><p>The fee grids can be reproduced as an experiment <strong>design</strong>. The paper's numerical curves cannot be recovered from the extracted text. Do not infer them from the headline profits, interpolate them from the listed percentages, or label a newly generated local curve as the paper's result.</p><p>A local result should instead be described as:</p><p>performance under the repository's documented reconstruction and local dataset</p><p>That wording preserves the value of sensitivity analysis while distinguishing it from empirical reproduction. The experiment layer's current contribution is reproducible configuration and result recording; direct end-to-end trading orchestration remains a repair task rather than an implemented fact.</p><h2>17. Run the Offline Workflow from the Command Line</h2><p>The project defines an offline command-line interface (CLI) for inspecting local data, generating synthetic demonstrations, and requesting research experiments. It does not connect to an exchange, download prices, store credentials, or submit orders.</p><p>The CLI should therefore be understood as a documented interface around the research modules, not as a verified end-to-end trading application. Several generated integrations currently have known incompatibilities. The examples below describe the intended command contracts and identify where the generated files still require repair. No command execution is claimed.</p><h3>17.1 Command registration and optional dependencies</h3><p>The shell command is registered in <code>pyproject.toml</code>:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;toml&quot;,&quot;nodeId&quot;:&quot;85be3289-1549-443a-898b-2f5fd367f6af&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-toml">[project.scripts]
quant-trading-model = "quantitative_trading_model.cli:main"</code></pre></div><p>After installation, this maps <code>quant-trading-model</code> to <code>main()</code> in <code>src/quantitative_trading_model/cli.py</code>. The parser selects a subcommand and the handler produces local output.</p><p>Modeling dependencies are optional:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;toml&quot;,&quot;nodeId&quot;:&quot;8aedac15-6a0d-4194-9e51-0642e8dc15ae&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-toml">[project.optional-dependencies]
deep-learning = ["torch&gt;=2.0"]
classical = [
    "statsmodels&gt;=0.14",
    "pmdarima&gt;=2.0",
    "hmmlearn&gt;=0.3",
    "xgboost&gt;=2.0"
]</code></pre></div><p>The core data, streak, sizing, risk, and accounting modules use the lighter dependencies declared in the main project configuration. PyTorch and classical forecasting packages are imported only by features that need them.</p><p>The intended behavior is to report a missing optional dependency clearly. However, the generated <code>main()</code> catches <code>CLIError</code>, <code>FileNotFoundError</code>, <code>ValueError</code>, and <code>TypeError</code>, but not <code>ImportError</code>. Consequently, a missing PyTorch installation may still produce an uncaught traceback when a forecast command reaches the model layer. The CLI needs an explicit <code>except ImportError</code> branch before this behavior can be described as reliable shell-level error handling.</p><h3>17.2 Defined subcommands</h3><p><code>build_parser()</code> defines these subcommands:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!FofB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d470a25-4a0c-4eb9-bda6-7ae64fe3e465_1412x578.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!FofB!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d470a25-4a0c-4eb9-bda6-7ae64fe3e465_1412x578.png 424w, /__u/substackcdn.com/image/fetch/$s_!FofB!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d470a25-4a0c-4eb9-bda6-7ae64fe3e465_1412x578.png 848w, /__u/substackcdn.com/image/fetch/$s_!FofB!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d470a25-4a0c-4eb9-bda6-7ae64fe3e465_1412x578.png 1272w, /__u/substackcdn.com/image/fetch/$s_!FofB!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d470a25-4a0c-4eb9-bda6-7ae64fe3e465_1412x578.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!FofB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d470a25-4a0c-4eb9-bda6-7ae64fe3e465_1412x578.png" width="1412" height="578" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8d470a25-4a0c-4eb9-bda6-7ae64fe3e465_1412x578.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:578,&quot;width&quot;:1412,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:149631,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d470a25-4a0c-4eb9-bda6-7ae64fe3e465_1412x578.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!FofB!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d470a25-4a0c-4eb9-bda6-7ae64fe3e465_1412x578.png 424w, /__u/substackcdn.com/image/fetch/$s_!FofB!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d470a25-4a0c-4eb9-bda6-7ae64fe3e465_1412x578.png 848w, /__u/substackcdn.com/image/fetch/$s_!FofB!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d470a25-4a0c-4eb9-bda6-7ae64fe3e465_1412x578.png 1272w, /__u/substackcdn.com/image/fetch/$s_!FofB!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d470a25-4a0c-4eb9-bda6-7ae64fe3e465_1412x578.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Thus, the commands are <strong>statically specified interfaces</strong>, not all working workflows. A complete runtime path requires repairing the interfaces identified below.</p><p>The parser excerpt illustrates the command structure:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;63cb07db-1d50-4604-814e-15e952cb9c84&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def build_parser() -&gt; argparse.ArgumentParser:
    parser = argparse.ArgumentParser(
        prog="quant-trading-model",
        description="Offline educational tools for the quantitative trading model paper.",
    )
    subparsers = parser.add_subparsers(dest="command", required=True)

    demo = subparsers.add_parser("demo")
    demo.add_argument("--rows", type=int, default=500)
    demo.add_argument("--seed", type=int, default=7)
    demo.add_argument("--asset", choices=("gold", "bitcoin"), default="gold")
    demo.set_defaults(handler=run_demo)

    # validate-data, forecast, backtest, and fee-sensitivity are
    # registered later in the same function.
    return parser</code></pre></div><p>This is an explanatory excerpt, not a complete replacement for <code>cli.py</code>.</p><h3>17.3 Local-only paths and data provenance</h3><p>The CLI includes path checks intended to reject URLs and network paths:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;35691b31-601c-4d35-8f1c-cbc5d53bf3a0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def _safe_input_path(value: str) -&gt; Path:
    if "://" in value:
        raise CLIError("Only local files are supported; URLs and network paths are not accepted")
    path = Path(value).expanduser()
    if not path.exists():
        raise CLIError(f"Input file does not exist: {path}")
    if not path.is_file():
        raise CLIError(f"Input path is not a regular file: {path}")
    return path.resolve()


def _safe_output_path(value: str) -&gt; Path:
    if "://" in value:
        raise CLIError("Output must be a local path, not a URL")
    path = Path(value).expanduser()
    if path.exists() and path.is_dir():
        raise CLIError(f"Output path is a directory: {path}")
    return path.resolve()</code></pre></div><p>These checks constrain file locations; they do not establish that a dataset matches the paper. Every local result should also record the source file, instrument, currency, frequency, date range, timezone, cleaning rules, and configuration snapshot.</p><p>There is a further generated-file inconsistency in <code>_load_local_data()</code>: it passes <code>asset_name=args.asset</code>, while the generated <code>load_price_csv()</code> function accepts <code>asset</code>, not <code>asset_name</code>. Depending on the compatibility wrapper, this can fail or cause the loader to retain its default asset label. The keyword must be aligned before asset-specific provenance can be trusted.</p><h3>17.4 Synthetic demonstration</h3><p>The intended synthetic command is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;9fdaf686-9fc4-4276-808a-53faf0b86831&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">quant-trading-model demo \
  --rows 500 \
  --asset gold \
  --seed 7 \
  --output results/demo.json</code></pre></div><p><code>run_demo()</code> calls <code>generate_synthetic_prices()</code> and records a compact summary. Synthetic data are useful for teaching timestamp handling, window construction, tensor shapes, finite sizing, VaR decisions, and portfolio accounting. They are not historical gold or Bitcoin data and cannot reproduce the paper's reported outcomes.</p><p>The optional <code>--run-experiment</code> flag is currently not a complete path. The generator returns a <code>PriceData</code> object, and <code>run_demo()</code> passes that object to <code>run_forecast_experiment()</code>. In the generated <code>experiments.py</code>, <code>_price_array()</code> handles a pandas <code>Series</code>, a pandas <code>DataFrame</code>, or an array-like object, but does not handle <code>PriceData</code>. The experiment can therefore fail before window construction. The repair must either teach <code>_price_array()</code> to extract <code>PriceData.frame</code> or make the CLI pass a compatible DataFrame or Series.</p><p>This same <code>PriceData</code> mismatch affects the <code>forecast</code> handler, because <code>_load_local_data()</code> also returns <code>PriceData</code> and passes it to <code>run_forecast_experiment()</code>. These commands are therefore documented as intended interfaces until that data contract is repaired.</p><h3>17.5 Validate a local CSV</h3><p>The generated parser uses a positional CSV argument. A statically accurate example is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;845cb841-4c7c-4098-bcf3-c98a78b15496&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">quant-trading-model validate-data data/gold.csv \
  --timestamp-column timestamp \
  --price-column close \
  --asset gold \
  --output results/gold_data_summary.json</code></pre></div><p><code>run_validate_data()</code> is intended to:</p><p>check the local path;</p><p>call <code>load_price_csv()</code>;</p><p>summarize the resulting object; and</p><p>print or write local JSON.</p><p>The loader requires parseable timestamps and finite, positive prices, and it sorts and deduplicates observations. The summary should be saved with a provenance record rather than treated as evidence that the data are the same data used by the paper.</p><p>Some older README examples use <code>--input data/gold.csv</code>, but the generated parser defines <code>csv</code> positionally and does not define an <code>--input</code> option. The positional form above matches the generated CLI more closely.</p><h3>17.6 Forecast command: intended interface and current limitation</h3><p>The intended command is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;2250484b-b9e0-4476-a7d3-796abf07961b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">quant-trading-model forecast data/gold.csv \
  --timestamp-column timestamp \
  --price-column close \
  --asset gold \
  --output results/gold_forecast.json</code></pre></div><p>Optional benchmarks are requested with:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;e8507ab8-743d-4af2-94f5-1c8c1e5be647&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">quant-trading-model forecast data/gold.csv \
  --timestamp-column timestamp \
  --price-column close \
  --asset gold \
  --benchmarks \
  --output results/gold_forecast_with_benchmarks.json</code></pre></div><p>The intended data flow is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;516b58ec-034c-462b-997d-9228fe98245a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">local CSV
  -&gt; local loader
  -&gt; standardized prices
  -&gt; 100-observation windows
  -&gt; chronological split
  -&gt; Att-BiLSTM or selected forecaster
  -&gt; forecast metrics
  -&gt; local result artifact</code></pre></div><p>The current generated path is not yet a verified runtime workflow. The CLI passes <code>PriceData</code> into <code>run_forecast_experiment()</code>, while <code>_price_array()</code> in <code>experiments.py</code> does not extract <code>PriceData</code>. This must be repaired before the command can reliably reach the forecasting model.</p><p>Even after that data mismatch is fixed, a missing PyTorch installation may escape through <code>main()</code> as an uncaught <code>ImportError</code>. The CLI should catch that exception and convert it into an actionable message. Until then, the optional-dependency behavior is a model-layer intention, not a guaranteed CLI behavior.</p><p>A repaired forecast artifact should include the asset, model, window length, horizon, split protocol, seed, preprocessing decisions, forecast metrics, package versions, and warnings. Forecast metrics remain separate from portfolio metrics and do not measure fees, turnover, drawdown, or liquidation.</p><h3>17.7 Backtest command: distinguish forecasting from portfolio simulation</h3><p>The intended command is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;c8c0b2d7-0e96-4a60-8610-a249f99419c8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">quant-trading-model backtest data/bitcoin.csv \
  --timestamp-column timestamp \
  --price-column close \
  --asset bitcoin \
  --output results/bitcoin_backtest.json</code></pre></div><p>In the generated CLI, <code>run_backtest_command()</code> calls <code>run_walk_forward_experiment()</code>. That function is intended to evaluate sequential forecasts; it does not, by itself, compose forecast signals, exponential sizing, VaR, greedy action selection, fills, fees, and portfolio valuation.</p><p>The repository does contain separate strategy and accounting modules. A complete portfolio backtest must explicitly connect them:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;5a3df806-a2a9-4bb0-856b-41f303c6f63b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">walk-forward forecast
  -&gt; signal generation
  -&gt; finite sizing
  -&gt; VaR filter
  -&gt; greedy action selection
  -&gt; next-bar order
  -&gt; portfolio simulator</code></pre></div><p>Therefore, the current <code>backtest</code> command is a statically described walk-forward entry point, not proof that a complete paper trading strategy is running. Its name should not be interpreted as evidence that the paper's missing execution protocol has been recovered.</p><h3>17.8 Fee-sensitivity command: parser exists, callback wiring is incomplete</h3><p>The paper-inspired fee grids are:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!7022!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5b9cc7-4a51-4275-a7df-0bcfb0f338ef_1096x224.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!7022!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5b9cc7-4a51-4275-a7df-0bcfb0f338ef_1096x224.png 424w, /__u/substackcdn.com/image/fetch/$s_!7022!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5b9cc7-4a51-4275-a7df-0bcfb0f338ef_1096x224.png 848w, /__u/substackcdn.com/image/fetch/$s_!7022!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5b9cc7-4a51-4275-a7df-0bcfb0f338ef_1096x224.png 1272w, /__u/substackcdn.com/image/fetch/$s_!7022!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5b9cc7-4a51-4275-a7df-0bcfb0f338ef_1096x224.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!7022!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5b9cc7-4a51-4275-a7df-0bcfb0f338ef_1096x224.png" width="1096" height="224" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4e5b9cc7-4a51-4275-a7df-0bcfb0f338ef_1096x224.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:224,&quot;width&quot;:1096,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:37524,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5b9cc7-4a51-4275-a7df-0bcfb0f338ef_1096x224.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!7022!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5b9cc7-4a51-4275-a7df-0bcfb0f338ef_1096x224.png 424w, /__u/substackcdn.com/image/fetch/$s_!7022!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5b9cc7-4a51-4275-a7df-0bcfb0f338ef_1096x224.png 848w, /__u/substackcdn.com/image/fetch/$s_!7022!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5b9cc7-4a51-4275-a7df-0bcfb0f338ef_1096x224.png 1272w, /__u/substackcdn.com/image/fetch/$s_!7022!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5b9cc7-4a51-4275-a7df-0bcfb0f338ef_1096x224.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>The intended command is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;b7c2d34c-a93a-4742-a9bf-d6fac483c694&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">quant-trading-model fee-sensitivity data/gold.csv \
  --timestamp-column timestamp \
  --price-column close \
  --asset gold \
  --output results/gold_fee_sensitivity.json</code></pre></div><p>The command is currently only a parser-level interface. <code>run_fee_sensitivity()</code> in <code>experiments.py</code> requires a positional <code>run_backtest</code> callback, followed by the asset and fee configuration. However, <code>run_fee_sensitivity_command()</code> calls it without supplying that required callback. As written, the handler cannot successfully invoke the fee experiment.</p><p>A repair must construct or inject a callback that accepts <code>fee_rate</code>, runs the same strategy and data under that fee, and returns portfolio results. Only then can the CLI perform a sensitivity sweep. The callback should hold all non-fee assumptions fixed: data, dates, forecasts, signal rules, sizing, slippage, initial cash, and liquidation. If fees change feasibility or greedy ranking, decisions must be regenerated and that dependency recorded.</p><p>The missing numeric values in the paper's fee figures must not be inferred from the fee grid itself.</p><h3>17.9 Optional dependency installation</h3><p>The intended installation commands are:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;bash&quot;,&quot;nodeId&quot;:&quot;399e026e-839d-4252-8720-044aae7306ba&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-bash">python -m pip install -e .
python -m pip install -e '.[deep-learning]'
python -m pip install -e '.[classical]'</code></pre></div><p>The generated <code>pyproject.toml</code> defines the <code>classical</code> extra. Some README text refers to a <code>benchmarks</code> extra, which is inconsistent; <code>classical</code> is the authoritative generated metadata name.</p><p>A missing optional package should ideally produce a message such as:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;ea9f3dc5-5311-4f28-8b0e-79ac7db234eb&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">PyTorch is required for recurrent forecasters. Install the optional deep-learning dependencies for this project.</code></pre></div><p>At present, the model layer can raise that message, but the CLI's <code>main()</code> does not catch <code>ImportError</code>. Repairing the exception handling is necessary for the message to be consistently presented as a normal command-line error.</p><h3>17.10 Archive results with provenance</h3><p>The CLI's output helper is intended to print JSON or write it to a local file:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;15e4c981-a435-4456-856f-c50f81a3bd60&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">def _write_output(value: Any, output: str | None) -&gt; None:
    rendered = _dump(value)
    if output is None:
        print(rendered)
        return
    path = _safe_output_path(output)
    path.parent.mkdir(parents=True, exist_ok=True)
    path.write_text(rendered + "\n", encoding="utf-8")
    print(f"Wrote {path}")</code></pre></div><p>A useful research directory might contain:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;3e2d5614-b063-40e6-a7ce-128318e9756b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">results/
&#9500;&#9472;&#9472; gold_data_summary.json
&#9500;&#9472;&#9472; gold_forecast.json
&#9500;&#9472;&#9472; gold_fee_sensitivity.json
&#9500;&#9472;&#9472; config_snapshot.json
&#9500;&#9472;&#9472; package_versions.txt
&#9492;&#9472;&#9472; provenance.md</code></pre></div><p>The configuration snapshot should include the window length, target horizon, scaler policy, model settings, fee rate, slippage, VaR settings, sizing schedule, signal thresholds, execution timing, and liquidation policy. The provenance file should identify the local data file and instrument metadata.</p><p>The paper's approximate <code>$646</code> gold value and <code>$215,487</code> Bitcoin value remain unverified claims. A local JSON file without dates, data provenance, and execution assumptions cannot validate either number.</p><h3>17.11 Offline workflow checklist</h3><p>The intended research workflow is:</p><p>Install only the dependencies required for the selected module.</p><p>Inspect <code>quant-trading-model --help</code> and the relevant subcommand help.</p><p>Use <code>demo</code> to inspect synthetic data flow, without treating it as empirical evidence.</p><p>Use <code>validate-data</code> on each local asset file.</p><p>Save the data summary and provenance metadata.</p><p>Run a chronological forecast experiment after repairing the <code>PriceData</code> handoff.</p><p>Record model status, target horizon, metrics, seed, and configuration.</p><p>Run walk-forward evaluation for trading-oriented conclusions.</p><p>Connect forecasts explicitly to signals, sizing, VaR, greedy selection, and next-bar execution for a portfolio backtest.</p><p>Supply a valid backtest callback before using fee sensitivity.</p><p>Compare portfolio results with buy-and-hold and report drawdown, volatility, turnover, and fees.</p><p>Label every result as implemented, reconstructed, illustrative, unavailable, or reported claim as appropriate.</p><p>The generated CLI and orchestration files also have known static-review findings: placeholder detection was reported for <code>cli.py</code>, and semantic code verification was skipped. Some generated tests and orchestration calls are API-inconsistent with the generated modules. These findings reinforce the correct interpretation of this section: it documents intended offline interfaces and their current repair requirements; it does not claim that any command ran successfully.</p><p>The central lesson is that a CLI can make a research workflow repeatable in form, but it cannot supply missing data or algorithmic details. Reliable interpretation still requires local data provenance, explicit configuration, causal signal timing, repaired module interfaces, and a strict distinction between computed results and the paper's unverified claims.</p><h2>18. Verification: Static, Semantic, and Invariant Checks</h2><p>A paper-to-code project needs more than a collection of source files. It also needs a disciplined way to check whether the implementation is structurally coherent, whether its assumptions are visible, and whether its financial logic avoids obvious look-ahead errors. At the same time, verification must not be overstated. A file that parses successfully is not necessarily semantically correct, and a passing unit test would not prove that the paper's reported profits are reproducible.</p><p>For this project, verification has three layers:</p><p><strong>Static verification</strong> checks source structure without running the research pipeline.</p><p><strong>Semantic review</strong> checks whether the implementation matches the intended mathematics and data flow.</p><p><strong>Empirical validation</strong> would compare executed results with the paper, but it is currently unavailable because the original data and protocol are missing.</p><p>The supplied local verification performed the first layer and recorded a partial result. Semantic code verification was explicitly skipped. Therefore, the discussion below describes the intended checks and the reported findings; it does not claim that the generated code was executed or that the test suite passed.</p><h3>18.1 Static verification: checking structure without execution</h3><p>Static checks inspect files as text or syntax trees. They are useful because they can identify malformed Python, missing symbols, accidental placeholders, unsafe patterns, and documentation inconsistencies before a runtime experiment is attempted.</p><p>The local verification report applied checks such as:</p><ul><li><p>non-empty file validation;</p></li><li><p>detection of Markdown fences in generated code fields;</p></li><li><p>placeholder and TODO detection;</p></li><li><p>checks for live-trading patterns;</p></li><li><p>generic text checks;</p></li><li><p>Python AST parsing for Python files.</p></li></ul><p>AST parsing is stronger than looking for balanced parentheses with a text search. It asks Python's parser whether a file has valid syntax and produces a structured representation of imports, classes, functions, and statements. It does not, however, prove that imported names exist at runtime or that two modules agree on an API.</p><p>The static checks reported successful parsing and structural checks for the core modules, including:</p><ul><li><p><code>data.py</code>, which constructs windows and chronological partitions;</p></li><li><p><code>metrics.py</code>, which separates forecast and portfolio measures;</p></li><li><p><code>models/attention.py</code> and <code>models/forecasters.py</code>, which define the tensor-facing model interfaces;</p></li><li><p><code>streaks.py</code>, <code>sizing.py</code>, and <code>risk.py</code>, which implement the reconstructed analytical components;</p></li><li><p><code>strategy.py</code> and <code>backtest.py</code>, which define candidate actions and portfolio accounting; and</p></li><li><p>the generated test files.</p></li></ul><p>That result gives useful confidence that these files are non-empty and syntactically parseable. It does not establish that the modules can be imported together, that optional dependencies are available, or that the implementation behaves correctly for real data.</p><h3>18.2 Reported static-check findings</h3><p>The overall local static-verification result was <strong>not passing</strong>. Four generated files were reported with issues:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!dQY9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb85f769-b559-4068-9d43-58974c21fff8_1402x912.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!dQY9!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb85f769-b559-4068-9d43-58974c21fff8_1402x912.png 424w, /__u/substackcdn.com/image/fetch/$s_!dQY9!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb85f769-b559-4068-9d43-58974c21fff8_1402x912.png 848w, /__u/substackcdn.com/image/fetch/$s_!dQY9!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb85f769-b559-4068-9d43-58974c21fff8_1402x912.png 1272w, /__u/substackcdn.com/image/fetch/$s_!dQY9!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb85f769-b559-4068-9d43-58974c21fff8_1402x912.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!dQY9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb85f769-b559-4068-9d43-58974c21fff8_1402x912.png" width="1402" height="912" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/db85f769-b559-4068-9d43-58974c21fff8_1402x912.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:912,&quot;width&quot;:1402,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:207931,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb85f769-b559-4068-9d43-58974c21fff8_1402x912.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!dQY9!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb85f769-b559-4068-9d43-58974c21fff8_1402x912.png 424w, /__u/substackcdn.com/image/fetch/$s_!dQY9!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb85f769-b559-4068-9d43-58974c21fff8_1402x912.png 848w, /__u/substackcdn.com/image/fetch/$s_!dQY9!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb85f769-b559-4068-9d43-58974c21fff8_1402x912.png 1272w, /__u/substackcdn.com/image/fetch/$s_!dQY9!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb85f769-b559-4068-9d43-58974c21fff8_1402x912.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The Markdown findings are format warnings in the generated-file representation. A Markdown tutorial normally uses fenced code blocks, so removing every fence would make the document less useful to readers. The important distinction is between a validation rule for a generated <code>code</code> field and the normal syntax of a Markdown document.</p><p>The placeholder findings are more substantive. They should be reviewed manually because a placeholder can mean anything from an intentionally unfinished branch to an incomplete implementation. Until that review and repair occur, the experiment and CLI layers should be described as planned or partially integrated artifacts rather than as verified end-to-end interfaces.</p><h3>18.3 Semantic review: checking meaning, not just syntax</h3><p>Semantic verification asks whether the code's operations match the intended method. For this project, the most important semantic questions are temporal, mathematical, and accounting-related.</p><h4>Temporal and data invariants</h4><p>The windowing and data tests are intended to review the following properties:</p><ul><li><p>timestamps are strictly increasing after cleaning;</p></li><li><p>prices are finite and positive;</p></li><li><p>a target timestamp occurs after the final timestamp in its input window;</p></li><li><p>chronological training and test partitions remain ordered;</p></li><li><p>scaling parameters are fitted only on training observations; and</p></li><li><p>no forecast uses a target or realized return from the future.</p></li></ul><p>These checks matter because a time-series implementation can be syntactically valid while still leaking information. For example, fitting a scaler on the complete series or randomly mixing overlapping windows can make a forecast evaluation look stronger than a deployment-like evaluation.</p><p>The test file <code>tests/test_data_and_streaks.py</code> is designed around small hand-calculated paths. That is a good verification strategy for boundaries: a short sequence makes it possible to calculate the expected target values and return signs manually. It is more informative for temporal invariants than a large opaque fixture.</p><h4>Attention and model-shape invariants</h4><p>The model tests are intended to review tensor contracts rather than predictive quality. The key conditions are:</p><ul><li><p>recurrent inputs have shape <code>(batch, time, features)</code>;</p></li><li><p>attention scores have shape <code>(batch, time)</code>;</p></li><li><p>attention weights are nonnegative;</p></li><li><p>weights sum to one across the time dimension;</p></li><li><p>the context vector has shape <code>(batch, features)</code>;</p></li><li><p>one-step and three-step heads return the configured horizon; and</p></li><li><p>the primary model contains the configured 20% dropout layer.</p></li></ul><p>The attention normalization check is particularly important. Applying softmax across features instead of time would produce a different algorithm from the paper's temporal attention equations. Shape tests can catch that kind of error without claiming that the model has learned a useful forecast.</p><p>The model tests also need PyTorch-aware handling. PyTorch is optional in this project, so a minimal installation should skip tensor-execution tests cleanly rather than fail while importing unrelated data or accounting modules. A skipped optional test is not a passing model validation; it means that the environment did not perform that check.</p><h4>Streak and return invariants</h4><p>The streak tests review the reconstruction of the paper's consecutive-rise and decline analysis:</p><ul><li><p>the default simple return uses the previous price as denominator;</p></li><li><p>zero returns do not silently become rises or declines;</p></li><li><p>maximal runs and overlapping subsequences produce different, intentional counts;</p></li><li><p>streak counts are nonnegative integers; and</p></li><li><p>percentile calculations use only the supplied historical segment.</p></li></ul><p>These checks are important because the paper's use of the term &#8220;Apriori&#8221; is not operationally supported by frequent-itemset support or confidence calculations. The implementation should therefore be reviewed as streak analysis, with its chosen return and zero policies recorded explicitly.</p><h4>Position-sizing invariants</h4><p>The sizing tests should confirm that the normalized exponential reconstruction remains bounded:</p><ul><li><p>every score and allocation is finite and nonnegative;</p></li><li><p>normalized weights sum to one within tolerance;</p></li><li><p>the planned allocations do not exceed the schedule budget;</p></li><li><p>an addition level cannot be consumed twice accidentally;</p></li><li><p>the approved amount cannot exceed available cash; and</p></li><li><p>the exposure cap prevents unlimited averaging down.</p></li></ul><p>These are safety and consistency properties of the reconstruction. They do not prove that the paper's corrupted equations (5)&#8211;(8) were recovered correctly. In fact, exact verification of those equations remains unavailable until the original notation or equation images are supplied.</p><h4>VaR and strategy invariants</h4><p>The risk and strategy layers require a different kind of semantic review. The intended checks include:</p><ul><li><p>VaR uses only returns available before the decision timestamp;</p></li><li><p>insufficient history is represented explicitly;</p></li><li><p>a blocked trade receives zero approved quantity;</p></li><li><p>risk scaling never increases the proposed order;</p></li><li><p>candidate scores use forecasts, current portfolio state, fees, and constraints only;</p></li><li><p>infeasible actions are excluded before ranking; and</p></li><li><p>greedy tie-breaking is deterministic.</p></li></ul><p>These conditions prevent a risk filter from becoming decorative. A VaR calculation that uses future returns, or a strategy score that accidentally includes realized profit, would invalidate the backtest even if every function returns a number.</p><p>The paper does not define its VaR parameters or greedy objective. Consequently, semantic verification can check the internal consistency of the chosen reconstruction, but it cannot check fidelity to an unavailable original algorithm.</p><h4>Portfolio-accounting invariants</h4><p>The backtest tests are designed around hand-computable transactions. The central identity is:</p><p>\[ V<em>t = \operatorname{cash}</em>t + \operatorname{quantity}<em>t p</em>t. \]</p><p>The intended accounting checks include:</p><ul><li><p>every fill occurs after its signal timestamp;</p></li><li><p>buy and sell fees are applied exactly once to executed notional;</p></li><li><p>unaffordable buys do not create negative cash under no-leverage settings;</p></li><li><p>holdings do not become negative when shorting is disabled;</p></li><li><p>marked portfolio value equals cash plus marked holdings;</p></li><li><p>liquidation is explicit; and</p></li><li><p>the buy-and-hold benchmark follows a documented fee and timing convention.</p></li></ul><p>These tests verify the simulator's stated accounting contract. They do not verify the paper's accounting because the paper does not specify its execution dates, fee-side convention, slippage, partial-unit rules, or liquidation policy.</p><h3>18.4 Why the skipped semantic verification matters</h3><p>The semantic verification report states that code semantic verification was skipped using the <code>--skip-code-semantic-verification</code> option. Its overall assessment was therefore &#8220;not semantically reviewed.&#8221; This is an important limitation, not a minor footnote.</p><p>Static AST parsing can show that Python syntax is valid. It cannot reliably detect every issue that a semantic review would investigate, such as:</p><ul><li><p>a caller passing <code>split_config</code> to a function that accepts only <code>config</code>;</p></li><li><p>a test constructing a class with fields that differ from the generated class;</p></li><li><p>a benchmark adapter assuming a metric method that another module does not expose;</p></li><li><p>a forecast output using a different shape from the target array; or</p></li><li><p>an experiment function describing end-to-end backtesting while only orchestrating forecasts.</p></li></ul><p>The supplied project notes also warn that some generated tests and orchestration code appear API-inconsistent with the generated modules. That means the tests should be treated as intended invariant specifications and static artifacts until their interfaces are reconciled. No claim should be made that the test suite passes.</p><h3>18.5 Verification is not execution and execution is not reproduction</h3><p>There are three separate statements that should not be conflated:</p><p><strong>The source parses.</strong> Python's parser accepted the file structure.</p><p><strong>The implementation is semantically coherent.</strong> Functions, types, formulas, and timestamps agree with one another under review.</p><p><strong>The paper's experiment is reproduced.</strong> The implementation matches the original data, protocol, and reported results.</p><p>The available evidence supports only parts of the first statement. The second was not completed by the supplied semantic verifier. The third is impossible at present because the paper's data and several algorithmic details are missing.</p><p>In particular, this tutorial does not claim that:</p><ul><li><p>the generated code was executed;</p></li><li><p>all dependencies were installed;</p></li><li><p>the tests passed;</p></li><li><p>the CLI completed a forecast or backtest;</p></li><li><p>the reported <code>$646</code> gold value or <code>$215,487</code> Bitcoin value was reproduced; or</p></li><li><p>the paper's benchmark metrics or fee-sensitivity curves were confirmed.</p></li></ul><p>The correct interpretation is narrower: the project contains intended implementations, tests, and verification rules, while the local static report identifies both successful structural checks and unresolved findings.</p><h3>18.6 A practical review order</h3><p>When repairing or extending the project, review the layers in this order:</p><p><strong>Resolve format and placeholder findings.</strong> Inspect <code>experiments.py</code> and <code>cli.py</code>; decide whether each detected placeholder is intentional or incomplete. Treat the Markdown-fence findings as representation warnings where fences are normal Markdown syntax.</p><p><strong>Reconcile public APIs.</strong> Compare function signatures and dataclass fields across source modules and tests. In particular, check the experiment-to-data, strategy-to-risk, and test-to-sizer interfaces.</p><p><strong>Run syntax and import checks in the target environment.</strong> Confirm that core imports do not require optional PyTorch or benchmark packages, and that optional failures are explicit.</p><p><strong>Review data boundaries.</strong> Hand-check window endpoints, target timestamps, scaler fitting, and chronological splits.</p><p><strong>Review tensor shapes.</strong> Confirm attention normalization across time and forecast-target horizon agreement.</p><p><strong>Review pure analytical components.</strong> Check return denominators, streak policies, normalized sizing, and VaR sign conventions with small numerical examples.</p><p><strong>Review portfolio accounting.</strong> Reconcile every fill, fee, cash change, holding quantity, and mark-to-market value.</p><p><strong>Only then run local experiments.</strong> Preserve configuration, data provenance, software versions, warnings, and result metadata.</p><p><strong>Compare with the paper cautiously.</strong> Label any mismatch as unresolved unless the original data and protocol are available.</p><p>This order moves from inexpensive structural checks to more demanding semantic and empirical work. It also prevents a striking backtest number from distracting attention from basic timestamp or accounting errors.</p><h3>18.7 Reproduction status remains a separate question</h3><p>The reproduction matrix in <code>docs/reproduction_matrix.md</code> is the appropriate place to track paper fidelity. It distinguishes direct mechanics&#8212;such as temporal attention equations, window construction, and portfolio identities&#8212;from reconstructions such as normalized exponential sizing, historical VaR, and greedy action selection.</p><p>A useful final status summary is:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!BWID!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F669e00cd-16dc-4efb-a7de-5d2d8087fd0a_1412x684.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!BWID!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F669e00cd-16dc-4efb-a7de-5d2d8087fd0a_1412x684.png 424w, /__u/substackcdn.com/image/fetch/$s_!BWID!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F669e00cd-16dc-4efb-a7de-5d2d8087fd0a_1412x684.png 848w, /__u/substackcdn.com/image/fetch/$s_!BWID!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F669e00cd-16dc-4efb-a7de-5d2d8087fd0a_1412x684.png 1272w, /__u/substackcdn.com/image/fetch/$s_!BWID!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F669e00cd-16dc-4efb-a7de-5d2d8087fd0a_1412x684.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!BWID!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F669e00cd-16dc-4efb-a7de-5d2d8087fd0a_1412x684.png" width="1412" height="684" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/669e00cd-16dc-4efb-a7de-5d2d8087fd0a_1412x684.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:684,&quot;width&quot;:1412,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:134565,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F669e00cd-16dc-4efb-a7de-5d2d8087fd0a_1412x684.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!BWID!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F669e00cd-16dc-4efb-a7de-5d2d8087fd0a_1412x684.png 424w, /__u/substackcdn.com/image/fetch/$s_!BWID!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F669e00cd-16dc-4efb-a7de-5d2d8087fd0a_1412x684.png 848w, /__u/substackcdn.com/image/fetch/$s_!BWID!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F669e00cd-16dc-4efb-a7de-5d2d8087fd0a_1412x684.png 1272w, /__u/substackcdn.com/image/fetch/$s_!BWID!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F669e00cd-16dc-4efb-a7de-5d2d8087fd0a_1412x684.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The purpose of this verification process is therefore not to manufacture certainty. It is to make uncertainty inspectable: readers can see which files parse, which invariants are intended, which findings remain unresolved, and which paper claims cannot yet be tested. That is the appropriate standard for an educational implementation of an incompletely specified quantitative-trading paper.</p><h2>19. Worked Synthetic Example from Price Path to Portfolio Record</h2><p>This section connects the data, streak, sizing, risk, strategy, and accounting layers with a short synthetic price path:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;6daf763a-79a6-40a3-9583-6226e0b8ce9b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">prices: 100, 102, 101, 99, 100, 103</code></pre></div><p>The example follows this causal sequence:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;d370ae09-6fda-462f-a940-24f3f0006883&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">prices
  &#8594; standard returns and streaks
  &#8594; manually supplied forecast
  &#8594; finite position-size schedule
  &#8594; VaR decision
  &#8594; greedy candidate selection
  &#8594; next-bar fill
  &#8594; portfolio valuation</code></pre></div><p>This is an <strong>illustrative interface example</strong>, not an empirical backtest. The prices are synthetic and the forecast is supplied manually rather than produced by a trained model. The paper's exact position-sizing equations, VaR procedure, and greedy policy are unavailable, so those components remain documented reconstructions. The code was not executed here; the calculations below describe the intended contracts and hand-checkable accounting.</p><h3>19.1 Prepare the synthetic prices</h3><p>The data layer accepts local historical data. For a small example, create a pandas <code>Series</code> and convert it into the standard timestamp-and-price table:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7b0df02e-0e8c-4fdb-99dd-06e03f008585&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import pandas as pd

prices = pd.Series(
    [100.0, 102.0, 101.0, 99.0, 100.0, 103.0],
    index=pd.date_range("2024-01-01", periods=6, freq="D"),
    name="price",
)

from quantitative_trading_model.data import standardize_prices

clean = standardize_prices(
    prices.rename_axis("timestamp").reset_index(),
    timestamp_column="timestamp",
    price_column="price",
    asset="synthetic",
)</code></pre></div><p><code>standardize_prices</code> parses timestamps, sorts observations, removes invalid rows when configured to do so, resolves duplicate timestamps, and returns a <code>PriceData</code> object. Prices must be finite and strictly positive. The function does not identify a market-data source or make network requests.</p><p>Six observations are far shorter than the paper-inspired 100-observation forecasting window. This path is therefore unsuitable for training the Att-BiLSTM. It is only long enough to demonstrate return alignment, decision timing, and portfolio accounting.</p><h3>19.2 Compute standard returns</h3><p>The default reconstruction uses the conventional simple return:</p><p>\[ r<em>t = \frac{p</em>t-p<em>{t-1}}{p</em>{t-1}}. \]</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;337dcd80-1cfd-4dba-b8cd-e6e982d9797c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from quantitative_trading_model.streaks import compute_returns

returns = compute_returns(clean.prices, denominator="previous")</code></pre></div><p>The hand calculations are:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!uC99!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc247d11c-ac66-460d-b842-32ff8bd25123_558x412.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!uC99!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc247d11c-ac66-460d-b842-32ff8bd25123_558x412.png 424w, /__u/substackcdn.com/image/fetch/$s_!uC99!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc247d11c-ac66-460d-b842-32ff8bd25123_558x412.png 848w, /__u/substackcdn.com/image/fetch/$s_!uC99!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc247d11c-ac66-460d-b842-32ff8bd25123_558x412.png 1272w, /__u/substackcdn.com/image/fetch/$s_!uC99!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc247d11c-ac66-460d-b842-32ff8bd25123_558x412.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!uC99!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc247d11c-ac66-460d-b842-32ff8bd25123_558x412.png" width="558" height="412" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c247d11c-ac66-460d-b842-32ff8bd25123_558x412.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:412,&quot;width&quot;:558,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:40950,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc247d11c-ac66-460d-b842-32ff8bd25123_558x412.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!uC99!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc247d11c-ac66-460d-b842-32ff8bd25123_558x412.png 424w, /__u/substackcdn.com/image/fetch/$s_!uC99!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc247d11c-ac66-460d-b842-32ff8bd25123_558x412.png 848w, /__u/substackcdn.com/image/fetch/$s_!uC99!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc247d11c-ac66-460d-b842-32ff8bd25123_558x412.png 1272w, /__u/substackcdn.com/image/fetch/$s_!uC99!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc247d11c-ac66-460d-b842-32ff8bd25123_558x412.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Each return is aligned to the later price. Thus, the return from 101 to 99 becomes known at the 99-price observation. It must not be assigned to the earlier 101 timestamp when constructing an online risk history.</p><p>The extracted paper equation appears to use the current price as the denominator. That alternative is available through <code>denominator="current"</code> or <code>denominator="paper"</code>, but one convention must be used consistently throughout a given experiment.</p><h3>19.3 Classify movements and identify streaks</h3><p>The sign sequence is:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;0248ac5f-6072-457e-8c0c-2154ed0d985a&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">+  -  -  +  +</code></pre></div><p>The streak layer separates calculation, classification, and counting:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4bf63827-3689-4505-9925-d29af15c2b5f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from quantitative_trading_model.streaks import classify_returns, count_streaks

labels = classify_returns(returns, zero_policy="neutral")

rise_counts = count_streaks(
    labels,
    value=1,
    mode="maximal",
    minimum_length=1,
)
decline_counts = count_streaks(
    labels,
    value=-1,
    mode="maximal",
    minimum_length=1,
)</code></pre></div><p>The maximal runs are:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;aa252bbe-72e6-4f95-99c6-4cc1ffa1d443&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">positive run of length 1
negative run of length 2
positive run of length 2</code></pre></div><p>Conceptually, the results are:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;81e52354-a740-410a-aeb6-f694384bbd05&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">rise_streaks:    {1: 1, 2: 1}
decline_streaks: {2: 1}</code></pre></div><p><code>count_streaks</code> also supports <code>mode="overlapping"</code>. A run of four positive returns contains three overlapping two-return subsequences, whereas maximal mode records one complete run of length four. These statistics are different, and the paper does not specify which one it used.</p><p>A complete summary can be requested with:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;67f618d7-a202-4996-82e3-efd2e67b60e3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from quantitative_trading_model.streaks import compute_streak_statistics

stats = compute_streak_statistics(
    clean.prices,
    denominator="previous",
    zero_policy="neutral",
    streak_mode="maximal",
)

print(stats.gain_percentile_90)
print(stats.decline_percentile_10)
print(stats.rise_streaks)
print(stats.decline_streaks)</code></pre></div><p>The summary exposes the 90th-percentile gain, signed 10th-percentile decline, median gain, and decline magnitudes. With only five returns, these estimates are not statistically meaningful. In a real walk-forward experiment, the statistics must be calculated from observations available before the relevant decision.</p><p>This implementation calls the procedure <strong>streak analysis</strong> or <strong>sequential-pattern analysis</strong>, not standard Apriori. The paper does not provide Apriori itemsets, support, or confidence calculations.</p><h3>19.4 Build and consume a finite sizing level</h3><p>The paper's equations (5)&#8211;(8) are incomplete in the extracted source. The implementation therefore uses a normalized exponential reconstruction. Suppose the schedule has a <code>$300</code> budget and three addition levels:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ccc6f3d0-da63-4695-969d-21bdd1eccade&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from quantitative_trading_model.sizing import (
    ExponentialPositionSizer,
    build_exponential_schedule,
)

schedule = build_exponential_schedule(
    budget=300.0,
    max_additions=3,
    growth=0.50,
    amplitude=1.0,
)
sizer = ExponentialPositionSizer(schedule=schedule)

print(schedule.weights)
print(schedule.allocations)</code></pre></div><p>The reconstructed scores have the form:</p><p>si=Aexp&#8289;(Bi),</p><p>followed by normalization:</p><p>\[ w<em>i=\frac{s</em>i}{\sum<em>j s</em>j}, \qquad \operatorname{allocation}<em>i=300w</em>i. \]</p><p>For indices <code>i = 0, 1, 2</code>, <code>A = 1</code>, and <code>B = 0.5</code>, the scores are approximately:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;f428c096-70f3-4beb-86aa-6997efe77fce&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">1.0000, 1.6487, 2.7183</code></pre></div><p>The resulting planned allocations are approximately:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;0e17515d-23e7-4cba-b0b2-cdcaaea1841b&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">level 1: $55.89
level 2: $92.16
level 3: $151.95</code></pre></div><p>The first schedule level is not merely described as selected: it is connected to the strategy through <code>ExponentialPositionSizer</code>. The strategy uses one-based addition levels in its public action record, while the schedule stores zero-based Python indices internally. After the action is approved, the selected level is consumed exactly once:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ec987b11-fc28-4ebd-a72a-4a9f9dac779f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">allocation = sizer.allocate_next_addition(
    available_cash=500.0,
    current_exposure=0.0,
)

print(allocation.level)          # first internal level is 0
print(allocation.approved_amount)</code></pre></div><p>The allocation cannot exceed available cash or the configured exposure capacity. This is a finite schedule, not unlimited averaging down. The formula and any calibration from gains, declines, or fees are reconstructions rather than faithful transcriptions of the paper's unreadable recursive break-even equations.</p><h3>19.5 Produce a forecast-derived signal</h3><p>Assume that when the observed price is 99, a forecasting model predicts a one-step price of 101. The forecast is manually supplied for this example:</p><p>rforecast=10199&#8722;1&#8776;0.020202.</p><p>That is an expected increase of about 2.02%.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;369a2173-c4ea-43b9-a1bd-31d80b57327f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from quantitative_trading_model.strategy import SignalGenerator

signal_generator = SignalGenerator(
    buy_threshold=0.01,
    sell_threshold=0.01,
    aggregation="terminal",
)

decision = signal_generator.decide(
    current_price=99.0,
    forecasts=[101.0],
)

print(decision.signal)
print(decision.expected_return)</code></pre></div><p>The expected increase exceeds the 1% buy threshold, so the intended signal is <code>Signal.BUY</code>. <code>SignalGenerator.decide</code> uses only the current price and forecast. It does not receive the later realized prices of 100 or 103.</p><p>For a three-day forecast, the same class can aggregate values such as <code>[100.0, 101.0, 102.0]</code>. The default <code>terminal</code> policy uses the final forecast; <code>mean</code> and <code>max</code> are also available. The paper mentions three-day forecasts but does not define this aggregation rule, so it remains an explicit reconstruction choice.</p><h3>19.6 Apply historical VaR before ranking actions</h3><p>The paper names VaR but does not specify its confidence level, horizon, lookback, or action rule. This example uses a rolling historical estimator with enough pre-decision observations to produce an estimate:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4d5436f2-074c-4ed0-8911-c040fbd74a67&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from quantitative_trading_model.config import VaRConfig
from quantitative_trading_model.risk import HistoricalVaR, VaRTradeFilter

var_config = VaRConfig(
    enabled=True,
    confidence=0.95,
    horizon=1,
    lookback=100,
    min_observations=3,
    action="/__u/onepagecode.substack.com/scale",
    insufficient_history="allow",
    max_var_fraction=0.05,
)

var_estimator = HistoricalVaR(var_config)
var_filter = VaRTradeFilter(config=var_config)</code></pre></div><p>At the 99-price decision point, the returns known from the path are the first three values:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;ebd470b2-f2a2-4f21-8bdc-1d2f265b28b6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">pre_decision_returns = returns[:3]
estimate = var_estimator.estimate(pre_decision_returns)
print(estimate.var)
print(estimate.return_quantile)</code></pre></div><p>The estimator now has three observations, matching the configured minimum. This is still only a toy quantile estimate; it is not a meaningful risk model. If the lower-tail return quantile were <code>-1.8%</code>, an exposure of <code>$55.89</code> would imply an estimated loss of roughly:</p><p>55.89&#215;0.018&#8776;$1.01.</p><p>The actual estimate should be read from <code>estimate</code>, not inferred from this hypothetical number.</p><p>The risk decision uses currency notional units:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;01b7783b-c6e3-4c0e-b0e1-c0849e9eba96&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">risk_decision = var_filter.decide(
    proposed_quantity=55.89,
    current_exposure=0.0,
    portfolio_value=500.0,
    returns=pre_decision_returns,
)

print(risk_decision.action)
print(risk_decision.approved_quantity)
print(risk_decision.reason)</code></pre></div><p>Here, <code>proposed_quantity</code> is the VaR module's name for a proposed currency notional. It is not asset quantity. A positive approved amount means the risk filter accepted or scaled the proposed notional; a blocked trade has an approved amount of zero. The filter uses only the supplied pre-decision returns.</p><h3>19.7 Rank the forecast and sizing candidates</h3><p>The sizing schedule and risk decision are now connected explicitly. The risk-approved amount is passed as the maximum buy notional, while the sizer supplies the next addition level:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5699f371-5329-4d34-937e-ea38813908b0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from quantitative_trading_model.strategy import GreedyPolicy, PortfolioSnapshot

approved_notional = float(risk_decision.approved_quantity)

snapshot = PortfolioSnapshot(
    cash=500.0,
    holdings=0.0,
    price=99.0,
    addition_level=0,
    portfolio_value=500.0,
)

policy = GreedyPolicy(fee_rate=0.002)
selected = policy.choose(
    snapshot,
    forecasts=[101.0],
    signal_generator=signal_generator,
    sizer=sizer,
    max_additions=3,
    buy_notional=approved_notional,
    returns=pre_decision_returns,
)

print(selected.signal)
print(selected.notional)
print(selected.addition_level)</code></pre></div><p>This call does not pass <code>VaRTradeFilter</code> into <code>GreedyPolicy</code>. The VaR decision has already been applied directly with the compatible <code>VaRTradeFilter.decide</code> interface. That explicit separation avoids confusing a currency-notional risk result with an asset-unit order.</p><p>The candidate set conceptually contains:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;66dd9b9f-6d04-4a6c-87fb-ad0b91f829ea&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">hold:  score 0
buy:   risk-approved notional, using current cash
add:   next unused finite sizing level
sell:  unavailable because holdings are zero</code></pre></div><p>The reconstructed greedy score for a buy or add action is:</p><p>score(a)=N(rforecast&#8722;f),</p><p>where <code>N</code> is currency notional, <code>r_forecast</code> is the predicted return, and <code>f</code> is the fee rate. If <code>N = 55.89</code>, <code>r_forecast &#8776; 0.020202</code>, and <code>f = 0.002</code>, then:</p><p>score&#8776;55.89(0.020202&#8722;0.002)&#8776;1.02.</p><p>The hold candidate has score zero, so a positive approved buy or add candidate can outrank it. The actual selected amount must still be read from <code>selected.notional</code>; the paper does not define its own greedy objective or candidate set.</p><p>The <code>GreedyPolicy</code> call uses forecast information, current portfolio state, fees, and the already-approved risk amount. It does not use the later realized 99-to-100 return. That restriction prevents look-ahead bias.</p><h3>19.8 Consume the selected schedule level</h3><p>The strategy's selected addition level is an action description. The sizing schedule should consume its corresponding level only after the action has been accepted for execution:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;50e3ded3-bfb3-4d5e-9ea8-3629f3ef7ce4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">if selected.signal.value == "add" and selected.addition_level is not None:
    allocation = sizer.allocate_next_addition(
        available_cash=snapshot.cash,
        current_exposure=snapshot.holdings * snapshot.price,
        level=selected.addition_level - 1,
    )
    print(allocation.approved_amount)</code></pre></div><p>The subtraction converts the strategy's one-based public level to the sizer's zero-based schedule index. In a production orchestration layer, the approved amount returned by the sizer and the amount sent to the order should be compared explicitly. If risk scaling or cash caps reduce the amount, the order must use the reduced approved notional rather than the original planned allocation.</p><h3>19.9 Execute on the next bar and calculate portfolio value</h3><p>The signal is formed at the 99-price observation. The next bar has price 100, so the order executes at 100 rather than at the signal price of 99.</p><p>The example uses these units:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!8NNB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc89285b3-4e72-438a-b08f-105d0bf92120_1154x412.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!8NNB!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc89285b3-4e72-438a-b08f-105d0bf92120_1154x412.png 424w, /__u/substackcdn.com/image/fetch/$s_!8NNB!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc89285b3-4e72-438a-b08f-105d0bf92120_1154x412.png 848w, /__u/substackcdn.com/image/fetch/$s_!8NNB!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc89285b3-4e72-438a-b08f-105d0bf92120_1154x412.png 1272w, /__u/substackcdn.com/image/fetch/$s_!8NNB!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc89285b3-4e72-438a-b08f-105d0bf92120_1154x412.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!8NNB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc89285b3-4e72-438a-b08f-105d0bf92120_1154x412.png" width="1154" height="412" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c89285b3-4e72-438a-b08f-105d0bf92120_1154x412.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:412,&quot;width&quot;:1154,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:87048,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc89285b3-4e72-438a-b08f-105d0bf92120_1154x412.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!8NNB!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc89285b3-4e72-438a-b08f-105d0bf92120_1154x412.png 424w, /__u/substackcdn.com/image/fetch/$s_!8NNB!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc89285b3-4e72-438a-b08f-105d0bf92120_1154x412.png 848w, /__u/substackcdn.com/image/fetch/$s_!8NNB!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc89285b3-4e72-438a-b08f-105d0bf92120_1154x412.png 1272w, /__u/substackcdn.com/image/fetch/$s_!8NNB!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc89285b3-4e72-438a-b08f-105d0bf92120_1154x412.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Assume the approved notional is <code>$55.89</code>, execution price is <code>$100</code>, the fee is <code>0.002</code>, and slippage is zero. The hand calculation is:</p><p>q=55.89100=0.5589&nbsp;units,</p><p>fee=55.89&#215;0.002=$0.11178,</p><p>cash after buy=500&#8722;55.89&#8722;0.11178=$443.99822.</p><p>Immediately after the fill, the marked portfolio value is:</p><p>V=443.99822+0.5589&#215;100=$499.88822.</p><p>The difference from <code>$500</code> is the purchase fee.</p><p>The corresponding accounting objects are:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;80d96d2b-6278-459e-84ae-18774593e17c&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from quantitative_trading_model.backtest import Order, PortfolioSimulator

simulator = PortfolioSimulator(
    initial_cash=500.0,
    fee_rate=0.002,
    slippage_rate=0.0,
    allow_short=False,
    allow_leverage=False,
)

order = Order(
    side="buy",
    notional=55.89,
    signal_timestamp="2024-01-04",   # the bar priced at 99
    execution_timestamp="2024-01-05",
    addition_level=1,
    reason="forecast recovery signal",
)

fill = simulator.execute(order, "2024-01-05", 100.0)
record = simulator.mark_to_market("2024-01-05", 100.0)

print(fill.notional, fill.fee)
print(record.portfolio_value)</code></pre></div><p><code>Order.notional</code> is deliberately used here because the strategy and VaR layers work in currency-notional terms. The simulator converts that notional into asset units at the execution price. If an order instead supplies <code>Order.quantity</code>, the simulator treats it as asset units directly.</p><p><code>PortfolioSimulator.execute</code> applies configured slippage, calculates executed notional and fees, updates cash and holdings, and records a <code>Fill</code>. <code>mark_to_market</code> creates a <code>BacktestRecord</code> satisfying:</p><p>\[ V<em>t=\text{cash}</em>t+\text{quantity}<em>t p</em>t. \]</p><p>The default simulator is long-only and non-leveraged. It clips an unaffordable buy to the quantity affordable with current cash; the current generated execution path does not implement a separate rejection branch for the <code>reject_insufficient_cash</code> configuration field. That behavior should therefore not be described as configurable rejection without a corresponding backtest-code change.</p><p>If the price later reaches 103, the marked value before selling is approximately:</p><p>443.99822+0.5589&#215;103&#8776;$501.675.</p><p>A later sale would incur a second fee under the default two-sided fee convention. The final realized value would depend on the sale price, slippage, and whether end-of-test liquidation is enabled.</p><h3>19.10 What the example demonstrates</h3><p>The example makes the following contracts visible:</p><ul><li><p>returns are aligned with the later price observation;</p></li><li><p>streaks are calculated from the selected return convention;</p></li><li><p>the sizing schedule is finite, normalized, and connected to the action level;</p></li><li><p>VaR is applied before the action is ranked;</p></li><li><p>risk and strategy values are currency notionals, while backtest quantities are asset units;</p></li><li><p>the greedy score uses forecasts and estimated fees, not realized future returns;</p></li><li><p>the signal at price 99 executes at the later price of 100;</p></li><li><p>the purchase fee is deducted exactly once;</p></li><li><p>portfolio value equals cash plus marked holdings.</p></li></ul><p>It does <strong>not</strong> demonstrate profitability. The forecast is manual, the price path is synthetic, and the paper's exact sizing, VaR, and greedy methods are unavailable. The example is a causal and accounting demonstration, not a reproduction of the paper.</p><h3>19.11 Relation to the paper's reported outcomes</h3><p>The paper reports approximately <code>$646</code> from a <code>$500</code> gold allocation and approximately <code>$215,487</code> from a <code>$500</code> Bitcoin allocation. Neither result can be inferred from this six-price example. Reproducing either claim would require the original datasets, dates, target definition, complete position-sizing equations, VaR settings, greedy rules, fee treatment, execution timing, and liquidation policy.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!2MbP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F922cdd0b-bf8b-484c-a94f-b4a5c4bf509c_948x470.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!2MbP!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F922cdd0b-bf8b-484c-a94f-b4a5c4bf509c_948x470.png 424w, /__u/substackcdn.com/image/fetch/$s_!2MbP!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F922cdd0b-bf8b-484c-a94f-b4a5c4bf509c_948x470.png 848w, /__u/substackcdn.com/image/fetch/$s_!2MbP!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F922cdd0b-bf8b-484c-a94f-b4a5c4bf509c_948x470.png 1272w, /__u/substackcdn.com/image/fetch/$s_!2MbP!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F922cdd0b-bf8b-484c-a94f-b4a5c4bf509c_948x470.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!2MbP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F922cdd0b-bf8b-484c-a94f-b4a5c4bf509c_948x470.png" width="948" height="470" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/922cdd0b-bf8b-484c-a94f-b4a5c4bf509c_948x470.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:470,&quot;width&quot;:948,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:75219,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/206399283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F922cdd0b-bf8b-484c-a94f-b4a5c4bf509c_948x470.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!2MbP!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F922cdd0b-bf8b-484c-a94f-b4a5c4bf509c_948x470.png 424w, /__u/substackcdn.com/image/fetch/$s_!2MbP!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F922cdd0b-bf8b-484c-a94f-b4a5c4bf509c_948x470.png 848w, /__u/substackcdn.com/image/fetch/$s_!2MbP!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F922cdd0b-bf8b-484c-a94f-b4a5c4bf509c_948x470.png 1272w, /__u/substackcdn.com/image/fetch/$s_!2MbP!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F922cdd0b-bf8b-484c-a94f-b4a5c4bf509c_948x470.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The example is useful because it makes causality, units, and accounting explicit. It should not be used as evidence about expected returns, investment quality, or the validity of the paper's headline performance.</p><h2>20. Limitations, Interpretation, and What a True Reproduction Would Need</h2><p>The implementation now expresses the paper&#8217;s main ideas as separate, inspectable components: an attention-enhanced BiLSTM for forecasting, benchmark forecasters, streak analysis, a normalized exponential position-sizing reconstruction, historical VaR, a reconstructed greedy policy, portfolio accounting, and transaction-fee sensitivity experiments. That modular structure is useful, but it should not be mistaken for proof that the paper&#8217;s numerical results or trading claims have been reproduced.</p><p>The central limitation is not Python syntax. It is missing experimental information. A genuine reproduction requires the same observations, preprocessing, model settings, decision rules, execution assumptions, and evaluation dates. Several of those details are absent or corrupted in the source material. The generated project therefore distinguishes what is implemented mechanically from what is reconstructed or merely reported as a claim.</p><h3>20.1 Interpretation labels</h3><p>Use the following vocabulary when describing results from this project:</p><p>| Label | Meaning | | --- | --- | | <strong>Implemented</strong> | The paper describes enough behavior to encode a concrete mechanism, such as a sliding window or dot-product attention operation. | | <strong>Reconstructed</strong> | The paper names a method but omits equations, parameters, or operational rules, so the project supplies a documented interpretation. | | <strong>Illustrative</strong> | The result demonstrates an interface, invariant, or workflow using synthetic or user-provided local data. It is not evidence of paper reproduction. | | <strong>Unavailable</strong> | The paper does not provide enough information or data to validate the method or number. | | <strong>Reported claim</strong> | A value or conclusion stated by the paper and preserved for comparison, but not independently verified. |</p><p>This distinction applies across all major components:</p><ul><li><p><strong>Att-BiLSTM:</strong> the broad sequence-model and attention design is implemented, but hidden dimensions, layer count, query construction, target definition, and training schedule are reconstructed choices.</p></li><li><p><strong>Benchmarks:</strong> recurrent, statistical, HMM, and boosted-tree adapters are available, but comparable numerical results require the original data and exact configurations.</p></li><li><p><strong>Position sizing:</strong> the finite normalized exponential schedule is a reconstruction because equations (5)&#8211;(8) cannot be recovered reliably.</p></li><li><p><strong>VaR:</strong> the rolling historical estimator and trade filter are operational reconstructions because the paper gives no confidence level, horizon, lookback, or action rule.</p></li><li><p><strong>Greedy selection:</strong> the fee-aware candidate-ranking policy is reconstructed because the paper does not define its candidates, objective, constraints, or tie-breaking.</p></li><li><p><strong>Backtesting:</strong> cash, holdings, fees, slippage, next-bar execution, and liquidation are explicit implementation assumptions, not verified details of the original backtest.</p></li><li><p><strong>Fee sensitivity:</strong> the reported fee grids are preserved, but the figures&#8217; numeric outputs are unavailable.</p></li></ul><p>The status table in <code>docs/reproduction_matrix.md</code> is the project&#8217;s main audit record. It maps each paper claim to its implementation location, validation strength, and missing information. It should be updated whenever a source equation, dataset, or protocol detail becomes available.</p><h3>20.2 Why forecast accuracy does not establish trading profitability</h3><p>A forecast metric measures the distance between a prediction and its target. A trading result depends on a much longer chain:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;ce5fa8c3-4f72-4350-962b-45a4a36d48d0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">observed history
    -&gt; forecast
    -&gt; signal threshold
    -&gt; position size
    -&gt; risk filter
    -&gt; order timing
    -&gt; execution price
    -&gt; fees and slippage
    -&gt; holdings and cash
    -&gt; portfolio value</code></pre></div><p>An Att-BiLSTM can have a lower RMSE than a benchmark and still fail to produce a profitable strategy. For example:</p><ul><li><p>the predicted movement may be smaller than the round-trip transaction cost;</p></li><li><p>a small timing error can turn a buy into a loss;</p></li><li><p>a position-sizing rule can amplify an incorrect forecast;</p></li><li><p>the forecast may be evaluated at a different timestamp from the order execution;</p></li><li><p>slippage, spread, liquidity, and partial fills can change the economics;</p></li><li><p>the asset&#8217;s price process may change after the historical training period.</p></li></ul><p>This is why <code>metrics.py</code> keeps forecast metrics separate from portfolio metrics. RMSE, MAE, MAPE, and R-squared should be reported alongside, but not confused with, cumulative return, maximum drawdown, volatility, Sharpe ratio, turnover, and total fees. A buy-and-hold benchmark is also necessary: a strategy&#8217;s final value has little meaning without knowing what passive ownership of the same asset and period would have produced.</p><h3>20.3 Financial and market limitations</h3><p>Historical financial data are not a stationary laboratory process. The relationship between past prices and later prices can change because of market structure, liquidity, regulation, macroeconomic conditions, participant behavior, or technological changes. A model that performs well in one period may not generalize to another.</p><p>The backtest also simplifies real execution. The generated <code>backtest.py</code> simulator supports explicit fees, optional proportional slippage, fractional units, exposure limits, and next-bar execution, but it does not model every market constraint. In particular, it does not provide a full treatment of:</p><ul><li><p>bid-ask spreads and changing spreads;</p></li><li><p>market impact and order-book depth;</p></li><li><p>liquidity limits or volume participation;</p></li><li><p>partial fills and order cancellation;</p></li><li><p>latency and stale signals;</p></li><li><p>exchange outages or data revisions;</p></li><li><p>custody, borrowing, margin, taxes, or funding costs;</p></li><li><p>instrument-specific trading calendars;</p></li><li><p>corporate actions or contract rolls where relevant.</p></li></ul><p>These omissions matter especially for a strategy that repeatedly adds to positions. Transaction costs are not a cosmetic adjustment: they can change whether an action is feasible, whether a greedy candidate has a positive score, and whether the entire position schedule remains solvent.</p><h3>20.4 Why the large Bitcoin result should not be generalized</h3><p>The paper reports an approximate final value of <code>$215,487</code> from a <code>$500</code> Bitcoin allocation, alongside an approximate <code>$646</code> value from a <code>$500</code> gold allocation. Those values are <strong>reported claims</strong>, not reproduced outputs of this project.</p><p>The Bitcoin number may depend on an unusually favorable historical interval, aggressive reinvestment, repeated position additions, leverage-like exposure, unrealistic fills, or an accounting convention that is not described. Without the exact dates, instrument definition, price series, signals, sizing parameters, transaction rules, and liquidation policy, the number cannot be separated into market-period effects and strategy effects.</p><p>It should therefore not be presented as expected future performance, evidence of general profitability, or a target that a synthetic demonstration ought to match. The same caution applies to the reported combined value of approximately <code>$216,133</code> from <code>$1,000</code> and to the paper&#8217;s benchmark forecast metrics.</p><h3>20.5 What a true empirical reproduction would require</h3><p>To replace the current reconstructions with verified implementations, obtain and preserve the following artifacts.</p><h4>Data and timestamps</h4><p>The original gold and Bitcoin files.</p><p>The precise instrument or index represented by each file.</p><p>Currency, timezone, sampling frequency, and date range.</p><p>The meaning of each price column: close, adjusted close, settlement, or another field.</p><p>Rules for missing observations, duplicates, invalid prices, and corporate or market-event adjustments.</p><p>The exact train, validation, and test boundaries.</p><p>Evidence that scaling and feature construction were fitted without using future observations.</p><p>The data layer in <code>data.py</code> can then be adapted to the original schema rather than relying on generic local CSV assumptions.</p><h4>Forecast target and Att-BiLSTM architecture</h4><p>Whether the model predicts raw prices, normalized prices, returns, or percentage changes.</p><p>Whether the forecast horizon is one day, three days, or a sequence of separately generated one-step predictions.</p><p>The number of LSTM layers and hidden units.</p><p>Whether the bidirectional outputs are concatenated, averaged, or projected.</p><p>The exact construction of the attention query <code>q</code>.</p><p>Dense layers, output activation, initialization, and regularization.</p><p>Learning rate, batch size, epoch count, stopping rule, random seed, and optimizer options.</p><p>The exact feature set, including whether volume or technical variables were used.</p><p>The current <code>models/attention.py</code> and <code>models/forecasters.py</code> implement the visible mathematics and a learned-query interpretation, but they cannot infer these missing architectural details.</p><h4>Benchmark configurations</h4><p>For each LSTM, GRU, BiLSTM, Holt-Winters, ARIMA, HMM, and XGBoost model, obtain the input representation, hyperparameters, tuning procedure, forecast horizon, and test-time procedure. This is particularly important for HMM and XGBoost:</p><ul><li><p>Viterbi decoding identifies latent states; it does not by itself define a future price forecast.</p></li><li><p>XGBoost is a gradient-boosted tree method, not a random forest.</p></li><li><p>A fair benchmark comparison requires aligned targets, timestamps, preprocessing, and untouched test data.</p></li></ul><p>Without those details, a locally computed benchmark table can be informative, but it is not a verified reconstruction of the paper&#8217;s table.</p><h4>Position management, VaR, and greedy policy</h4><p>The original source equations or high-resolution figures are required for equations (5)&#8211;(8). In addition, clarify:</p><ul><li><p>whether position variables represent cash, asset units, or notional exposure;</p></li><li><p>how the recovery percentile and decline percentile enter each recursive equation;</p></li><li><p>when an additional position is triggered;</p></li><li><p>the maximum number of additions and any leverage rules;</p></li><li><p>whether fees apply to both sides of each transaction;</p></li><li><p>the VaR confidence level, horizon, lookback, distribution, and aggregation method;</p></li><li><p>whether VaR blocks trades, scales them, or only reports risk;</p></li><li><p>the greedy algorithm&#8217;s candidate actions, objective, constraints, and tie-breaking.</p></li></ul><p>Until those materials are supplied, <code>sizing.py</code>, <code>risk.py</code>, and <code>strategy.py</code> should be described as explicit, auditable reconstructions rather than faithful source translations.</p><h4>Backtest and fee-sensitivity protocol</h4><p>A verified portfolio comparison also needs the backtest start and end dates, initial holdings, cash allocation, signal timestamp, execution timestamp, order type, partial-unit policy, fee convention, spread, slippage, liquidation rule, and reinvestment behavior. For fee sensitivity, the numeric data behind the paper&#8217;s figures are required; the fee grid alone does not determine the curve.</p><p>The experiment layer can hold non-fee assumptions fixed and rerun a supplied backtest callback, but it cannot infer the original callback from the paper. A sensitivity table generated locally should therefore be labeled as a new experiment under documented assumptions.</p><h3>20.6 Educational research code versus deployable trading software</h3><p>This project is intentionally an offline educational framework. It demonstrates how to make data boundaries, model contracts, risk rules, execution timing, and accounting explicit. It is not a deployable trading system.</p><p>A production system would require substantially more work, including independent data validation, secure configuration and credential handling, exchange or broker integration, order-state reconciliation, monitoring, failure recovery, compliance review, capacity analysis, and extensive operational testing. None of those capabilities should be inferred from the presence of a backtest simulator or a command-line interface.</p><p>The safest use of the generated project is as a research scaffold:</p><p>inspect the assumptions;</p><p>replace synthetic data with identified local data;</p><p>preserve chronological boundaries;</p><p>validate each component independently;</p><p>record configuration and provenance;</p><p>compare against passive and simpler baselines;</p><p>treat every result as conditional on its protocol.</p><h3>20.7 Verification confidence and unresolved project findings</h3><p>The available verification evidence is limited. Static checks parsed the generated Python files and inspected basic structure, but the overall static stage did not pass: <code>experiments.py</code> and <code>cli.py</code> triggered placeholder detection, while the generated <code>README.md</code> and <code>docs/tutorial.md</code> fields were flagged because they contain Markdown fences. These findings are recorded issues, not proof that the underlying ideas are correct or incorrect.</p><p>Semantic code verification was explicitly skipped. Consequently, this article must not claim that the generated code was executed, that the test suite passed, or that the model trained successfully. The tests and modules document intended interfaces and invariants, but unresolved API inconsistencies may remain&#8212;for example, between some generated tests, orchestration calls, and implementation signatures.</p><p>This is an important distinction:</p><ul><li><p><strong>Static verification</strong> asks whether source files parse and meet structural checks.</p></li><li><p><strong>Semantic review</strong> asks whether functions, types, formulas, and data flow agree.</p></li><li><p><strong>Runtime testing</strong> would execute the code on controlled fixtures.</p></li><li><p><strong>Empirical reproduction</strong> would compare results with the paper using the original data and protocol.</p></li></ul><p>Only the first category was partially recorded here, and semantic verification was skipped. None of these checks, even if fully successful, would establish that the paper&#8217;s financial claims are valid.</p><h3>20.8 Final interpretation policy</h3><p>When presenting future results, use precise language:</p><ul><li><p>Say <strong>&#8220;the implementation uses&#8221;</strong> for explicit code behavior.</p></li><li><p>Say <strong>&#8220;we reconstruct&#8221;</strong> when interpreting incomplete equations or policies.</p></li><li><p>Say <strong>&#8220;the synthetic example illustrates&#8221;</strong> for generated demonstrations.</p></li><li><p>Say <strong>&#8220;the paper reports&#8221;</strong> for unverified numerical claims.</p></li><li><p>Say <strong>&#8220;locally measured under this protocol&#8221;</strong> for results obtained on identified local data.</p></li><li><p>Do not say <strong>&#8220;reproduced&#8221;</strong> unless the data, dates, preprocessing, model, decision rules, execution protocol, and evaluation outputs have been independently matched.</p></li></ul><p>The project&#8217;s strongest contribution is therefore methodological rather than evidentiary. It shows how to translate an incomplete quantitative-trading description into modular Python components while preserving temporal causality, finite allocation, explicit risk assumptions, transaction-cost accounting, and honest uncertainty. A stronger scientific conclusion must wait for the original data and missing protocol details.</p><h2>Complete Generated Code Files</h2><p>Use the button below to download the source code: </p>
      <p>
          <a href="/__u/onepagecode.substack.com/p/quantitative-trading-model-from-paper">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Advanced Text Generation Techniques and Tools]]></title><description><![CDATA[Chapter 7 of Hands-On Large Language Models: Advanced text generation pipelines, orchestration with LangChain, memory management, and autonomous agents]]></description><link>https://onepagecode.substack.com/p/advanced-text-generation-techniques</link><guid isPermaLink="false">https://onepagecode.substack.com/p/advanced-text-generation-techniques</guid><dc:creator><![CDATA[Onepagecode]]></dc:creator><pubDate>Thu, 09 Jul 2026 19:51:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!qHOK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c98ed2a-e79d-49df-b1e8-e3a5ee0e102c_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Why a single prompt stops being enough</h2><p>A single prompt is often enough when the task is simple and self-contained. If you need a brief answer, a rewrite, or a one-off classification, prompt design may be all you need. The limit appears when the same task must be repeated, routed, validated, or combined with other steps. At that point, the question is no longer only how to get a better response from the model. It is how to build a reusable process around the model so the result is consistent each time.</p><h2><strong>&#128216; Buy the Entire Book &#8211; Available on Amazon Worldwide (If you are paid subscriber, use the url at the end of this article to download the book for free!)</strong></h2><p>You can now purchase the complete book directly from Amazon in your country:</p><ul><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/us/dualbookshelf.marketplacelink/B0H6WZJ57S">United States (US)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/uk/dualbookshelf.marketplacelink/B0H6WZJ57S">United Kingdom (UK)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/de/dualbookshelf.marketplacelink/B0H6WZJ57S">Germany (DE)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/fr/dualbookshelf.marketplacelink/B0H6WZJ57S">France (FR)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/es/dualbookshelf.marketplacelink/B0H6WZJ57S">Spain (ES)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/it/dualbookshelf.marketplacelink/B0H6WZJ57S">Italy (IT)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/nl/dualbookshelf.marketplacelink/B0H6WZJ57S">Netherlands (NL)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/jp/dualbookshelf.marketplacelink/B0H6WZJ57S">Japan (JP)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/br/dualbookshelf.marketplacelink/B0H6WZJ57S">Brazil (BR):</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/ca/dualbookshelf.marketplacelink/B0H6WZJ57S">Canada (CA)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/mx/dualbookshelf.marketplacelink/B0H6WZJ57S">Mexico (MX)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/au/dualbookshelf.marketplacelink/B0H6WZJ57S">Australia (AU)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/in/dualbookshelf.marketplacelink/B0H6WZJ57S">India (IN)</a></strong></p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!qHOK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c98ed2a-e79d-49df-b1e8-e3a5ee0e102c_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!qHOK!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c98ed2a-e79d-49df-b1e8-e3a5ee0e102c_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!qHOK!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c98ed2a-e79d-49df-b1e8-e3a5ee0e102c_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!qHOK!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c98ed2a-e79d-49df-b1e8-e3a5ee0e102c_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!qHOK!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c98ed2a-e79d-49df-b1e8-e3a5ee0e102c_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!qHOK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c98ed2a-e79d-49df-b1e8-e3a5ee0e102c_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2c98ed2a-e79d-49df-b1e8-e3a5ee0e102c_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7707739,&quot;alt&quot;:&quot;Figure 7.1: A single prompt is sufficient for simple tasks, but repeated, routed, and validated work requires an orchestrated workflow with explicit steps and checks.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204107502?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c98ed2a-e79d-49df-b1e8-e3a5ee0e102c_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 7.1: A single prompt is sufficient for simple tasks, but repeated, routed, and validated work requires an orchestrated workflow with explicit steps and checks." title="Figure 7.1: A single prompt is sufficient for simple tasks, but repeated, routed, and validated work requires an orchestrated workflow with explicit steps and checks." srcset="/__u/substackcdn.com/image/fetch/$s_!qHOK!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c98ed2a-e79d-49df-b1e8-e3a5ee0e102c_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!qHOK!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c98ed2a-e79d-49df-b1e8-e3a5ee0e102c_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!qHOK!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c98ed2a-e79d-49df-b1e8-e3a5ee0e102c_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!qHOK!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c98ed2a-e79d-49df-b1e8-e3a5ee0e102c_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>That shift matters because LLM applications become systems, not just prompts. A support workflow, for example, may need a ticket summary, a tone check, a response draft, and a final formatting pass. One prompt can produce each piece, but it does not by itself ensure the order, the handoffs, or the checks that make the workflow dependable. In that sense, prompt engineering is about behavior you ask for, while orchestration is about behavior you design for.</p><h2>Orchestration as the layer around model calls</h2><p>Orchestration is the control layer that sits around generation. It decides how inputs move through the application, when model calls happen, what happens to intermediate output, and which modules can intervene before the final response is returned. Prompt design shapes the model&#8217;s input. Orchestration shapes the workflow that surrounds the model.</p><p>A useful metaphor is a recipe versus an assembly line. A prompt can tell the model what to make, but orchestration determines how the work is sequenced, checked, and assembled. That is why orchestration improves quality, repeatability, and structure in real applications: it makes the process explicit instead of leaving every response to a single pass.</p><h2>LangChain as the organizing framework</h2><p>LangChain is the framework used in this chapter to keep that system view organized. It is not required for every LLM application, but it gives the chapter a shared language for connecting model I/O, templates, chains, memory, and agents. Thinking in that framework makes it easier to see which part of the application is responsible for input shaping, which part manages workflow, and which part handles longer-lived context or tool use.</p><p>This is also where the distinction between prompt engineering and orchestration becomes useful. Prompt engineering improves the wording of a request. Orchestration arranges the reusable structure around that request. LangChain helps separate those concerns so the application does not collapse into one oversized prompt that is hard to understand or maintain.</p><h2>Chains, memory, and agents in the workflow stack</h2><p>Chains are the most predictable orchestration pattern. They connect model calls and other modules into a known sequence, which makes them well suited to workflows with clear steps. Memory extends the interaction across turns so earlier context can remain available. Agents add flexibility by letting the system choose actions at runtime when the next step is not fully known in advance.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!4ZdU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7962923-7427-4d13-bb72-f5cce4b2dc97_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!4ZdU!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7962923-7427-4d13-bb72-f5cce4b2dc97_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!4ZdU!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7962923-7427-4d13-bb72-f5cce4b2dc97_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!4ZdU!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7962923-7427-4d13-bb72-f5cce4b2dc97_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!4ZdU!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7962923-7427-4d13-bb72-f5cce4b2dc97_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!4ZdU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7962923-7427-4d13-bb72-f5cce4b2dc97_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b7962923-7427-4d13-bb72-f5cce4b2dc97_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:6966571,&quot;alt&quot;:&quot;Figure 7.2: Prompt, chain, memory, and agent represent increasing levels of orchestration, from a single call to runtime decision-making.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204107502?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7962923-7427-4d13-bb72-f5cce4b2dc97_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 7.2: Prompt, chain, memory, and agent represent increasing levels of orchestration, from a single call to runtime decision-making." title="Figure 7.2: Prompt, chain, memory, and agent represent increasing levels of orchestration, from a single call to runtime decision-making." srcset="/__u/substackcdn.com/image/fetch/$s_!4ZdU!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7962923-7427-4d13-bb72-f5cce4b2dc97_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!4ZdU!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7962923-7427-4d13-bb72-f5cce4b2dc97_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!4ZdU!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7962923-7427-4d13-bb72-f5cce4b2dc97_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!4ZdU!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7962923-7427-4d13-bb72-f5cce4b2dc97_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A simple decision rule follows from that spectrum. Use a prompt when one model call is enough. Use a chain when the steps are predictable. Use an agent when the next action depends on what happens at runtime. Chains give control, agents give flexibility, and memory supports continuity when the conversation itself carries important context.</p><h2>How orchestration improves real systems</h2><p>The practical value of orchestration shows up in maintainability and repeatability. A workflow built from explicit steps is easier to test, easier to revise, and easier to extend than a single prompt that tries to do everything at once. It is also easier to improve output quality because validation, routing, and enrichment can happen before the final generation step instead of being left to chance.</p><p>That is why this chapter moves beyond prompt engineering without abandoning it. Prompt design still matters, but the larger goal is to design the system around the model. LangChain provides one practical way to do that, and the patterns introduced here will prepare you for memory and agents later in the chapter. Retrieval will come later as a complement to orchestration by supplying external context, but the foundation is the same: structure the workflow so the model is part of a reliable application, not an isolated call.</p><h2>What to expect next in the chapter</h2><p>The next sections move from this motivation into the mechanics of model loading, prompt templates, chains, memory, and agents, with retrieval reserved for the following chapter.</p><h2>Loading local and hosted models for chain-based generation</h2><h3>Start with the deployment decision: local vs hosted</h3><p>Before a prompt ever reaches a chain, the model itself has to be available somewhere. In practice, that means choosing between local inference and a hosted API, and the choice affects far more than convenience. A local model gives you control over data, predictable access, and the ability to experiment offline, but it asks for hardware, model files, and a little more patience during setup. A hosted model removes most of that friction, but shifts the trade-off toward network dependence, usage limits, and external cost. For this chapter, the local path is useful because it makes the mechanics visible, yet the hosted fallback keeps the workflow accessible on a laptop or in a constrained environment.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!ssdo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe09aaf54-2092-4018-b386-3d6931f99031_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!ssdo!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe09aaf54-2092-4018-b386-3d6931f99031_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!ssdo!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe09aaf54-2092-4018-b386-3d6931f99031_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!ssdo!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe09aaf54-2092-4018-b386-3d6931f99031_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ssdo!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe09aaf54-2092-4018-b386-3d6931f99031_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!ssdo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe09aaf54-2092-4018-b386-3d6931f99031_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e09aaf54-2092-4018-b386-3d6931f99031_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:8295255,&quot;alt&quot;:&quot;Figure 7.3: Local and hosted model deployment choices differ in hardware, cost, and access constraints, but the surrounding orchestration layer can stay the same.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204107502?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe09aaf54-2092-4018-b386-3d6931f99031_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 7.3: Local and hosted model deployment choices differ in hardware, cost, and access constraints, but the surrounding orchestration layer can stay the same." title="Figure 7.3: Local and hosted model deployment choices differ in hardware, cost, and access constraints, but the surrounding orchestration layer can stay the same." srcset="/__u/substackcdn.com/image/fetch/$s_!ssdo!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe09aaf54-2092-4018-b386-3d6931f99031_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!ssdo!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe09aaf54-2092-4018-b386-3d6931f99031_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!ssdo!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe09aaf54-2092-4018-b386-3d6931f99031_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!ssdo!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe09aaf54-2092-4018-b386-3d6931f99031_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A good rule of thumb is simple. If you are iterating on prompt structure, chain design, or message formatting, a hosted model is often the fastest way to get moving. If you want to inspect deployment constraints, work offline, or understand how much model size and context budget shape behavior, a local model is worth the extra setup. The important point is that the orchestration layer should stay conceptually stable even when the backend changes; only the engine underneath needs to swap.</p><h3>What quantization changes in practice</h3><p>The local model in this chapter uses a quantized GGUF file. Quantization compresses the model weights by storing them with fewer bits than a full-precision model would use. You can think of it as packing a suitcase more tightly: the contents are still there, but they have been folded into a smaller space so the model can fit on more modest hardware. That smaller footprint usually means lower VRAM usage and often faster inference, especially on consumer machines. The cost is that the model no longer preserves every detail of the original parameters, so a little fidelity is traded away for practicality.</p><p>That trade-off is exactly why quantized models matter. A large model in full precision may be inconvenient or impossible to run locally, while a compressed variant can become usable without changing the surrounding application logic. The same idea underlies many modern local deployment choices: reduce precision enough to make the model portable, but not so much that the quality drops below what the task needs.</p><h3>Fetch and load the local GGUF model in LangChain</h3><p>The example in this chapter uses a GGUF artifact downloaded from Hugging Face. The specific file matters because the precision variant is part of the contract between the model and your hardware. Choosing a different variant changes memory use, speed, and sometimes output quality, so the file name is not just a label; it is part of the deployment decision.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;5a5eeedb-ab7b-4eec-b706-9fa2de4ca90f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">!wget https://huggingface.co/microsoft/Phi-3-mini-4k-instruct-gguf/resolve/main/Phi-3-mini-4k-instruct-fp16.gguf</code></pre></div><p>Once the file is available locally, LangChain can wrap it through <code>LlamaCpp</code>. This is the local model interface for the chapter, and the parameters you pass here shape the generation behavior before any chain is built on top.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e7d3c062-70db-42c5-9dce-35518f4f0f51&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from langchain import LlamaCpp

# Make sure the model path is correct for your system!
llm = LlamaCpp(
    model_path="Phi-3-mini-4k-instruct-fp16.gguf",
    n_gpu_layers=-1,
    max_tokens=500,
    n_ctx=2048,
    seed=42,
    verbose=False
)</code></pre></div><p>The practical meaning of those settings is straightforward. <code>model_path</code> points to the GGUF file you downloaded. <code>n_gpu_layers</code> controls how much of the model is offloaded to the GPU when that is available. <code>max_tokens</code> caps the length of the generated answer, which is part of your token budget. <code>n_ctx</code> sets the context window the model can condition on, and that matters because every extra instruction, conversation turn, or retrieved snippet has to fit alongside the answer you hope to generate. <code>seed</code> helps make experiments repeatable, and <code>verbose</code> keeps the output quieter while you are testing.</p><p>That context limit is easy to underestimate. Think of it as a suitcase that must hold both the prompt and the response. If you keep stuffing history into the prompt, you leave less room for the answer, and eventually the model has to truncate, ignore, or compress something. This is why context size becomes a design constraint later in the chapter when chains and memory start accumulating text.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!akEk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841ad40b-8274-48f7-a388-4ac925d6e4c8_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!akEk!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841ad40b-8274-48f7-a388-4ac925d6e4c8_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!akEk!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841ad40b-8274-48f7-a388-4ac925d6e4c8_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!akEk!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841ad40b-8274-48f7-a388-4ac925d6e4c8_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!akEk!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841ad40b-8274-48f7-a388-4ac925d6e4c8_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!akEk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841ad40b-8274-48f7-a388-4ac925d6e4c8_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/841ad40b-8274-48f7-a388-4ac925d6e4c8_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7443994,&quot;alt&quot;:&quot;Figure 7.4: The context window is a shared budget for prompt text, history, retrieval, and generated output, so input growth reduces available response space.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204107502?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841ad40b-8274-48f7-a388-4ac925d6e4c8_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 7.4: The context window is a shared budget for prompt text, history, retrieval, and generated output, so input growth reduces available response space." title="Figure 7.4: The context window is a shared budget for prompt text, history, retrieval, and generated output, so input growth reduces available response space." srcset="/__u/substackcdn.com/image/fetch/$s_!akEk!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841ad40b-8274-48f7-a388-4ac925d6e4c8_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!akEk!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841ad40b-8274-48f7-a388-4ac925d6e4c8_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!akEk!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841ad40b-8274-48f7-a388-4ac925d6e4c8_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!akEk!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F841ad40b-8274-48f7-a388-4ac925d6e4c8_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Why the raw model call can fail without the right prompt shape</h3><p>Once the model is loaded, it is tempting to call it directly and assume that any text input will work. With some instruct or chat-tuned models, that assumption is wrong. The model may be running correctly, but the request still fails because the input shape does not match the format it expects. That is less like a runtime crash and more like sending a valid HTTP request to the wrong endpoint schema.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;text&quot;,&quot;nodeId&quot;:&quot;e4fcfef7-90f6-4b17-a21d-34361e4ca161&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-text">llm.invoke("Hi! My name is Maarten. What is 1 + 1?")

''</code></pre></div><p>The empty result is the important lesson. The model is not broken; the prompt contract is incomplete. Some models need a specific instruction wrapper, a chat template, or other formatting cues before they respond sensibly. In other words, prompt formatting is not an optional refinement layered on top of the model. For certain backends, it is part of the interface definition.</p><p>That distinction matters for the rest of the chapter. If the request format is unstable, every downstream chain inherits that fragility. Once you know the model&#8217;s expected shape, you can stop hand-assembling prompts and move toward reusable templates that enforce the right structure every time.</p><h3>Use a hosted chat model when local inference is not practical</h3><p>The hosted fallback follows the same overall pattern but uses an API-backed chat model instead of a local GGUF file. Operationally, that changes where the computation happens and what constraints matter most. You no longer manage model files or local GPU placement, but you do need credentials, network access, and a model service that fits your use case.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;053c9e3c-f441-4abc-8295-a0b4d45b9609&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from langchain.chat_models import ChatOpenAI

# Create a chat-based LLM
chat_model = ChatOpenAI(openai_api_key="MY_KEY")</code></pre></div><p>This is the same interface idea in a different engine. The point is not that one path is universally better than the other, but that the surrounding generation workflow can remain portable when the model-loading layer is abstracted cleanly. That portability is exactly why the chapter moves next from model invocation to chains: once the backend is in place, the real productivity gain comes from composing prompts, inputs, and outputs into repeatable workflows rather than hand-calling the model one request at a time.</p><h2>Why templates beat one-off prompts</h2><p>A one-off prompt is fine when you are experimenting, but it becomes brittle the moment the same task needs to run again with different inputs. In practice, the useful unit is not a single prompt string; it is a prompt template, which is a stable prompt frame with named slots for the parts that change. That distinction matters because the reusable part is the workflow contract, while the variable part is just the data you feed into it.</p><p>For the ReviewDesk workflow, that difference is easy to see. A support summary might be rewritten into a clean answer draft today, then reused tomorrow for a different ticket, a different tone, or a different product line. If the structure of the request stays mostly the same, a template gives you consistency without forcing you to copy and edit prose by hand. If the task itself needs to be split into distinct decisions, such as naming a title before drafting a response, then the problem has outgrown a single prompt and is ready for chaining.</p><h2>Defining variables and model-specific prompt format</h2><p>A prompt template is useful because it makes the changing parts explicit. Instead of hiding the input inside a long string, you name the fields that the chain expects. That turns prompt writing into a small interface design exercise: the template defines what information must be supplied, and the model receives the formatted text that results from filling those slots.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!8vv9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06de3df2-9d87-4e78-a9f4-308a0adcbc94_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!8vv9!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06de3df2-9d87-4e78-a9f4-308a0adcbc94_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!8vv9!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06de3df2-9d87-4e78-a9f4-308a0adcbc94_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!8vv9!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06de3df2-9d87-4e78-a9f4-308a0adcbc94_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!8vv9!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06de3df2-9d87-4e78-a9f4-308a0adcbc94_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!8vv9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06de3df2-9d87-4e78-a9f4-308a0adcbc94_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/06de3df2-9d87-4e78-a9f4-308a0adcbc94_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:6614112,&quot;alt&quot;:&quot;Figure 7.5: A prompt template defines a stable text frame with named slots, which are filled with variable inputs before the formatted prompt is sent to the model.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204107502?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06de3df2-9d87-4e78-a9f4-308a0adcbc94_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 7.5: A prompt template defines a stable text frame with named slots, which are filled with variable inputs before the formatted prompt is sent to the model." title="Figure 7.5: A prompt template defines a stable text frame with named slots, which are filled with variable inputs before the formatted prompt is sent to the model." srcset="/__u/substackcdn.com/image/fetch/$s_!8vv9!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06de3df2-9d87-4e78-a9f4-308a0adcbc94_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!8vv9!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06de3df2-9d87-4e78-a9f4-308a0adcbc94_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!8vv9!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06de3df2-9d87-4e78-a9f4-308a0adcbc94_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!8vv9!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06de3df2-9d87-4e78-a9f4-308a0adcbc94_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Some models also care about exact formatting. A chat-oriented model may expect role markers, delimiters, or other special tokens around the user and assistant turns. In that case, the template is not only about wording; it is also about preserving the structure the model was trained to recognize. If that structure is wrong, the model can still produce text, but the chain may become less reliable because the model is being asked to read a format it does not expect.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8c827a97-340f-45ee-a1e1-8315e539e274&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from langchain import PromptTemplate

# Create a prompt template with the "input_prompt" variable
template = """&lt;s&gt;&lt;|user|&gt;
{input_prompt}&lt;|end|&gt;
&lt;|assistant|&gt;"""
prompt = PromptTemplate(
    template=template,
    input_variables=["input_prompt"]
)</code></pre></div><h2>Wiring a prompt directly into an LLM</h2><p>Once the prompt is packaged as a template, the next step is to connect it to the model. LangChain&#8217;s pipe syntax makes that composition explicit: the prompt formats the input, and the model consumes the formatted result. Conceptually, this is the smallest useful chain. It is still just one step, but the step is now reusable and named instead of being a hand-written string glued to a model call.</p><p>The invocation also becomes clearer. You do not send the model a free-form blob of text; you provide a mapping that matches the template&#8217;s variable names. That small discipline pays off quickly, because it makes the call site readable and reduces the chance of silently passing the wrong value into the prompt.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;1c2aa124-9ffb-4eea-b225-14e2c12d1d97&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">basic_chain = prompt | llm

# Use the chain
basic_chain.invoke(
    {
        "input_prompt": "Hi! My name is Maarten. What is 1 + 1?",
    }
)</code></pre></div><h2>Reusing the same pattern for different jobs</h2><p>The real value of templates shows up when the same structure is reused for a different task. A business-name generator, a ticket summarizer, and a response drafter can all follow the same pattern even though they solve different problems. The only thing that changes is the variable name and the wording around it. That is the hallmark of a good template: when the nouns change but the workflow stays stable, you have separated the prompt&#8217;s public interface from its implementation.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;40eb6c2e-8a97-4ec3-a64c-a63c99b60271&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># Create a Chain that creates our business' name
template = "Create a funny name for a business that sells {product}."
name_prompt = PromptTemplate(
    template=template,
    input_variables=["product"]
)</code></pre></div><h2>Breaking a generation task into linked steps</h2><p>A larger generation task often becomes easier to control when it is decomposed into smaller stages. The ReviewDesk analogue is straightforward. Suppose you have a short ticket summary and want first a concise title, then a character sketch for a response style, and finally a full draft. Each step can focus on one subgoal, which makes the output easier to inspect and easier to correct.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!yZzD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2a28f65-3fc7-4f49-acbb-0cd1b68d9816_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!yZzD!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2a28f65-3fc7-4f49-acbb-0cd1b68d9816_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!yZzD!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2a28f65-3fc7-4f49-acbb-0cd1b68d9816_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!yZzD!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2a28f65-3fc7-4f49-acbb-0cd1b68d9816_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!yZzD!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2a28f65-3fc7-4f49-acbb-0cd1b68d9816_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!yZzD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2a28f65-3fc7-4f49-acbb-0cd1b68d9816_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f2a28f65-3fc7-4f49-acbb-0cd1b68d9816_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:9089237,&quot;alt&quot;:&quot;Figure 7.6: A sequential chain breaks a larger generation task into labeled intermediate steps, where each output becomes input to the next prompt.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204107502?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2a28f65-3fc7-4f49-acbb-0cd1b68d9816_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 7.6: A sequential chain breaks a larger generation task into labeled intermediate steps, where each output becomes input to the next prompt." title="Figure 7.6: A sequential chain breaks a larger generation task into labeled intermediate steps, where each output becomes input to the next prompt." srcset="/__u/substackcdn.com/image/fetch/$s_!yZzD!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2a28f65-3fc7-4f49-acbb-0cd1b68d9816_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!yZzD!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2a28f65-3fc7-4f49-acbb-0cd1b68d9816_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!yZzD!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2a28f65-3fc7-4f49-acbb-0cd1b68d9816_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!yZzD!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2a28f65-3fc7-4f49-acbb-0cd1b68d9816_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This is where sequential chaining becomes more than syntactic convenience. The title is not just another string; it is a labeled artifact that later prompts can use. The character description is likewise not a final answer but a structured intermediate result. By giving each stage a specific role, you make the hidden subgoals visible, much as a kitchen keeps prep, plating, and finishing separate so each station knows exactly what to hand off.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c3bcf66c-368b-4207-b69f-a0063fdb2c85&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from langchain import LLMChain

# Create a chain for the title of our story
template = """&lt;s&gt;&lt;|user|&gt;
Create a title for a story about {summary}. Only return the title.&lt;|end|&gt;
&lt;|assistant|&gt;"""
title_prompt = PromptTemplate(template=template, input_variables=["summary"])
title = LLMChain(llm=llm, prompt=title_prompt, output_key="title")

title.invoke({"summary": "a girl that lost her mother"})

# Create a chain for the character description using the summary and title
template = """&lt;s&gt;&lt;|user|&gt;
Describe the main character of a story about {summary} with the title {title}. Use only two sentences.&lt;|end|&gt;
&lt;|assistant|&gt;"""
character_prompt = PromptTemplate(
    template=template, input_variables=["summary", "title"]
)
character = LLMChain(llm=llm, prompt=character_prompt, output_key="character")

# Create a chain for the story using the summary, title, and character description
template = """&lt;s&gt;&lt;|user|&gt;
Create a story about {summary} with the title {title}. The main character is: {character}. Only return the story and it cannot be longer than one paragraph. &lt;|end|&gt;
&lt;|assistant|&gt;"""
story_prompt = PromptTemplate(
    template=template, input_variables=["summary", "title", "character"]
)
story = LLMChain(llm=llm, prompt=story_prompt, output_key="story")

# Combine all three components to create the full chain
llm_chain = title | character | story

llm_chain.invoke("a girl that lost her mother")</code></pre></div><h2>Reading, passing, and constraining intermediate outputs</h2><p>Output keys are the shipping labels of a chain. They tell you what each stage produced and what the next stage should expect. In the example above, the first step stores its result as <code>title</code>, the second as <code>character</code>, and the third as <code>story</code>. That naming is not cosmetic; it is how the pipeline keeps track of which artifact should flow into which prompt.</p><p>This is also where many failures start. If a downstream prompt expects <code>{title}</code> but the upstream step saved its result under a different key, the chain breaks or receives the wrong field. The same issue appears when an intermediate output is too verbose. A title prompt that ignores the instruction to return only the title can pollute the next stage with extra text, which is why generation constraints belong inside the prompt rather than as an afterthought.</p><h2>Choosing between a single prompt and a chain</h2><p>Use a single prompt when the task is simple enough that one well-shaped request can handle it. Use a chain when the task has genuinely separable substeps, when you want intermediate artifacts you can inspect, or when the same structure will be reused across many inputs. A prompt template is the reusable interface; a chain is the orchestration around that interface. Keeping that boundary clear will matter even more in the next sections, where stateful memory and tool-using agents build on the same idea of explicit handoffs.</p><h2>Conversation memory and the cost of keeping context</h2><h3>Why a plain chain does not remember yesterday</h3><p>A plain LLM call is stateless. It answers from whatever text you include in the current prompt, and then the interaction ends. That is easy to miss because the first reply can feel conversational: if you tell the model your name in one turn and ask a related question immediately afterward, the answer may seem context-aware. But that impression can be misleading. The model is not carrying a private notebook between calls; unless you pass the earlier exchange back in, the next invocation starts fresh.</p><p>That distinction matters in a support workflow like ReviewDesk. Suppose a user says, &#8220;Hi, I&#8217;m Maarten, and I need help with a billing issue,&#8221; and the assistant responds helpfully. If a second call asks, &#8220;What is my name?&#8221; a basic chain has no reason to know. The missing ingredient is not a better prompt in the abstract, but a mechanism for carrying forward conversation state. In LangChain, that mechanism is memory.</p><h3>Wiring chat history into the prompt</h3><p>Memory does not replace the prompt. It feeds the prompt. The prompt template reserves a slot, often called <code>chat_history</code> or something equivalent, and the memory object fills that slot before the model is called. That separation is important: the chain orchestrates the call, the memory object stores and loads state, and the prompt template defines where that state should appear.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!pwCU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203e055c-e4ad-4ad9-a02a-3041a90e63e3_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!pwCU!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203e055c-e4ad-4ad9-a02a-3041a90e63e3_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!pwCU!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203e055c-e4ad-4ad9-a02a-3041a90e63e3_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!pwCU!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203e055c-e4ad-4ad9-a02a-3041a90e63e3_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!pwCU!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203e055c-e4ad-4ad9-a02a-3041a90e63e3_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!pwCU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203e055c-e4ad-4ad9-a02a-3041a90e63e3_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/203e055c-e4ad-4ad9-a02a-3041a90e63e3_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7909710,&quot;alt&quot;:&quot;Figure 7.7: A memory object loads prior conversation state into a prompt template slot such as `chat_history` before each LLM call.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204107502?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203e055c-e4ad-4ad9-a02a-3041a90e63e3_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 7.7: A memory object loads prior conversation state into a prompt template slot such as `chat_history` before each LLM call." title="Figure 7.7: A memory object loads prior conversation state into a prompt template slot such as `chat_history` before each LLM call." srcset="/__u/substackcdn.com/image/fetch/$s_!pwCU!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203e055c-e4ad-4ad9-a02a-3041a90e63e3_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!pwCU!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203e055c-e4ad-4ad9-a02a-3041a90e63e3_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!pwCU!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203e055c-e4ad-4ad9-a02a-3041a90e63e3_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!pwCU!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203e055c-e4ad-4ad9-a02a-3041a90e63e3_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A useful mental model is chat-state compression. The chain is the delivery system, the prompt is the envelope, and memory decides how much of the prior conversation to pack into it. If the envelope has no history field, nothing from earlier turns can be injected, no matter how conversational the wording sounds.</p><h3>Full-buffer memory as the simplest persistent transcript</h3><p>The simplest baseline is <code>ConversationBufferMemory</code>. It keeps the full conversation transcript and sends it back into the prompt on each turn. Because nothing is discarded, it is the easiest strategy to reason about. If the user said it earlier, it is still available now. That makes buffer memory a good starting point for short, high-stakes conversations where exact recall matters more than efficiency.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2e4e832c-46c4-4f09-a19e-6f01e71b5f60&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">template = """&lt;s&gt;&lt;|user|&gt;Current conversation:{chat_history}\n\n{input_prompt}&lt;|end|&gt;\n&lt;|assistant|&gt;"""

prompt = PromptTemplate(
    template=template,
    input_variables=["input_prompt", "chat_history"]
)

memory = ConversationBufferMemory(memory_key="chat_history")

llm_chain = LLMChain(
    prompt=prompt,
    llm=llm,
    memory=memory
)</code></pre></div><p>With that wiring in place, the same two-turn ReviewDesk exchange now behaves like a real conversation. The first invocation stores the user&#8217;s name, and the second can recover it because the entire transcript is still attached to the chain.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;209da89f-6915-4395-8daf-ec20599cfa88&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">llm_chain.invoke({"input_prompt": "Hi! My name is Maarten. What is 1 + 1?"})
llm_chain.invoke({"input_prompt": "What is my name?"})</code></pre></div><p>The advantage is clarity. The downside appears quickly in longer dialogs. Every new turn expands the prompt, which increases token usage, slows generation, and raises cost. Buffer memory is faithful, but it is also greedy. Like carrying every receipt in your wallet, it works until the accumulation becomes the problem.</p><h3>Windowed memory: keep the recent turns, drop the rest</h3><p><code>ConversationBufferWindowMemory</code> trims that growth by keeping only the most recent <code>k</code> turns. The idea is a sliding window over the dialogue. Recent context stays available, older context falls away. This is often a better fit when the current exchange depends on the last few messages but not on the entire history.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4d23fc0a-15fc-4153-bd64-50a54c4a8e54&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">memory = ConversationBufferWindowMemory(k=2, memory_key="chat_history")

llm_chain = LLMChain(
    prompt=prompt,
    llm=llm,
    memory=memory
)</code></pre></div><p>With a small window, the chain can still answer follow-up questions tied to the recent exchange, but it eventually loses earlier facts. In a ReviewDesk conversation, that can be enough if the assistant only needs to maintain the active thread of one support case. If the user says their name and issue in one or two turns, the model can continue to refer to them correctly. If the discussion drifts long enough, earlier commitments, counts, or names may drop out of scope.</p><p>This is the most common trap with too-small windows: the model sounds coherent until it is asked about a detail that has already slid past the edge of the buffer. Windowed memory is therefore a budgeting tool, not a weaker version of buffer memory. It is an explicit decision to value recency over completeness.</p><h3>Summary memory: compress the conversation instead of replaying it</h3><p><code>ConversationSummaryMemory</code> takes a different path. Instead of preserving the full transcript, it keeps a running summary of the interaction. After each turn, the earlier conversation is compressed into a shorter form, and that condensed state is what gets injected later. In effect, the model is asked to remember the story, not the transcript.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;91274cf5-265b-4da6-96df-219912131b6d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">summary_prompt_template = """&lt;s&gt;&lt;|user|&gt;Summarize the conversations and update with the new lines.\n\nCurrent summary:\n{summary}\n\nnew lines of conversation:\n{new_lines}\n\nNew summary:&lt;|end|&gt;\n&lt;|assistant|&gt;"""
summary_prompt = PromptTemplate(
    input_variables=["new_lines", "summary"],
    template=summary_prompt_template
)

memory = ConversationSummaryMemory(
    llm=llm,
    memory_key="chat_history",
    prompt=summary_prompt
)

llm_chain = LLMChain(
    prompt=prompt,
    llm=llm,
    memory=memory
)</code></pre></div><p>Because the summary is updated incrementally, the chain can stay conversational across longer sessions without replaying everything. That is the main appeal: lower token cost and better scalability. But compression is never free. A summary tends to preserve the gist while softening exact wording, counts, and small commitments. The most subtle failure mode is summary drift, where the condensed history remains plausible but slowly becomes less precise. It is a little like a game of telephone played over many turns: the storyline survives, but the details can blur.</p><p>A direct inspection call makes the behavior less mysterious. You can ask the memory object what it currently holds and see that it returns the compressed state rather than the raw transcript.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;136f588d-6852-43dc-b8e1-b982d5158d3d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">memory.load_memory_variables({})</code></pre></div><p>That inspection step is useful because it reminds you that memory is just another component in the chain. The chain does not magically &#8220;remember&#8221;; it asks memory what should be inserted, then the prompt receives that state like any other input variable.</p><h3>How to choose a memory strategy in practice</h3><p>The choice among buffer, windowed, and summary memory is a trade-off among retention, token usage, latency, and accuracy. If the conversation is short and every detail matters, full buffer memory is the safest choice. If the assistant only needs the latest turns, a small window keeps the system lean. If the discussion is long-lived and the broad arc matters more than exact phrasing, summary memory is usually the most efficient option.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!zc35!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffef380b7-46b2-4569-aab1-e36da8e5b829_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!zc35!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffef380b7-46b2-4569-aab1-e36da8e5b829_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!zc35!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffef380b7-46b2-4569-aab1-e36da8e5b829_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!zc35!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffef380b7-46b2-4569-aab1-e36da8e5b829_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!zc35!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffef380b7-46b2-4569-aab1-e36da8e5b829_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!zc35!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffef380b7-46b2-4569-aab1-e36da8e5b829_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fef380b7-46b2-4569-aab1-e36da8e5b829_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:8990378,&quot;alt&quot;:&quot;Figure 7.8: Buffer, windowed, and summary memory trade off transcript retention against token cost, latency, and recall fidelity.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204107502?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffef380b7-46b2-4569-aab1-e36da8e5b829_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 7.8: Buffer, windowed, and summary memory trade off transcript retention against token cost, latency, and recall fidelity." title="Figure 7.8: Buffer, windowed, and summary memory trade off transcript retention against token cost, latency, and recall fidelity." srcset="/__u/substackcdn.com/image/fetch/$s_!zc35!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffef380b7-46b2-4569-aab1-e36da8e5b829_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!zc35!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffef380b7-46b2-4569-aab1-e36da8e5b829_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!zc35!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffef380b7-46b2-4569-aab1-e36da8e5b829_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!zc35!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffef380b7-46b2-4569-aab1-e36da8e5b829_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>One useful rule is to think in terms of continuity versus knowledge. Memory is for continuity of the conversation, not for discovering facts outside the chat. A support assistant may remember that the user already described a billing issue, but it should not be treated as a database of policy documents or past tickets. When the task shifts from remembering the conversation to finding external information, memory is no longer enough on its own. That is the point where retrieval becomes the right tool, not a bigger buffer.</p><h2>Agents, tools, and ReAct-style control flow</h2><p>A fixed chain is like an assembly line. The same stations run in the same order every time, which is ideal when the task is predictable. An agent is different. It behaves more like a dispatcher that decides which specialist to call next, then uses that result to choose the following step. The change is not just that the agent is smarter. The change is that control flow is no longer predetermined. A chain follows a route; an agent selects the route as it goes.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!zilo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff284d28c-cbac-41d1-bdcd-06900a753d1b_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!zilo!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff284d28c-cbac-41d1-bdcd-06900a753d1b_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!zilo!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff284d28c-cbac-41d1-bdcd-06900a753d1b_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!zilo!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff284d28c-cbac-41d1-bdcd-06900a753d1b_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!zilo!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff284d28c-cbac-41d1-bdcd-06900a753d1b_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!zilo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff284d28c-cbac-41d1-bdcd-06900a753d1b_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f284d28c-cbac-41d1-bdcd-06900a753d1b_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:8252503,&quot;alt&quot;:&quot;Figure 7.9: Fixed chains execute the same steps in the same order, while agents choose the next action dynamically by routing through tools or specialists at runtime.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204107502?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff284d28c-cbac-41d1-bdcd-06900a753d1b_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 7.9: Fixed chains execute the same steps in the same order, while agents choose the next action dynamically by routing through tools or specialists at runtime." title="Figure 7.9: Fixed chains execute the same steps in the same order, while agents choose the next action dynamically by routing through tools or specialists at runtime." srcset="/__u/substackcdn.com/image/fetch/$s_!zilo!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff284d28c-cbac-41d1-bdcd-06900a753d1b_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!zilo!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff284d28c-cbac-41d1-bdcd-06900a753d1b_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!zilo!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff284d28c-cbac-41d1-bdcd-06900a753d1b_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!zilo!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff284d28c-cbac-41d1-bdcd-06900a753d1b_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>That matters when the answer depends on capabilities outside the model itself. If the task is stable and self-contained, a chain is usually simpler and safer. If the task needs current facts, a calculation, or a choice among several external actions, an agent becomes useful because it can orchestrate those steps dynamically. The model is still just generating text, but the runtime lets it decide when to search, when to calculate, and when to stop.</p><h2>What counts as a tool, and why tools change the game</h2><p>A tool is an external function that the agent can call through a defined interface. It is not a model weight, and it is not hidden reasoning. It is an explicit capability exposed to the agent so the model can use it as part of its plan. Search is one common tool type because it can fetch current or obscure facts. Math is another because it can support deterministic calculation even when the model should not be trusted to carry the arithmetic in its own text output.</p><p>This separation is the key idea. The model handles selection and coordination, while the tools handle execution. That makes the system compositional: search can be swapped for another retrieval service, and math can be replaced with a different calculator, without changing the basic control pattern. It also gives a practical heuristic for design. Use a chain when the steps are fixed. Use an agent when the next step depends on what just happened. Add tools when the model needs to reach beyond its own text generation.</p><h2>ReAct as the loop that binds thinking to acting</h2><p>ReAct turns that control pattern into a repeating loop of thought, action, and observation. The thought step is the model deciding what to do next. The action step names the tool and the input to send to it. The observation is the tool output that comes back into the system. That observation is then folded into the next decision, which is why the loop can adapt instead of merely replaying a script.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!j7Rm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4c8a097-7190-44a4-84f3-33fe1d462df5_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!j7Rm!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4c8a097-7190-44a4-84f3-33fe1d462df5_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!j7Rm!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4c8a097-7190-44a4-84f3-33fe1d462df5_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!j7Rm!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4c8a097-7190-44a4-84f3-33fe1d462df5_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!j7Rm!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4c8a097-7190-44a4-84f3-33fe1d462df5_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!j7Rm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4c8a097-7190-44a4-84f3-33fe1d462df5_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e4c8a097-7190-44a4-84f3-33fe1d462df5_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:9310332,&quot;alt&quot;:&quot;Figure 7.10: ReAct agents iterate through thought, action, and observation, while the scratchpad preserves history so each new decision can use prior tool results.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204107502?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4c8a097-7190-44a4-84f3-33fe1d462df5_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 7.10: ReAct agents iterate through thought, action, and observation, while the scratchpad preserves history so each new decision can use prior tool results." title="Figure 7.10: ReAct agents iterate through thought, action, and observation, while the scratchpad preserves history so each new decision can use prior tool results." srcset="/__u/substackcdn.com/image/fetch/$s_!j7Rm!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4c8a097-7190-44a4-84f3-33fe1d462df5_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!j7Rm!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4c8a097-7190-44a4-84f3-33fe1d462df5_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!j7Rm!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4c8a097-7190-44a4-84f3-33fe1d462df5_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!j7Rm!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4c8a097-7190-44a4-84f3-33fe1d462df5_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The scratchpad, or intermediate-step state, is what preserves that history. It records prior actions and observations so the agent can see what it has already tried before choosing the next move. Without that state, the model would have to make each decision in isolation. With it, each tool call becomes part of the context for the following turn.</p><p>The prompt has to make the available action space visible as well. Tool names and tool descriptions are injected into the template so the model knows what it is allowed to call and what each tool is for. That is a major difference from a free-form prompt. The agent is not inventing arbitrary actions; it is selecting from the tools that the runtime has exposed.</p><h2>How LangChain wires tools, prompts, and execution</h2><p>The implementation begins with the model. In the example, a hosted chat model is loaded with a zero temperature setting so the agent behaves more consistently when deciding among tools.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;fa85d313-d234-4546-9ff3-f7eefabfa507&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import os
from langchain_openai import ChatOpenAI

# Load OpenAI's LLMs with LangChain
os.environ["OPENAI_API_KEY"] = "MY_KEY"
openai_llm = ChatOpenAI(model_name="gpt-3.5-turbo", temperature=0)</code></pre></div><p>Next comes the ReAct prompt. Its job is to define the interaction format and reserve placeholders for the tool list, tool names, the user input, and the scratchpad. That is what allows the same template to support many different runs while keeping the intermediate state attached to the current one.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d2a2ebbf-d624-4afd-a61e-e5b0995bfeb8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># Create the ReAct template
react_template = """Answer the following questions as best you can. You have access to the following tools:

{tools}

Use the following format:

Question: the input question you must answer
Thought: you should always think about what to do
Action: the action to take, should be one of [{tool_names}]
Action Input: the input to the action
Observation: the result of the action
... (this Thought/Action/Action Input/Observation can repeat N times)
Thought: I now know the final answer
Final Answer: the final answer to the original input question

Begin!

Question: {input}
Thought:{agent_scratchpad}"""

prompt = PromptTemplate(
    template=react_template,
    input_variables=["tools", "tool_names", "input", "agent_scratchpad"]
)</code></pre></div><p>The tools are then assembled into a shared list. Here the agent gets both a web search tool and an LLM-backed math tool. That combination is enough to show why agents matter: the model can fetch a live fact, then hand a numeric value to another capability for a deterministic follow-up step.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;dc1c6620-0869-471e-a141-d5e6f4b74743&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from langchain.agents import load_tools, Tool
from langchain.tools import DuckDuckGoSearchResults

# You can create the tool to pass to an agent
search = DuckDuckGoSearchResults()
search_tool = Tool(
    name="duckduck",
    description="A web search engine. Use this to as a search engine for general queries.",
    func=search.run,
)

# Prepare tools
tools = load_tools(["llm-math"], llm=openai_llm)
tools.append(search_tool)</code></pre></div><p>The agent itself is created from the model, the tools, and the prompt. The executor is the runtime layer that makes the whole pattern work. It runs the loop, calls tools, gathers observations, and can print intermediate steps when verbose tracing is enabled. It also has to handle malformed action strings and parsing failures, because model output is not always perfectly formatted.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8db114d9-6731-485b-95b2-fd5c5edca26d&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from langchain.agents import AgentExecutor, create_react_agent

# Construct the ReAct agent
agent = create_react_agent(openai_llm, tools, prompt)
agent_executor = AgentExecutor(
    agent=agent, tools=tools, verbose=True, handle_parsing_errors=True
)</code></pre></div><h2>A worked multi-step example: search plus arithmetic</h2><p>The benefit of agentic control flow becomes clear with a query that mixes current information and computation. Suppose you want the current price of a MacBook Pro in USD and then want to convert that amount to EUR using an exchange rate. A fixed prompt can only guess at the live price. An agent can search for the current price first, then use the math tool on the retrieved value.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;7ac7581c-d72d-4273-8635-d58e05ed27e2&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># What is the price of a MacBook Pro?
agent_executor.invoke(
    {
        "input": "What is the current price of a MacBook Pro in USD? How much would it cost in EUR if the exchange rate is 0.85 EUR for 1 USD."
    }
)</code></pre></div><p>A typical run starts with the agent choosing search because it needs a live fact. The search tool returns an observation, which becomes input to the next decision. If the result is a price in USD, the agent can pass that value into the math tool to compute the EUR equivalent. The executor exposes those intermediate steps, so you can see both the tool call and the observation that followed it. That trace is valuable because it shows not just the final answer, but the path the agent took to get there.</p><p>This example also makes the division of labor obvious. The model is not doing the retrieval or the arithmetic itself. It is coordinating the search tool, the math tool, and the loop that connects them. That is what makes agent systems more flexible than chains, but also more fragile than chains, because each step depends on the quality of the previous one.</p><h2>Where agents fail, and why retrieval comes next</h2><p>Tool use adds power, but it also adds risk. A tool can return irrelevant or stale output. The model can produce malformed action text that fails parsing. A loop can continue taking actions without converging on a final answer. These are the kinds of failures that fixed chains largely avoid because their execution path is predetermined.</p><p>That trade-off is why retrieval-oriented systems come next. Agents are one way to orchestrate access to information, but they are not the only way, and they are not a replacement for retrieval architecture. The next chapter builds on this idea by showing how search and retrieval can be structured more deliberately, so the model can work with information in a more controlled way.</p><p><em>Use the url below to download the enitre book as pdf:</em></p>
      <p>
          <a href="/__u/onepagecode.substack.com/p/advanced-text-generation-techniques">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Prompt Engineering]]></title><description><![CDATA[Chapter 6 of Hands-On Large Language Models: Iterative prompt design, controlling text generation with temperature and top-p, and structured JSON output parsing]]></description><link>https://onepagecode.substack.com/p/prompt-engineering</link><guid isPermaLink="false">https://onepagecode.substack.com/p/prompt-engineering</guid><dc:creator><![CDATA[Onepagecode]]></dc:creator><pubDate>Wed, 08 Jul 2026 19:48:51 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Oj1o!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98c658d-0d01-4743-8613-d3e023fcc448_2816x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>From Representation Models to Generative Prompts</h2><p>Earlier chapters treated text models mainly as representation engines. A review could become an embedding, a support ticket could become a feature vector, and a document could become a label. That perspective is still valuable, but prompt engineering starts when the goal changes. Instead of asking a model to reveal information already encoded in text, we ask it to produce new text that serves a task. The mental model shifts from extracting a signal to directing a generation.</p><p>That shift is important because generation is open ended. A classifier usually lives inside a small answer space, while a generative model can move in many plausible directions. The prompt is the operating surface for that process. It is the control layer through which we frame the task, set the tone, and indicate what kind of output should come next. Prompt engineering is not a writing trick and it is not a substitute for model capability. It is a practical engineering discipline for making the model legible to the task.</p><h2><strong>&#128216; Buy the Entire Book &#8211; Available on Amazon Worldwide (If you are paid subscriber, use the url at the end of this article to download the book for free!)</strong></h2><p>You can now purchase the complete book directly from Amazon in your country:</p><ul><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/us/dualbookshelf.marketplacelink/B0H6WZJ57S">United States (US)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/uk/dualbookshelf.marketplacelink/B0H6WZJ57S">United Kingdom (UK)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/de/dualbookshelf.marketplacelink/B0H6WZJ57S">Germany (DE)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/fr/dualbookshelf.marketplacelink/B0H6WZJ57S">France (FR)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/es/dualbookshelf.marketplacelink/B0H6WZJ57S">Spain (ES)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/it/dualbookshelf.marketplacelink/B0H6WZJ57S">Italy (IT)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/nl/dualbookshelf.marketplacelink/B0H6WZJ57S">Netherlands (NL)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/jp/dualbookshelf.marketplacelink/B0H6WZJ57S">Japan (JP)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/br/dualbookshelf.marketplacelink/B0H6WZJ57S">Brazil (BR):</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/ca/dualbookshelf.marketplacelink/B0H6WZJ57S">Canada (CA)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/mx/dualbookshelf.marketplacelink/B0H6WZJ57S">Mexico (MX)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/au/dualbookshelf.marketplacelink/B0H6WZJ57S">Australia (AU)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/in/dualbookshelf.marketplacelink/B0H6WZJ57S">India (IN)</a></strong></p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Oj1o!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98c658d-0d01-4743-8613-d3e023fcc448_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Oj1o!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98c658d-0d01-4743-8613-d3e023fcc448_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!Oj1o!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98c658d-0d01-4743-8613-d3e023fcc448_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!Oj1o!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98c658d-0d01-4743-8613-d3e023fcc448_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Oj1o!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98c658d-0d01-4743-8613-d3e023fcc448_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Oj1o!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98c658d-0d01-4743-8613-d3e023fcc448_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f98c658d-0d01-4743-8613-d3e023fcc448_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:6641920,&quot;alt&quot;:&quot;Figure 6.1: Figure: Prompt engineering shifts the model from extracting a representation to directing open-ended generation through a prompt control layer.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204106890?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98c658d-0d01-4743-8613-d3e023fcc448_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 6.1: Figure: Prompt engineering shifts the model from extracting a representation to directing open-ended generation through a prompt control layer." title="Figure 6.1: Figure: Prompt engineering shifts the model from extracting a representation to directing open-ended generation through a prompt control layer." srcset="/__u/substackcdn.com/image/fetch/$s_!Oj1o!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98c658d-0d01-4743-8613-d3e023fcc448_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!Oj1o!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98c658d-0d01-4743-8613-d3e023fcc448_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!Oj1o!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98c658d-0d01-4743-8613-d3e023fcc448_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Oj1o!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98c658d-0d01-4743-8613-d3e023fcc448_2816x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Seen that way, prompt engineering becomes an iterative workflow. You define the intent, examine what the model returns, notice where it drifts or overreaches, and revise the wording or structure. A weak prompt often behaves like an underspecified test case: it leaves too much room for the model to guess. Better prompts reduce that uncertainty by tightening the task boundary, making the expected behavior clearer, and improving consistency across runs and users. Stronger models help, but they do not remove the need for this control loop. Ambiguity, variability, and task-specific constraints still affect whether the output is useful in practice.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!pgqY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ffcb4b-b19b-4f19-8804-326700f3e908_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!pgqY!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ffcb4b-b19b-4f19-8804-326700f3e908_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!pgqY!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ffcb4b-b19b-4f19-8804-326700f3e908_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!pgqY!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ffcb4b-b19b-4f19-8804-326700f3e908_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!pgqY!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ffcb4b-b19b-4f19-8804-326700f3e908_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!pgqY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ffcb4b-b19b-4f19-8804-326700f3e908_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/52ffcb4b-b19b-4f19-8804-326700f3e908_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7405562,&quot;alt&quot;:&quot;Figure 6.2: Figure: Prompt engineering is an iterative feedback loop that refines intent, reduces ambiguity, and improves output consistency.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204106890?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ffcb4b-b19b-4f19-8804-326700f3e908_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 6.2: Figure: Prompt engineering is an iterative feedback loop that refines intent, reduces ambiguity, and improves output consistency." title="Figure 6.2: Figure: Prompt engineering is an iterative feedback loop that refines intent, reduces ambiguity, and improves output consistency." srcset="/__u/substackcdn.com/image/fetch/$s_!pgqY!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ffcb4b-b19b-4f19-8804-326700f3e908_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!pgqY!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ffcb4b-b19b-4f19-8804-326700f3e908_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!pgqY!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ffcb4b-b19b-4f19-8804-326700f3e908_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!pgqY!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ffcb4b-b19b-4f19-8804-326700f3e908_2816x1536.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>ReviewDesk makes the change in mindset easy to see. In a representation-first flow, a support item might be assigned a category for analysis. In a generation-first flow, the same item becomes a request for the next text the system should produce, such as a concise reply draft or a short internal summary. The business task is no longer just, what label is this, but what should the model produce next.</p><p>That is why prompt structure continues to matter even as models improve. Good prompts help the system stay consistent, easier to inspect, and easier to use downstream. They also set up the next concerns in this chapter: how to reason with generated outputs, how to verify them, and how to evaluate whether the result is actually fit for the job. Prompt engineering is the first control mechanism, not the last one.</p><h2>Loading a Chat Model and Controlling Sampling</h2><h3>Why Start with a Small Open Chat Model</h3><p>Prompt engineering becomes practical only when you can observe how a model changes as the prompt and decoding settings change. That is why a small open chat model is a sensible starting point instead of a large closed service hidden behind an API. A compact model is easier to run on modest hardware, cheaper to experiment with, and fast enough to support repeated trial and error. In other words, it gives you a low-friction place to learn the workflow before you move on to larger systems.</p><p>A model such as Phi-3-mini is especially useful in this learning phase because it is compact enough to be realistic on limited VRAM while still behaving like a modern instruct-tuned chat model. That makes it a good fit for iterative work: you can revise a prompt, rerun it, and inspect the difference without waiting on heavyweight infrastructure. The goal here is not to claim that this model is universally best. The goal is to choose a practical baseline that keeps experimentation simple.</p><h3>Loading the Model, Tokenizer, and Generation Pipeline</h3><p>The standard Transformers workflow is to load a causal language model, load its tokenizer, and then wrap both in a text-generation pipeline. Each object has a distinct role. The model contains the learned weights. The tokenizer contains the vocabulary and special tokens needed to turn text into token IDs and back again. The pipeline is a convenience layer that packages inference into a single interface.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!yFJY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1fd7be2-aa12-49f3-9ee2-638ae5b89255_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!yFJY!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1fd7be2-aa12-49f3-9ee2-638ae5b89255_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!yFJY!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1fd7be2-aa12-49f3-9ee2-638ae5b89255_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!yFJY!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1fd7be2-aa12-49f3-9ee2-638ae5b89255_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!yFJY!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1fd7be2-aa12-49f3-9ee2-638ae5b89255_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!yFJY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1fd7be2-aa12-49f3-9ee2-638ae5b89255_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b1fd7be2-aa12-49f3-9ee2-638ae5b89255_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5701231,&quot;alt&quot;:&quot;Figure 6.3: The standard generation stack: tokenizer, causal language model, and pipeline work together to turn prompts into decoded text under explicit generation settings.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204106890?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1fd7be2-aa12-49f3-9ee2-638ae5b89255_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 6.3: The standard generation stack: tokenizer, causal language model, and pipeline work together to turn prompts into decoded text under explicit generation settings." title="Figure 6.3: The standard generation stack: tokenizer, causal language model, and pipeline work together to turn prompts into decoded text under explicit generation settings." srcset="/__u/substackcdn.com/image/fetch/$s_!yFJY!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1fd7be2-aa12-49f3-9ee2-638ae5b89255_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!yFJY!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1fd7be2-aa12-49f3-9ee2-638ae5b89255_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!yFJY!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1fd7be2-aa12-49f3-9ee2-638ae5b89255_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!yFJY!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1fd7be2-aa12-49f3-9ee2-638ae5b89255_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;398a76a8-bf73-4e14-85db-1876c3fbeaf0&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline

# Load model and tokenizer
model = AutoModelForCausalLM.from_pretrained(
    "microsoft/Phi-3-mini-4k-instruct",
    device_map="cuda",
    torch_dtype="auto",
    trust_remote_code=True,
)
tokenizer = AutoTokenizer.from_pretrained("microsoft/Phi-3-mini-4k-instruct")

# Create a pipeline
pipe = pipeline(
    "text-generation",
    model=model,
    tokenizer=tokenizer,
    return_full_text=False,
    max_new_tokens=500,
    do_sample=False,
)</code></pre></div><p>This configuration establishes a deterministic baseline. <code>return_full_text=False</code> keeps the output focused on the newly generated continuation instead of echoing the original prompt. <code>max_new_tokens</code> limits how far the response can grow. <code>do_sample=False</code> disables sampling, so generation follows the most likely next token at each step. That makes the result repeatable, which is exactly what you want when you are checking whether a prompt change caused a different answer.</p><p>The pipeline also makes the code easier to read. Instead of wiring generation calls by hand, you pass in a prompt or a message list and let the wrapper handle the inference details. That simplicity matters early in the workflow, because it lets you focus on prompt behavior rather than on low-level generation mechanics.</p><h3>How Chat Messages Become a Serialized Prompt</h3><p>Chat models are usually written against a message list rather than a single string. The list is the form you control as the developer, but it is not yet the exact input the model sees. Before generation, the tokenizer applies a model-specific chat template that serializes the messages into a prompt string and inserts the control tokens the model expects.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;8dc690b5-1ac6-4f0f-9e0e-52f659e5a701&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># Prompt
messages = [
    {"role": "user", "content": "Create a funny joke about chickens."}
]

# Generate the output
output = pipe(messages)
print(output[0]["generated_text"])</code></pre></div><p>The structure above is simple on purpose. The key idea is that the message object is only the starting point. The pipeline turns that structure into the actual text representation the model conditions on.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d871ad69-43d1-4200-b592-2044ae739b58&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># Apply prompt template
prompt = pipe.tokenizer.apply_chat_template(messages, tokenize=False)
print(prompt)</code></pre></div><p>That printed string is the serialized prompt. It may include role tags, separators, and end markers that do not appear in the original message list. Those details are model-specific, which is why chat templates are not portable in a universal way. A prompt that is formatted correctly for one model can be serialized differently for another, even if the user-visible content looks identical.</p><p>A useful way to think about the distinction is that the message list is your instruction, while the chat template is the translation into the model&#8217;s dialect. When a response looks odd, too verbose, or unexpectedly truncated, the first thing to inspect is often the serialized prompt rather than the message object alone.</p><h3>Turning On Sampling: Temperature, Top-p, and Output Variability</h3><p>Deterministic decoding is the baseline, but it is not the only useful behavior. When sampling is enabled, the model has more freedom in how it selects the next token. Temperature and top-p are the two main controls for shaping that freedom.</p><p>Temperature adjusts how sharply the model favors the most likely tokens. Lower values concentrate probability mass on the top candidates and make outputs more predictable. Higher values spread probability more widely and make less likely tokens more available. In practice, that means a low temperature usually gives steadier wording, while a higher temperature usually produces more variation.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!gAvB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b19c5f0-d52a-4220-8c32-a23ea835117a_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!gAvB!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b19c5f0-d52a-4220-8c32-a23ea835117a_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!gAvB!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b19c5f0-d52a-4220-8c32-a23ea835117a_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!gAvB!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b19c5f0-d52a-4220-8c32-a23ea835117a_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!gAvB!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b19c5f0-d52a-4220-8c32-a23ea835117a_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!gAvB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b19c5f0-d52a-4220-8c32-a23ea835117a_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0b19c5f0-d52a-4220-8c32-a23ea835117a_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7045545,&quot;alt&quot;:&quot;Figure 6.4: Temperature reshapes the probability distribution, while top-p selects the smallest token set whose cumulative probability reaches the threshold.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204106890?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b19c5f0-d52a-4220-8c32-a23ea835117a_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 6.4: Temperature reshapes the probability distribution, while top-p selects the smallest token set whose cumulative probability reaches the threshold." title="Figure 6.4: Temperature reshapes the probability distribution, while top-p selects the smallest token set whose cumulative probability reaches the threshold." srcset="/__u/substackcdn.com/image/fetch/$s_!gAvB!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b19c5f0-d52a-4220-8c32-a23ea835117a_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!gAvB!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b19c5f0-d52a-4220-8c32-a23ea835117a_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!gAvB!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b19c5f0-d52a-4220-8c32-a23ea835117a_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!gAvB!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b19c5f0-d52a-4220-8c32-a23ea835117a_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;fe0a81b4-94ba-453a-a1a5-9c9639000fa4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># Using a high temperature
output = pipe(messages, do_sample=True, temperature=1)
print(output[0]["generated_text"])</code></pre></div><p>Top-p, also called nucleus sampling, uses cumulative probability mass to limit the candidate set. Instead of considering every token, the model keeps the smallest set of tokens whose combined probability reaches the chosen threshold. This makes the pool adaptive to the prompt and the model state. A related concept is top-k, which keeps a fixed number of highest-probability tokens, but temperature and top-p are the main controls you are likely to use most often.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;66625d0c-0703-46da-830d-e218c0c955a5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># Using a high top_p
output = pipe(messages, do_sample=True, top_p=1)
print(output[0]["generated_text"])</code></pre></div><p>Sampling should be treated as a reliability choice, not just a creativity feature. The more freedom you allow, the more likely the output is to vary from run to run. That variability can be helpful, but it also means you may need more validation if the output feeds a customer-facing or workflow-critical task. Deterministic generation is often the best debugging mode because fixed output makes it easier to tell whether a change came from the prompt or from the decoding settings.</p><h3>Reading the Same Prompt Three Ways: Deterministic, Tempered, and Nucleus-Sampled</h3><p>In ReviewDesk, imagine a support request asking for a polite reply to a customer who has failed to log in several times. With deterministic decoding, the response is usually stable and conservative, which makes it easier to review and compare. With a higher temperature, the wording may become more varied and expressive. With top-p sampling, the answer remains focused on plausible continuations but has more room to choose among them. The prompt stays the same while the decoding policy changes.</p><p>That difference matters because support, triage, and drafting workflows usually need reliability before style. Sampled outputs are useful examples, but they are not guaranteed to repeat exactly on another run, especially when temperature or top-p is enabled. For that reason, deterministic settings are a good place to start whenever you want to debug a prompt or isolate the effect of a formatting change.</p><h2>From Single Prompt to Prompt Assembly</h2><p>A prompt is easier to design when you stop treating it as one uninterrupted string and start treating it as an assembly of reusable parts. That shift matters because prompt engineering is not just about writing well once. It is about building an interface to the model that can be edited, tested, and reused across tasks. In practice, the useful question is not whether a prompt is long or short, but which parts are doing the work and how those parts interact.</p><p>A modular prompt makes that structure visible. If a response is too vague, the likely cause may be a missing instruction, weak context, an unclear format cue, or a mismatch between the audience and the tone. Thinking this way turns prompt writing into a debugging exercise. You can toggle one component at a time, observe the output, and learn which part actually changes the model&#8217;s behavior. That is a more reliable habit than rewriting everything from scratch whenever the result disappoints you.</p><h2>Anatomy of a Modular Prompt</h2><p>A prompt can be built from several reusable design units. Persona sets the role the model should inhabit, such as a support agent, analyst, or editor. Instruction states the job to perform. Context supplies the background that makes the job meaningful. Format tells the model how the answer should be shaped. Audience specifies who the output is for. Tone guides voice and level of formality. Data is the concrete content the task acts on. These pieces overlap in practice, but each one gives a different kind of control.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!QIRn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e65f60-045b-4ab9-89a0-2f846e2872f6_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!QIRn!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e65f60-045b-4ab9-89a0-2f846e2872f6_1408x768.png 424w, /__u/substackcdn.com/image/fetch/$s_!QIRn!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e65f60-045b-4ab9-89a0-2f846e2872f6_1408x768.png 848w, /__u/substackcdn.com/image/fetch/$s_!QIRn!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e65f60-045b-4ab9-89a0-2f846e2872f6_1408x768.png 1272w, /__u/substackcdn.com/image/fetch/$s_!QIRn!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e65f60-045b-4ab9-89a0-2f846e2872f6_1408x768.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!QIRn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e65f60-045b-4ab9-89a0-2f846e2872f6_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/57e65f60-045b-4ab9-89a0-2f846e2872f6_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3238639,&quot;alt&quot;:&quot;Figure 6.5: A modular prompt is composed of reusable fields&#8212;persona, instruction, context, format, audience, tone, and data&#8212;each contributing a different control signal to the model.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204106890?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e65f60-045b-4ab9-89a0-2f846e2872f6_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 6.5: A modular prompt is composed of reusable fields&#8212;persona, instruction, context, format, audience, tone, and data&#8212;each contributing a different control signal to the model." title="Figure 6.5: A modular prompt is composed of reusable fields&#8212;persona, instruction, context, format, audience, tone, and data&#8212;each contributing a different control signal to the model." srcset="/__u/substackcdn.com/image/fetch/$s_!QIRn!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e65f60-045b-4ab9-89a0-2f846e2872f6_1408x768.png 424w, /__u/substackcdn.com/image/fetch/$s_!QIRn!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e65f60-045b-4ab9-89a0-2f846e2872f6_1408x768.png 848w, /__u/substackcdn.com/image/fetch/$s_!QIRn!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e65f60-045b-4ab9-89a0-2f846e2872f6_1408x768.png 1272w, /__u/substackcdn.com/image/fetch/$s_!QIRn!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e65f60-045b-4ab9-89a0-2f846e2872f6_1408x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>They are not a universal schema. A prompt does not need every field, and the fields do not need to appear in one fixed order. The point is to make the design choices explicit so they can be reused or removed deliberately. That is why prompt composition is best understood as a set of modules rather than a rigid template. If persona says &#8220;brief and direct&#8221; while tone says &#8220;detailed and explanatory,&#8221; the prompt sends mixed signals. If context is added late after the main instruction, it may still help, but the model may attend to the earlier framing more strongly. Ordering is not magic, but it does affect which cue feels dominant.</p><p>Consider a support-style summarization prompt. The persona can establish an expert voice, the instruction can ask for a concise summary, the context can explain that the goal is to extract the most important points, the format can request a labeled output, the audience can point to busy researchers, the tone can keep the result professional, and the data field can hold the text to summarize. Each component narrows the space of acceptable completions in a different way. If one part is removed, the output often changes in a visible and useful way. If the format cue disappears, the answer may become less structured. If the audience cue disappears, the wording may become less tailored. If the tone cue changes, the same content may feel more casual or more formal.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d664b503-d7a0-463c-8007-9a6023826bbb&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># Prompt components
persona = "You are an expert in Large Language models. You excel at breaking down complex papers into digestible summaries.\n"
instruction = "Summarize the key findings of the paper provided.\n"
context = "Your summary should extract the most crucial points that can help researchers quickly understand the most vital information of the paper.\n"
data_format = "Create a bullet-point summary that outlines the method. Follow this up with a concise paragraph that encapsulates the main results.\n"
audience = "The summary is designed for busy researchers that quickly need to grasp the newest trends in Large Language Models.\n"
tone = "The tone should be professional and clear.\n"
text = "MY TEXT TO SUMMARIZE"
data = f"Text to summarize: {text}"

# The full prompt - remove and add pieces to view its impact on the generated output
query = persona + instruction + context + data_format + audience + tone + data</code></pre></div><p>The engineering value here is not the string concatenation itself. It is the ability to edit one field at a time and compare outputs. That makes prompt work feel closer to A/B testing than to one-shot writing. You can remove the audience clause, shorten the context, or move the format cue earlier and inspect what changes. This is especially useful when a prompt needs to be maintained over time, because each component can be revised without losing the whole structure.</p><h2>Prompt Tuning by Editing the Pieces</h2><p>Prompt refinement is usually iterative. Add a component, remove one, or reorder them, then compare the result. If the answer becomes more precise after adding format, that field is earning its place. If a persona statement causes unnecessary verbosity, it may be too strong for the task. If a context block is helpful only when placed near the instruction, that ordering choice is worth keeping. The point is to isolate cause and effect so the prompt behaves like a controlled design, not a guess.</p><p>A good rule of thumb is to change one thing at a time when possible. That makes the output easier to interpret. It also helps you distinguish a genuinely useful constraint from a cosmetic one. Prompt design becomes more stable when you treat each field as a testable feature rather than as decoration.</p><h2>One-Shot Examples as Behavior Demonstrations</h2><p>Another way to steer a model is to show it a worked example inside the prompt. This is one-shot in-context learning. It is not training, and it does not update the model&#8217;s weights. The example only changes the context the model is currently conditioning on. That distinction matters because a prompt example can demonstrate a pattern, but it does not rewrite the model&#8217;s underlying knowledge.</p><p>One-shot prompting is best understood as behavior demonstration. It can teach style, response shape, label mapping, or task format. It is useful when you want the model to imitate a pattern rather than infer it from a long rule description. It is less reliable as a way to inject brand-new facts the model never learned. A single example can show how a reply should sound, but it cannot guarantee durable factual learning the way training would.</p><p>With chat-based models, the example is usually written as role-structured messages. Those messages are then serialized through the model&#8217;s chat template into the actual input format expected by that model. The distinction is important. The message list is the structured prompt design, while the template serialization is the model-specific conversion step that turns that structure into tokens. In other words, the example is not just text; it is a sequence of roles and contents that the template packages for generation.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5ffcedc2-9599-4baa-be3a-18ad34cf85ea&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># Use a single example of using the made-up word in a sentence
one_shot_prompt = [
    {
        "role": "user",
        "content": "A 'Gigamuru' is a type of Japanese musical instrument. An example of a sentence that uses the word Gigamuru is:"
    },
    {
        "role": "assistant",
        "content": "I have a Gigamuru that my uncle gave me as a gift. I love to play it at home."
    },
    {
        "role": "user",
        "content": "To 'screeg' something is to swing a sword at it. An example of a sentence that uses the word screeg is:"
    }
]
print(tokenizer.apply_chat_template(one_shot_prompt, tokenize=False))

# Generate the output
outputs = pipe(one_shot_prompt)
print(outputs[0]["generated_text"])</code></pre></div><p>This pattern is useful when the task is mostly about format or style. A single example can show how a ticket summary should be phrased, how a classification label should appear, or how a handoff note should be structured. It is not a shortcut to knowledge acquisition. It is a way of showing the model the pattern you want it to continue.</p><h2>Chaining Prompts into a Workflow</h2><p>Prompting also becomes more powerful when it is split across multiple calls. Chain prompting uses the output of one call as input to the next. Instead of asking the model to do everything in one pass, you let it perform one subtask, capture the intermediate output, and then reuse that result downstream. This is a workflow pattern, not a new model capability.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!NVIf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F260846da-267b-492c-b1ee-e9e608177373_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!NVIf!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F260846da-267b-492c-b1ee-e9e608177373_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!NVIf!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F260846da-267b-492c-b1ee-e9e608177373_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!NVIf!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F260846da-267b-492c-b1ee-e9e608177373_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!NVIf!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F260846da-267b-492c-b1ee-e9e608177373_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!NVIf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F260846da-267b-492c-b1ee-e9e608177373_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/260846da-267b-492c-b1ee-e9e608177373_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:6270695,&quot;alt&quot;:&quot;Figure 6.6: Chain prompting breaks a task into multiple model calls, where each intermediate output becomes the input to the next stage.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204106890?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F260846da-267b-492c-b1ee-e9e608177373_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 6.6: Chain prompting breaks a task into multiple model calls, where each intermediate output becomes the input to the next stage." title="Figure 6.6: Chain prompting breaks a task into multiple model calls, where each intermediate output becomes the input to the next stage." srcset="/__u/substackcdn.com/image/fetch/$s_!NVIf!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F260846da-267b-492c-b1ee-e9e608177373_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!NVIf!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F260846da-267b-492c-b1ee-e9e608177373_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!NVIf!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F260846da-267b-492c-b1ee-e9e608177373_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!NVIf!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F260846da-267b-492c-b1ee-e9e608177373_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>That separation is valuable when a task naturally breaks into stages. One call can draft a name, summary, or internal note. A later call can transform that draft into a customer-facing message, a shorter pitch, or a cleaner revision. The first output becomes a controlled input for the second prompt. That makes the pipeline easier to reason about because each step has a narrower purpose.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6d93e4a6-b8d0-4c97-8c32-b20d21822f3e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python"># Create name and slogan for a product
product_prompt = [
    {"role": "user", "content": "Create a name and slogan for a chatbot that leverages LLMs."}
]
outputs = pipe(product_prompt)
product_description = outputs[0]["generated_text"]
print(product_description)

# Based on a name and slogan for a product, generate a sales pitch
sales_prompt = [
    {"role": "user", "content": f"Generate a very short sales pitch for the following product: '{product_description}'"}
]
outputs = pipe(sales_prompt)
sales_pitch = outputs[0]["generated_text"]
print(sales_pitch)</code></pre></div><p>In this example, the first call produces the product name and slogan, and the second call consumes that output to produce a sales pitch. That handoff is the key idea. Chain prompting is useful when you want decomposition, drafting, or refinement. It is often preferable to a single prompt when the full task is too broad, when the first pass is exploratory, or when the second pass needs a cleaner input than the raw source text.</p><p>It also has failure modes. Errors can accumulate across steps, vague intermediate outputs can become worse in the next call, and brittle handoffs can cause the later prompt to lose the original intent. For that reason, chain prompting works best when each intermediate output is named clearly and is specific enough to support the next stage. Think of it as an assembly line: each station does a smaller job, but each station also depends on the quality of the previous one.</p><h2>What Modular Prompting Can and Cannot Do</h2><p>Modular prompting improves steering, clarity, and reuse, but it does not retrain the model. Prompt examples do not update weights. They do not create permanent knowledge. They only shape the behavior of the current interaction. That is why prompt design and model training are related but fundamentally different. Prompting works by changing inputs; training works by changing parameters.</p><p>The practical takeaway is simple. Use modular prompts when you want to control behavior, demonstrate style, or decompose a task into stages. Use chain prompting when a single pass is too crowded and intermediate results will help. Use examples when you want to demonstrate a pattern. Keep in mind that these methods improve structure, not certainty. The next section extends that idea by moving toward techniques that ask the model to work through tasks more deliberately.</p><h2>What reasoning-oriented prompting is trying to change</h2><p>Ordinary prompt refinement mainly changes what the model says. Reasoning-oriented prompting tries to change the process the model is nudged to produce before it settles on an answer. That difference matters when the task has hidden structure, overlapping labels, or underspecified details. In those situations, a prompt that asks for intermediate work can make the output more stable than a prompt that simply asks for a polished response.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Rzwu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8f156b-7b95-4fa5-ad9d-aff9173eff5b_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Rzwu!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8f156b-7b95-4fa5-ad9d-aff9173eff5b_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!Rzwu!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8f156b-7b95-4fa5-ad9d-aff9173eff5b_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!Rzwu!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8f156b-7b95-4fa5-ad9d-aff9173eff5b_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Rzwu!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8f156b-7b95-4fa5-ad9d-aff9173eff5b_2752x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Rzwu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8f156b-7b95-4fa5-ad9d-aff9173eff5b_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7b8f156b-7b95-4fa5-ad9d-aff9173eff5b_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:8266307,&quot;alt&quot;:&quot;Figure 6.7: Reasoning-oriented prompting changes the model&#8217;s process, not just the wording of the final answer, which is especially helpful on ambiguous tasks.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204106890?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8f156b-7b95-4fa5-ad9d-aff9173eff5b_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 6.7: Reasoning-oriented prompting changes the model&#8217;s process, not just the wording of the final answer, which is especially helpful on ambiguous tasks." title="Figure 6.7: Reasoning-oriented prompting changes the model&#8217;s process, not just the wording of the final answer, which is especially helpful on ambiguous tasks." srcset="/__u/substackcdn.com/image/fetch/$s_!Rzwu!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8f156b-7b95-4fa5-ad9d-aff9173eff5b_2752x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!Rzwu!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8f156b-7b95-4fa5-ad9d-aff9173eff5b_2752x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!Rzwu!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8f156b-7b95-4fa5-ad9d-aff9173eff5b_2752x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Rzwu!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8f156b-7b95-4fa5-ad9d-aff9173eff5b_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The important point is not that the model becomes a person-like reasoner. It does not suddenly think the way a human does. The practical effect is narrower and more useful: the prompt can steer generation toward a stepwise style that exposes intermediate choices, which often improves average reliability on messy tasks. For ReviewDesk, that is helpful when a ticket mentions several plausible categories at once and the model needs to compare evidence instead of guessing from the first clue.</p><h2>Chain-of-thought prompting as single-path guided reasoning</h2><p>Chain-of-thought prompting is the simplest version of this idea. You show the model one worked example that includes intermediate reasoning, then ask a new question in the same format. The example acts as a pattern for the next completion, so the model is more likely to produce a visible reasoning trace before giving the final label. This is useful when the answer depends on a hidden condition or on weighing more than one clue.</p><p>Here is a compact ReviewDesk-style example. The first ticket is answered with a brief justification, and the second ticket is left for the model to solve in the same style.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b5e8fbe2-f6b9-475c-940e-c94e00cab5ce&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">cot_prompt = [
    {"role": "user", "content": "Ticket: 'I was charged twice, but I also can't access my account after the password reset.' Classify the main issue and briefly explain why."},
    {"role": "assistant", "content": "The billing problem is the clearest actionable issue because the duplicate charge is explicit. Access trouble is present, but the payment issue is the strongest label. Final answer: billing."},
    {"role": "user", "content": "Ticket: 'My subscription renewed today, but the app says my plan expired and the invoice looks wrong.' Classify the main issue and briefly explain why."}
]

outputs = pipe(cot_prompt)
print(outputs[0]["generated_text"])</code></pre></div><p>The worked example matters because it demonstrates not just the final format but the expected treatment of evidence. That is why chain-of-thought often helps on ambiguous support tickets, policy triage, or any other task where the label depends on comparing several plausible interpretations. It is still only a prompt-level nudge, though. A reasoning trace can look convincing and still be wrong if the task statement is unclear or if the model latches onto the wrong clue early.</p><h2>Zero-shot chain-of-thought: triggering reasoning without examples</h2><p>Zero-shot chain-of-thought keeps the same goal but removes the example. Instead of demonstrating reasoning, it asks for it directly with a short cue such as &#8220;think step by step.&#8221; This lighter version is useful when you want to save context, avoid crafting a demonstration, or quickly test whether a task benefits from explicit reasoning guidance.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f50e5146-0e6b-4083-a813-62259cc7a947&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">zeroshot_cot_prompt = [
    {"role": "user", "content": "Ticket: 'My subscription renewed today, but the app says my plan expired and the invoice looks wrong.' Think step by step, then give the main issue."}
]

outputs = pipe(zeroshot_cot_prompt)
print(outputs[0]["generated_text"])</code></pre></div><p>A short cue can still change generation style because it alters what the model is trying to produce next, even without examples. In practice, this is often a good first step. If it works, the prompt stays compact. If it does not, that is a clue that the task may need a better framing, a few-shot example, or a clearer decision rule rather than just a larger model.</p><h2>Self-consistency: sample several reasoning traces and aggregate the answer</h2><p>Self-consistency goes beyond a single reasoning path. Instead of trusting one completion, you sample several completions under stochastic decoding and then aggregate the final answers. The idea is simple: if the prompt is robust, different sampled traces should arrive at the same label. If the prompt is fragile, the variation between samples exposes that instability, and voting over the answers often improves the final decision.</p><p>This method depends on diversity in the samples, so decoding settings matter. If temperature is too low or top_p is too restrictive, the samples may be nearly identical and self-consistency loses much of its value. If sampling is too loose, the outputs can become noisy in a different way. The practical goal is enough diversity to reveal alternative reasoning paths while keeping the completions plausible.</p><p>For ReviewDesk, this is useful when a ticket could reasonably fit more than one category and the first answer is unreliable. A single pass might lock onto the most obvious word in the text. Several sampled passes make it more likely that the dominant pattern emerges, and a simple majority vote over the final labels often gives a better result.</p><p>The tradeoff is cost. Multiple samples mean more tokens, more latency, and more compute expense. That makes self-consistency a good fit for unstable or high-value decisions, but a poor fit for easy tasks where a single answer is already good enough.</p><h2>Tree-of-thought: branching into alternative solution paths</h2><p>Tree-of-thought extends the same general idea, but it changes the unit of exploration. Instead of sampling several final answers, it branches into several intermediate solution paths, compares them, and keeps the promising ones while discarding the weak ones. The key difference is that this is about path exploration, not answer voting.</p><p>That distinction is important. Self-consistency asks, in effect, which complete answer shows up most reliably across samples. Tree-of-thought asks which partial line of reasoning appears most promising before the answer is finalized. It is a drafting-and-narrowing process rather than a voting process.</p><p>A simplified ReviewDesk prompt can imitate this by asking for a few candidate interpretations and then a final choice.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;2561ab67-7836-4bbd-b833-05ee1961ea32&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">tot_prompt = [
    {"role": "user", "content": "Ticket: 'I was charged after canceling, and the dashboard still says my payment failed.' Consider a few possible interpretations, compare them briefly, then choose the best category and explain why."}
]

outputs = pipe(tot_prompt)
print(outputs[0]["generated_text"])</code></pre></div><p>You can think of the method as asking the model to draft several candidate explanations, debate which one fits best, and then narrow to the strongest branch. That is enough to capture the practical idea without turning the prompt into a full search algorithm. Tree-of-thought is most useful when the task is open-ended or when the model needs to choose among several plausible paths rather than retrieve one obvious fact.</p><h2>Choosing a reasoning strategy in practice</h2><p>A simple mental model helps. Use a single-pass prompt when the task is straightforward. Use self-consistency when the task is unstable and the cost of a wrong answer is higher than the cost of extra sampling. Use tree-of-thought when the task requires comparing several competing interpretations or exploring a small decision space before choosing.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!qo_S!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08cecff9-514d-4884-b5c2-02f95223d63d_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!qo_S!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08cecff9-514d-4884-b5c2-02f95223d63d_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!qo_S!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08cecff9-514d-4884-b5c2-02f95223d63d_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!qo_S!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08cecff9-514d-4884-b5c2-02f95223d63d_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!qo_S!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08cecff9-514d-4884-b5c2-02f95223d63d_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!qo_S!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08cecff9-514d-4884-b5c2-02f95223d63d_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/08cecff9-514d-4884-b5c2-02f95223d63d_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7271415,&quot;alt&quot;:&quot;Figure 6.8: A practical rule of thumb: use single-pass prompts for easy tasks, self-consistency for unstable decisions, and tree-of-thought when the model must explore competing interpretations.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204106890?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08cecff9-514d-4884-b5c2-02f95223d63d_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 6.8: A practical rule of thumb: use single-pass prompts for easy tasks, self-consistency for unstable decisions, and tree-of-thought when the model must explore competing interpretations." title="Figure 6.8: A practical rule of thumb: use single-pass prompts for easy tasks, self-consistency for unstable decisions, and tree-of-thought when the model must explore competing interpretations." srcset="/__u/substackcdn.com/image/fetch/$s_!qo_S!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08cecff9-514d-4884-b5c2-02f95223d63d_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!qo_S!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08cecff9-514d-4884-b5c2-02f95223d63d_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!qo_S!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08cecff9-514d-4884-b5c2-02f95223d63d_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!qo_S!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08cecff9-514d-4884-b5c2-02f95223d63d_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>All three methods improve prompt behavior and often improve average outcomes, but none of them guarantees correctness. They reduce uncertainty; they do not eliminate it. If the ticket is ambiguous, underspecified, or high stakes, you should still verify the result with downstream checks or human review. The strongest practical benefit of reasoning-oriented prompting is not perfect logic. It is a better chance of getting a careful answer when the problem itself is not already easy.</p><h2>Structured Output, Verification, and Constrained Decoding</h2><p>Structured output becomes a production requirement the moment a model response stops being a reading experience and starts being input for something else. A support triage system cannot &#8220;mostly&#8221; receive JSON. A database loader cannot guess where the category field ended. An API client cannot safely consume prose that looks organized to a human but fails at the border checkpoint of a parser. This is the practical distinction that matters here: a response can look right on screen and still not run right in software.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!4iLt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fde96c5-b704-41d3-a3e3-81db026c766a_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!4iLt!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fde96c5-b704-41d3-a3e3-81db026c766a_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!4iLt!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fde96c5-b704-41d3-a3e3-81db026c766a_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!4iLt!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fde96c5-b704-41d3-a3e3-81db026c766a_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!4iLt!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fde96c5-b704-41d3-a3e3-81db026c766a_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!4iLt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fde96c5-b704-41d3-a3e3-81db026c766a_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1fde96c5-b704-41d3-a3e3-81db026c766a_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7192827,&quot;alt&quot;:&quot;Figure 6.9: A response can appear correct to a person while still failing at the parser boundary; downstream software requires machine-valid structure, not just visual plausibility.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204106890?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fde96c5-b704-41d3-a3e3-81db026c766a_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 6.9: A response can appear correct to a person while still failing at the parser boundary; downstream software requires machine-valid structure, not just visual plausibility." title="Figure 6.9: A response can appear correct to a person while still failing at the parser boundary; downstream software requires machine-valid structure, not just visual plausibility." srcset="/__u/substackcdn.com/image/fetch/$s_!4iLt!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fde96c5-b704-41d3-a3e3-81db026c766a_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!4iLt!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fde96c5-b704-41d3-a3e3-81db026c766a_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!4iLt!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fde96c5-b704-41d3-a3e3-81db026c766a_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!4iLt!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fde96c5-b704-41d3-a3e3-81db026c766a_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The simplest way to improve structure is to ask for it directly. A zero-shot prompt can request a fixed schema, but the model is still free to drift, add commentary, or return partial objects. In ReviewDesk, you might ask for a ticket classification result containing a category, a priority, and a short draft reply. That is enough to guide the model, but not enough to guarantee anything. If the task is forgiving, a prompt alone may be acceptable. If the output is going straight into another system, the same request becomes fragile.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;df31863f-f89d-48a0-8fdb-f659ebfba0d3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">prompt = [
    {
        "role": "user",
        "content": (
            "Classify this ReviewDesk ticket and return valid JSON with exactly these keys: "
            "category, priority, draft_response. "
            "Ticket: 'My login works on mobile but not on desktop after the last update.'"
        ),
    }
]

output = llm.create_chat_completion(messages=prompt, temperature=0)["choices"][0]["message"]["content"]
print(output)</code></pre></div><p>A one-shot or few-shot template usually improves adherence because the model sees the target shape rather than merely being told about it. The example acts like a small map of the destination. You are not training the model, and you are not changing its weights; you are narrowing the space of likely completions by showing a format that should be imitated. That distinction is easy to miss. In-context examples are temporary steering signals, not new knowledge stored by the model.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e4d8d577-ee7f-4ff3-8709-f7c73a2a5744&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">prompt = [
    {
        "role": "user",
        "content": (
            "Return ReviewDesk tickets in this JSON format only:\n"
            '{"category":"...","priority":"...","draft_response":"..."}\n\n'
            "Example:\n"
            '{"category":"authentication","priority":"high","draft_response":"Please try resetting your password..."}\n\n'
            "Ticket: 'The billing page shows an error when I open it from Safari.'"
        ),
    }
]

output = llm.create_chat_completion(messages=prompt, temperature=0)["choices"][0]["message"]["content"]
print(output)</code></pre></div><p>This is where the contrast between guidance and enforcement becomes important. Prompt shaping is persuasive; constrained decoding is restrictive. Prompting tells the model what kind of answer is wanted. Grammar-constrained decoding or JSON-enforced generation limits the next token choices so the model can only produce output that fits the allowed format. In other words, prompting steers the car, while constrained decoding puts rails on the road. The prompt still matters, but the runtime is now doing part of the reliability work.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Q4oG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c7b9ff-1549-4d1c-a619-fb2b9feb23f5_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Q4oG!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c7b9ff-1549-4d1c-a619-fb2b9feb23f5_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!Q4oG!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c7b9ff-1549-4d1c-a619-fb2b9feb23f5_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!Q4oG!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c7b9ff-1549-4d1c-a619-fb2b9feb23f5_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Q4oG!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c7b9ff-1549-4d1c-a619-fb2b9feb23f5_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Q4oG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c7b9ff-1549-4d1c-a619-fb2b9feb23f5_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e6c7b9ff-1549-4d1c-a619-fb2b9feb23f5_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7486089,&quot;alt&quot;:&quot;Figure 6.10: Prompting nudges the model toward the desired structure, while constrained decoding limits token choices so only schema-valid output can be produced.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204106890?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c7b9ff-1549-4d1c-a619-fb2b9feb23f5_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 6.10: Prompting nudges the model toward the desired structure, while constrained decoding limits token choices so only schema-valid output can be produced." title="Figure 6.10: Prompting nudges the model toward the desired structure, while constrained decoding limits token choices so only schema-valid output can be produced." srcset="/__u/substackcdn.com/image/fetch/$s_!Q4oG!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c7b9ff-1549-4d1c-a619-fb2b9feb23f5_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!Q4oG!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c7b9ff-1549-4d1c-a619-fb2b9feb23f5_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!Q4oG!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c7b9ff-1549-4d1c-a619-fb2b9feb23f5_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Q4oG!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c7b9ff-1549-4d1c-a619-fb2b9feb23f5_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>With runtimes that support structured generation, the same ReviewDesk task can be asked for as JSON-only output. A model such as a quantized GGUF chat model loaded through <code>llama-cpp-python</code> can be configured so the response must satisfy a JSON grammar or a JSON response format. That does not make the answer correct in the semantic sense, but it does strongly improve the guarantee that the string is machine-readable.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6fee3632-9f7b-40d7-aedd-7e71bd0bc732&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from llama_cpp.llama import Llama

llm = Llama.from_pretrained(
    repo_id="microsoft/Phi-3-mini-4k-instruct-gguf",
    filename="*fp16.gguf",
    n_gpu_layers=-1,
    n_ctx=2048,
    verbose=False,
)

output = llm.create_chat_completion(
    messages=[
        {
            "role": "user",
            "content": (
                "For this ReviewDesk ticket, return JSON with category, priority, and draft_response. "
                "Ticket: 'My invoice downloaded as a blank PDF.'"
            ),
        }
    ],
    response_format={"type": "json_object"},
    temperature=0,
)["choices"][0]["message"]["content"]

print(output)</code></pre></div><p>The runtime can enforce shape, but it cannot enforce truth. That is why verification remains necessary. A visually plausible blob of text is not the same thing as valid JSON, and valid JSON is not the same thing as a correct or useful answer. Parsing is the first real test. If <code>json.loads</code> fails, the output is not ready. If parsing succeeds, the next step is to check that the expected fields are present and that their types make sense for the downstream consumer.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;45ddb9b1-2786-4191-9d00-5a0c3daac3a4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import json

parsed = json.loads(output)
print(json.dumps(parsed, indent=2))

required_keys = {"category", "priority", "draft_response"}
if not required_keys.issubset(parsed):
    raise ValueError(f"Missing keys: {required_keys - set(parsed)}")

if not isinstance(parsed["category"], str):
    raise TypeError("category must be a string")
if not isinstance(parsed["priority"], str):
    raise TypeError("priority must be a string")
if not isinstance(parsed["draft_response"], str):
    raise TypeError("draft_response must be a string")</code></pre></div><p>That small parse-and-validate step is more than a convenience. It is the border checkpoint between generated text and usable data. If the model returns a malformed object, a truncated string, or a response with commentary wrapped around the JSON, the failure mode should be classified rather than ignored. A malformed object usually calls for regeneration. A truncated object often points to token limits or overly long outputs. Commentary-laden output may be fixed by stronger instructions or by constrained decoding. In a brittle pipeline, the right response is often to retry with lower temperature, a tighter schema, or a stricter decoder. In a less brittle workflow, it may be acceptable to reject the result and fall back to a human review path.</p><p>The operational rule is simple. Shape with examples when failure is tolerable, enforce with decoding when the downstream consumer is brittle, and validate whenever the output is reused programmatically. That sequence keeps prompt engineering from becoming guesswork. It treats malformed output as a workflow incident rather than as a cosmetic annoyance, which is exactly the mindset needed when a model stops being a chatbot and starts being a component.</p><p><em>Use the url below to download the enitre book as pdf:</em></p>
      <p>
          <a href="/__u/onepagecode.substack.com/p/prompt-engineering">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Text Clustering and Topic Modeling]]></title><description><![CDATA[Chapter 5 of Hands-On Large Language Models: Leveraging dense embeddings, dimensionality reduction (UMAP), density-based clustering (HDBSCAN), and BERTopic workflows for topic discovery]]></description><link>https://onepagecode.substack.com/p/text-clustering-and-topic-modeling</link><guid isPermaLink="false">https://onepagecode.substack.com/p/text-clustering-and-topic-modeling</guid><dc:creator><![CDATA[Onepagecode]]></dc:creator><pubDate>Tue, 07 Jul 2026 19:44:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2V5g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d208012-f50a-455a-9f73-9532b2a62632_2816x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1>Why Cluster Documents Before You Label Topics?</h1><p>When a text collection gets large enough, the first problem is usually not prediction but orientation. A support team may have thousands of ReviewDesk tickets, and reading them one by one is like opening a crowded inbox after a long weekend: you can see that there is structure, but you cannot yet see the structure clearly enough to act on it. Unsupervised text clustering gives you a way to sort that inbox by similarity before you decide what the folders should be called. The goal is not to assign final categories on day one. The goal is to reveal recurring shapes in the data so that a human can inspect them with much less effort.</p><p>That distinction matters. A cluster is a discovered grouping of documents that appear to belong together; a topic label is the human-readable name you attach after you inspect that group. Those are related tasks, but they are not the same task. If you collapse them too early, you end up treating an interpretive label as if it were an objective property of the corpus. In practice, clustering is the intermediate representation that helps you get from a messy archive to a manageable set of candidate themes.</p><p>The reason embedding-based methods work so well here is that they group by meaning rather than by surface overlap. Two ReviewDesk tickets can talk about the same billing issue using very different words, and a keyword-matching approach may miss that relationship entirely. A semantic model can still place them near each other because it has learned that the underlying intent is similar, even when the vocabulary is not. That is the crucial advantage of dense representations: they make latent structure visible when the obvious words are misleading.</p><h2><strong>&#128216; Buy the Entire Book &#8211; Available on Amazon Worldwide (If you are paid subscriber, use the url at the end of this article to download the book for free!)</strong></h2><p>You can now purchase the complete book directly from Amazon in your country:</p><ul><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/us/dualbookshelf.marketplacelink/B0H6WZJ57S">United States (US)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/uk/dualbookshelf.marketplacelink/B0H6WZJ57S">United Kingdom (UK)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/de/dualbookshelf.marketplacelink/B0H6WZJ57S">Germany (DE)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/fr/dualbookshelf.marketplacelink/B0H6WZJ57S">France (FR)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/es/dualbookshelf.marketplacelink/B0H6WZJ57S">Spain (ES)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/it/dualbookshelf.marketplacelink/B0H6WZJ57S">Italy (IT)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/nl/dualbookshelf.marketplacelink/B0H6WZJ57S">Netherlands (NL)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/jp/dualbookshelf.marketplacelink/B0H6WZJ57S">Japan (JP)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/br/dualbookshelf.marketplacelink/B0H6WZJ57S">Brazil (BR):</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/ca/dualbookshelf.marketplacelink/B0H6WZJ57S">Canada (CA)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/mx/dualbookshelf.marketplacelink/B0H6WZJ57S">Mexico (MX)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/au/dualbookshelf.marketplacelink/B0H6WZJ57S">Australia (AU)</a></strong></p></li><li><p><strong><a href="https://kdp.amazon.com/amazon-dp-action/in/dualbookshelf.marketplacelink/B0H6WZJ57S">India (IN)</a></strong></p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!2V5g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d208012-f50a-455a-9f73-9532b2a62632_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!2V5g!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d208012-f50a-455a-9f73-9532b2a62632_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!2V5g!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d208012-f50a-455a-9f73-9532b2a62632_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!2V5g!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d208012-f50a-455a-9f73-9532b2a62632_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!2V5g!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d208012-f50a-455a-9f73-9532b2a62632_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!2V5g!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d208012-f50a-455a-9f73-9532b2a62632_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4d208012-f50a-455a-9f73-9532b2a62632_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5434419,&quot;alt&quot;:&quot;Figure 5.1: Dense embeddings can place documents near each other by meaning, even when the wording is different, making latent structure visible to clustering algorithms.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204106410?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d208012-f50a-455a-9f73-9532b2a62632_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 5.1: Dense embeddings can place documents near each other by meaning, even when the wording is different, making latent structure visible to clustering algorithms." title="Figure 5.1: Dense embeddings can place documents near each other by meaning, even when the wording is different, making latent structure visible to clustering algorithms." srcset="/__u/substackcdn.com/image/fetch/$s_!2V5g!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d208012-f50a-455a-9f73-9532b2a62632_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!2V5g!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d208012-f50a-455a-9f73-9532b2a62632_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!2V5g!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d208012-f50a-455a-9f73-9532b2a62632_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!2V5g!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d208012-f50a-455a-9f73-9532b2a62632_2816x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This is especially useful in unlabeled corpora, where classification is not available or not worth the cost. In many teams, labels are incomplete, stale, or too coarse to explain the patterns hidden in the archive. Clustering gives you a fast way to explore the shape of the dataset before you commit to a taxonomy. In a support setting, that might surface separate piles for product bugs, billing questions, feature requests, and operational issues. In a media workflow, the same idea could expose groups of similar assets, recurring content briefs, or unusual one-off submissions that deserve special handling.</p><p>Outliers are part of the value, not a nuisance to be ignored. A density-based method may leave some documents unassigned because they do not fit any sufficiently coherent group. That is often exactly what you want. The odd ticket that references two unrelated problems, the duplicate message copied from another channel, or the rare edge case from a premium customer can be more informative than a clean cluster. These documents tell you where the corpus is noisy, where workflows are messy, and where your categories may need revision.</p><p>Once you have clusters, topic modeling becomes a second layer of interpretation rather than a replacement for discovery. The cluster tells you that a set of documents belong together; the topic label tells you what to call that set in a way a person can read quickly. In that sense, topic modeling is like writing captions for a photo album after the photos have already been grouped. The grouping comes first because it reduces the search space. The naming comes later because it is easier to label a coherent pile than to assign meaning to thousands of isolated documents.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!bzJk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5fc44823-aabd-46b8-ab41-7348b08d48cc_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!bzJk!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5fc44823-aabd-46b8-ab41-7348b08d48cc_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!bzJk!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5fc44823-aabd-46b8-ab41-7348b08d48cc_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!bzJk!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5fc44823-aabd-46b8-ab41-7348b08d48cc_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!bzJk!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5fc44823-aabd-46b8-ab41-7348b08d48cc_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!bzJk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5fc44823-aabd-46b8-ab41-7348b08d48cc_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5fc44823-aabd-46b8-ab41-7348b08d48cc_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5610316,&quot;alt&quot;:&quot;Figure 5.2: Topic labeling is a later interpretive step: clustering first discovers coherent document groups, then humans name those groups after inspection.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204106410?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5fc44823-aabd-46b8-ab41-7348b08d48cc_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 5.2: Topic labeling is a later interpretive step: clustering first discovers coherent document groups, then humans name those groups after inspection." title="Figure 5.2: Topic labeling is a later interpretive step: clustering first discovers coherent document groups, then humans name those groups after inspection." srcset="/__u/substackcdn.com/image/fetch/$s_!bzJk!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5fc44823-aabd-46b8-ab41-7348b08d48cc_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!bzJk!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5fc44823-aabd-46b8-ab41-7348b08d48cc_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!bzJk!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5fc44823-aabd-46b8-ab41-7348b08d48cc_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!bzJk!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5fc44823-aabd-46b8-ab41-7348b08d48cc_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>That is the workflow this chapter develops. We will start with embeddings, use them to group documents semantically, and then refine those groups into readable topics. The later BERTopic-style steps will add keyword summaries, representation methods, and labeling tools, but their usefulness depends on this first move: discovering clusters before pretending to know their names.</p><h2>Anchor the Running Corpus</h2><p>The chapter&#8217;s shared corpus is the ArXiv cs.CL collection, a convenient stand-in for a real document stream that already contains a broad spread of computation and language research abstracts. Treat it as the chapter&#8217;s working file cabinet: the documents are already collected, but they still need to be arranged into a clean structure before any semantic grouping can happen. That structure matters because clustering does not begin with labels; it begins with a corpus that can be read consistently, inspected later, and passed unchanged through the next stages of the pipeline.</p><h2>Load the Dataset from Hugging Face</h2><p>The starting point is intentionally simple. The dataset is fetched from Hugging Face, and the chapter uses the training split as the source corpus. That gives us a reproducible pool of documents without needing to assemble the collection by hand. In practice, this kind of loading step is the bridge between a hosted dataset and an analysis-ready document table.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Tbeu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90a1e6c2-8381-422f-8c84-ad29c8a5afb7_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Tbeu!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90a1e6c2-8381-422f-8c84-ad29c8a5afb7_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!Tbeu!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90a1e6c2-8381-422f-8c84-ad29c8a5afb7_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!Tbeu!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90a1e6c2-8381-422f-8c84-ad29c8a5afb7_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Tbeu!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90a1e6c2-8381-422f-8c84-ad29c8a5afb7_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Tbeu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90a1e6c2-8381-422f-8c84-ad29c8a5afb7_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/90a1e6c2-8381-422f-8c84-ad29c8a5afb7_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:6745411,&quot;alt&quot;:&quot;Figure 5.3: Loading a hosted dataset and separating abstracts from titles produces an analysis-ready corpus table for later embedding and clustering.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204106410?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90a1e6c2-8381-422f-8c84-ad29c8a5afb7_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 5.3: Loading a hosted dataset and separating abstracts from titles produces an analysis-ready corpus table for later embedding and clustering." title="Figure 5.3: Loading a hosted dataset and separating abstracts from titles produces an analysis-ready corpus table for later embedding and clustering." srcset="/__u/substackcdn.com/image/fetch/$s_!Tbeu!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90a1e6c2-8381-422f-8c84-ad29c8a5afb7_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!Tbeu!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90a1e6c2-8381-422f-8c84-ad29c8a5afb7_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!Tbeu!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90a1e6c2-8381-422f-8c84-ad29c8a5afb7_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Tbeu!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90a1e6c2-8381-422f-8c84-ad29c8a5afb7_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b48f2647-328f-4964-b9be-c2af8d65606e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from datasets import load_dataset

dataset = load_dataset("maartengr/arxiv_nlp")["train"]

abstracts = dataset["Abstracts"]
titles = dataset["Titles"]</code></pre></div><p>The code does not do any modeling yet. Its job is to make the corpus explicit. If the public dataset is unavailable in another setting, the same pattern still applies to a domain corpus such as ReviewDesk support tickets: load the source table, choose the slice you want to analyze, and separate the text you want to embed from the metadata you want to keep for interpretation.</p><h2>Separate Text from Metadata</h2><p>Here the abstracts are the primary documents. They are the text that will later be embedded, compared, and grouped. Titles play a different role. They are lightweight metadata that help you inspect a cluster, understand a document at a glance, and later assign or evaluate a topic label. Keeping those two fields separate is a small design choice with large downstream consequences. If you merge the title into the abstract too early, you blur the boundary between content and annotation, and later inspection becomes noisier than it needs to be.</p><p>A useful way to think about it is that the abstract is the folder contents, while the title is the tab on the folder. Clustering works on the contents; the tab helps you understand what ended up together after the fact. That separation also makes error analysis easier, because you can ask whether a document landed in the right group without having to untangle extra metadata from the text itself.</p><h2>Sanity-Check the Corpus Shape</h2><p>Before moving on, it is worth confirming that the corpus is usable as a one-document-per-row collection. The abstract and title arrays should stay aligned, and missing entries should be understood before they propagate into later steps. Even when the dataset is already curated, this is the point where a quick structural check pays off: if the document table is misaligned now, every cluster inspection afterward will inherit the mistake.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!jg7-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fa064a-b9dc-4f4d-be2a-248fba4f580a_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!jg7-!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fa064a-b9dc-4f4d-be2a-248fba4f580a_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!jg7-!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fa064a-b9dc-4f4d-be2a-248fba4f580a_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!jg7-!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fa064a-b9dc-4f4d-be2a-248fba4f580a_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!jg7-!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fa064a-b9dc-4f4d-be2a-248fba4f580a_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!jg7-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fa064a-b9dc-4f4d-be2a-248fba4f580a_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d2fa064a-b9dc-4f4d-be2a-248fba4f580a_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:6864659,&quot;alt&quot;:&quot;Figure 5.4: A quick alignment and missing-value check prevents corpus-table errors from propagating into downstream clustering and inspection.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204106410?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fa064a-b9dc-4f4d-be2a-248fba4f580a_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 5.4: A quick alignment and missing-value check prevents corpus-table errors from propagating into downstream clustering and inspection." title="Figure 5.4: A quick alignment and missing-value check prevents corpus-table errors from propagating into downstream clustering and inspection." srcset="/__u/substackcdn.com/image/fetch/$s_!jg7-!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fa064a-b9dc-4f4d-be2a-248fba4f580a_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!jg7-!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fa064a-b9dc-4f4d-be2a-248fba4f580a_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!jg7-!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fa064a-b9dc-4f4d-be2a-248fba4f580a_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!jg7-!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fa064a-b9dc-4f4d-be2a-248fba4f580a_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Why This Structure Supports Later Clustering</h2><p>This preparation step does not yet cluster anything, but it sets up the exact input shape that clustering expects: a clean collection of documents with stable metadata attached beside them. That makes the later embedding stage straightforward, and it gives you a reliable way to read the results once groups begin to emerge. In unsupervised text analysis, the quality of the final clusters is only part of the story; the quality of the document table determines whether those clusters can be interpreted with confidence.</p><h2>From Embeddings to Clusters and Visual Maps</h2><h3>Why embeddings are the right starting point for text groups</h3><p>If you want to sort a messy inbox, you do not begin by counting shared words and hoping the piles organize themselves. That kind of sparse lexical view is useful for some problems, but it is a weak way to group whole documents by meaning. Two support tickets can describe the same issue with different product names and different wording, yet still belong together. Semantic embeddings handle that better because they turn each document into a dense vector where related texts end up near one another.</p><p>That is why this workflow starts with embeddings rather than bag-of-words or TF-IDF. The goal is not to preserve exact vocabulary. The goal is to preserve enough semantic neighborhood structure that similar documents can be grouped even when surface wording differs. In practice, this makes a mixed corpus easier to sort before you assign human-readable topic names.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!4htA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57573f8a-dfa3-4452-b869-34e1b17020b0_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!4htA!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57573f8a-dfa3-4452-b869-34e1b17020b0_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!4htA!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57573f8a-dfa3-4452-b869-34e1b17020b0_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!4htA!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57573f8a-dfa3-4452-b869-34e1b17020b0_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!4htA!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57573f8a-dfa3-4452-b869-34e1b17020b0_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!4htA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57573f8a-dfa3-4452-b869-34e1b17020b0_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/57573f8a-dfa3-4452-b869-34e1b17020b0_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5741465,&quot;alt&quot;:&quot;Figure 5.5: Dense semantic embeddings preserve neighborhood structure, making it possible to group documents by meaning even when they share few surface words.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204106410?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57573f8a-dfa3-4452-b869-34e1b17020b0_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 5.5: Dense semantic embeddings preserve neighborhood structure, making it possible to group documents by meaning even when they share few surface words." title="Figure 5.5: Dense semantic embeddings preserve neighborhood structure, making it possible to group documents by meaning even when they share few surface words." srcset="/__u/substackcdn.com/image/fetch/$s_!4htA!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57573f8a-dfa3-4452-b869-34e1b17020b0_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!4htA!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57573f8a-dfa3-4452-b869-34e1b17020b0_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!4htA!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57573f8a-dfa3-4452-b869-34e1b17020b0_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!4htA!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57573f8a-dfa3-4452-b869-34e1b17020b0_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Think of embeddings as a smarter sorting surface. The model has not named the folders yet, but it has already arranged the documents so likely neighbors sit close together. That is the foundation for the rest of the pipeline.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a027ed44-03bc-45c7-9133-c7a2096ec9dc&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from sentence_transformers import SentenceTransformer

embedding_model = SentenceTransformer("thenlper/gte-small")
embeddings = embedding_model.encode(titles_and_texts, show_progress_bar=True)</code></pre></div><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;fbfe2116-cb39-4fbb-95cf-f3786c25b848&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">embeddings.shape</code></pre></div><h3>From high-dimensional vectors to a clusterable representation</h3><p>The embedding matrix is still high-dimensional, which is good for representation but awkward for clustering. Many clustering methods become less effective when the space is wide and sparse, especially when the useful structure lives in local neighborhoods rather than across every dimension. Reducing the vectors first makes the geometry more tractable.</p><p>This reduction has two distinct jobs in the workflow. One reduced space is used as the working representation for clustering. A separate two-dimensional projection is used later for plotting. The first helps the algorithm organize the documents. The second helps you inspect the result. Those are related, but they are not the same step.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!fvVH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf1b3c2a-0f21-459d-938b-cd2bbfce9fc0_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!fvVH!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf1b3c2a-0f21-459d-938b-cd2bbfce9fc0_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!fvVH!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf1b3c2a-0f21-459d-938b-cd2bbfce9fc0_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!fvVH!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf1b3c2a-0f21-459d-938b-cd2bbfce9fc0_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fvVH!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf1b3c2a-0f21-459d-938b-cd2bbfce9fc0_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!fvVH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf1b3c2a-0f21-459d-938b-cd2bbfce9fc0_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bf1b3c2a-0f21-459d-938b-cd2bbfce9fc0_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7166513,&quot;alt&quot;:&quot;Figure 5.6: Dimensionality reduction serves two different roles here: a reduced working space for clustering and a separate 2D projection for visualization.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204106410?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf1b3c2a-0f21-459d-938b-cd2bbfce9fc0_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 5.6: Dimensionality reduction serves two different roles here: a reduced working space for clustering and a separate 2D projection for visualization." title="Figure 5.6: Dimensionality reduction serves two different roles here: a reduced working space for clustering and a separate 2D projection for visualization." srcset="/__u/substackcdn.com/image/fetch/$s_!fvVH!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf1b3c2a-0f21-459d-938b-cd2bbfce9fc0_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!fvVH!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf1b3c2a-0f21-459d-938b-cd2bbfce9fc0_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!fvVH!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf1b3c2a-0f21-459d-938b-cd2bbfce9fc0_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!fvVH!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf1b3c2a-0f21-459d-938b-cd2bbfce9fc0_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The distance measure matters too. Cosine distance fits embedding-based text similarity well because semantic embeddings are usually interpreted by direction rather than raw magnitude. If two documents point in a similar semantic direction, they should be treated as close even if their lengths differ or their vector norms do not match exactly.</p><p>A useful way to picture this is a working geometry, not a true map. The reducer is trying to preserve the neighborhoods that matter for grouping, not to reveal some perfect universal coordinate system for language.</p><h3>Using UMAP to preserve neighborhood structure</h3><p>UMAP is a practical reducer for this task because it aims to preserve local neighborhoods while compressing the data into fewer dimensions. Documents that are close in the original embedding space should remain close after reduction, at least well enough for clustering to use the result. That makes UMAP useful both as a preparatory step for clustering and later as a projection for display.</p><p>The main parameters here are easy to read in conceptual terms. <code>n_components</code> sets the size of the reduced working space. A small value such as five often gives a clustering algorithm enough room to separate groups without keeping the full embedding width. <code>metric</code> is set to <code>cosine</code> so the reduction respects embedding geometry. <code>min_dist</code> controls how tightly neighboring points may pack together, and a small value encourages compact local structure. <code>random_state</code> fixes the stochastic parts of the process so repeated runs are easier to compare.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b2dffb8e-7f9b-423f-b684-0d76555fd028&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from umap import UMAP

umap_model = UMAP(
    n_components=5,
    min_dist=0.0,
    metric="cosine",
    random_state=42,
)
reduced_embeddings = umap_model.fit_transform(embeddings)</code></pre></div><p>That reproducibility matters because cluster results are not absolute facts. They depend on the model, the reducer, and the seed. If the seed changes, the exact arrangement can change too. Holding it fixed makes interpretation much easier when you are trying to understand whether a group is stable or just an artifact of a particular run.</p><h3>Letting HDBSCAN discover groups and reject noise</h3><p>Once the embeddings have been reduced, HDBSCAN can discover clusters without requiring you to choose <code>k</code> in advance. That is one reason it works well for open-ended text corpora. You do not always know how many themes are present, and forcing a fixed cluster count can be awkward when the corpus is messy or uneven.</p><p>HDBSCAN looks for dense regions in the reduced space. Where points form a compact neighborhood, it assigns a cluster. Where points are sparse, ambiguous, or transitional, it can leave them unassigned. That unassigned behavior is not a bug. It is often the most useful part of the method, because real corpora usually contain items that do not belong cleanly anywhere.</p><p>The label <code>-1</code> marks those noise points. In practice, that means borderline documents, vague requests, or one-off items can be separated from the main clusters instead of being forced into a misleading group. If you expect messy real-world text, this is a valuable feature, not a failure.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e342df11-3e6f-4abe-8e32-968a95be3f79&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from hdbscan import HDBSCAN

hdbscan_model = HDBSCAN(
    min_cluster_size=50,
    metric="euclidean",
    cluster_selection_method="eom",
).fit(reduced_embeddings)
clusters = hdbscan_model.labels_

len(set(clusters))</code></pre></div><p>The exact cluster labels are parameter-dependent. Change the embedding model, the UMAP settings, or the HDBSCAN thresholds, and the cluster assignment can shift. That is normal. The point is not to treat one run as permanent truth. The point is to get a usable grouping that fits the corpus and the task.</p><h3>Checking cluster meaning by reading sample documents</h3><p>A cluster becomes meaningful only after you read some of the documents inside it. The simplest validation habit is to sample a few members from one cluster and ask whether they share a recognizable theme. In a support-ticket corpus, that might reveal a cluster about login failures, another about billing corrections, or another about export requests.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;55f28549-6e91-448f-8582-1675f4389d0e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import numpy as np

cluster = 0
for index in np.where(clusters == cluster)[0][:3]:
    print(titles_and_texts[index][:300] + "...\n")</code></pre></div><p>This step catches an important failure mode. A cluster can look neat numerically and still mix several subthemes that happen to share vocabulary. Human reading is what tells you whether the cluster is genuinely coherent or merely compact in reduced space. Numerical grouping gives you structure, but human inspection gives you confidence.</p><h3>Projecting into two dimensions for diagnosis, not proof</h3><p>After clustering, you can make a two-dimensional projection for plotting. This is a diagnostic view, not the clustering step itself. It helps you spot broad structure, noise points, and overlap, but it is not a ground-truth map of meaning. A nice scatterplot can be misleading if you treat it as proof that the clusters are perfectly pure.</p><p>A practical plotting convention is to keep outliers separate from clustered points. Noise points are often drawn in gray so they recede visually, while clustered points are colored by cluster ID. The axes are usually hidden because the exact coordinates matter less than the relative arrangement.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;5b997b2f-ff66-4712-8d11-6da8251659f8&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import pandas as pd
from umap import UMAP

plot_embeddings = UMAP(
    n_components=2,
    min_dist=0.0,
    metric="cosine",
    random_state=42,
).fit_transform(embeddings)

df = pd.DataFrame(plot_embeddings, columns=["x", "y"])
df["title"] = titles
df["cluster"] = [str(c) for c in clusters]

to_plot = df.loc[df.cluster != "-1", :]
outliers = df.loc[df.cluster == "-1", :]</code></pre></div><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6e179bb7-1063-4044-93bf-de5ebcf43cd3&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import matplotlib.pyplot as plt

plt.scatter(outliers.x, outliers.y, alpha=0.05, s=2, c="grey")
plt.scatter(
    to_plot.x,
    to_plot.y,
    c=to_plot.cluster.astype(int),
    alpha=0.6,
    s=2,
    cmap="tab20b",
)
plt.axis("off")</code></pre></div><p>Read that plot as an X-ray, not as a map. It can show where clusters sit, where outliers collect, and where the method may be uncertain. It cannot prove semantic purity. Visual separation and semantic purity are related, but they are not identical. Two clusters can look far apart and still share content, and two clusters can sit near each other while still representing distinct themes.</p><p>This is also where the workflow connects to topic modeling. Clustering has organized the corpus into semantically related groups, but those groups still need readable topic descriptions and representative terms. The next section builds on that structure rather than starting over from scratch.</p><h2>From Cluster IDs to Topic Representations</h2><p>Once you have clustered document embeddings, you have already found something useful, but the result is still unnamed. Clustering tells you which documents belong together; topic modeling asks what those groups are about. BERTopic treats those as separate layers. It can reuse the embeddings, UMAP reduction, and HDBSCAN clusters from the previous step, then build a topic representation on top of them. The cluster assignment is the backbone, and the topic name is the label that comes later.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!3vBC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F108e2575-e3ab-4198-962a-0bc197728f10_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!3vBC!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F108e2575-e3ab-4198-962a-0bc197728f10_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!3vBC!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F108e2575-e3ab-4198-962a-0bc197728f10_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!3vBC!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F108e2575-e3ab-4198-962a-0bc197728f10_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!3vBC!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F108e2575-e3ab-4198-962a-0bc197728f10_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!3vBC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F108e2575-e3ab-4198-962a-0bc197728f10_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/108e2575-e3ab-4198-962a-0bc197728f10_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:6374898,&quot;alt&quot;:&quot;Figure 5.7: BERTopic keeps clustering and topic naming separate: embeddings and clusters are reused, while topic labels are built and refined on top of the fixed cluster assignments.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204106410?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F108e2575-e3ab-4198-962a-0bc197728f10_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 5.7: BERTopic keeps clustering and topic naming separate: embeddings and clusters are reused, while topic labels are built and refined on top of the fixed cluster assignments." title="Figure 5.7: BERTopic keeps clustering and topic naming separate: embeddings and clusters are reused, while topic labels are built and refined on top of the fixed cluster assignments." srcset="/__u/substackcdn.com/image/fetch/$s_!3vBC!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F108e2575-e3ab-4198-962a-0bc197728f10_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!3vBC!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F108e2575-e3ab-4198-962a-0bc197728f10_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!3vBC!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F108e2575-e3ab-4198-962a-0bc197728f10_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!3vBC!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F108e2575-e3ab-4198-962a-0bc197728f10_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>That separation is what makes BERTopic modular. You can keep the same document groups while changing only how they are described. The inbox analogy is a good one: clustering sorts the messages into folders, and representation models rename the folders after you look inside. In ReviewDesk, one folder might collect billing failures, another mobile crashes, and another feature requests. The groups already exist before they are named, and BERTopic gives you a structured way to turn those groups into searchable topics.</p><p>Outliers carry over from the clustering step as well. If HDBSCAN leaves some documents unassigned, BERTopic preserves that behavior with topic <code>-1</code> instead of forcing every document into a topic that may not fit. That is useful when the corpus contains short, ambiguous, or genuinely unusual documents.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;febbd642-4dcf-4083-ab83-90c6aa6cdaba&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from bertopic import BERTopic

# Train our model with our previously defined models
topic_model = BERTopic(
    embedding_model=embedding_model,
    umap_model=umap_model,
    hdbscan_model=hdbscan_model,
    verbose=True
).fit(abstracts, embeddings)</code></pre></div><h2>How c-TF-IDF Builds the First Topic Keywords</h2><p>BERTopic turns each cluster into an initial keyword description with c-TF-IDF, or class-based TF-IDF. The key idea is to pool all documents in a cluster and treat that pool as one class-level text. The model then scores words by asking a simple question: what appears often in this folder, and how distinctive is it compared with the other folders?</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!7AkB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75f6ff4b-985b-41b4-99fb-8a6929d40b61_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!7AkB!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75f6ff4b-985b-41b4-99fb-8a6929d40b61_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!7AkB!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75f6ff4b-985b-41b4-99fb-8a6929d40b61_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!7AkB!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75f6ff4b-985b-41b4-99fb-8a6929d40b61_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!7AkB!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75f6ff4b-985b-41b4-99fb-8a6929d40b61_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!7AkB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75f6ff4b-985b-41b4-99fb-8a6929d40b61_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/75f6ff4b-985b-41b4-99fb-8a6929d40b61_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7068293,&quot;alt&quot;:&quot;Figure 5.8: c-TF-IDF pools documents within each cluster and ranks words by how common they are inside that cluster and how rare they are across the others.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204106410?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75f6ff4b-985b-41b4-99fb-8a6929d40b61_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 5.8: c-TF-IDF pools documents within each cluster and ranks words by how common they are inside that cluster and how rare they are across the others." title="Figure 5.8: c-TF-IDF pools documents within each cluster and ranks words by how common they are inside that cluster and how rare they are across the others." srcset="/__u/substackcdn.com/image/fetch/$s_!7AkB!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75f6ff4b-985b-41b4-99fb-8a6929d40b61_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!7AkB!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75f6ff4b-985b-41b4-99fb-8a6929d40b61_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!7AkB!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75f6ff4b-985b-41b4-99fb-8a6929d40b61_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!7AkB!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75f6ff4b-985b-41b4-99fb-8a6929d40b61_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>That is the important difference from ordinary document-level TF-IDF. Standard TF-IDF works on individual documents. c-TF-IDF shifts the unit of analysis to the cluster, so the words are chosen to represent the group as a whole rather than any single document. In practice, this is what makes the topic keywords feel topic-like instead of document-like.</p><p>A compact version of the scoring idea is</p><p>c-TF-IDF(w,c)=tf(w,c)&#8901;log&#8289;(Ndf(w)),</p><p>where w is a word, c is a cluster, tf(w,c) measures how often the word appears in the pooled cluster text, and df(w) counts how many clusters contain that word. The formula is only a shorthand for the intuition: topic words should be common inside one cluster and relatively uncommon across the others.</p><p>This class-level weighting also explains why topic <code>-1</code> can appear in the output. If some documents remain outliers, their pooled text becomes its own class-level representation. The resulting words are a reminder that density-based clustering can preserve uncertainty instead of inventing a confident label where none is justified.</p><h2>Working with the Topic API in Practice</h2><p>After fitting the model, the fastest way to understand what it found is to inspect the topic summary table. This shows topic IDs, document counts, and the first pass at representative keywords. It is a quick audit of whether the clusters look meaningful and whether the outlier topic is large enough to deserve attention.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;1882c03d-3953-46bb-b9b9-3565cb4e6228&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">topic_model.get_topic_info()</code></pre></div><p>To inspect one topic in more detail, you can ask for its ranked keywords directly. That is often the most practical way to verify whether a cluster named by c-TF-IDF really corresponds to the subject you expect.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f3504386-430b-41a8-91ea-372b9218c516&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">topic_model.get_topic(0)</code></pre></div><p>You can also look up the topic assigned to a specific document. In a support workflow, that is how you check whether a ticket about payment failures, app crashes, or account access landed in the right group.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f0e606fd-2940-4fcc-99c0-ae776edbaffd&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">topic_model.topics_[titles.index("BERTopic: Neural topic modeling with a class-based TF-IDF procedure")]</code></pre></div><p>Query-based search works a little differently. It asks which discovered topic is most similar to a text prompt, such as <code>"topic modeling"</code>. This is topic retrieval, not corpus search. You are not finding documents that contain the phrase; you are asking which folder in the learned topic space best matches the phrase.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;6dc1e9bb-f76e-4db7-882f-fa69c03fa5c4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">topic_model.find_topics("topic modeling")</code></pre></div><h2>Improving Topic Word Lists Without Reclustering</h2><p>BERTopic keeps cluster membership separate from topic wording, which means you can refine the labels without changing the groups themselves. That is a major advantage of the framework. The same cluster can be named in several different ways, and those names can be updated without rerunning the whole embedding and clustering pipeline.</p><p>The baseline topic words come from c-TF-IDF, but representation models can rerank or diversify them afterward. This is the post-processing layer. It is cheaper than retraining and often produces clearer labels, although not always more faithful ones. Cleaner wording is useful, but it can also flatten nuance if a reranker removes a distinctive term.</p><p>A small helper makes the before-and-after comparison easier to inspect.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;53760b56-42fe-491d-847f-ec5079d351ee&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from copy import deepcopy
original_topics = deepcopy(topic_model.topic_representations_)

def topic_differences(model, original_topics, nr_topics=5):
    df = pd.DataFrame(columns=["Topic", "Original", "Updated"])
    for topic in range(nr_topics):
        og_words = " | ".join(list(zip(*original_topics[topic]))[0][:5])
        new_words = " | ".join(list(zip(*model.get_topic(topic)))[0][:5])
        df.loc[len(df)] = [topic, og_words, new_words]
    return df</code></pre></div><p><code>KeyBERTInspired</code> is a similarity-focused reranker. It prefers words that sit closer to the semantic center of the topic, so the resulting keyword list often reads more coherently than raw c-TF-IDF output.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f460b2b7-133c-441a-aabb-f58d2ec6b3f4&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from bertopic.representation import KeyBERTInspired

representation_model = KeyBERTInspired()
topic_model.update_topics(abstracts, representation_model=representation_model)
topic_differences(topic_model, original_topics)</code></pre></div><p><code>MaximalMarginalRelevance</code>, or MMR, pushes in a different direction. It still wants relevant words, but it also tries to avoid redundancy. If KeyBERTInspired looks for the best-fitting tracks on a playlist, MMR tries to keep the playlist from sounding repetitive. The benefit is a more varied label; the limitation is that too much diversification can hide useful detail.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d8581e76-408c-4f9b-94af-9d8fd6259119&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from bertopic.representation import MaximalMarginalRelevance

representation_model = MaximalMarginalRelevance(diversity=0.2)
topic_model.update_topics(abstracts, representation_model=representation_model)
topic_differences(topic_model, original_topics)</code></pre></div><p>In practice, reranking is best seen as label editing rather than topic discovery. It can make a topic easier to read, but it does not change which documents belong to that topic. That is why the cluster boundaries stay fixed while the sign on the folder changes.</p><h2>Generating Human-Readable Labels with LLMs</h2><p>Representation models can also use text generation. In that setup, BERTopic gives the model representative documents and topic keywords, then asks it to produce a short label. This feels closer to summarizing a cluster than to training a new model. The prompt acts like a captioning template for the topic.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;b20397d6-00d6-4ecd-9d37-d08136738d55&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from transformers import pipeline
from bertopic.representation import TextGeneration

prompt = """I have a topic that contains the following documents: 
[DOCUMENTS]

The topic is described by the following keywords: '[KEYWORDS]'.

Based on the documents and keywords, what is this topic about?"""

generator = pipeline("text2text-generation", model="google/flan-t5-small")
representation_model = TextGeneration(
    generator, prompt=prompt, doc_length=50, tokenizer="whitespace"
)
topic_model.update_topics(abstracts, representation_model=representation_model)
topic_differences(topic_model, original_topics)</code></pre></div><p>A chat model can do the same job with a slightly different prompt format.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;41813690-0906-4f80-9d07-bae0d7059fe5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import openai
from bertopic.representation import OpenAI

prompt = """
I have a topic that contains the following documents:
[DOCUMENTS]

The topic is described by the following keywords: [KEYWORDS]

Based on the information above, extract a short topic label in the following format:
topic: &lt;short topic label&gt;
"""

client = openai.OpenAI(api_key="YOUR_KEY_HERE")
representation_model = OpenAI(
    client, model="gpt-3.5-turbo", exponential_backoff=True, chat=True, prompt=prompt
)
topic_model.update_topics(abstracts, representation_model=representation_model)
topic_differences(topic_model, original_topics)</code></pre></div><p>Generated labels are flexible, but that flexibility comes with caveats. They can be broad, corpus-dependent, and slightly different across runs or models. A label like <code>Billing Issues</code> may be fluent, yet a more specific label such as <code>Refund Delays and Invoice Problems</code> may be more useful for support triage. The generated name should therefore be treated as a suggestion that still needs human review.</p><h2>Visualizing Topics as an Inspection Tool</h2><p>Visualization helps you audit the learned topic structure, but it does not prove that a topic is correct. A document map can show whether documents within a topic are close together, whether two topics overlap, or whether outliers sit far from the main groups. A keyword bar chart can show which words dominate a topic, and a heatmap or hierarchy can reveal similarity between topics. These plots are diagnostic tools, not final evidence.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f2d98348-d8c1-45b2-a546-6cd159477247&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">fig = topic_model.visualize_documents(
    titles, 
    reduced_embeddings=reduced_embeddings, 
    width=1200, 
    hide_annotations=True
)

fig.update_layout(font=dict(size=16))

topic_model.visualize_barchart()
topic_model.visualize_heatmap(n_clusters=30)
topic_model.visualize_hierarchy()</code></pre></div><p>A datamap-style view is another way to inspect the same structure at a glance.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;f1975735-77c4-4c38-907e-7a734370591f&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">fig = topic_model.visualize_document_datamap(
    titles,
    topics=list(range(20)),
    reduced_embeddings=reduced_embeddings,
    width=1200,
    label_font_size=11,
    label_wrap_width=20,
    use_medoids=True,
)</code></pre></div><p>The practical value of these visualizations is that they help you ask better questions. Do similar topics sit near each other? Are the largest topics internally consistent? Do outliers suggest noise or a missing theme? In that sense, the plots support interpretation, but the topic meaning still comes from the cluster structure and the words chosen to describe it.</p><h2>Recap: Choosing and Interpreting Topic Workflows</h2><p>The chapter&#8217;s workflow is best remembered as a translation pipeline: embeddings capture semantic proximity, dimensionality reduction makes that structure easier to work with, clustering discovers document groups, and topic representation turns those groups into language a person can inspect. That sequence matters because each stage solves a different problem. The geometry in embedding space tells you which documents belong together; it does not yet tell you what to call the cluster or whether the grouping deserves to be treated as a coherent theme.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!lZ4r!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F000e8dbd-9845-419f-a208-743d432b8224_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!lZ4r!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F000e8dbd-9845-419f-a208-743d432b8224_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!lZ4r!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F000e8dbd-9845-419f-a208-743d432b8224_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!lZ4r!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F000e8dbd-9845-419f-a208-743d432b8224_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!lZ4r!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F000e8dbd-9845-419f-a208-743d432b8224_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!lZ4r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F000e8dbd-9845-419f-a208-743d432b8224_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/000e8dbd-9845-419f-a208-743d432b8224_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5892497,&quot;alt&quot;:&quot;Figure 5.9: Topic workflows are a translation pipeline: embeddings and clustering discover structure, while topic representation converts that structure into human-readable labels.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204106410?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F000e8dbd-9845-419f-a208-743d432b8224_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 5.9: Topic workflows are a translation pipeline: embeddings and clustering discover structure, while topic representation converts that structure into human-readable labels." title="Figure 5.9: Topic workflows are a translation pipeline: embeddings and clustering discover structure, while topic representation converts that structure into human-readable labels." srcset="/__u/substackcdn.com/image/fetch/$s_!lZ4r!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F000e8dbd-9845-419f-a208-743d432b8224_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!lZ4r!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F000e8dbd-9845-419f-a208-743d432b8224_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!lZ4r!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F000e8dbd-9845-419f-a208-743d432b8224_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!lZ4r!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F000e8dbd-9845-419f-a208-743d432b8224_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>That separation between discovery and interpretation is the central idea to carry forward. Clustering is like sorting a crowded inbox into piles by similarity. Topic work is what writes the folder tabs. The piles can be useful even before they are named, but the names are what make the output communicate well enough for analysis, reporting, or downstream workflows. In practice, a dense cluster may correspond to a real theme, a loose mixture of related complaints, or simply a repeated phrasing pattern. Outlier handling is part of that reality, not a sign that the method failed.</p><p>BERTopic sits on top of this discovery process rather than replacing it. It reuses the clustered documents and then applies c-TF-IDF to surface the terms that best characterize each group. That gives the first readable version of a topic. Representation models then refine that view by changing how the cluster is summarized, which can improve keyword quality, reduce repetition, or emphasize more informative phrases. Methods such as maximal marginal relevance and KeyBERTInspired are useful here because they help topic descriptions become less redundant and more legible without reclustering the corpus.</p><p>Generative topic labels add another layer of usefulness, especially when the keyword view is too sparse or too technical. But these labels should be treated as assistive summaries, not authoritative truth. A model can produce a crisp label that sounds plausible while still flattening an important distinction in the documents underneath. For that reason, manual inspection remains the final check before a topic is used operationally. A good practical rule is simple: if you only need exploration, raw clusters may be enough; if you need communication, topic labels help; if you need accountability, human review is still required.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!_j7P!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4478358-a9d4-41ec-82f7-8094a6d9b76e_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!_j7P!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4478358-a9d4-41ec-82f7-8094a6d9b76e_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!_j7P!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4478358-a9d4-41ec-82f7-8094a6d9b76e_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!_j7P!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4478358-a9d4-41ec-82f7-8094a6d9b76e_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!_j7P!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_webp, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4478358-a9d4-41ec-82f7-8094a6d9b76e_2816x1536.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!_j7P!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4478358-a9d4-41ec-82f7-8094a6d9b76e_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b4478358-a9d4-41ec-82f7-8094a6d9b76e_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7265244,&quot;alt&quot;:&quot;Figure 5.10: The right interpretation layer depends on the use case: exploration can stop at clusters, communication benefits from labels, and operational decisions require human validation.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://onepagecode.substack.com/i/204106410?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4478358-a9d4-41ec-82f7-8094a6d9b76e_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 5.10: The right interpretation layer depends on the use case: exploration can stop at clusters, communication benefits from labels, and operational decisions require human validation." title="Figure 5.10: The right interpretation layer depends on the use case: exploration can stop at clusters, communication benefits from labels, and operational decisions require human validation." srcset="/__u/substackcdn.com/image/fetch/$s_!_j7P!, /__u/onepagecode.substack.com/w_424, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4478358-a9d4-41ec-82f7-8094a6d9b76e_2816x1536.png 424w, /__u/substackcdn.com/image/fetch/$s_!_j7P!, /__u/onepagecode.substack.com/w_848, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4478358-a9d4-41ec-82f7-8094a6d9b76e_2816x1536.png 848w, /__u/substackcdn.com/image/fetch/$s_!_j7P!, /__u/onepagecode.substack.com/w_1272, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4478358-a9d4-41ec-82f7-8094a6d9b76e_2816x1536.png 1272w, /__u/substackcdn.com/image/fetch/$s_!_j7P!, /__u/onepagecode.substack.com/w_1456, /__u/onepagecode.substack.com/c_limit, /__u/onepagecode.substack.com/f_auto, /__u/onepagecode.substack.com/q_auto:good, /__u/onepagecode.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4478358-a9d4-41ec-82f7-8094a6d9b76e_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This chapter closes by moving from unsupervised discovery to controlled generation. In the next chapter, prompts become the main way to steer what a model says, which is the natural complement to the unsupervised methods here.</p><p><em>Use the url below to download the enitre book as pdf:</em></p>
      <p>
          <a href="/__u/onepagecode.substack.com/p/text-clustering-and-topic-modeling">
              Read more
          </a>
      </p>
   ]]></content:encoded></item></channel></rss>