<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Josh Becker]]></title><description><![CDATA[Josh Becker]]></description><link>https://senatorbecker.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!mkmy!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fd10f7074-d368-445c-bac4-1021b13adffc_4032x3024.jpeg</url><title>Josh Becker</title><link>https://senatorbecker.substack.com</link></image><generator>Substack</generator><lastBuildDate>Thu, 03 Sep 2026 01:12:10 GMT</lastBuildDate><atom:link href="/__u/senatorbecker.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Josh Becker]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[senatorbecker@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[senatorbecker@substack.com]]></itunes:email><itunes:name><![CDATA[Josh Becker]]></itunes:name></itunes:owner><itunes:author><![CDATA[Josh Becker]]></itunes:author><googleplay:owner><![CDATA[senatorbecker@substack.com]]></googleplay:owner><googleplay:email><![CDATA[senatorbecker@substack.com]]></googleplay:email><googleplay:author><![CDATA[Josh Becker]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Welcome: AI Transparency in California]]></title><description><![CDATA[Legislation Keeping Pace With Innovation]]></description><link>https://senatorbecker.substack.com/p/welcome-ai-transparency-in-california</link><guid isPermaLink="false">https://senatorbecker.substack.com/p/welcome-ai-transparency-in-california</guid><dc:creator><![CDATA[Josh Becker]]></dc:creator><pubDate>Tue, 11 Aug 2026 18:25:16 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!mkmy!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fd10f7074-d368-445c-bac4-1021b13adffc_4032x3024.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hi, I&#8217;m California State Senator Josh Becker. I represent a million people in Silicon Valley. Over the last few years, we&#8217;ve seen unprecedented investments in artificial intelligence and its related infrastructure. Alongside that investment myself and others have raised concerns around safety, misinformation, and cybersecurity. It&#8217;s clear we need to regulate AI, but to do so effectively is complex. I&#8217;ll be writing a series of posts to be transparent about the process behind tackling these issues, including decision points, key matters of debate, and technical questions we still need to answer.</p><p>I represent a district whose constituents work for many of the companies building the models, compute, benchmarks, and data pipelines driving a lot of these conversations. This may come as a surprise to many, but they are also the people I hear from who are most concerned about the future of AI. That&#8217;s why I&#8217;ll be writing this series on my Substack, to discuss how we can legislate to meet the moment, and what I&#8217;m hearing from experts across industry, academia, and civil society. I want to start by talking about some progress we&#8217;ve already made on providing more transparency to consumers around AI-generated content, and how we got here.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://senatorbecker.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe to stay up to date on my work here in California</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p>In 2024, I passed the California AI Transparency Act - the first law in the nation to impose content transparency requirements on AI developers, and as of August 2nd, came into effect on the same day as the EU AI Act&#8217;s obligations under Article 50. The goal was straightforward: ensure that Californians understand where the content they view online comes from, so they can know what to trust. But there&#8217;s a lot of complexity in regulating an industry that is evolving almost daily.</p><p>AI has made remarkable advances to the benefit of businesses and consumers alike. We&#8217;ve also continued to see harmful deepfakes and AI-powered misinformation spread across the internet, undermining public institutions, news outlets, and democracy. This year, I&#8217;m authoring SB 1000 to adapt the law to reflect the technical realities that have come to the forefront in the years since the original law was signed. I&#8217;ve engaged with a wide set of subject matter experts across industry, civil society, and international authorities to ensure that the legislation we pass here in California helps push the world towards a more cohesive ecosystem of provenance. This includes meeting with those in the EU who are currently drafting the Code of Practice, as well as the French Ambassador for Digital Affairs, Clara Chappaz, around their thinking on content labeling.</p><div><hr></div><h3><strong>How the law works</strong></h3><p>In practice, CAITA will mandate that providers of AI systems embed a common set of information into the metadata of content they generate - similar to a nutrition label that discloses:</p><ol><li><p>The name of the developer,</p></li><li><p>Whether or not the content is generated or modified by artificial intelligence,</p></li><li><p>The name and version information of the AI system that made or altered the content,</p></li><li><p>The time and date of the content&#8217;s creation,</p></li><li><p>A unique identifier of the content, and</p></li><li><p>In 2029, whether the system that created the content is designed to function as an assistive technology for individuals with disabilities.</p></li></ol><p>This information will travel with the content, and when posted to a large online platform like Instagram or YouTube, they&#8217;ll be required to disclose whether or not the content was generated or substantially altered by AI. The law also mandates that the tools used to detect the information have an accessible API so websites, online platforms, apps, or any other web-based software can call on the tool to check the origin of content.</p><div><hr></div><h3>Standardization - Benefits and Trade-offs</h3><p>One of the major decisions we made this year was to include language in the bill that would encourage the adoption of technical standards. This was significant because - as I&#8217;ll outline below - content provenance technology is still developing and is by no means perfect. Many developers have adopted an open-spec standard called C2PA,  created by an industry-led coalition which serves as a possible baseline for how provenance data can be embedded in content. It&#8217;s designed to be extensible so that developers can add measures to increase the resilience of that information. For the past year I&#8217;ve been in negotiations with the steering committee members of the standard in collaboration with the advocacy groups who originally helped draft the bill. Think of it like a certified mail system for digital content. When you send certified mail, the postal service stamps it with a verified record of who sent it, when, and where - and that record travels with the envelope. Under C2PA, every piece of content gets a machine-readable stamp that records its origin and any modifications, and that stamp is cryptographically tied to a verified account so it can&#8217;t be forged without detection - known as a &#8220;hard-binding.&#8221; </p><div><hr></div><h3>Let&#8217;s Get Technical</h3><p>The standard is structured so that content generators (ChatGPT, Midjourney, or camera manufacturers like Nikon) apply to join and are issued certificates to sign the disclosures they embed in their content using cryptographic keys. Validators (<a href="https://support.google.com/gemini/answer/16722517?hl=en-AU&amp;ref_topic=13194540">Google SynthID</a>, <a href="https://verify.contentauthenticity.org/">C2PA&#8217;s own verifier</a>, or any other tool used to read C2PA disclosures) are also approved by the standards body, and verify those disclosures by checking the signer&#8217;s certificate against a registry of trusted certificate authorities.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> Those certificates expire after a set validity period for security purposes. If a signing key is compromised or stolen, the issuing certificate authority can revoke the certificate before it expires. For users, when metadata is lost in a piece of C2PA conformant content, the validator system that checks for the content&#8217;s information simply shows that the content&#8217;s origin can&#8217;t be verified, unless the content has what&#8217;s called a &#8220;soft-binding&#8221; - a type of disclosure that allows a validator tool to recover lost metadata.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> If the signed disclosures in the metadata are like stamps on a piece of mail, you can think of soft-bindings like invisible ink written on the envelope itself that only machines can read. These aren&#8217;t yet required by SB 1000 as these kinds of watermarks are still being refined, although some tools like Google&#8217;s SynthID have begun incorporating them more frequently. However, I think it&#8217;s important to acknowledge the shortcomings of this technology generally, particularly because of C2PA&#8217;s position as the leading standard.</p><p>First, consistency and scalability of techniques used to ensure provenance data is resilient to transformations is an ongoing challenge. Another problem we&#8217;re seeing even in the basic &#8220;hard-binding&#8221;  disclosures (the stamp on the envelope) of C2PA is that the tools used to validate disclosures aren&#8217;t always consistent with one another, which can lead to the same piece of content being portrayed in two different ways across disclosure verification tools. Second, a major issue I have been pushing C2PA to address is that a validator tool can see a perfectly intact disclosure, confirm the certificate chain, and declare the disclosure &#8220;valid&#8221; without ever checking whether the certificate authority has since invalidated the key that made that disclosure. If a business is C2PA conformant, but its signing key becomes compromised, there is no requirement that a validator check whether those certificates have been revoked. Due to these limitations, we included language which would allow for new technologies that are interoperable with leading standards to comply with the law so that innovation can continue without fragmenting the ecosystem beyond what&#8217;s useful to accomplish the policy goals, and we don&#8217;t settle for the first solution presented to us.</p><p>Another significant addition we made this year was to allow people to embed personal information in the content disclosures required under the bill if they explicitly consent to doing so. It&#8217;s important to be clear that personal information can be embedded in all manner of content with C2PA - synthetic or otherwise. Many researchers and governments are looking for ways to leverage content provenance infrastructure for the purposes of digital rights management to help ensure that people have a say on whether the content they publish can be used for AI training, and explore ways to fairly compensate creators for licensed use. While that is out of the scope of my work this year, it is certainly something I&#8217;m considering for the future.</p><div><hr></div><h3>Visible Markings</h3><p>When most people think of marking AI content, they envision visible marks on images. When we originally passed the California AI Transparency Act (SB 942) we included a provision that mandated companies provide users with the option to include them in image, video, and audio. However, SB 1000 removes that obligation, and here&#8217;s why. Perceptible watermarks are legible and frictionless, but that same ease of access has real trade-offs. First, there is not a single, monolithic mark that is universally applied to content. There would be nothing stopping people from creating their own labels which could confuse or mislead consumers. Additionally, visible marks are extremely easy to remove or misappropriate with little evidence that the content was tampered with. This is part of the reason why the EU has backed away from visible disclosures, requiring them only for specific kinds of content, in specific contexts. Finally, as awareness of the amount of AI content on the internet grows, people may come to assume that content <em>without</em> a perceptible watermark is inherently <em>more</em> credible.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> That assumption would be wrong.</p><p>Whether content is factually accurate and whether it was synthetically generated are largely unrelated questions. A photo can be real and misleading when put in the wrong context, and an AI-generated image can be benign. The real policy problem we&#8217;re trying to solve is <em>misinformation</em>, which requires two things: 1) Increased media literacy so people know to critically examine what they see online, and 2) Reliable information people can trust about where content comes from. The AI Transparency Act addresses the latter of these issues head-on. It pushes companies, device manufacturers, and media organizations to conform to standardized provenance approaches so that there is a consistent stream of information being embedded in the content we see online. That information will be required to be made available on large online platforms so people can easily assess where content comes from, and whether or not to trust it.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://senatorbecker.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/senatorbecker.substack.com/subscribe"><span>Subscribe now</span></a></p><div><hr></div><h3>The Open-Source/Open-Weight Question</h3><p>One of the harder problems we&#8217;ve been wrestling with is what to do about open-weight or open-source models that anyone can download and run locally on their own hardware.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a> While these models are often smaller (13B-30B parameters), the performance gap between them and those on the frontier is closing.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a> This is important given the rise of AI-generated CSAM and NCII, much of which is created using these kinds of systems.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a> </p><p>This presents two problems. The first is that it&#8217;s difficult to trace content from these systems back to specific individuals - since the user prompt and model response are contained to their system. The second is that regulating the upstream provision of these models could harm the ecosystem of open systems, which have also been beneficial to academia, small businesses, and law-abiding users. While there are currently technical solutions being developed around baking provenance into the internals of these models, none are yet mature enough to make it mandatory by law. It&#8217;s a space I will continue to monitor as the bill goes into effect. Under SB 1000, open-source systems are required to embed provenance to the extent technically feasible, and if a licensor discovers someone is out of compliance, they have to notify the licensee that they&#8217;re breaking the law and request they modify the system to comply. If the licensee doesn&#8217;t alert the licensor of what action they&#8217;re taking, the licensor reports them to the Attorney General.</p><div><hr></div><h3>What comes next</h3><p>Lawmakers can&#8217;t approach AI content labeling through a binary &#8220;AI or not AI&#8221; lens. The technical limitations and the potential for unintended consequences demand that we get the details right. Legislation must innovate along with technology. I am a firm believer in our government being capable of finding innovative solutions for complex problems. Achieving that goal requires nuance, and a willingness to adapt. The reality is that no technology is perfect, and we need to make sure that the laws we pass stand the test of time. The California AI Transparency Act will move us toward a more trustworthy information ecosystem, and I&#8217;m committed to being as nimble as the Silicon Valley startups I represent as each new wave of tech comes. California has led on complex issues before, and I&#8217;m confident we&#8217;ll lead here too.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://senatorbecker.substack.com/p/welcome-ai-transparency-in-california?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Feel free to subscribe to stay up to date with my work here in California. </p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://senatorbecker.substack.com/p/welcome-ai-transparency-in-california?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="/__u/senatorbecker.substack.com/p/welcome-ai-transparency-in-california?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share</span></a></p></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>See the list of conformant generators and validators <a href="https://spec.c2pa.org/conformance-explorer/">here</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>See <a href="https://spec.c2pa.org/specifications/specifications/2.4/index.html">C2PA&#8217;s specification</a> for more information.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p><a href="https://pubsonline.informs.org/doi/abs/10.1287/mnsc.2019.3478">The Implied Truth Effect: Attaching Warnings to a Subset of Fake News Headlines Increases Perceived Accuracy of Headlines Without Warnings</a></p><p><span>Gordon Pennycook</span>, <span>Adam Bear</span>, <span>Evan T. Collins</span>, and <span>David G. Rand</span></p><p><span>Management Science202066:11,4944-495710.1287/mnsc.2019.3478 </span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>Even defining what &#8220;open-source&#8221; means in AI has been the subject of a much debate. For that reason, we did not define it in SB 1000 and instead focused on the nature of licensor&#8217;s relationships to licensees, and the level of technical oversight over these kinds of systems generally.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p><a href="https://www.aisi.gov.uk/frontier-ai-trends-report">The AI Security Institute estimates</a> that open models are between 4 and 8 months behind the frontier.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>See the Internet Watch Foundation&#8217;s <a href="https://www.iwf.org.uk/about-us/why-we-exist/our-research/how-ai-is-being-abused-to-create-child-sexual-abuse-imagery">2026 AI CSAM Report</a>.</p></div></div>]]></content:encoded></item></channel></rss>