Newly unredacted material in The New York Times' three-year-old copyright suit against OpenAI and Microsoft shows Microsoft director of Applied Science Brent Hecht wrote in a January 2023 internal memo that the companies' AI training practices were "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history"; OpenAI's head of ChatGPT, Nick Turley, separately called the threat to publishers "existential." The filings detail OpenAI's mid-training datasets containing 91,692+ copies of NYT/Daily News/Center for Investigative Reporting works, a Common Crawl dataset with 2M+ nytimes.com documents, and a "Project Mango" dataset with 160,903+ unique publisher works — plus an OpenAI researcher's paywall-bypass "hack" (Sam Altman co-founder Greg Brockman's reply: "ah nice") and deliberate stripping of copyright notices before training.
Microsoft's own data shows its Copilot "answer engine" cut NYT click-through rates by up to 93% versus plain Bing search — internally flagged by Hecht as a "doom loop" threatening "the entire web" — and CEO Satya Nadella testified he'd have forced OpenAI to retrain its models had he known paywalled content was used. The admissions land three weeks after the Trump administration filed a brief backing OpenAI's fair-use defense, and cut directly against that defense's market-substitution requirement.
NewsAPI Step 1b main query returned a genuine 0 results again this run — confirms news-agent/CLAUDE.md line 62 remains unresolved since 2026-07-07 (the + should be OR between phrase-quoted terms). Escalation to Ayush is still the right path since this agent's write scope doesn't extend to CLAUDE.md.
NewsAPI competitor query (1c) returned 3 raw hits, 0 genuine competitor hits — pure keyword-collision noise (a Nigerian aviation PR piece, two Hoover.org geopolitics pieces); no useful stray hit today.
Firecrawl: all 3 targeted scrapes succeeded cleanly — both TechCrunch pages returned a Cloudflare challenge banner at the top of the scraped markdown but the full article content followed underneath in both cases; the Inc42 UPI-MDR scrape was clean throughout. First 3/3-clean run logged in over a week.
Anthropic Newsroom rss.xml still 404s — unresolved; no Anthropic-specific standalone news broke today beyond its role in the Emerald AI grid coalition (Must Know).
shopify.dev's changelog feed.xml remains HTTP 500 — still unresolved; Shopify Changelog RSS continues to cover primary updates.
memory.md size discipline: single-previous-day rule maintained — replaced the Sep 17 "Last Report" full narrative with today's (kept Sep 17's Top 3 Stories in full); trimmed the competitor moves log's Jun 19 row (crossed the 90-day window as of today), leaving Jun 20 as the earliest entry — next trim due once Jun 20 crosses 90 days back (around Sep 19).