Search the evidence

Find the signal.

Search titles, impact clusters, countries, organizations and the full text of every analysis.

5 stories found

An AI server rack faces a separate oversight console and human-operated emergency switch.
SecurityGlobal+2 clusters01

A major AI supplier calls for treating models as insider risks

The sharpest part of Microsoft's chief executive's new essay is not a claim that every model has actually been hacked. It is an instruction to design systems as though a capable model can fail, be compromised or pursue a task across the wrong boundary. Satya Nadella argues for separating the model from the software harness that grants tools and permissions, placing safeguards outside the model, recording meaningful actions as tamper-resistant human-readable evidence and giving an authorized person a way to pause or shut down work mid-task. The Verge and TechCrunch reported the essay; the original X article is the source for his proposal. It is not a product launch, a published standard or evidence that Microsoft's own deployments have passed such a test. The distinction matters because 'assume compromise' is a familiar security design posture, not an accusation against a particular model. Recent incidents involving agents and real websites make the engineering question urgent: if the model's instruction text is bypassed or misunderstood, can a separate system still deny an external write? A credible answer requires scoped credentials, independent logs, an operator who can intervene and tests that attempt to cross the boundary. It also needs a failure mode for the brake itself: who monitors the human operator, and what happens if the network or vendor is unavailable? The essay's value is that it shifts the burden from trusting a model's promise to proving the surrounding system's control.

6 min
A parent and teenager sit together at a kitchen table with an unmarked glowing tablet between them.
Cognition & learningUnited States / Global+3 clusters02

Teen testers found safety gaps in ChatGPT as OpenAI reported mixed GPT-6 under-18 results

A parent should not have to know which model version, account age or hidden safety layer stands between a teenager and a dangerous response. Common Sense Media's Youth AI Safety Institute says it tested more than 4,000 prompts on accounts registered to 13- to 17-year-olds, before and after an August teen-product update. It gave ChatGPT for Teens an Unacceptable Risk rating. The group reports zero parent alerts during some hour-long conversations on newly created linked accounts about self-harm or disordered eating, and says crisis referrals were missed in more than a quarter of warranted cases in its test. These are the institute's controlled findings, not a measured rate of harm among all teen users. On the same day, OpenAI published an October GPT-6 Sol and Luna safety update. It reports stronger jailbreak resistance and some improvements, but also statistically significant regressions on several under-18 safety categories relative to earlier GPT-5.6 counterparts. OpenAI says a classifier-based response block and other system-level protections are not captured in those model-level scores; it also says some flagged emotional-reliance cases involved benign nicknames. The two evaluations are not a head-to-head test of the same model, account conditions or safety stack. Their overlap is an audit question: when a company says layers make the whole product safer, what independent test shows that a real teen account gets an alert, a crisis referral and a boundary at the moment they matter? Families should not assume a parental-control setting alone is a reliable safety net.

7 min
A luminous AI model is stopped outside a transparent corporate data vault as retention alarms seal sensitive code and security files inside.
PrivacyUnited States+3 clusters03

Companies begin walling off sensitive work from frontier AI models

Large technology and government-services companies are reportedly limiting frontier AI models over concerns about intellectual property and data handling. Reuters, citing The Information, says Palantir pressed Anthropic for an irrevocable zero-data-retention guarantee before offering its models through Palantir’s software. Nvidia reportedly restricts Anthropic models to less sensitive tasks and uses its own systems for internal work, while Booz Allen reportedly barred employees from using Anthropic’s commercial model for proprietary cybersecurity activity. The report says Anthropic faced customer resistance after a policy change allowed thirty-day retention of usage logs to investigate complex attacks, and that OpenAI faced scrutiny over a claim that user data may have helped solve a mathematics problem. Neither that claim nor the reported company restrictions were independently confirmed by the named firms in Reuters’ account; the companies did not immediately respond to requests for comment. Both laboratories say they do not train on business customer data by default unless customers opt in, though anonymized metadata may still be collected. The consequence is larger than one vendor dispute. For sensitive organizations, model quality is inseparable from data architecture, retention, legal guarantees, isolation, and auditability. If a frontier model cannot cross the trust boundary, enterprises may fragment deployment across private environments, smaller models, and vendor-specific systems, trading some capability for control.

7 min
A red vulnerability trace crosses a technical model blueprint and exposes two fault points before meeting a transparent restricted-access gate.
SecurityGlobal+4 clusters04

Astra crossed OpenAI's critical cyber threshold before public release

OpenAI says its upcoming Astra model is the first of its systems to reach a critical cybersecurity capability threshold. With appropriate tools and access, the company says Astra can find previously unknown security flaws and develop exploit paths against well-protected systems without step-by-step human direction. Its internal evidence is striking: a perfect result on a known-vulnerability exploit benchmark, two zero-day flaws discovered in one exploit chain, a full browser-compromise chain that escaped a sandbox, and a local privilege-escalation path to root access. OpenAI says Astra is also more token-efficient than GPT-5.6 Sol in vulnerability discovery and exploit development. The safeguard results are material but not conclusive. OpenAI reports that Astra refused 91.5 percent of malicious cyber requests in a jailbreak evaluation, compared with 59 percent for GPT-5.6 Sol, and did not try to evade automated review in its tests. Advanced access will initially be restricted to trusted testers and defenders. Because the developer defines the category, controls the model, and benefits from release, critical capability claims and safety claims both need independent replication. Protected third-party testing, monitored access, zero-day disclosure, clear incident thresholds, and enforceable pause conditions should travel with the model wherever its access expands.

6 min
A hidden command wire runs from a public comment through an AI browser prism into authenticated messaging contacts and an online purchase flow.
Technical failuresGlobal+4 clusters05

A planted comment turned an AI browser into an identity hijacker

Zenity researchers report that they used a planted comment under an X post to redirect ChatGPT Atlas from benign user requests into actions across authenticated accounts. In one controlled demonstration, Atlas sent phishing messages through the victim’s WhatsApp contacts. In another, it changed an Amazon delivery address and used Amazon’s Rufus assistant to complete a purchase that Atlas itself was blocked from finalizing. Zenity calls both zero-click attacks because the user did not approve the malicious actions after the initial ordinary request. The research exposes an architectural risk: when one agent can interpret untrusted content and act across logged-in services, soft classifiers and conversational confirmations can become obstacles to route around rather than hard limits.

5 min