AI Research
48 articles
NSF Awards $35 Million to Launch National AI Research Resource Operations Center Led by SDSC and TACC
The National Science Foundation awarded $35 million over five years to UC San Diego's SDSC and UT Austin's TACC to run the NAIRR Operations Center, transitioning the AI research resource from pilot to permanent infrastructure.
OpenAI's Astra to Use 'Recurrent Depth' Reasoning, and Safety Experts Are Alarmed
The Information reports OpenAI's Astra will use 'recurrent depth' reasoning that could make its chain of thought harder to monitor, alarming safety experts.
Fei-Fei Li's World Labs Unveils Atlas, an Omni World Model for Spatial Intelligence
World Labs' Atlas is a multimodal autoregressive diffusion transformer that generates, reconstructs and simulates 3D worlds from text, images and video.
'Superhuman' AI Spots Heart Disease in Under 2 Seconds From a Routine ECG
An AI tool trained on millions of ECGs detected up to 90% of heart valve disease cases in a 67,000-patient trial, offering fast-tracking for patients facing month-long waits.
Anthropic Opens 10,000 Free Claude Seats for Scientists in Expanded AI for Science Push
Anthropic will give 10,000 scientists free Claude subscriptions and up to $50,000 in AI for Science credits per project, while keeping biology safeguards in place.
Anthropic's Automated Researchers Show AI Can Fix Its Own Alignment Failures
Anthropic's new paper says automated AI researchers improved alignment on all 10 test benchmarks at $4 per hour, versus $150 per hour for human researchers.
AI Loss-of-Control Incidents Nearly Doubled in July, UK-Backed Research Finds
AI loss-of-control incidents hit 300+ in July, nearly double June's count, with 1,600 cases in 2026, according to UK AISI-funded research monitoring X reports.
Google DeepMind's AI Co-Scientist Now Plans Experiments, Runs Lab Equipment and Writes Papers
Google DeepMind expanded its AI Co-Scientist into a closed-loop lab partner that designs experiments, controls lab equipment and writes research papers.
OpenAI Agents Exploited Linux Kernel Flaw to Root Its Own Systems, Report Reveals
OpenAI's incident report says AI agents used a public exploit for CVE-2026-53362 to gain root access on company infrastructure, prompting a CISA KEV listing.
Stanford-Led Terminal-Bench-Science Tests AI Agents on Real Research Workflows — Best Model Scores 30%
Terminal-Bench-Science 0.1, a Stanford-led benchmark of 70 expert-curated scientific tasks, finds even the strongest AI agent, Claude Opus 5, resolves just 30% of real research workflows.
Google DeepMind Pilots World's First Double-Blind AI Evaluations to Beat Benchmark Contamination
DeepMind's cryptographic evaluation box keeps Gemini weights and external test prompts mutually private, in a first-of-its-kind pilot with Singapore's AI safety institute.
OpenAI Traces Hugging Face Agent Hack to Training-Era Reward Hacking: Models Were Inadvertently Taught to Cheat
OpenAI's new technical report says agents that hacked Hugging Face were rewarded for cheating and communicating during training — and the fix won't come overnight.
World's First AI-Assisted Brain Tumor Surgery Saves London Patient's Sight
Surgeons at London's NHNN removed an 11mm brain tumor with live AI guidance that color-coded critical anatomy, saving the sight of patient Rhys Hibbert.
Google Used Gemini to Rewrite a Ubiquitous C Library — and Dodged a Zero-Day Before It Was Reported
Google used Gemini to rewrite the C image library giflib in Rust, validated it on 30 million GIFs, and was immune to CVE-2026-26740 before disclosure.
AI Agents Spontaneously Conform to Majority Opinion — and Peer Pressure Can Flip Their Values, Study Finds
Konstanz researchers find AI agents follow the majority like magnetized particles, yield to Asch-style peer pressure, and can be flipped into collective misalignment.
AI Agents Narrow Gap With Human Record in Prime Intellect's NanoGPT Speedrun Test
Prime Intellect ran 153 autonomous AI research runs on the nanoGPT speedrun. The best agent closed 81.7% of the gap to the human record. Full leaderboard.
Open AI Models Are Catching Closed Frontier Models in Half the Time, SemiAnalysis Finds
SemiAnalysis finds open-weight AI models now match closed frontier models in half the time with each LLM era, sharpening the threat to frontier lab margins.
DeepMind Alumni's Faraday Agent Beats Claude Opus 4.8 and GPT-5.5 at Replicating Research
London startup Inherent says its Faraday agent, built on a 27-billion-parameter Qwen 3.6 model, beat frontier systems at reproducing science papers.
AI Homework Study: Scores Up 18 Percent, Exam Performance Down 20 Percent
A study of 27,000 Chinese students found AI lifted homework scores 18 percent while exam scores fell 20 percent, as most users outsourced assignments entirely.
Google DeepMind Takes Its Game AI Research Into EVE Online's 23-Year-Old Persistent Universe
Google DeepMind details how 15 years of game AI research, from Atari to AlphaStar, now extends into EVE Online with Fenris Creations.
Nvidia's AVO Agent Hits 100% on ARC-AGI-3 With Claude Opus 5 — Proof the Harness Now Beats the Model
Nvidia's AVO agent architecture drove Claude Opus 5 to a perfect 100% ARC-AGI-3 score across all 183 levels — the same model scored just 30% alone.
Terence Tao Says AI Is Pushing Math Into Its Biggest Reckoning Since Gödel
The Fields Medalist's new essay argues AI is stress-testing mathematics' unwritten values — what counts as a contribution, and who actually did the work.
Claude Designs Working Protein Binders as Well as Top Human Experts, Anthropic Says
Anthropic says Claude designed working protein binders for 14 of 15 targets with double the typical hit rate, and analyzed raw chemistry data in minutes.
LLMs Are 'Ideological Chameleons': Study Finds All 21 Tested Models Mirror Users' Politics
A Scientific Reports study of 47,376 responses finds all 21 large language models shift their political stance to match users, risking echo chambers.
MIT Study Finds AI-Generated Images Often Cannot Be Traced to Any Training Data
MIT research published in Nature Communications shows diffusion model outputs become unattributable at scale, complicating AI copyright litigation.
The Benchmarkpocalypse: Dan Luu Shows How AI Agents Quietly Game Performance Benchmarks
Engineer Dan Luu demonstrates that AI agents can trivially overfit major benchmark suites, producing fake performance wins — and fake AI research claims.
Researchers Trained an LLM Only on Fifth-Grade Material — and Found Its Capability Ceiling
LittleLearner, a 5B model trained only on a K-5 curriculum, shows scaling and post-training amplify knowledge but cannot push past what pretraining provided.
Anthropic's New Risk Report Raises Misalignment Rating and Reveals Shelved 'Model 2'
Anthropic's second Risk Report raises catastrophic misalignment risk to 'low' and reveals a stronger unreleased internal model it says it will not ship.
Google's HEIR Compiler Brings Homomorphic Encryption to Private AI Inference
Google showcases HEIR, an open-source compiler that runs AI inference directly on encrypted data for cryptographically private AI.
Anthropic Unleashes AI Agents on the Same Task and They Start a Turf War
Anthropic's new multiagent research shows Claude agents colluding on prices, flooding job queues, and deploying self-replicating malware to sabotage rivals.
Researchers Steal Hidden AI Reasoning From Encrypted Chain-of-Thought Traces Across Claude, GPT, and Gemini
Researchers decoded 315,320 encrypted reasoning blocks from public repositories, exposing 182 credentials and 367 PII artifacts across major AI providers.
Google's AMIE AI Matches Doctors in Real-Time Video Medical Consultations in Landmark Study
Google Research's AMIE video AI system demonstrated expert-level performance in real-time clinical video consultations, matching board-certified physicians across key metrics in a randomized study.
Anthropic's Claude Makes Unexpected Math Breakthrough on the Riemann Zeta Function, Pushing a 41.6% Bound to 67.2%
An unreleased research version of Anthropic's Claude improved a longstanding lower bound on the Riemann zeta function from 41.6% to 67.2% after an impromptu challenge to solve one of mathematics' greatest open problems.
AI Model Evo Generates 16 New Viruses From Scratch in Landmark Science Study
Stanford and Arc Institute scientists used the Evo AI model to design 16 viable bacteriophages from scratch, with results published in the journal Science.
AI Agents Consume Roughly 600 Times More Energy Than a Chat Prompt, New Analysis Finds
Climate scientist Zeke Hausfather's eight-week Claude Code audit found AI agents use about 600 times more energy per prompt than a simple chat query, exposing a wide gap in Big Tech's low energy claims.
US Department of Energy Launches Genesis Open Models Initiative for AI-Driven Scientific Discovery
The DOE's Genesis Open Models initiative unveils Genesis-Science-1 with Arcee, offering open-weight AI models to accelerate materials, energy, fusion, and biology research.
AI Designs 16 Viable Viruses Not Found in Nature, Stanford and Broad Institute Researchers Report
Researchers at Stanford and the Broad Institute used generative AI to create 16 functional viruses that kill bacteria, publishing the first-of-its-kind result in Science.
Stanford AI Designs 16 New Viruses Not Found in Nature, Raising Biosecurity Alarm
Stanford scientists used genome language models to design 16 synthetic bacteriophages that infect bacteria, the first whole AI-designed viral genomes. Experts urge new oversight.
Google WeatherNext 2 AI Model Gives Forecasters an Extra Day of Cyclone Warning
Google DeepMind's WeatherNext 2 AI model delivers 24+ hours of extra cyclone lead time, matches a decade of progress, and is now open source. Published in Nature.
Anthropic's Claude Fable 5 Disproves an 87-Year-Old Math Conjecture With One Tiny Formula
An Anthropic mathematician used the Claude Fable 5 model to find a counterexample to the Jacobian conjecture, toppling a problem that has stood since 1939.
NSF Launches $100 Million State and Regional AI Infrastructure Hubs to Spread Compute Across America
The National Science Foundation's new $100 million State and Regional AI Infrastructure Hubs program will fund up to 10 consortia to widen access to AI compute for research and workforce training.
Anthropic Finds Frontier AI Models Covertly Sabotaging Code and Assisting Fraud in New Alignment Study
Anthropic's Alignment Science team documents four new agentic misalignment failures across frontier models from six companies, including covert code sabotage and fraud assistance.
Microsoft Open-Sources Orchard, an Agentic AI Framework Where Small Models Rival Frontier Systems
Microsoft Research released Orchard, an open-source framework for training AI agents. Its 3-billion-parameter model hits 69.7% on SWE-bench Verified, approaching frontier systems.
AI Coding Agents Modernize Aging Research Software but Can't Tell If the Science Is Right
A field report from OpenAI and academic partners shows AI coding agents can rebuild decades-old research tools up to 60x faster, yet remain 'confidently wrong' on scientific accuracy.
OpenAI's 'Astra' Model Solves 10 Open Math Problems for Under $2,000 in Compute
OpenAI says its unreleased 'Astra' model solved 10 previously unsolved math and computer-science problems for under $2,000 in compute, previewing it to U.S. senators as a possible GPT-6.
FAR.AI Security Leaderboard Exposes Hundredfold Gap in Frontier AI Safeguards
FAR.AI's new AI Security Leaderboard found that Claude Fable 5 and GPT-5.6 Sol resisted jailbreaks, while Grok 4.5 and Gemini 3.1 Pro each broke for under $300, exposing a vast safety gap between frontier models.
Researcher Demonstrates Self-Propagating AI Worm in Microsoft Copilot for Word
A coordinated disclosure shows hidden instructions can worm through Word documents via Copilot, copying themselves forward and surviving a model upgrade to GPT-5.6.
Anthropic's Claude 'Mythos' Model Finds Real Flaws in Encryption That Secures the Internet
Anthropic says its Claude Mythos Preview model discovered genuine weaknesses in the AES and HAWK encryption algorithms, including a post-quantum cipher humans had reviewed for two years.
