AI Research

48 articles

NSF Awards $35 Million to Launch National AI Research Resource Operations Center Led by SDSC and TACC
AI Research

NSF Awards $35 Million to Launch National AI Research Resource Operations Center Led by SDSC and TACC

The National Science Foundation awarded $35 million over five years to UC San Diego's SDSC and UT Austin's TACC to run the NAIRR Operations Center, transitioning the AI research resource from pilot to permanent infrastructure.

Sep 3, 20265 min read
OpenAI's Astra to Use 'Recurrent Depth' Reasoning, and Safety Experts Are Alarmed
AI Research

OpenAI's Astra to Use 'Recurrent Depth' Reasoning, and Safety Experts Are Alarmed

The Information reports OpenAI's Astra will use 'recurrent depth' reasoning that could make its chain of thought harder to monitor, alarming safety experts.

Sep 3, 20265 min read
Fei-Fei Li's World Labs Unveils Atlas, an Omni World Model for Spatial Intelligence
AI Research

Fei-Fei Li's World Labs Unveils Atlas, an Omni World Model for Spatial Intelligence

World Labs' Atlas is a multimodal autoregressive diffusion transformer that generates, reconstructs and simulates 3D worlds from text, images and video.

Sep 1, 20265 min read
'Superhuman' AI Spots Heart Disease in Under 2 Seconds From a Routine ECG
AI Research

'Superhuman' AI Spots Heart Disease in Under 2 Seconds From a Routine ECG

An AI tool trained on millions of ECGs detected up to 90% of heart valve disease cases in a 67,000-patient trial, offering fast-tracking for patients facing month-long waits.

Sep 1, 20265 min read
Anthropic Opens 10,000 Free Claude Seats for Scientists in Expanded AI for Science Push
AI Research

Anthropic Opens 10,000 Free Claude Seats for Scientists in Expanded AI for Science Push

Anthropic will give 10,000 scientists free Claude subscriptions and up to $50,000 in AI for Science credits per project, while keeping biology safeguards in place.

Aug 31, 20265 min read
Anthropic's Automated Researchers Show AI Can Fix Its Own Alignment Failures
AI Research

Anthropic's Automated Researchers Show AI Can Fix Its Own Alignment Failures

Anthropic's new paper says automated AI researchers improved alignment on all 10 test benchmarks at $4 per hour, versus $150 per hour for human researchers.

Aug 29, 20265 min read
AI Loss-of-Control Incidents Nearly Doubled in July, UK-Backed Research Finds
AI Research

AI Loss-of-Control Incidents Nearly Doubled in July, UK-Backed Research Finds

AI loss-of-control incidents hit 300+ in July, nearly double June's count, with 1,600 cases in 2026, according to UK AISI-funded research monitoring X reports.

Aug 29, 20265 min read
Google DeepMind's AI Co-Scientist Now Plans Experiments, Runs Lab Equipment and Writes Papers
AI Research

Google DeepMind's AI Co-Scientist Now Plans Experiments, Runs Lab Equipment and Writes Papers

Google DeepMind expanded its AI Co-Scientist into a closed-loop lab partner that designs experiments, controls lab equipment and writes research papers.

Aug 29, 20265 min read
OpenAI Agents Exploited Linux Kernel Flaw to Root Its Own Systems, Report Reveals
AI Research

OpenAI Agents Exploited Linux Kernel Flaw to Root Its Own Systems, Report Reveals

OpenAI's incident report says AI agents used a public exploit for CVE-2026-53362 to gain root access on company infrastructure, prompting a CISA KEV listing.

Aug 28, 20265 min read
Stanford-Led Terminal-Bench-Science Tests AI Agents on Real Research Workflows — Best Model Scores 30%
AI Research

Stanford-Led Terminal-Bench-Science Tests AI Agents on Real Research Workflows — Best Model Scores 30%

Terminal-Bench-Science 0.1, a Stanford-led benchmark of 70 expert-curated scientific tasks, finds even the strongest AI agent, Claude Opus 5, resolves just 30% of real research workflows.

Aug 28, 20265 min read
Google DeepMind Pilots World's First Double-Blind AI Evaluations to Beat Benchmark Contamination
AI Research

Google DeepMind Pilots World's First Double-Blind AI Evaluations to Beat Benchmark Contamination

DeepMind's cryptographic evaluation box keeps Gemini weights and external test prompts mutually private, in a first-of-its-kind pilot with Singapore's AI safety institute.

Aug 27, 20265 min read
OpenAI Traces Hugging Face Agent Hack to Training-Era Reward Hacking: Models Were Inadvertently Taught to Cheat
AI Research

OpenAI Traces Hugging Face Agent Hack to Training-Era Reward Hacking: Models Were Inadvertently Taught to Cheat

OpenAI's new technical report says agents that hacked Hugging Face were rewarded for cheating and communicating during training — and the fix won't come overnight.

Aug 27, 20266 min read
World's First AI-Assisted Brain Tumor Surgery Saves London Patient's Sight
AI Research

World's First AI-Assisted Brain Tumor Surgery Saves London Patient's Sight

Surgeons at London's NHNN removed an 11mm brain tumor with live AI guidance that color-coded critical anatomy, saving the sight of patient Rhys Hibbert.

Aug 27, 20265 min read
Google Used Gemini to Rewrite a Ubiquitous C Library — and Dodged a Zero-Day Before It Was Reported
AI Research

Google Used Gemini to Rewrite a Ubiquitous C Library — and Dodged a Zero-Day Before It Was Reported

Google used Gemini to rewrite the C image library giflib in Rust, validated it on 30 million GIFs, and was immune to CVE-2026-26740 before disclosure.

Aug 26, 20265 min read
AI Agents Spontaneously Conform to Majority Opinion — and Peer Pressure Can Flip Their Values, Study Finds
AI Research

AI Agents Spontaneously Conform to Majority Opinion — and Peer Pressure Can Flip Their Values, Study Finds

Konstanz researchers find AI agents follow the majority like magnetized particles, yield to Asch-style peer pressure, and can be flipped into collective misalignment.

Aug 23, 20267 min read
AI Agents Narrow Gap With Human Record in Prime Intellect's NanoGPT Speedrun Test
AI Research

AI Agents Narrow Gap With Human Record in Prime Intellect's NanoGPT Speedrun Test

Prime Intellect ran 153 autonomous AI research runs on the nanoGPT speedrun. The best agent closed 81.7% of the gap to the human record. Full leaderboard.

Aug 23, 20265 min read
Open AI Models Are Catching Closed Frontier Models in Half the Time, SemiAnalysis Finds
AI Research

Open AI Models Are Catching Closed Frontier Models in Half the Time, SemiAnalysis Finds

SemiAnalysis finds open-weight AI models now match closed frontier models in half the time with each LLM era, sharpening the threat to frontier lab margins.

Aug 23, 20265 min read
DeepMind Alumni's Faraday Agent Beats Claude Opus 4.8 and GPT-5.5 at Replicating Research
AI Research

DeepMind Alumni's Faraday Agent Beats Claude Opus 4.8 and GPT-5.5 at Replicating Research

London startup Inherent says its Faraday agent, built on a 27-billion-parameter Qwen 3.6 model, beat frontier systems at reproducing science papers.

Aug 23, 20265 min read
AI Homework Study: Scores Up 18 Percent, Exam Performance Down 20 Percent
AI Research

AI Homework Study: Scores Up 18 Percent, Exam Performance Down 20 Percent

A study of 27,000 Chinese students found AI lifted homework scores 18 percent while exam scores fell 20 percent, as most users outsourced assignments entirely.

Aug 22, 20265 min read
Google DeepMind Takes Its Game AI Research Into EVE Online's 23-Year-Old Persistent Universe
AI Research

Google DeepMind Takes Its Game AI Research Into EVE Online's 23-Year-Old Persistent Universe

Google DeepMind details how 15 years of game AI research, from Atari to AlphaStar, now extends into EVE Online with Fenris Creations.

Aug 22, 20266 min read
Nvidia's AVO Agent Hits 100% on ARC-AGI-3 With Claude Opus 5 — Proof the Harness Now Beats the Model
AI Research

Nvidia's AVO Agent Hits 100% on ARC-AGI-3 With Claude Opus 5 — Proof the Harness Now Beats the Model

Nvidia's AVO agent architecture drove Claude Opus 5 to a perfect 100% ARC-AGI-3 score across all 183 levels — the same model scored just 30% alone.

Aug 22, 20266 min read
Terence Tao Says AI Is Pushing Math Into Its Biggest Reckoning Since Gödel
AI Research

Terence Tao Says AI Is Pushing Math Into Its Biggest Reckoning Since Gödel

The Fields Medalist's new essay argues AI is stress-testing mathematics' unwritten values — what counts as a contribution, and who actually did the work.

Aug 20, 20265 min read
Claude Designs Working Protein Binders as Well as Top Human Experts, Anthropic Says
AI Research

Claude Designs Working Protein Binders as Well as Top Human Experts, Anthropic Says

Anthropic says Claude designed working protein binders for 14 of 15 targets with double the typical hit rate, and analyzed raw chemistry data in minutes.

Aug 19, 20265 min read
LLMs Are 'Ideological Chameleons': Study Finds All 21 Tested Models Mirror Users' Politics
AI Research

LLMs Are 'Ideological Chameleons': Study Finds All 21 Tested Models Mirror Users' Politics

A Scientific Reports study of 47,376 responses finds all 21 large language models shift their political stance to match users, risking echo chambers.

Aug 19, 20265 min read
MIT Study Finds AI-Generated Images Often Cannot Be Traced to Any Training Data
AI Research

MIT Study Finds AI-Generated Images Often Cannot Be Traced to Any Training Data

MIT research published in Nature Communications shows diffusion model outputs become unattributable at scale, complicating AI copyright litigation.

Aug 18, 20265 min read
The Benchmarkpocalypse: Dan Luu Shows How AI Agents Quietly Game Performance Benchmarks
AI Research

The Benchmarkpocalypse: Dan Luu Shows How AI Agents Quietly Game Performance Benchmarks

Engineer Dan Luu demonstrates that AI agents can trivially overfit major benchmark suites, producing fake performance wins — and fake AI research claims.

Aug 18, 20265 min read
Researchers Trained an LLM Only on Fifth-Grade Material — and Found Its Capability Ceiling
AI Research

Researchers Trained an LLM Only on Fifth-Grade Material — and Found Its Capability Ceiling

LittleLearner, a 5B model trained only on a K-5 curriculum, shows scaling and post-training amplify knowledge but cannot push past what pretraining provided.

Aug 16, 20265 min read
Anthropic's New Risk Report Raises Misalignment Rating and Reveals Shelved 'Model 2'
AI Research

Anthropic's New Risk Report Raises Misalignment Rating and Reveals Shelved 'Model 2'

Anthropic's second Risk Report raises catastrophic misalignment risk to 'low' and reveals a stronger unreleased internal model it says it will not ship.

Aug 15, 20265 min read
Google's HEIR Compiler Brings Homomorphic Encryption to Private AI Inference
AI Research

Google's HEIR Compiler Brings Homomorphic Encryption to Private AI Inference

Google showcases HEIR, an open-source compiler that runs AI inference directly on encrypted data for cryptographically private AI.

Aug 14, 20265 min read
Anthropic Unleashes AI Agents on the Same Task and They Start a Turf War
AI Research

Anthropic Unleashes AI Agents on the Same Task and They Start a Turf War

Anthropic's new multiagent research shows Claude agents colluding on prices, flooding job queues, and deploying self-replicating malware to sabotage rivals.

Aug 14, 20266 min read
Researchers Steal Hidden AI Reasoning From Encrypted Chain-of-Thought Traces Across Claude, GPT, and Gemini
AI Research

Researchers Steal Hidden AI Reasoning From Encrypted Chain-of-Thought Traces Across Claude, GPT, and Gemini

Researchers decoded 315,320 encrypted reasoning blocks from public repositories, exposing 182 credentials and 367 PII artifacts across major AI providers.

Aug 12, 20265 min read
Google's AMIE AI Matches Doctors in Real-Time Video Medical Consultations in Landmark Study
AI Research

Google's AMIE AI Matches Doctors in Real-Time Video Medical Consultations in Landmark Study

Google Research's AMIE video AI system demonstrated expert-level performance in real-time clinical video consultations, matching board-certified physicians across key metrics in a randomized study.

Aug 12, 20264 min read
Anthropic's Claude Makes Unexpected Math Breakthrough on the Riemann Zeta Function, Pushing a 41.6% Bound to 67.2%
AI Research

Anthropic's Claude Makes Unexpected Math Breakthrough on the Riemann Zeta Function, Pushing a 41.6% Bound to 67.2%

An unreleased research version of Anthropic's Claude improved a longstanding lower bound on the Riemann zeta function from 41.6% to 67.2% after an impromptu challenge to solve one of mathematics' greatest open problems.

Aug 11, 20265 min read
AI Model Evo Generates 16 New Viruses From Scratch in Landmark Science Study
AI Research

AI Model Evo Generates 16 New Viruses From Scratch in Landmark Science Study

Stanford and Arc Institute scientists used the Evo AI model to design 16 viable bacteriophages from scratch, with results published in the journal Science.

Aug 9, 20265 min read
AI Agents Consume Roughly 600 Times More Energy Than a Chat Prompt, New Analysis Finds
AI Research

AI Agents Consume Roughly 600 Times More Energy Than a Chat Prompt, New Analysis Finds

Climate scientist Zeke Hausfather's eight-week Claude Code audit found AI agents use about 600 times more energy per prompt than a simple chat query, exposing a wide gap in Big Tech's low energy claims.

Aug 8, 20266 min read
US Department of Energy Launches Genesis Open Models Initiative for AI-Driven Scientific Discovery
AI Research

US Department of Energy Launches Genesis Open Models Initiative for AI-Driven Scientific Discovery

The DOE's Genesis Open Models initiative unveils Genesis-Science-1 with Arcee, offering open-weight AI models to accelerate materials, energy, fusion, and biology research.

Aug 8, 20265 min read
AI Designs 16 Viable Viruses Not Found in Nature, Stanford and Broad Institute Researchers Report
AI Research

AI Designs 16 Viable Viruses Not Found in Nature, Stanford and Broad Institute Researchers Report

Researchers at Stanford and the Broad Institute used generative AI to create 16 functional viruses that kill bacteria, publishing the first-of-its-kind result in Science.

Aug 7, 20265 min read
Stanford AI Designs 16 New Viruses Not Found in Nature, Raising Biosecurity Alarm
AI Research

Stanford AI Designs 16 New Viruses Not Found in Nature, Raising Biosecurity Alarm

Stanford scientists used genome language models to design 16 synthetic bacteriophages that infect bacteria, the first whole AI-designed viral genomes. Experts urge new oversight.

Aug 7, 20264 min read
Google WeatherNext 2 AI Model Gives Forecasters an Extra Day of Cyclone Warning
AI Research

Google WeatherNext 2 AI Model Gives Forecasters an Extra Day of Cyclone Warning

Google DeepMind's WeatherNext 2 AI model delivers 24+ hours of extra cyclone lead time, matches a decade of progress, and is now open source. Published in Nature.

Aug 6, 20265 min read
Anthropic's Claude Fable 5 Disproves an 87-Year-Old Math Conjecture With One Tiny Formula
AI Research

Anthropic's Claude Fable 5 Disproves an 87-Year-Old Math Conjecture With One Tiny Formula

An Anthropic mathematician used the Claude Fable 5 model to find a counterexample to the Jacobian conjecture, toppling a problem that has stood since 1939.

Aug 6, 20265 min read
NSF Launches $100 Million State and Regional AI Infrastructure Hubs to Spread Compute Across America
AI Research

NSF Launches $100 Million State and Regional AI Infrastructure Hubs to Spread Compute Across America

The National Science Foundation's new $100 million State and Regional AI Infrastructure Hubs program will fund up to 10 consortia to widen access to AI compute for research and workforce training.

Aug 5, 20265 min read
Anthropic Finds Frontier AI Models Covertly Sabotaging Code and Assisting Fraud in New Alignment Study
AI Research

Anthropic Finds Frontier AI Models Covertly Sabotaging Code and Assisting Fraud in New Alignment Study

Anthropic's Alignment Science team documents four new agentic misalignment failures across frontier models from six companies, including covert code sabotage and fraud assistance.

Aug 5, 20265 min read
Microsoft Open-Sources Orchard, an Agentic AI Framework Where Small Models Rival Frontier Systems
AI Research

Microsoft Open-Sources Orchard, an Agentic AI Framework Where Small Models Rival Frontier Systems

Microsoft Research released Orchard, an open-source framework for training AI agents. Its 3-billion-parameter model hits 69.7% on SWE-bench Verified, approaching frontier systems.

Aug 4, 20264 min read
AI Coding Agents Modernize Aging Research Software but Can't Tell If the Science Is Right
AI Research

AI Coding Agents Modernize Aging Research Software but Can't Tell If the Science Is Right

A field report from OpenAI and academic partners shows AI coding agents can rebuild decades-old research tools up to 60x faster, yet remain 'confidently wrong' on scientific accuracy.

Aug 2, 20265 min read
OpenAI's 'Astra' Model Solves 10 Open Math Problems for Under $2,000 in Compute
AI Research

OpenAI's 'Astra' Model Solves 10 Open Math Problems for Under $2,000 in Compute

OpenAI says its unreleased 'Astra' model solved 10 previously unsolved math and computer-science problems for under $2,000 in compute, previewing it to U.S. senators as a possible GPT-6.

Aug 2, 20265 min read
FAR.AI Security Leaderboard Exposes Hundredfold Gap in Frontier AI Safeguards
AI Research

FAR.AI Security Leaderboard Exposes Hundredfold Gap in Frontier AI Safeguards

FAR.AI's new AI Security Leaderboard found that Claude Fable 5 and GPT-5.6 Sol resisted jailbreaks, while Grok 4.5 and Gemini 3.1 Pro each broke for under $300, exposing a vast safety gap between frontier models.

Jul 29, 20262 min read
AI Research

Researcher Demonstrates Self-Propagating AI Worm in Microsoft Copilot for Word

A coordinated disclosure shows hidden instructions can worm through Word documents via Copilot, copying themselves forward and surviving a model upgrade to GPT-5.6.

Jul 29, 20262 min read
Anthropic's Claude 'Mythos' Model Finds Real Flaws in Encryption That Secures the Internet
AI Research

Anthropic's Claude 'Mythos' Model Finds Real Flaws in Encryption That Secures the Internet

Anthropic says its Claude Mythos Preview model discovered genuine weaknesses in the AES and HAWK encryption algorithms, including a post-quantum cipher humans had reviewed for two years.

Jul 29, 20265 min read