When AI looks all-powerful
AI can already write code, solve math problems, generate images, and is entering hospitals and labs. Is it replacing people across the board, or just unusually good at a few tasks? Stanford Institute for Human-Centered Artificial Intelligence’s AI Index Report 2026 makes this year’s sharpest tension clear: model capabilities race ahead, while governance, evaluation, education, and data systems fall further behind.
The index spans R&D, performance, the economy, science, medicine, education, governance, and public opinion. What lingers is not AI improving at a steady pace, but several things at once: capability jumps, real utility appearing, measurement tools going askew, and society lagging in readiness.
Brilliant—and still spectacularly wrong
What misleads most is uneven intelligence. The same system can look expert on hard problems and fail at seemingly simple ones. Gemini Deep Think can reach gold-medal level at the International Mathematical Olympiad, yet reads traditional analog clocks at only 50.1% accuracy versus about 90.1% for humans. One impressive score does not prove stable, general understanding.
Worse, the scoring yardsticks can be wrong. Some traditional benchmark items have error rates as high as 42%, and leaderboards can be distorted by training-data contamination and targeted gaming. Converging model scores mostly show intensifying competition—not that systems are equally reliable in the real world.
Supply is concentrating among a few players. In 2025, among 102 notable models released, 93 came from industry—about 91.2%—and academia contributed only 2; many models also withhold training code. Compute and capital decide who trains the next systems, making independent verification harder for outside researchers.
AI’s impact on work is uneven
The economic ledger is mixed too. Generative AI created about $172 billion in annual consumer surplus in the U.S.—the extra value users feel from tools worth keeping. Many already find chat, writing, and search more convenient, but that does not mean corporate profits rose in step, or that every industry shared the same gains.
Efficiency gains show up most where tasks are well-bounded and results easy to check. In customer service and software development, AI brings roughly 14% to 26% productivity gains. The evidence better supports helping people finish a piece of work than rewriting entire careers.
Job effects also split by experience. Among software developers aged 22 to 25, employment compared with 2024 fell nearly 20%; senior developers held up better. Pressure on junior roles is a warning, but interest rates, tech layoffs, and hiring freezes may be mixed in. Available data still cannot separate how much of that 20% is truly from AI.
Process uses in medicine and research
In healthcare, the more solid use is letting AI handle recording, organizing, and alerts first—not diagnosing alone. Multiple institutions have deployed ambient AI clinical documentation tools, cutting documentation time by up to 83%, with return on investment as high as 112%. Those figures show specific workflows sped up; they do not mean AI can independently make complex medical decisions.
Research has similar boundaries. AI can help propose hypotheses, process materials, and design experiments; but on PaperArena, which requires reading literature, using tools, and finishing end-to-end research, the best AI agent scored only 38.8% accuracy versus 83.5% for human PhD experts. The better role remains collaborator, not a reliable autonomous scientist taking over full research.
Tech outruns social readiness
Capabilities race ahead; rules and evaluation lag. Models may pass standard safety tests, then collapse under deliberate jailbreaks; stronger privacy protection can also pull against accuracy and fairness. More deployments do not automatically mean controllable risk.
Public feeling also diverges from expert expectations. A Pew survey found 73% of AI experts think AI will improve how people work, versus only 23% of the general public—a 50-point gap. Builders and researchers more often see efficiency; people facing job change more often feel uncertainty.
The so-called global picture still has large blanks. Some enterprise surveys omit China; international education data miss China, India, and many African countries; GitHub also misses other code platforms. This index is good for broad trends, not for every country, industry, or household’s shared experience.
Rewrite tasks first, then careers
A safer reading: AI first rewrites concrete steps in work—organizing, generating, searching, and some analysis get faster, and assistive flows in medicine and research show gains; full jobs that need long-term planning, real-world verification, accountability, and complex collaboration still lag humans badly.
Capability, adoption, and capital are accelerating together; productivity evidence, job attribution, evaluation reliability, and governance readiness still do not. It can show which tasks are already hit, but not whether a given person will lose a job next year or how much a company will definitely save.
Source institutions:Stanford Institute for Human-Centered Artificial Intelligence
This content is for reading and understanding research reports. It does not constitute investment advice or trading signals.
Read in App
Read global research reports on mobile.
This content is for research reading and does not constitute investment advice.