AI is getting stronger. Are we ready?

When AI looks all-powerful

AI can already write code, solve math problems, generate images, and is entering hospitals and labs. Is it replacing people across the board, or just unusually good at a few tasks? Stanford Institute for Human-Centered Artificial Intelligence’s AI Index Report 2026 makes this year’s sharpest tension clear: model capabilities race ahead, while governance, evaluation, education, and data systems fall further behind.

The index spans R&D, performance, the economy, science, medicine, education, governance, and public opinion. What lingers is not AI improving at a steady pace, but several things at once: capability jumps, real utility appearing, measurement tools going askew, and society lagging in readiness.

Brilliant—and still spectacularly wrong

What misleads most is uneven intelligence. The same system can look expert on hard problems and fail at seemingly simple ones. Gemini Deep Think can reach gold-medal level at the International Mathematical Olympiad, yet reads traditional analog clocks at only 50.1% accuracy versus about 90.1% for humans. One impressive score does not prove stable, general understanding.

Worse, the scoring yardsticks can be wrong. Some traditional benchmark items have error rates as high as 42%, and leaderboards can be distorted by training-data contamination and targeted gaming. Converging model scores mostly show intensifying competition—not that systems are equally reliable in the real world.

Supply is concentrating among a few players. In 2025, among 102 notable models released, 93 came from industry—about 91.2%—and academia contributed only 2; many models also withhold training code. Compute and capital decide who trains the next systems, making independent verification harder for outside researchers.

AI’s impact on work is uneven

The economic ledger is mixed too. Generative AI created about $172 billion in annual consumer surplus in the U.S.—the extra value users feel from tools worth keeping. Many already find chat, writing, and search more convenient, but that does not mean corporate profits rose in step, or that every industry shared the same gains.

Efficiency gains show up most where tasks are well-bounded and results easy to check. In customer service and software development, AI brings roughly 14% to 26% productivity gains. The evidence better supports helping people finish a piece of work than rewriting entire careers.

Job effects also split by experience. Among software developers aged 22 to 25, employment compared with 2024 fell nearly 20%; senior developers held up better. Pressure on junior roles is a warning, but interest rates, tech layoffs, and hiring freezes may be mixed in. Available data still cannot separate how much of that 20% is truly from AI.

Process uses in medicine and research

In healthcare, the more solid use is letting AI handle recording, organizing, and alerts first—not diagnosing alone. Multiple institutions have deployed ambient AI clinical documentation tools, cutting documentation time by up to 83%, with return on investment as high as 112%. Those figures show specific workflows sped up; they do not mean AI can independently make complex medical decisions.

Research has similar boundaries. AI can help propose hypotheses, process materials, and design experiments; but on PaperArena, which requires reading literature, using tools, and finishing end-to-end research, the best AI agent scored only 38.8% accuracy versus 83.5% for human PhD experts. The better role remains collaborator, not a reliable autonomous scientist taking over full research.

Tech outruns social readiness

Capabilities race ahead; rules and evaluation lag. Models may pass standard safety tests, then collapse under deliberate jailbreaks; stronger privacy protection can also pull against accuracy and fairness. More deployments do not automatically mean controllable risk.

Public feeling also diverges from expert expectations. A Pew survey found 73% of AI experts think AI will improve how people work, versus only 23% of the general public—a 50-point gap. Builders and researchers more often see efficiency; people facing job change more often feel uncertainty.

The so-called global picture still has large blanks. Some enterprise surveys omit China; international education data miss China, India, and many African countries; GitHub also misses other code platforms. This index is good for broad trends, not for every country, industry, or household’s shared experience.

Rewrite tasks first, then careers

A safer reading: AI first rewrites concrete steps in work—organizing, generating, searching, and some analysis get faster, and assistive flows in medicine and research show gains; full jobs that need long-term planning, real-world verification, accountability, and complex collaboration still lag humans badly.

Capability, adoption, and capital are accelerating together; productivity evidence, job attribution, evaluation reliability, and governance readiness still do not. It can show which tasks are already hit, but not whether a given person will lose a job next year or how much a company will definitely save.


Source institutions:Stanford Institute for Human-Centered Artificial Intelligence

This content is for reading and understanding research reports. It does not constitute investment advice or trading signals.

Read in App

Read global research reports on mobile.

This content is for research reading and does not constitute investment advice.


了解 InCosmos Vision 的更多信息

订阅后即可通过电子邮件收到最新文章。

本文内容基于公开信息整理与数据分析,不构成投资建议,不构成任何金融产品的买卖要约。大宗商品投资涉及显著风险,历史表现不预示未来结果。

了解 InCosmos Vision 的更多信息

立即订阅以继续阅读并访问完整档案。

继续阅读