
New benchmark exposes how badly AI struggles with real knowledge work
A new benchmark demonstrates that even the most advanced AI models perform poorly on realistic knowledge work tasks,...

A new benchmark demonstrates that even the most advanced AI models perform poorly on realistic knowledge work tasks,...

Liquid AI has introduced LFM2.5-Embedding-350M and LFM2.5-ColBERT-350M, two retrieval models combining dense bi-encoder...

Enterprises are reassessing their AI spending after an initial period of aggressive AI adoption ('tokenmaxxing')....

Zhipu AI released GLM-5.2, an open-source large language model with a 1-million-token context window under the MIT...

MiniMax released MSA, a sparse attention mechanism built on Grouped Query Attention that reduces per-token attention...
LifeSciBench is a new expert-authored and expert-reviewed benchmark designed to evaluate how AI systems perform on...

The Institute of the Estonian Language has released a benchmark that measures how vulnerable AI language models are to...

The article examines how companies are dealing with rising token usage costs as they deploy AI models, revealing...

Cohere has released North Mini, a code model designed to give AI developers more control and transparency compared to...

Z.ai launched GLM-5.2 on June 13, 2026, featuring a 1-million-token context window and two thinking-effort levels...

KPMG retracted a report on AI usage after discovering it contained hallucinations generated by AI tools. The incident...

Moonshot AI released Kimi K2.7 Code, an open-weights trillion-parameter model for programming that costs up to 12x less...

Zyphra released Zamba2-VL, a family of open vision-language models in sizes from 1.2B to 7B parameters using a hybrid...

A tutorial on building a code dataset pipeline using NVIDIA's Nemotron-Pretraining-Code-v3 metadata index for code...

The article discusses how tech companies are exploring the use of cheaper AI models that can maintain quality...

Google has ordered over three million AI chips from Intel for 2028, while Nvidia tests Intel's manufacturing technology...

Xiaomi's MiMo team and TileRT have released MiMo-V2.5-Pro-UltraSpeed, a serving mode that achieves over 1000 tokens per...

This tutorial demonstrates using GEPA, a reflective prompt-evolution framework, to optimize how small language models...

A new open-source voice model called Audio Interaction enables continuous listening and can decide every 0.4 seconds...