LLM · Other Companies

New benchmark exposes how badly AI struggles with real knowledge work
LLM

New benchmark exposes how badly AI struggles with real knowledge work

A new benchmark demonstrates that even the most advanced AI models perform poorly on realistic knowledge work tasks,...

The Decoder
Liquid AI Introduces LFM2.5-Embedding-350M and LFM2.5-ColBERT-350M: Dense Bi-Encoder and Late-Interaction Models for Fast Multilingual Search Across 11 Languages
LLM

Liquid AI Introduces LFM2.5-Embedding-350M and LFM2.5-ColBERT-350M: Dense Bi-Encoder and Late-Interaction Models for Fast Multilingual Search Across 11 Languages

Liquid AI has introduced LFM2.5-Embedding-350M and LFM2.5-ColBERT-350M, two retrieval models combining dense bi-encoder...

MarkTechPost
NEA’s Tiffany Luck says enterprises are still figuring out their AI ROI
LLM

NEA’s Tiffany Luck says enterprises are still figuring out their AI ROI

Enterprises are reassessing their AI spending after an initial period of aggressive AI adoption ('tokenmaxxing')....

TechCrunch AI
Zhipu AI's GLM-5.2 closes in on closed-source leaders in coding marathons
LLM

Zhipu AI's GLM-5.2 closes in on closed-source leaders in coding marathons

Zhipu AI released GLM-5.2, an open-source large language model with a 1-million-token context window under the MIT...

The Decoder
MiniMax Sparse Attention (MSA): a Two-Branch Block-Sparse Attention Trained on a 109B-Parameter MoE With a 3T-Token Budget
LLM

MiniMax Sparse Attention (MSA): a Two-Branch Block-Sparse Attention Trained on a 109B-Parameter MoE With a 3T-Token Budget

MiniMax released MSA, a sparse attention mechanism built on Grouped Query Attention that reduces per-token attention...

MarkTechPost
LLM

Introducing LifeSciBench

LifeSciBench is a new expert-authored and expert-reviewed benchmark designed to evaluate how AI systems perform on...

OpenAI Blog
How easily can Russian propaganda fool AI models? A new benchmark finds out
LLM

How easily can Russian propaganda fool AI models? A new benchmark finds out

The Institute of the Estonian Language has released a benchmark that measures how vulnerable AI language models are to...

The Decoder
‘Pretty Crazy’ Token Usage Is Testing Bosses’ Bet on AI
LLM

‘Pretty Crazy’ Token Usage Is Testing Bosses’ Bet on AI

The article examines how companies are dealing with rising token usage costs as they deploy AI models, revealing...

Wired AI
Cohere North Mini Code Gives AI Developers More Control
LLM

Cohere North Mini Code Gives AI Developers More Control

Cohere has released North Mini, a code model designed to give AI developers more control and transparency compared to...

AI Business
Z.ai Launches GLM-5.2 With a Usable 1M-Token Context, Two Thinking-Effort Levels, and No Benchmarks at Launch
LLM

Z.ai Launches GLM-5.2 With a Usable 1M-Token Context, Two Thinking-Effort Levels, and No Benchmarks at Launch

Z.ai launched GLM-5.2 on June 13, 2026, featuring a 1-million-token context window and two thinking-effort levels...

MarkTechPost
KPMG pulls report on AI usage due to apparent hallucinations
LLM

KPMG pulls report on AI usage due to apparent hallucinations

KPMG retracted a report on AI usage after discovering it contained hallucinations generated by AI tools. The incident...

TechCrunch AI
Open model Kimi K2.7 Code undercuts GPT-5.5 and Claude by up to 12x on price per token
LLM

Open model Kimi K2.7 Code undercuts GPT-5.5 and Claude by up to 12x on price per token

Moonshot AI released Kimi K2.7 Code, an open-weights trillion-parameter model for programming that costs up to 12x less...

The Decoder
Zyphra Release Zamba2-VL: Hybrid Mamba2–Transformer Vision-Language Models That Cut Time-to-First-Token by About an Order of Magnitude
LLM

Zyphra Release Zamba2-VL: Hybrid Mamba2–Transformer Vision-Language Models That Cut Time-to-First-Token by About an Order of Magnitude

Zyphra released Zamba2-VL, a family of open vision-language models in sizes from 1.2B to 7B parameters using a hybrid...

MarkTechPost
Building a Code Dataset Pipeline from NVIDIA Nemotron-Pretraining-Code-v3 Metadata with Streaming, Pandas, and tiktoken
LLM

Building a Code Dataset Pipeline from NVIDIA Nemotron-Pretraining-Code-v3 Metadata with Streaming, Pandas, and tiktoken

A tutorial on building a code dataset pipeline using NVIDIA's Nemotron-Pretraining-Code-v3 metadata index for code...

MarkTechPost
Can tech companies learn to love cheaper AI models? 
LLM

Can tech companies learn to love cheaper AI models? 

The article discusses how tech companies are exploring the use of cheaper AI models that can maintain quality...

TechCrunch AI
Intel gets a second life as Google and Nvidia explore it as a TSMC backup for AI chips
LLM

Intel gets a second life as Google and Nvidia explore it as a TSMC backup for AI chips

Google has ordered over three million AI chips from Intel for 2028, while Nvidia tests Intel's manufacturing technology...

The Decoder
Xiaomi MiMo and TileRT Push a 1-Trillion-Parameter Model Past 1000 Tokens Per Second on Commodity GPUs
LLM

Xiaomi MiMo and TileRT Push a 1-Trillion-Parameter Model Past 1000 Tokens Per Second on Commodity GPUs

Xiaomi's MiMo team and TileRT have released MiMo-V2.5-Pro-UltraSpeed, a serving mode that achieves over 1000 tokens per...

MarkTechPost
Building Reflective Prompt Optimization with GEPA: Multi-Component Prompts, Structured Feedback, and Held-Out Validation
LLM

Building Reflective Prompt Optimization with GEPA: Multi-Component Prompts, Structured Feedback, and Held-Out Validation

This tutorial demonstrates using GEPA, a reflective prompt-evolution framework, to optimize how small language models...

MarkTechPost
New open-source voice model listens nonstop and decides every 0.4 seconds whether to speak or stay silent
LLM

New open-source voice model listens nonstop and decides every 0.4 seconds whether to speak or stay silent

A new open-source voice model called Audio Interaction enables continuous listening and can decide every 0.4 seconds...

The Decoder