LLM · Other Companies

GameStop offers $56 billion for eBay, struggles to explain how it'll pay for it
LLM

GameStop offers $56 billion for eBay, struggles to explain how it'll pay for it

Amid falling revenue and store closures, GameStop wants to buy the much larger eBay.

Ars Technica AI
A Developer’s Guide to Systematic Prompting: Mastering Negative Constraints, Structured JSON Outputs, and Multi-Hypothesis Verbalized Sampling
LLM

A Developer’s Guide to Systematic Prompting: Mastering Negative Constraints, Structured JSON Outputs, and Multi-Hypothesis Verbalized Sampling

This article provides a developer's guide to systematic prompting techniques for large language models, focusing on...

MarkTechPost
Xiaomi's open-weight MiMo-V2.5-Pro takes aim at Claude Opus with hours-long autonomous coding
LLM

Xiaomi's open-weight MiMo-V2.5-Pro takes aim at Claude Opus with hours-long autonomous coding

Xiaomi released MiMo-V2.5-Pro, an open-weight model that nearly matches Anthropic's Claude Opus 4.6 on coding...

The Decoder
What is Tokenization Drift and How to Fix It?
LLM

What is Tokenization Drift and How to Fix It?

The article discusses tokenization drift, a subtle but critical issue where AI models can degrade in performance due to...

MarkTechPost
Even the latest AI models make three systematic reasoning errors, ARC-AGI-3 analysis shows
LLM

Even the latest AI models make three systematic reasoning errors, ARC-AGI-3 analysis shows

Analysis of OpenAI's GPT-5.5 and Anthropic's Opus 4.7 on the ARC-AGI-3 benchmark reveals three systematic reasoning...

The Decoder
A Coding Guide on LLM Post Training with TRL from Supervised Fine Tuning to DPO and GRPO Reasoning
LLM

A Coding Guide on LLM Post Training with TRL from Supervised Fine Tuning to DPO and GRPO Reasoning

A hands-on tutorial guide for post-training large language models using the TRL library, covering four key techniques:...

MarkTechPost
Qwen AI Releases Qwen-Scope: An Open-Source Sparse AutoEncoders (SAE) Suite That Turns LLM Internal Features into Practical Development Tools
LLM

Qwen AI Releases Qwen-Scope: An Open-Source Sparse AutoEncoders (SAE) Suite That Turns LLM Internal Features into Practical Development Tools

Qwen AI has released Qwen-Scope, an open-source Sparse Autoencoder suite that provides tools to interpret and utilize...

MarkTechPost
Moonshot AI Open-Sources FlashKDA: CUTLASS Kernels for Kimi Delta Attention with Variable-Length Batching and H20 Benchmarks
LLM

Moonshot AI Open-Sources FlashKDA: CUTLASS Kernels for Kimi Delta Attention with Variable-Length Batching and H20 Benchmarks

Moonshot AI has open-sourced FlashKDA, a high-performance implementation of Kimi Delta Attention optimized for the...

MarkTechPost
Tencent's 440 MB AI model translates 33 languages offline on your phone
LLM

Tencent's 440 MB AI model translates 33 languages offline on your phone

Tencent released a compact 440 MB AI translation model as an open-weight model that supports 33 languages and runs...

The Decoder
IBM Releases Two Granite Speech 4.1 2B Models: Autoregressive ASR with Translation and Non-Autoregressive Editing for Fast Inference
LLM

IBM Releases Two Granite Speech 4.1 2B Models: Autoregressive ASR with Translation and Non-Autoregressive Editing for Fast Inference

IBM has released two versions of the Granite Speech 4.1 2B model for automatic speech recognition (ASR), featuring both...

MarkTechPost
Qwen Team Releases FlashQLA: a High-Performance Linear Attention Kernel Library That Achieves Up to 3× Speedup on NVIDIA Hopper GPUs
LLM

Qwen Team Releases FlashQLA: a High-Performance Linear Attention Kernel Library That Achieves Up to 3× Speedup on NVIDIA Hopper GPUs

The QwenLM team released FlashQLA, a high-performance kernel library that accelerates linear attention mechanisms,...

MarkTechPost
Sanctioned Chinese AI Firm SenseTime Releases Image Model Built for Speed
LLM

Sanctioned Chinese AI Firm SenseTime Releases Image Model Built for Speed

SenseTime, a sanctioned Chinese AI firm, has released a new image generation model optimized to run on Chinese-made...

Wired AI
With Nemotron 3 Nano Omni, Nvidia reveals what really goes into a modern multimodal model
LLM

With Nemotron 3 Nano Omni, Nvidia reveals what really goes into a modern multimodal model

Nvidia releases Nemotron 3 Nano Omni, an open multimodal model capable of processing text, image, video, and audio. The...

The Decoder
Here is what an LLM that knows nothing after 1930 thinks our world looks like in 2026
LLM

Here is what an LLM that knows nothing after 1930 thinks our world looks like in 2026

A 13-parameter language model called 'Talkie' trained exclusively on texts from before 1931 generates predictions for...

The Decoder
OpenMOSS Releases MOSS-Audio: An Open-Source Foundation Model for Speech, Sound, Music, and Time-Aware Audio Reasoning
LLM

OpenMOSS Releases MOSS-Audio: An Open-Source Foundation Model for Speech, Sound, Music, and Time-Aware Audio Reasoning

OpenMOSS released MOSS-Audio, an open-source foundation model that unifies speech, sound, music, and temporal reasoning...

MarkTechPost
The company with a monopoly on AI's most critical machine is racing to build more
LLM

The company with a monopoly on AI's most critical machine is racing to build more

ASML, which holds a monopoly on EUV lithography machines essential for AI chip production, is significantly increasing...

The Decoder
The LoRA Assumption That Breaks in Production 
LLM

The LoRA Assumption That Breaks in Production 

LoRA, a popular efficient fine-tuning method for large models, relies on the assumption that all model updates are...

MarkTechPost
How to Build a Fully Searchable AI Knowledge Base with OpenKB, OpenRouter, and Llama
LLM

How to Build a Fully Searchable AI Knowledge Base with OpenKB, OpenRouter, and Llama

This tutorial demonstrates how to build a searchable AI knowledge base using OpenKB with Llama models accessed through...

MarkTechPost
500 investment bankers review AI outputs and find none ready for client delivery
LLM

500 investment bankers review AI outputs and find none ready for client delivery

A benchmark study where 500 investment bankers evaluated outputs from leading AI models like GPT-5.4 and Claude Opus...

The Decoder