New Open Decision Models Handle Text, Images, and Audio Fast
Liquid AI released new open decision models called d1-3B and d1-omni-600M. These tools make fast choices using text, images, and audio.
In short: Liquid AI released new open decision models called d1-3B and d1-omni-600M. These tools make fast choices using text, images, and audio.
Imagine a computer program that looks at a picture or reads a message and immediately gives you a structured answer without having to chat back and forth.
What happened, in plain words
Liquid AI announced two new open-weight decision models named d1-3B and d1-omni-600M. Unlike typical chat systems that write out long sentences word by word, these models provide answers in a single step called a forward pass. The team tested the models across seven public topics, such as reading comprehension and medical questions, and checked their speed on various hardware devices.
Key points
- Top scores in size category The d1-3B model scored 48.57 on the Decision Index 0.2.1, beating out several larger models.
- Handling different types of media While d1-3B accepts text and images, the d1-omni-600M model can handle text combined with either images or audio.
- High-speed responses On tested hardware, d1-3B answers a question in under 50 milliseconds across all measured edge devices.
- Early research stage The d1-omni-600M model is currently an early research release undergoing further development.
Terms explained
- Decision Models — Computer tools designed to pick an answer or make a choice instead of writing long paragraphs. Example: A program that looks at a customer complaint and sorts it into the correct support category.
- Multimodal — The ability of a system to understand more than one type of input, such as both pictures and words. Example: An app that can read a text caption and look at a photograph at the same time.
- Parameters — The internal settings of a computer model that help it learn patterns and make choices. Example: The adjustable dials on a sound mixing board that shape the final audio.
Why it matters
These models can help developers build tools that sort information, check images, or analyze text very quickly on local hardware devices.
What we still don't know
The d1-omni-600M model is an early release, and audio decision benchmarks remain an open problem without official scores reported in the source.
Based on reporting from Hugging Face Blog. This is an independent explainer, written in our own words with AI assistance; Hugging Face Blog has not reviewed or endorsed it. Read the original for the full details.