🔍 Read the full analysis: What Makes LFM2.5-VL-DSpark Useful For Vision-Language AI? on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Liquid AI has released LFM2.5-VL-DSpark, an experimental 280 million-parameter drafter for its LFM2.5-VL-3B vision-language model. The company reports decoding speedups of up to 3.13x on Apple silicon and 2.66x on an NVIDIA H100, while greedy outputs remain the same because the target model verifies proposed tokens.
Liquid AI has released LFM2.5-VL-DSpark, an experimental add-on intended to speed up its LFM2.5-VL-3B vision-language model, as detailed in the original analysis. The company says the 280 million-parameter drafter raises decoding speed by as much as 3.13x on Apple silicon and 2.66x on an NVIDIA H100, while preserving the target model’s output under greedy decoding.
DSpark uses speculative decoding: a smaller drafter proposes a block of tokens, then the larger target model checks those proposals. Matching tokens can be accepted in batches, reducing the number of sequential generation steps. Liquid AI says the target model verifies every proposed token, so greedy decoding produces the same output as running the target model alone. The drafter adds about 8.9% to the target model’s parameter count.
The company reports results across six vision-language task types under the MMSpec benchmark protocol, including general and text-based visual question answering, image captioning, chart questions, complex reasoning and multi-turn conversation. With MLX on an M5 Max, Liquid AI reports decoding gains of 2.30x to 3.13x and end-to-end latency improvements of 1.56x to 2.62x. With llama.cpp on an M3 Ultra, it reports 1.57x to 2.14x faster decoding and 1.30x to 1.77x end-to-end gains.
For an NVIDIA H100, Liquid AI reports decoding speedups reaching 2.66x and end-to-end gains of 1.64x to 2.27x. The company says those GPU measurements use a DSpark block size of eight. The drafter is available on Hugging Face in Safetensors and GGUF formats, and integrations are available for llama.cpp, MLX-VLM and SGLang, subject to the relevant support patches or builds.
Faster Local Vision-Language Inference
The release addresses a practical issue for local and edge AI: vision-language models process image information as well as text, which can make responses slower than text-only generation. A smaller drafter that speeds token generation could make a 3-billion-parameter model feel more responsive on supported consumer hardware, particularly when users interact with it directly.
The reported figures apply most strongly to the decoding phase, when a model generates its response. They do not mean the full image-to-answer process becomes three times faster. Image encoding and prompt prefill still take time, and the company’s own end-to-end results are lower than its decoding gains. The difference matters for deployments where total response time, rather than token generation alone, determines whether an application feels fast.
Availability in established inference tools may also lower the effort required to try the approach. Liquid AI says the integrations support common deployment paths, while the public model files let developers evaluate the drafter with their own hardware and workloads. The practical value will depend on whether the speed gains hold for those workloads and outweigh the added memory use.
DSpark Moves Beyond Text Models
Liquid AI’s DSpark recipe was previously applied to text-only LFM2.5 models. This release extends the technique to a vision-language target. According to the company, image patches and text tokens are brought into a shared representation before the tapped model layers. The drafter can therefore use hidden-state vectors with the same dimensions across modalities, while retaining the inference algorithm used for the text models.
The drafter is an attention-only model with four layers, selected after comparisons of three-, four- and five-layer versions. It drafts blocks of nine tokens in the described configuration. Liquid AI says it trained the drafter on a mixture of vision-language supervised fine-tuning data weighted toward anticipated serving workloads, and that acceptance improved over ten training epochs before gains began to level off. The company has not published a detailed breakdown of that mixture.
Speculative decoding does not make the target model smaller or replace its output decisions. Instead, the additional model predicts possible next tokens for the target to check. Liquid AI says its drafter contains about 279.5 million parameters, including a 193.0 million-parameter decoder stack, a 21.0 million-parameter hidden-state projection and a 65.5 million-parameter Markov head, plus small normalization and confidence components.
“It adds a speculative decoding path that trades a minimal increase in memory footprint for a larger speedup without changing output quality.”
— Liquid AI
Benchmark Limits and Open Questions
The speed figures are Liquid AI’s own measurements; independent benchmark results were not included in the supplied material. Performance can vary with hardware, prompt composition and image resolution. The release is labeled experimental, and the company has not said when it might become a stable release or whether it will do so.
One GPU benchmark range in the announcement is reported as “20.4x to 2.66x,” a range whose lower value exceeds its upper value. That figure appears internally inconsistent; it may contain a typo, but Liquid AI has not clarified it. The supplied results therefore do not establish the intended lower bound for that range.
Other open questions include how well the drafter performs on vision tasks unlike its training data, how it affects sampling-based generation rather than greedy decoding, and how much runtime memory it adds beyond its parameter count. The training-data mixture also remains undescribed in detail. The company’s claim that greedy output is unchanged should not be generalized to every generation method or deployment condition without further evidence.
Independent Tests Will Show Practical Gains
Developers can evaluate the release using the public model files and supported inference integrations. The next useful evidence will be independent measurements across devices, image sizes and task types, including full response latency and memory use rather than decoding speed alone. Such results could show where the drafter helps most and whether the company’s reported gains carry over to everyday applications.
Liquid AI has not announced a date for a stable release or provided a schedule for further benchmark clarifications. Users evaluating the experimental model will need to account for the available integration versions and test the exact workloads they plan to run. Further details on the inconsistent GPU range, sampling behavior and training data have not been provided in the supplied announcement.
Key Questions
What is LFM2.5-VL-DSpark?
It is an experimental speculative-decoding drafter intended to speed up Liquid AI’s LFM2.5-VL-3B vision-language model. It proposes candidate tokens for the target model to verify.
Does DSpark change the model’s answers?
Liquid AI says greedy decoding produces the same output because the target model verifies each proposed token. The supplied information does not establish the same behavior for sampling-based generation.
How much faster is it?
Liquid AI reports decoding speedups of up to 3.13x on an M5 Max and up to 2.66x on an H100. End-to-end gains are lower in its reported tests, and performance will vary by hardware and workload.
Where can developers get it?
The drafter is available on Hugging Face in Safetensors and GGUF formats. Liquid AI lists support for llama.cpp, MLX-VLM and SGLang, with specific patches or builds required for those integrations.
Are the benchmark results independently verified?
The figures described here are company-reported results. Independent confirmation and broader testing across prompts, image resolutions and deployment settings were not included in the supplied source material.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
