AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: What Makes LFM2.5-VL-DSpark Useful For Vision-Language AI? on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

Liquid AI has released LFM2.5-VL-DSpark, an experimental 280 million-parameter drafter for its LFM2.5-VL-3B vision-language model. The company reports decoding speedups of up to 3.13x on Apple silicon and 2.66x on an NVIDIA H100, while greedy outputs remain the same because the target model verifies proposed tokens.

Liquid AI has released LFM2.5-VL-DSpark, an experimental add-on intended to speed up its LFM2.5-VL-3B vision-language model, as detailed in the original analysis. The company says the 280 million-parameter drafter raises decoding speed by as much as 3.13x on Apple silicon and 2.66x on an NVIDIA H100, while preserving the target model’s output under greedy decoding.

DSpark uses speculative decoding: a smaller drafter proposes a block of tokens, then the larger target model checks those proposals. Matching tokens can be accepted in batches, reducing the number of sequential generation steps. Liquid AI says the target model verifies every proposed token, so greedy decoding produces the same output as running the target model alone. The drafter adds about 8.9% to the target model’s parameter count.

The company reports results across six vision-language task types under the MMSpec benchmark protocol, including general and text-based visual question answering, image captioning, chart questions, complex reasoning and multi-turn conversation. With MLX on an M5 Max, Liquid AI reports decoding gains of 2.30x to 3.13x and end-to-end latency improvements of 1.56x to 2.62x. With llama.cpp on an M3 Ultra, it reports 1.57x to 2.14x faster decoding and 1.30x to 1.77x end-to-end gains.

For an NVIDIA H100, Liquid AI reports decoding speedups reaching 2.66x and end-to-end gains of 1.64x to 2.27x. The company says those GPU measurements use a DSpark block size of eight. The drafter is available on Hugging Face in Safetensors and GGUF formats, and integrations are available for llama.cpp, MLX-VLM and SGLang, subject to the relevant support patches or builds.

At a glance
announcementWhen: Released in 2026; available now, with e…
The developmentLiquid AI released an experimental speculative-decoding drafter designed to speed up its LFM2.5-VL-3B vision-language model.
At a glance
announcementWhen: announced September 2026, available now
The developmentLiquid AI announced and released an experimental DSpark draft model for its LFM2.5-VL-3B vision-language model, available immediately on Hugging Face.

Faster Local Vision-Language Inference

The release addresses a practical issue for local and edge AI: vision-language models process image information as well as text, which can make responses slower than text-only generation. A smaller drafter that speeds token generation could make a 3-billion-parameter model feel more responsive on supported consumer hardware, particularly when users interact with it directly.

The reported figures apply most strongly to the decoding phase, when a model generates its response. They do not mean the full image-to-answer process becomes three times faster. Image encoding and prompt prefill still take time, and the company’s own end-to-end results are lower than its decoding gains. The difference matters for deployments where total response time, rather than token generation alone, determines whether an application feels fast.

Availability in established inference tools may also lower the effort required to try the approach. Liquid AI says the integrations support common deployment paths, while the public model files let developers evaluate the drafter with their own hardware and workloads. The practical value will depend on whether the speed gains hold for those workloads and outweigh the added memory use.

DSpark Moves Beyond Text Models

Liquid AI’s DSpark recipe was previously applied to text-only LFM2.5 models. This release extends the technique to a vision-language target. According to the company, image patches and text tokens are brought into a shared representation before the tapped model layers. The drafter can therefore use hidden-state vectors with the same dimensions across modalities, while retaining the inference algorithm used for the text models.

The drafter is an attention-only model with four layers, selected after comparisons of three-, four- and five-layer versions. It drafts blocks of nine tokens in the described configuration. Liquid AI says it trained the drafter on a mixture of vision-language supervised fine-tuning data weighted toward anticipated serving workloads, and that acceptance improved over ten training epochs before gains began to level off. The company has not published a detailed breakdown of that mixture.

Speculative decoding does not make the target model smaller or replace its output decisions. Instead, the additional model predicts possible next tokens for the target to check. Liquid AI says its drafter contains about 279.5 million parameters, including a 193.0 million-parameter decoder stack, a 21.0 million-parameter hidden-state projection and a 65.5 million-parameter Markov head, plus small normalization and confidence components.

“It adds a speculative decoding path that trades a minimal increase in memory footprint for a larger speedup without changing output quality.”

— Liquid AI

Benchmark Limits and Open Questions

The speed figures are Liquid AI’s own measurements; independent benchmark results were not included in the supplied material. Performance can vary with hardware, prompt composition and image resolution. The release is labeled experimental, and the company has not said when it might become a stable release or whether it will do so.

One GPU benchmark range in the announcement is reported as “20.4x to 2.66x,” a range whose lower value exceeds its upper value. That figure appears internally inconsistent; it may contain a typo, but Liquid AI has not clarified it. The supplied results therefore do not establish the intended lower bound for that range.

Other open questions include how well the drafter performs on vision tasks unlike its training data, how it affects sampling-based generation rather than greedy decoding, and how much runtime memory it adds beyond its parameter count. The training-data mixture also remains undescribed in detail. The company’s claim that greedy output is unchanged should not be generalized to every generation method or deployment condition without further evidence.

Independent Tests Will Show Practical Gains

Developers can evaluate the release using the public model files and supported inference integrations. The next useful evidence will be independent measurements across devices, image sizes and task types, including full response latency and memory use rather than decoding speed alone. Such results could show where the drafter helps most and whether the company’s reported gains carry over to everyday applications.

Liquid AI has not announced a date for a stable release or provided a schedule for further benchmark clarifications. Users evaluating the experimental model will need to account for the available integration versions and test the exact workloads they plan to run. Further details on the inconsistent GPU range, sampling behavior and training data have not been provided in the supplied announcement.

Key Questions

What is LFM2.5-VL-DSpark?

It is an experimental speculative-decoding drafter intended to speed up Liquid AI’s LFM2.5-VL-3B vision-language model. It proposes candidate tokens for the target model to verify.

Does DSpark change the model’s answers?

Liquid AI says greedy decoding produces the same output because the target model verifies each proposed token. The supplied information does not establish the same behavior for sampling-based generation.

How much faster is it?

Liquid AI reports decoding speedups of up to 3.13x on an M5 Max and up to 2.66x on an H100. End-to-end gains are lower in its reported tests, and performance will vary by hardware and workload.

Where can developers get it?

The drafter is available on Hugging Face in Safetensors and GGUF formats. Liquid AI lists support for llama.cpp, MLX-VLM and SGLang, with specific patches or builds required for those integrations.

Are the benchmark results independently verified?

The figures described here are company-reported results. Independent confirmation and broader testing across prompts, image resolutions and deployment settings were not included in the supplied source material.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Run macOS Software On Linux

Interest in running macOS applications on Linux is surging, driven by community experiments and unofficial tools, though official support remains absent.

Explanation Of Everything You Can See In Htop/top On Linux (2019)

Detailed explanation of all elements visible in htop and top commands on Linux, clarifying what each component represents and how to interpret system metrics.

When a Scanner Saves More Time Than a Printer

How can a scanner revolutionize your workflow and save time over a printer? Discover the surprising benefits that await you.

7 Best PC Routers for Prime Day Deals in 2026

Discover the best PC routers on Prime Day 2026, including WiFi 7, wired ports, and setup options. Find your ideal router today.