📊 Full opportunity report: Minerva. The opposite path. on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Italy’s Minerva, a large-scale European sovereign LLM trained from scratch, achieved impressive technical results but underperformed on Italian academic benchmarks. This raises questions about the scale of native-language investment needed for true country-specific language understanding.
Italy’s Minerva project, a large-scale sovereign language model trained entirely from scratch on 2.5 trillion tokens, scored just 4.9% on the INVALSI Italian school-exam benchmark, despite extensive investment and infrastructure. This outcome questions assumptions about the relationship between training scale and language understanding, and highlights the complexity of developing truly country-specific AI models.
The Minerva project, led by Sapienza University of Rome’s NLP group and supported by Italy’s national research infrastructure, trained models ranging from 350 million to 7 billion parameters using roughly 50% Italian data. The training involved 2.5 trillion tokens, with approximately 1.14 trillion Italian tokens, making it one of Europe’s most ambitious efforts to build a sovereign LLM from scratch.
Despite the scale and open publication of weights, data, and code, Minerva-3B scored only 4.9% on the INVALSI exam, a standardized Italian school assessment. Researchers concluded that, while dataset composition is important, the overall dataset size and model parameters are more crucial for handling complex language tasks. This suggests that even large-scale native-language training may be insufficient at current parameter levels to produce deep country-specific knowledge.
The results present a structural challenge: Italy’s investment produced a technically impressive model but not the language understanding needed for high-level academic tasks. The comparison with Portugal’s AMÁLIA, which layered specialization onto a multilingual foundation, underscores that the key issue is the scale of native-language investment required for meaningful country-knowledge depth.
Minerva.
The opposite
path.
Italy spent years building a European sovereign LLM from scratch. Then Minerva-3B scored 4.9% on the INVALSI Italian school exam.
Where AMÁLIA layered Portuguese specialization onto a multilingual foundation, Minerva trained from scratch on 2.5 trillion tokens with approximately 50% Italian content. Where AMÁLIA’s weights are not yet public, Minerva published weights, training data, and code as truly-open from day one. By every institutional measure, the Italian approach worked. But the empirical results contain a finding the press coverage has been quiet about — and it has implications that extend well beyond Italy.
Same problem. Opposite path.
European sovereign-LLM development has two primary architectural approaches. Italy chose from scratch with substantial native-language foundation. Portugal chose continuation pre-training of a multilingual model. The structural comparison surfaces what each commitment actually requires operationally.
The comparison is not “Italy did it better than Portugal.” Both projects respond to the same structural problem with different architectural strategies under different institutional and economic constraints. Italy’s national-AI investment is structurally larger by an order of magnitude — and Minerva is the visible artifact of that scale.

Advanced Language Tool Kit: Teaching the Structure of the English Language
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
4.9% on INVALSI. The bitter lesson surfaces.
In June 2024, researchers evaluated Minerva-3B on the Italian school-exam benchmark. The result was unambiguous. This is not a critique of Minerva — it is a critique of the public discourse around what Minerva’s empirical results actually demonstrate.

Large-Scale AI Engineering: Design, Train, and Optimize Foundation Models on NVIDIA GPU Clusters
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
350M to 7B. Four parameter scales, one architecture.
The Minerva model family covers four parameter tiers, each with specific training corpora. Each scale level reveals what the from-scratch path actually requires at different operating points.
Italian + English
100B English
~50% English
+ 200B code

AI Data Preparation Guide: Fuel AI With Quality Data | Labeling Tools Explained | Human-in-the-Loop Best Practices | Prepare to Train Smarter | Annotate for Success | Annotation Drives Intelligence
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Three answers. Same question.
Minerva, AMÁLIA, and OpenEuroLLM represent the three operational answers to the European sovereign-LLM question. Each makes different architectural and institutional bets. The strategic discourse benefits from treating all three as data points in the same empirical experiment.

AI Engineering: Building Applications with Foundation Models
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Three standards the movement should adopt.
The structural critique generalizes beyond Minerva. The European sovereign-LLM movement benefits from internalizing these lessons across every subsequent national project. Italy modeled the openness standard; the movement should adopt it as norm.
Minerva is one valid answer to the European sovereign-LLM question. AMÁLIA is another. OpenEuroLLM is potentially a third. The strategic discourse benefits from treating all three as data points in the same empirical experiment rather than as competing national-prestige projects. More analysis like this is needed. Not less.
Implications for European Sovereign-Language Models
The Minerva results demonstrate that simply increasing training data and model size may not be enough to develop models capable of understanding and performing complex tasks in a native language. For European countries pursuing sovereign LLMs, this suggests a need to reevaluate investment strategies, particularly regarding native-language data scale and model parameters. The findings challenge the narrative that ‘training from scratch’ alone guarantees high performance in country-specific tasks and highlight the importance of considering scale and resource commitments.
This has broader implications for European AI policy, emphasizing that building effective national models requires substantial, targeted investment beyond initial data collection. The results also raise questions about the feasibility of achieving true language and knowledge depth at current parameter scales, potentially influencing future AI infrastructure and funding decisions across Europe.
Background on European Sovereign-LLM Strategies
European efforts to develop sovereign language models have been characterized by contrasting approaches. Portugal’s AMÁLIA project, for example, layered Portuguese specialization onto a multilingual foundation, using a smaller proportion of native data. Conversely, Italy’s Minerva trained from scratch on a vastly larger dataset, with a focus on Italian content, and published its weights and data openly. The debate has centered on whether ‘from scratch’ training or continued pre-training offers better results for country-specific language understanding.
Prior to Minerva’s release, the general expectation was that larger native-language datasets and models would lead to superior performance. The Italian project represented a significant investment—both financially and infrastructurally—supported by Italy’s national AI strategy and supercomputing resources. However, the recent benchmark results challenge assumptions about the sufficiency of scale alone for complex language tasks.
“While dataset composition is important, the overall size of the dataset and the number of parameters are more crucial for handling complex language tasks.”
— Research team of Minerva
Unresolved Questions About Scale and Performance
It remains unclear whether increasing model size beyond current levels, or further targeted native-language training, will significantly improve performance on complex academic and reasoning tasks. The results are based on a single benchmark, and broader evaluations are ongoing. Additionally, the long-term impact of different training strategies on real-world applications is still being studied.
Next Steps in European Sovereign-Language AI Development
The Minerva team plans to continue iterating on their models, including experiments with larger parameter counts and different training data compositions. European policymakers and researchers are likely to reassess investment strategies, emphasizing the importance of scale, data quality, and specialized training. Further benchmarking and real-world testing are expected to clarify the path toward more effective country-specific language models.
Key Questions
Why did Minerva perform poorly on the Italian academic benchmark?
Despite extensive training on a large dataset, the results suggest that the current scale of parameters and data may still be insufficient to develop deep, country-specific knowledge necessary for high-level academic tasks.
How does Minerva compare to Portugal’s AMÁLIA project?
While Minerva trained from scratch on a larger native dataset, its performance on complex tasks was limited, contrasting with AMÁLIA’s approach of layered specialization, which has yet to publish comparable benchmarks.
What does this mean for future European AI projects?
It indicates that significant native-language investment, both in data and model size, is necessary to achieve country-specific language understanding, and that scale alone may not be enough.
Are there plans to improve Minerva’s performance?
Yes, the team plans ongoing experiments with larger models and different training strategies, aiming to enhance performance on complex language tasks.
What are the broader implications for AI policy in Europe?
European policymakers may need to reconsider funding allocations and strategic priorities, emphasizing the importance of scale, data quality, and targeted training for sovereign language models.
Source: ThorstenMeyerAI.com