📊 Full opportunity report: AMÁLIA · The Three Hard Questions. on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Portugal launched AMÁLIA, a €5.5 million European Portuguese LLM, which outperforms previous models on many benchmarks. However, critical questions about its openness, native data use, and goals remain unresolved, highlighting broader issues in European sovereign AI efforts.
Portugal’s €5.5 million AMÁLIA large language model is now operational, with a base version released in late September 2025, marking a significant step in the country’s AI ambitions. However, despite its technical achievements, key questions about its openness, native-language data, and strategic purpose remain unresolved, raising concerns about the broader European sovereign-LLM landscape.
AMÁLIA is a consortium project involving approximately 60 researchers from Portugal’s top institutions, including NOVA, IST, and IT, funded by the government and announced in December 2024. It is built as a continuation of the multilingual EuroLLM foundation, rather than training from scratch, with the base model handling Portuguese text and knowledge up to the end of 2023. The model outperforms previous open models on Portuguese benchmarks and beats Qwen 3-8B on most tests, though it still trails on specific tasks like ALBA, its primary benchmark.
Technical details reveal that only about 5.8 billion tokens from Portugal’s web archive were used in the extended pre-training phase, representing roughly 5.5% of the total tokens. The supervised fine-tuning phase increased the Portuguese data share to approximately 17-18%. The final version is expected by June 2026, with ongoing development and evaluation. The project exemplifies a strategic approach: leveraging existing multilingual models rather than training from scratch, contrasting with Italy’s Minerva, which trained from the ground up.
AMÁLIA
The three hard
questions.
Portugal spent €5.5M to build a European Portuguese LLM. The base version is operational, the benchmarks beat Qwen 3-8B on most pt-PT tasks. So why are the most important questions still unanswered?
Last month, Duarte O.Carmo published the sharpest public analysis of AMÁLIA — Portugal’s state-funded European Portuguese large language model. He prefaces his critique with the necessary diplomatic apparatus before doing what almost nobody else in the European-sovereign-LLM discourse has been willing to do publicly: asking hard questions about whether the work, as released, actually does what it set out to do. This piece is a structural extension of his analysis. The AMÁLIA case study exposes three hard questions every national LLM effort needs to answer publicly — and the broader European sovereign-LLM movement has been operating without explicit answers to any of them.
Three questions every national LLM effort needs to answer publicly.
Duarte O.Carmo’s framing maps cleanly onto the structural argument. Each question lands specifically in AMÁLIA — and the broader European sovereign-LLM movement has been operating without explicit answers to any of them.
The three questions form a structural feedback loop. Q3 (optimization target) determines Q2 (data volume needed) which conditions Q1 (openness sufficient for community contribution). The European sovereign-LLM movement collectively benefits from these questions becoming standard methodology disclosure, not exceptional critique.
107 billion tokens. 5.8 billion clearly pt-PT.
The structurally tractable question with a structurally surprising answer. For a model whose entire stated purpose is European Portuguese prioritization, the native-language share of extended pre-training is 5.5%. The implications cascade into every other question.
The Olmo standard. AMÁLIA’s current state.
Allen Institute for AI’s Olmo project defines what “fully open” operationally requires. Olmo doesn’t lead frontier benchmarks. That’s not the point. The point is to be the structural reference for openness. AMÁLIA’s “fully open source” claim should track to the operational standard.
Four strategic positions. AMÁLIA between two and three.
Approximately €100M+ in publicly disclosed European sovereign-LLM funding across the major initiatives. The structural question every project faces: what is the actual competitive position you’re staking? Four options — none mutually exclusive — but each requiring different commitments.
Three standards. For AMÁLIA and the movement.
The structural critique generalizes beyond AMÁLIA. Italy, France, Germany, Switzerland, the OpenEuroLLM consortium, and every subsequent national project benefit from public discourse holding national LLM efforts to operational standards on openness, data accounting, and strategic positioning.
The European sovereign-AI agenda is a serious strategic project that deserves serious public discourse. O.Carmo’s analysis is what serious public discourse looks like. Appropriately diplomatic. Structurally rigorous. Willing to ask the hard questions in public when the public investment justifies it. More of this is needed — across every European sovereign-LLM project, not just AMÁLIA.
Implications of AMÁLIA for European AI Sovereignty
The development of AMÁLIA underscores a broader pattern in European AI efforts: many countries are investing heavily in sovereign-language models but face persistent questions about transparency, data sufficiency, and strategic goals. These issues are crucial because they influence how national AI policies are shaped, how models are adopted in public sectors, and Europe’s competitiveness in AI innovation. The fact that Portugal’s model is publicly funded and involves national institutions makes these questions particularly impactful at a policy level, setting a precedent for other European countries.
Moreover, the unresolved questions about openness, native data, and objectives highlight potential risks: without clear answers, models might not meet transparency standards, may rely on limited native data, or fail to align with national priorities. This could hinder trust, adoption, and strategic sovereignty in AI development across Europe.
large language model development kit
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
European Sovereign LLM Efforts and the Structural Challenges
Portugal’s AMÁLIA is part of a wider movement across Europe, including Italy’s Minerva, Germany’s Aleph Alpha, France’s Mistral, and initiatives like OpenEuroLLM and AI Sweden. These efforts aim to develop sovereign language models to reduce reliance on US and Chinese AI giants. However, most of these projects share common structural issues: questions about how open their models truly are, how much native-language data is enough, and what objectives they should prioritize. These issues have been underexplored publicly but are central to evaluating the strategic value of these models.
Historically, European efforts have focused on technical benchmarks, but recent analyses, including Duarte O.Carmo’s critique, emphasize the importance of examining these structural questions to understand the real impact and sustainability of these models. AMÁLIA exemplifies this shift by highlighting the need for transparent discussions about openness, data sufficiency, and strategic goals in national AI initiatives.
“AMÁLIA is an impressive piece of work, but it raises fundamental questions about what it truly delivers and how open it really is.”
— Duarte O.Carmo
AI model training and fine-tuning software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About AMÁLIA’s Openness and Objectives
While the technical performance of AMÁLIA is documented, it remains unclear how open the model truly is, especially regarding access and transparency. The extent to which native Portuguese data suffices for strategic sovereignty is also uncertain, as the current data share appears limited. Additionally, the ultimate goals—whether to prioritize academic research, public service, or commercial deployment—are not yet clearly defined or communicated.
Further developments in the coming months may clarify these issues, but as of now, they remain open questions that influence the model’s strategic value and trustworthiness.
As an affiliate, we earn on qualifying purchases.
Next Steps in Evaluating and Developing AMÁLIA
The final version of AMÁLIA is scheduled for release by June 2026, with ongoing assessments expected to address current gaps. Researchers and policymakers will likely scrutinize its openness, native data use, and strategic objectives more closely. Additionally, Portugal is expected to publish more detailed transparency reports and engage in public discussions about the model’s purpose and governance. These steps will be crucial in determining whether AMÁLIA can serve as a sustainable, trustworthy foundation for Portugal’s AI ambitions and influence broader European strategies.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main purpose of AMÁLIA?
AMÁLIA aims to develop a high-performance European Portuguese language model to support academic, public, and potentially commercial applications, reducing reliance on non-European AI models.
How open is AMÁLIA really?
It is not yet clear how accessible the model will be to external researchers or developers, as transparency and openness policies are still being developed or clarified.
Why is native Portuguese data limited in AMÁLIA?
The training used only about 5.8 billion tokens from Portuguese sources, which may be insufficient for full linguistic sovereignty or advanced capabilities, raising questions about data sufficiency.
What are the broader implications for European AI development?
The questions raised by AMÁLIA reflect systemic issues across Europe’s sovereign AI efforts, emphasizing the need for transparency, clear objectives, and sufficient native data to ensure strategic independence and trust.
Source: ThorstenMeyerAI.com