Open source LLMs: how companies gain more control and flexibility

TL;DR
Open source LLMs are developing rapidly and increasingly reaching a level of performance comparable to closed source models such as GPT-4. Large models like DeepSeek R1 (671B) deliver impressive results, but they require powerful hardware. Smaller models such as Phi-4 (14B) are easier on resources but come with some limitations in quality. Even so, the performance is considerable and can now be described as production-ready for some use cases. Companies with high data protection requirements or long-term cost considerations should definitely include open source LLM alternatives in their evaluation.
Analysis: how capable are current open source LLMs?
The market for open source AI is growing steadily. While proprietary models such as GPT-4 or Claude are widely used, open source LLMs are becoming more significant. Recent benchmarks show that models such as DeepSeek R1 (671B) are competitive in terms of reasoning and answer quality, although there are challenges around hardware requirements and language support. At the same time there are always smaller model variants with fewer parameters that run faster but have to accept losses in quality.
A second important lever in balancing speed and quality is the question of whether or not to choose a dedicated reasoning model. Reasoning models invest more compute time in order to analyse context and connections more deeply, which usually leads to better results but also uses more resources. Non-reasoning models, by contrast, deliver fast answers but can reach their limits in complex fields of work. So the decisive point is always the trade-off between answer quality and answer time.
Performance: open source versus proprietary AI
Our test run of several open source LLMs, including DeepSeek R1, Phi-4 (14B) and various Llama models, showed:
- DeepSeek R1 (671B): delivers detailed and well-structured answers, but requires high-end GPUs.
- DeepSeek V3 (671B): not quite as capable as R1, but considerably faster, since it answers directly instead of thinking first.
- Phi-4 (14B): fast and easy on resources, though with slight grammar problems and limited accuracy. Currently our clear value-for-money winner.
- Llama 3.2 (3B) & Llama 3.1 (8B): fast models with lower memory requirements, but not at the same level in terms of quality. Not recommended.
Large models deliver convincing results but are not suitable for every use case because of the hardware they need. Smaller models are better suited to running on standard hardware but cannot keep up in quality.
Example 1: Phi-4 has minor grammar problems. The content is also simple. Still very impressive for the speed.

Example 2: DeepSeek R1 shows strong reasoning. Smaller models still have clear weaknesses in German, though, and sometimes mix in Chinese terms.

Example 3: smaller Llama models (Llama 3.1 and 3.2, for example) had clear difficulties with German grammar. This does not allow any conclusions about the quality of larger Llama variants, which were not tested here.

Hardware requirements: what do you need?
A decisive factor in using open source models is the compute power required. It determines not only whether a model can be run at all but also has a major influence on how fast answers are generated. Cloud-based AI services provide this compute power directly, whereas companies running local infrastructure have to invest in capable hardware accordingly:
- DeepSeek R1 (671B): requires high-end GPUs with a lot of VRAM (an NVIDIA A100 or 4090, for example).
- Phi-4 (14B): runs on capable consumer GPUs such as an RTX 3090 or 4080.
- Llama 3.2 (3B): can be run on mid-range GPUs, but offers only limited quality.
For your company that means: a one-off investment in hardware can be more cost-efficient in the long term than ongoing API fees, but it requires the corresponding know-how to run the models, as well as a certain commitment.
Where it makes sense: when is an open source LLM worth it?
Although proprietary models can currently be integrated into existing systems more easily, open source alternatives offer advantages:
Data protection and compliance: for industries with high data protection requirements in particular (healthcare, for example), running an AI locally can contribute to better compliance with regulatory rules.
Cost savings: while GPT-4 API costs can run to several hundred or thousand euros, a one-off investment in hardware makes long-term savings possible.
Adaptability: open source LLMs can be trained and optimised individually and chosen flexibly in terms of model size, which is ideal for getting faster results on simple tasks, for example.
On top of that, using open source models means companies avoid becoming heavily dependent on a single provider and can switch easily to whichever provider is currently the most capable.
There are challenges too, however:
Technical barrier to entry: running and maintaining your own model requires expertise and experience.
Limited language support: smaller open source LLMs in particular are trained primarily on English and Chinese. In German-language applications this often leads to losses in quality, as the examples above showed. Larger models are usually less affected by this problem.
Conclusion
Open source AI models are more capable today than ever and represent a serious alternative to closed source models for many companies. Anyone focusing on data protection and cost control in the long term should evaluate open source LLMs. A tailored model can be particularly advantageous for specific use cases with local processing.
Beyond that, deploying LLMs yourself opens up entirely new possibilities for you: model size can be adjusted flexibly to your requirements for performance and quality, data protection rules can be complied with more easily, and smaller models can even be run on consumer hardware when needed.
Sounds like a relevant use case for you? Let us talk and develop the best solution for you together.
Tech Newsletter
Join our 2,000+ subscribers and receive monthly updates on our latest articles, case studies, webinars, events, and industry news.




