R-03 Foundry & Azure
11,000 Models – So Which One Do I Actually Pick?
Why I am writing about this
It was towards the end of my last session at Microsoft. The slides were done, the coffee cold – and then the questions came. Not one, not two. The same topic over and over, from the audience, afterwards, on LinkedIn:
Over 11,000 models in the catalog – but which one is actually the right one for me?
Honestly, I had to swallow for a second. Because the question sounds so simple and at the same time is so damn legitimate. 11,000 models. That is no longer a choice, that is a shelf without labels in a supermarket that someone tripled in size overnight.
And the second thing that stuck with me: the people asking were not beginners. They were admins, architects, power users – people who know what they are doing. Still, the question was in the room. That tells me: it is not a knowledge problem. It is an orientation problem.
That is exactly what this article is for.
What is the Model Catalog in Azure AI Foundry anyway?
Azure AI Foundry (formerly Azure AI Studio) is Microsoft’s central platform for building, testing and deploying AI applications. The Model Catalog inside it is the control centre for choosing a model – and it is growing fast.
You will find models from Microsoft itself (the Phi family), OpenAI (GPT-4o, o1, o3…), Meta (Llama), Mistral, Cohere, Stability AI, xAI, Google (Gemma), Anthropic (Claude) – and many more. Curated, categorised, with benchmarks and descriptions.
The decisive advantage over direct API access: most models in the catalog run inside the Azure infrastructure. That means your data does not leave the Azure region your tenant is configured for. For enterprise environments with data protection and compliance requirements, that is not a nice-to-have – it is a prerequisite.
Anthropic models (the Claude family) are offered in the Foundry catalog, but as things stand, data processing still runs through US infrastructure – not through Azure directly. For GDPR-sensitive workloads or environments with an EU data localisation requirement, that is a blocker. It is not a permanent state, but as of today it is a relevant one. Check before you deploy.
The real stuff: how do I choose the right model?
1. Understand what you actually want
Sounds trivial. It isn’t. Before you open the catalog, answer three questions for yourself:
- What is the use case? Generating text, writing code, analysing images, summarising documents, understanding language?
- How much context do I need? Short prompts or long documents with lots of context?
- What are my compliance requirements? EU data residency, industry regulations, internal policies?
Only then does the catalog make sense.
2. Start small – seriously
This is my clear recommendation and I stand by it: start with the mini models.
GPT-4o mini, Phi-3.5-mini, Llama 3.2 (the smaller variants) – they are faster, cheaper per token and perfectly sufficient for most test scenarios. When you notice the model reaching its limits – quality fluctuates, context gets lost, answers become shallow – then you move up to the large variants.
Why that matters: larger models do not just cost more per token. They are also slower at inference, which becomes relevant in production scenarios. And: if you cannot get a concept to work with a mini model, it is usually not the model – it is the prompt or the design of your solution.
3. Models have specialities – use them
A few concrete examples that make the difference:
| Model | Strength | Typical use case |
|---|---|---|
| GPT-4o OpenAI | Multimodal, reasoning, long context | Document analysis, complex agents |
| Phi-3.5 / Phi-4 Microsoft | Efficient, strong at reasoning for its size | Edge scenarios, cost-sensitive workloads |
| Llama 3.x Meta | Open weights, flexible to adapt | Fine-tuning projects, on-prem scenarios |
| Mistral Large Mistral | Strong in European languages, code | Multilingual applications, coding assistance |
| Sora Not a text LLM | Video generation, visual understanding | Image and video workflows – not for text! |
| Deepseek Deepseek | Strong in maths, coding, analytics | Technical analysis, STEM domains |
Sora is not a language model in the classic sense. Anyone who puts Sora to work on text tasks is on the wrong track – and burns token budget for nothing. Deepseek is very powerful in analytical domains, but: check data processing and provenance carefully in a compliance context.
4. Read the model descriptions. Really.
I know, it sounds like “read the manual” – but it pays off. In the Foundry Model Catalog, every model lists:
- Benchmarks and comparison values
- Recommended use cases
- Context window size
- Licence model and terms of use
- Price per token
That last point is underestimated. Token costs between a mini model and the full model of the same family can differ by a factor of 10–20. For production workloads with high throughput, that is no longer an academic detail – that is budget planning.
5. Keep enterprise data protection in view
For Foundry projects in enterprise environments: use models that verifiably process on Azure infrastructure. That gives you:
- Data residency in your configured Azure region
- Integration into your existing Azure security architecture (private endpoints, VNet, managed identity)
- Provable compliance
Models you are not sure about can be deployed in the Foundry project, tested – and removed again if they don’t fit. The platform is built for that. Use it as a sandbox, not as a production environment.
Reality check
How you know it is working
- The model answers consistently and in the expected format
- The latency fits your use case
- Token costs stay within the planned budget
- Your compliance requirements are documented and met
If it doesn’t work – check these 3 things
- Wrong model for the use case? Did you use a language model for an image workflow, or the other way round?
- Context window too small? Long documents need models with a large context window – that is not the same for every model.
- Unclear data processing? Before production use: check explicitly in the Foundry catalog where the model is processed. When in doubt: ask, don’t assume.
Conclusion
The 11,000 models are not a threat – they are an opportunity. But as with any big toolbox: whoever reaches in without a plan loses time and money.
My recommendation is simple: do your research, read the descriptions, start with mini models, and only move up once you know what you need. Foundry lets you test models safely and remove them again – that is a feature, not an emergency exit.
And with all the enthusiasm for new models: compliance first. A model that sends your data to the wrong region is not a model for your company – no matter how good the benchmark numbers are.
The right model is not the strongest one – it is the one that solves your use case, respects your budget and leaves your data where it belongs.
Quick checklist: choosing a model in Azure AI Foundry
- Use case clearly defined before opening the catalog
- Start with mini models (e.g. GPT-4o mini, Phi-3.5-mini)
- Model description and benchmarks in the catalog read
- Specialisations considered (Sora ≠ text model, Deepseek = analytical/technical)
- Token costs compared before deployment
- Data processing location checked (Azure region vs. third country)
- Anthropic models: currently US processing – compliance check mandatory
- Deploy the model as a test in the Foundry project, evaluate, uninstall if needed
Did you run into the model selection problem on your last project too? Tell me about it – I am curious which use cases are keeping you busy right now.
