HomeNotesTalksThe roast curve
M365 Barista Talk · LinkedIn
Home /Notes /R-03

R-03 Foundry & Azure

11,000 Models – So Which One Do I Actually Pick?

17/03/20266 minMedium · Concept

Why I am writing about this

It was towards the end of my last session at Microsoft. The slides were done, the coffee cold – and then the questions came. Not one, not two. The same topic over and over, from the audience, afterwards, on LinkedIn:

Over 11,000 models in the catalog – but which one is actually the right one for me?

Honestly, I had to swallow for a second. Because the question sounds so simple and at the same time is so damn legitimate. 11,000 models. That is no longer a choice, that is a shelf without labels in a supermarket that someone tripled in size overnight.

And the second thing that stuck with me: the people asking were not beginners. They were admins, architects, power users – people who know what they are doing. Still, the question was in the room. That tells me: it is not a knowledge problem. It is an orientation problem.

That is exactly what this article is for.


What is the Model Catalog in Azure AI Foundry anyway?

Azure AI Foundry (formerly Azure AI Studio) is Microsoft’s central platform for building, testing and deploying AI applications. The Model Catalog inside it is the control centre for choosing a model – and it is growing fast.

You will find models from Microsoft itself (the Phi family), OpenAI (GPT-4o, o1, o3…), Meta (Llama), Mistral, Cohere, Stability AI, xAI, Google (Gemma), Anthropic (Claude) – and many more. Curated, categorised, with benchmarks and descriptions.

The decisive advantage over direct API access: most models in the catalog run inside the Azure infrastructure. That means your data does not leave the Azure region your tenant is configured for. For enterprise environments with data protection and compliance requirements, that is not a nice-to-have – it is a prerequisite.

IMPORTANT NOTE ON ANTHROPIC MODELS (AS OF NOW)

Anthropic models (the Claude family) are offered in the Foundry catalog, but as things stand, data processing still runs through US infrastructure – not through Azure directly. For GDPR-sensitive workloads or environments with an EU data localisation requirement, that is a blocker. It is not a permanent state, but as of today it is a relevant one. Check before you deploy.


The real stuff: how do I choose the right model?

1. Understand what you actually want

Sounds trivial. It isn’t. Before you open the catalog, answer three questions for yourself:

  • What is the use case? Generating text, writing code, analysing images, summarising documents, understanding language?
  • How much context do I need? Short prompts or long documents with lots of context?
  • What are my compliance requirements? EU data residency, industry regulations, internal policies?

Only then does the catalog make sense.

2. Start small – seriously

This is my clear recommendation and I stand by it: start with the mini models.

GPT-4o mini, Phi-3.5-mini, Llama 3.2 (the smaller variants) – they are faster, cheaper per token and perfectly sufficient for most test scenarios. When you notice the model reaching its limits – quality fluctuates, context gets lost, answers become shallow – then you move up to the large variants.

Why that matters: larger models do not just cost more per token. They are also slower at inference, which becomes relevant in production scenarios. And: if you cannot get a concept to work with a mini model, it is usually not the model – it is the prompt or the design of your solution.

3. Models have specialities – use them

A few concrete examples that make the difference:

Model Strength Typical use case
GPT-4o OpenAI Multimodal, reasoning, long context Document analysis, complex agents
Phi-3.5 / Phi-4 Microsoft Efficient, strong at reasoning for its size Edge scenarios, cost-sensitive workloads
Llama 3.x Meta Open weights, flexible to adapt Fine-tuning projects, on-prem scenarios
Mistral Large Mistral Strong in European languages, code Multilingual applications, coding assistance
Sora Not a text LLM Video generation, visual understanding Image and video workflows – not for text!
Deepseek Deepseek Strong in maths, coding, analytics Technical analysis, STEM domains
WATCH OUT FOR …

Sora is not a language model in the classic sense. Anyone who puts Sora to work on text tasks is on the wrong track – and burns token budget for nothing. Deepseek is very powerful in analytical domains, but: check data processing and provenance carefully in a compliance context.

4. Read the model descriptions. Really.

I know, it sounds like “read the manual” – but it pays off. In the Foundry Model Catalog, every model lists:

  • Benchmarks and comparison values
  • Recommended use cases
  • Context window size
  • Licence model and terms of use
  • Price per token

That last point is underestimated. Token costs between a mini model and the full model of the same family can differ by a factor of 10–20. For production workloads with high throughput, that is no longer an academic detail – that is budget planning.

5. Keep enterprise data protection in view

For Foundry projects in enterprise environments: use models that verifiably process on Azure infrastructure. That gives you:

  • Data residency in your configured Azure region
  • Integration into your existing Azure security architecture (private endpoints, VNet, managed identity)
  • Provable compliance

Models you are not sure about can be deployed in the Foundry project, tested – and removed again if they don’t fit. The platform is built for that. Use it as a sandbox, not as a production environment.


Reality check

How you know it is working

  • The model answers consistently and in the expected format
  • The latency fits your use case
  • Token costs stay within the planned budget
  • Your compliance requirements are documented and met

If it doesn’t work – check these 3 things

TROUBLESHOOTING
  • Wrong model for the use case? Did you use a language model for an image workflow, or the other way round?
  • Context window too small? Long documents need models with a large context window – that is not the same for every model.
  • Unclear data processing? Before production use: check explicitly in the Foundry catalog where the model is processed. When in doubt: ask, don’t assume.

Conclusion

The 11,000 models are not a threat – they are an opportunity. But as with any big toolbox: whoever reaches in without a plan loses time and money.

My recommendation is simple: do your research, read the descriptions, start with mini models, and only move up once you know what you need. Foundry lets you test models safely and remove them again – that is a feature, not an emergency exit.

And with all the enthusiasm for new models: compliance first. A model that sends your data to the wrong region is not a model for your company – no matter how good the benchmark numbers are.

ESPRESSO-MOMENT

The right model is not the strongest one – it is the one that solves your use case, respects your budget and leaves your data where it belongs.


Quick checklist: choosing a model in Azure AI Foundry

CHECKLIST
  • Use case clearly defined before opening the catalog
  • Start with mini models (e.g. GPT-4o mini, Phi-3.5-mini)
  • Model description and benchmarks in the catalog read
  • Specialisations considered (Sora ≠ text model, Deepseek = analytical/technical)
  • Token costs compared before deployment
  • Data processing location checked (Azure region vs. third country)
  • Anthropic models: currently US processing – compliance check mandatory
  • Deploy the model as a test in the Foundry project, evaluate, uninstall if needed

Did you run into the model selection problem on your last project too? Tell me about it – I am curious which use cases are keeping you busy right now.

Ferdi Lethen-Oellers
Ferdi Lethen-Oellers

Senior Modern Workplace Consultant at amexus, Microsoft MVP, author of “Microsoft 365 Administration für Dummies”.