---
title: "Meta Llama | Skills We Assign For | Azendo"
description: "Llama and self-hosted models — data control, fine-tuning, real cost, plus the Azendo roles assigned for it."
url: "https://azendo.co/skills/meta-llama/"
---

[Skills](https://azendo.co/skills/) AI and machine learning 

# Meta Llama.

Llama is Meta's family of open-weight language models, downloadable and self-hostable. Weights can be fine-tuned and the models run on your own infrastructure, which changes the data, cost and control position entirely.

## Where Meta Llama fits on a long engagement.

Self-hosting answers questions an API cannot. Data never leaves your infrastructure, which resolves most regulatory objections directly; the model version is fixed until you change it, so behaviour does not shift underneath you; and the weights can be fine-tuned on proprietary data.

The cost comparison is more nuanced than it appears. GPU instances are expensive and billed whether or not they are serving, so self-hosting is cheaper than an API only above a volume threshold that many workloads never reach. Below it, an API is both cheaper and considerably less work.

## What an assigned team does with Meta Llama.

Serving models is real infrastructure with real operational load: GPU capacity, batching, autoscaling that responds slowly because instances take minutes to start, and monitoring for a workload whose failure modes differ from a web service.

That is platform work as much as AI work, which is why it is scoped alongside [devops as a service](https://azendo.co/services/cloud-and-devops/) rather than treated as a model decision.

## What we use Meta Llama for.

* Data that cannot leave the estate Inference inside your own infrastructure, which resolves most residency objections.
* Behaviour that does not change underneath you A fixed model version, so a provider update cannot alter production output.
* Fine-tuning on proprietary data Adaptation to a domain where a general model performs poorly.

## How Meta Llama capacity is assigned.

Self-hosted model work is assigned under [managed ai services](https://azendo.co/services/ai-engineering/), with the cost crossover against an API established before infrastructure is committed.

## Service lines it sits in

* [AI engineering](https://azendo.co/services/ai-engineering/)

Capacity is agreed as a committed monthly capacity across a discipline, not per skill.

## Related in ai and machine learning

* [PyTorch — skill we assign for](https://azendo.co/skills/pytorch/)
* [TensorFlow — skill we assign for](https://azendo.co/skills/tensorflow/)
* [scikit-learn — skill we assign for](https://azendo.co/skills/scikit-learn/)
* [Pandas — skill we assign for](https://azendo.co/skills/pandas/)
* [NumPy — skill we assign for](https://azendo.co/skills/numpy/)
* [embeddings — skill we assign for](https://azendo.co/skills/embeddings/)
* [function calling — skill we assign for](https://azendo.co/skills/function-calling/)
* [vLLM — skill we assign for](https://azendo.co/skills/vllm/)

## Tell us what your roadmap needs Meta Llama for.

A service delivery manager replies with the disciplines we would assign, the monthly capacity and what the first month looks like.

[All skills we assign for](https://azendo.co/skills/)
