---
title: "vLLM | Skills We Assign For | Azendo"
description: "vLLM for self-hosted inference — when hosting your own model makes sense, and the Azendo roles assigned for it."
url: "https://azendo.co/skills/vllm/"
---

[Skills](https://azendo.co/skills/) AI and machine learning 

# vLLM.

vLLM is an inference server for running large language models efficiently on your own hardware, using memory techniques that raise throughput substantially over naive serving. It is the answer when self-hosting a model is the requirement.

## Where vLLM fits on a long engagement.

Self-hosting is chosen for a reason, not by default. Data that cannot leave your environment, cost at high sustained volume, or latency that an API round trip cannot meet. Absent one of those, a hosted API is cheaper and considerably less work.

Where it is justified, throughput per GPU decides the economics entirely. The difference between an efficient serving setup and a naive one is several times the hardware bill, which is why the serving layer deserves the same attention as the model.

## What an assigned team does with vLLM.

Self-hosted inference is an operational commitment. GPUs need capacity planning, model versions need rollout discipline, and throughput tuning is continuous as usage patterns change — none of which fits a project shape.

It is also frequently adopted for a reason that stops applying. Where hosted pricing has moved or volume has fallen, revisiting the decision honestly is part of the assignment, which sits under [AI engineering services](https://azendo.co/services/ai-engineering/) when the feature itself is also in scope.

## What we use vLLM for.

* Inference inside your own boundary Running models where regulation or contract prevents sending data to a third party.
* Cost at sustained volume Self-hosting where per-token pricing has passed the point of being the cheaper option.
* Getting more from existing GPUs Serving configuration tuned so the hardware already bought carries more load.

## How vLLM capacity is assigned.

Serving infrastructure is assigned under [MLOps engineering](https://azendo.co/services/mlops-engineering/), working on your own hardware or in your own cloud account.

## Roles we assign vLLM for

* [LLM Engineer AI engineering](https://azendo.co/services/ai-engineering/llm-engineer/)
* [Model Deployment Engineer MLOps engineering](https://azendo.co/services/mlops-engineering/model-deployment-engineer/)

## Service lines it sits in

* [AI engineering](https://azendo.co/services/ai-engineering/)
* [MLOps engineering](https://azendo.co/services/mlops-engineering/)

Capacity is agreed as a committed monthly capacity across a discipline, not per skill.

## Related in ai and machine learning

* [Weaviate — skill we assign for](https://azendo.co/skills/weaviate/)
* [LangChain — skill we assign for](https://azendo.co/skills/langchain/)
* [LlamaIndex — skill we assign for](https://azendo.co/skills/llamaindex/)
* [LangGraph — skill we assign for](https://azendo.co/skills/langgraph/)
* [agent frameworks — skill we assign for](https://azendo.co/skills/agent-frameworks/)
* [prompt and context engineering — skill we assign for](https://azendo.co/skills/prompt-and-context-engineering/)
* [Prompt evaluation — skill we assign for](https://azendo.co/skills/prompt-evaluation/)
* [LangSmith — skill we assign for](https://azendo.co/skills/langsmith/)

## Tell us what your roadmap needs vLLM for.

A service delivery manager replies with the disciplines we would assign, the monthly capacity and what the first month looks like.

[All skills we assign for](https://azendo.co/skills/)
