---
title: "Prompt evaluation | Skills We Assign For | Azendo"
description: "Evaluating prompts and LLM output — test sets, judges, regression detection, plus the Azendo roles assigned for it."
url: "https://azendo.co/skills/prompt-evaluation/"
---

[Skills](https://azendo.co/skills/) AI and machine learning 

# Prompt evaluation.

Prompt evaluation measures whether a change to a prompt, model or retrieval step improved output, using a fixed test set with defined criteria, scored automatically, by a model judge, or by people.

## Where Prompt evaluation fits on a long engagement.

Without evaluation, prompt work is guesswork with confidence. A change that improves one example often degrades three others, and nobody notices because nobody checked. A fixed test set with expected behaviour turns that from opinion into measurement.

Model-based judging scales but has its own biases: judges prefer longer answers, favour output from the same model family, and are inconsistent on borderline cases. Calibrating a judge against human scores on a sample is what makes its verdicts usable rather than merely available.

## What an assigned team does with Prompt evaluation.

The test set is the asset, and it should grow from production failures. Every case where the system got something wrong is a case that belongs in the suite, so the same failure cannot recur unnoticed.

Maintaining that set continuously is ongoing work rather than a setup task, held within an agreed [committed monthly capacity](https://azendo.co/pricing/).

## What we use Prompt evaluation for.

* Changes measured rather than assumed A fixed set scored before and after, so an improvement is evidenced.
* Regressions caught before release Evaluation in the pipeline, so a prompt change cannot quietly degrade other cases.
* A suite grown from real failures Production errors added as cases, so the same mistake cannot recur silently.

## How Prompt evaluation capacity is assigned.

Evaluation capacity is assigned as explicit scope under [ai engineering services](https://azendo.co/services/ai-engineering/), because an unmeasured system cannot be improved deliberately.

## Service lines it sits in

* [AI engineering](https://azendo.co/services/ai-engineering/)

Capacity is agreed as a committed monthly capacity across a discipline, not per skill.

## Related in ai and machine learning

* [PyTorch — skill we assign for](https://azendo.co/skills/pytorch/)
* [TensorFlow — skill we assign for](https://azendo.co/skills/tensorflow/)
* [scikit-learn — skill we assign for](https://azendo.co/skills/scikit-learn/)
* [Pandas — skill we assign for](https://azendo.co/skills/pandas/)
* [NumPy — skill we assign for](https://azendo.co/skills/numpy/)
* [embeddings — skill we assign for](https://azendo.co/skills/embeddings/)
* [function calling — skill we assign for](https://azendo.co/skills/function-calling/)
* [vLLM — skill we assign for](https://azendo.co/skills/vllm/)

## Tell us what your roadmap needs Prompt evaluation for.

A service delivery manager replies with the disciplines we would assign, the monthly capacity and what the first month looks like.

[All skills we assign for](https://azendo.co/skills/)
