---
title: "Observability | Skills We Assign For | Azendo"
description: "Observability in production systems — beyond dashboards, and the Azendo roles assigned for it."
url: "https://azendo.co/skills/observability/"
---

[Skills](https://azendo.co/skills/) Observability 

# Observability.

Observability is the property of a system that lets you understand its internal state from its outputs — logs, metrics and traces — including for conditions nobody anticipated when it was built.

## Where Observability fits on a long engagement.

The distinction from monitoring is about unanticipated questions. Monitoring answers questions you thought of in advance: is CPU high, is the error rate up. Observability is being able to ask a new question — why is this specific customer's request slow on Tuesdays — without deploying new instrumentation first.

That capability comes from high-cardinality structured data. Logs and traces carrying the identifiers that matter — user, tenant, request, version — allow arbitrary slicing after the fact. Aggregate metrics cannot answer those questions no matter how many dashboards are built from them.

## What an assigned team does with Observability.

AI systems raise the stakes because they are non-deterministic. Without the prompt, the retrieved context and the response recorded for each call, a reported failure cannot be investigated at all — there is no reproduction to run.

Designing that in from the start rather than adding it after an unexplainable incident is part of what is scoped under [ai engineering services](https://azendo.co/services/ai-engineering/).

## What we use Observability for.

* New questions answered without redeploying High-cardinality data allowing arbitrary slicing after an incident starts.
* Non-deterministic systems made investigable Prompt, context and response recorded, because there is no reproduction otherwise.
* Correlation across services Identifiers carried through, so one request is followable end to end.

## How Observability capacity is assigned.

Observability is assigned across platform and application capacity, designed in at the start because retrofitting it after an incident is when it is least available.

## Service lines it sits in

* [AI engineering](https://azendo.co/services/ai-engineering/)

Capacity is agreed as a committed monthly capacity across a discipline, not per skill.

## Industries that ask for it

* [Logistics and transportation](https://azendo.co/industries/logistics-and-transportation/)

## Related in observability

* [SLO and error budget design — skill we assign for](https://azendo.co/skills/slo-and-error-budget-design/)
* [Prometheus — skill we assign for](https://azendo.co/skills/prometheus/)
* [OpenTelemetry — skill we assign for](https://azendo.co/skills/opentelemetry/)
* [CloudWatch — skill we assign for](https://azendo.co/skills/cloudwatch/)
* [PagerDuty — skill we assign for](https://azendo.co/skills/pagerduty/)
* [Sentry — skill we assign for](https://azendo.co/skills/sentry/)
* [Logging — skill we assign for](https://azendo.co/skills/logging/)
* [audit logging — skill we assign for](https://azendo.co/skills/audit-logging/)

## Tell us what your roadmap needs Observability for.

A service delivery manager replies with the disciplines we would assign, the monthly capacity and what the first month looks like.

[All skills we assign for](https://azendo.co/skills/)
