Where Kubeflow fits on a long engagement.

The argument for Kubeflow is consistency. If the organisation already runs Kubernetes, putting ML workloads there means one platform to secure, monitor and operate rather than a separate stack with its own conventions.

The argument against is that it inherits Kubernetes complexity and adds its own. A team with three models and no platform engineer will spend more time operating Kubeflow than doing machine learning. It earns its place at a scale most organisations have not reached.

What an assigned team does with Kubeflow.

Running ML on Kubernetes means inheriting two operational surfaces, and organisations reach for it earlier than they should. The honest question at scoping is whether the number of models justifies the platform, and the answer is often no.

Where it genuinely is justified, the capacity is platform capacity rather than data science capacity. Where it is not, the same outcome is usually available under data science with far less to operate.

What we use Kubeflow for.

  • ML on the platform you already run Training and serving inside the existing cluster, with the same access control and monitoring.
  • Repeatable training pipelines Multi-step workflows defined once and re-run on a schedule or on new data.
  • Deciding whether you need it An honest assessment against simpler alternatives before committing to the operational cost.

How Kubeflow capacity is assigned.

Platform work of this kind is assigned under MLOps engineering, and we will say when a simpler stack would serve you better.

Tell us what your roadmap needs Kubeflow for.

A service delivery manager replies with the disciplines we would assign, the monthly capacity and what the first month looks like.

Loading the contact form… You can also email hello@azendo.co.

We reply within one working day. No obligation, and no newsletter.