Kubernetes AI Conformance¶
Welkin based on Cluster API meets all MUST requirements in the Kubernetes AI Conformance v1.35 checklist. This page describes how Welkin fulfills these requirements and provides links to supporting documentation and test evidence.
Note
This page is a work in progress. More requirements will be listed as we confirm that Welkin fulfills them.
Accelerators¶
DRA Support (dra_support)¶
Welkin supports the Dynamic Resource Allocation (DRA) APIs, which enable fine-grained resource requests beyond simple device counts. DRA is enabled by default in Kubernetes starting with v1.34.
The Kubernetes v1.35 conformance results for Welkin based on Cluster API include passing API tests for DeviceClass, ResourceClaim, ResourceClaimTemplate, and ResourceSlice. These tests provide evidence of DRA API support.
Scheduling and Orchestration¶
Gang Scheduling (gang_scheduling)¶
The Kueue scheduler can be installed on top of Welkin, ensuring all-or-nothing scheduling for distributed AI workloads. Automated, evidence-gathering tests have been implemented for the following scenarios:
| Scenario | Description |
|---|---|
successful_admission |
Admit and complete two workers within 200m CPU quota. |
insufficient_quota |
Keep three workers blocked by 200m CPU quota for 20 seconds. |
quota_recovery |
Unblock the same three-worker Job by raising quota to 300m. |
Cluster Autoscaling (cluster_autoscaling)¶
Welkin based on Cluster API uses the Cluster Autoscaler to add Nodes when Pods cannot be scheduled due to insufficient resources and remove Nodes that are no longer needed. Scaling is based on Pod resource requests rather than current resource utilization. For AI conformance, this capability must support scaling Node groups containing specific accelerator types based on Pods requesting those accelerators.
See the Cluster autoscaling documentation for scaling behavior and Infrastructure Provider availability.
Observability¶
AI Service Metrics (ai_service_metrics)¶
Welkin includes Prometheus and the Prometheus Operator to discover and collect metrics from workloads exposing metrics in the Prometheus format, including AI frameworks and model servers. Applications configure metrics collection through ServiceMonitor or PodMonitor resources.
The metrics collection documentation describes how these resources are discovered and includes a working example of scraping an application's metrics endpoint.