Dev News Daily ENDE

Amazon EKS adds Kubernetes 1.37 with a GA Metrics API and HPA scale-to-zero on by default

Amazon Elastic Kubernetes Service and Amazon EKS Distro now support Kubernetes 1.37, AWS announced this week. New clusters can be created on 1.37 and existing ones upgraded through the EKS console, eksctl or infrastructure-as-code tools, in every region where EKS runs, including the AWS GovCloud (US) regions. EKS Distro builds of 1.37 are published on the ECR Public Gallery and on GitHub.

AWS singles out three changes in the release.

The Metrics API is generally available as metrics.k8s.io/v1. It is the API that supplies pod and node CPU and memory usage to the Horizontal Pod Autoscaler and to kubectl top, so the version clusters have relied on for years is now a stable one.

DRA device taints and tolerations are generally available. Dynamic Resource Allocation drivers and administrators can taint a device such as a GPU, and the scheduler keeps workloads off it unless they tolerate the taint. That gives operators a way to drain or quarantine a single accelerator without cordoning the whole node.

HPA scale-to-zero has reached beta and is enabled by default. An autoscaler with minReplicas: 0 that scales on object or external metrics can now take a workload down to zero pods when it is idle and bring it back when demand returns. The announcement limits it to those two metric types: an autoscaler driven by the pods' own CPU or memory has nothing to measure once no pods are running.

Before upgrading, AWS points to EKS cluster insights, which flag issues that could affect a cluster upgrade, and to its documentation on the EKS version lifecycle.

Amazon EKS adds Kubernetes 1.37 with a GA Metrics API and HPA scale-to-zero on by default
Amazon EKS adds Kubernetes 1.37 with a GA Metrics API and HPA scale-to-zero on by default — Dev News Daily

What it means

The scale-to-zero default is the change that can surprise someone. Any existing autoscaler that already declares minReplicas: 0 with an external or object metric will start doing exactly that after the upgrade. Teams should search their manifests for that combination and confirm that the first request after a quiet period can tolerate a cold start.