AI EngineerVideoSource Linked
Infra behind Krea 2: How to train and serve at scale — Gabriel Jorge Menezes, Krea.ai
Krea’s K2 infrastructure shows why GPU utilization alone misleads: tensor-core, temperature, interconnect, and checkpoint metrics are what keep large training runs productive.
AI Engineer · Aug 18, 2026
Source Summary
Krea trained K2 from scratch on **thousands of GPUs**, yet reported GPU utilization stayed at 100% even when the cluster was inefficient. The team instead watched tensor-core use, removed GPUs above **78°C**, and checkpointed roughly every 20–30 minutes.
Practical Implication
Builders operating GPU agents or training jobs should instrument useful work, thermals, NVLink, InfiniBand, and storage throughput. Krea also uses queued gang scheduling and shifts inference to external capacity when training preempts production GPUs.
Agent-Ready Context
Krea trained K2 from scratch on **thousands of GPUs**, yet reported GPU utilization stayed at 100% even when the cluster was inefficient. The team instead watched tensor-core use, removed GPUs above **78°C**, and checkpointed roughly every 20–30 minutes. Builders operating GPU agents or training jobs should instrument useful work, thermals, NVLink, InfiniBand, and storage throughput. Krea also uses queued gang scheduling and shifts inference to external capacity when training preempts production GPUs. This design assumes substantial Kubernetes, networking, and storage investment. Krea’s thresholds and routing system reflect its own diffusion workload and should be validated before reuse elsewhere.
Context Map
infraimage#observabilityUncertainty
This design assumes substantial Kubernetes, networking, and storage investment. Krea’s thresholds and routing system reflect its own diffusion workload and should be validated before reuse elsewhere.