High-performance computing is unforgiving: thousands of nodes, parallel filesystems, job schedulers, and users whose research stalls every hour the system misbehaves. Broadwing engineers have stabilized a Top50 hybrid-cloud supercomputer with more than 10,000 nodes — co-developing automated health checks, resolving issues across Lustre, Kubernetes, and job scheduling, and coordinating hundreds of hardware replacements to meet every critical benchmark within six months.
We also build the software that runs these systems — including a high-availability, cloud-native hypervisor management plane with auto-discovery and Kubernetes-hosted microservices, installable in under an hour, built for a leading HPC manufacturer. And we modernize the engineering around HPC software: on one platform we took software test cycles from one month to one hour with CI/CD and artifact-management modernization.
If your compute is the product — research, simulation, AI training — we keep it fast, stable, and operable by your own team.