Broadwing
WHO WE ARE WHAT WE DO WORK CONNECT
WHAT WE DO / HPC

HPC.

We build and stabilize research-scale compute — from 10k+-node supercomputers to the software platforms that manage them.

CLUSTERS SCHEDULERS LUSTRE KUBERNETES HYBRID CLOUD
Our engineers have run some of the largest systems on Earth.

High-performance computing is unforgiving: thousands of nodes, parallel filesystems, job schedulers, and users whose research stalls every hour the system misbehaves. Broadwing engineers have stabilized a Top50 hybrid-cloud supercomputer with more than 10,000 nodes — co-developing automated health checks, resolving issues across Lustre, Kubernetes, and job scheduling, and coordinating hundreds of hardware replacements to meet every critical benchmark within six months.

We also build the software that runs these systems — including a high-availability, cloud-native hypervisor management plane with auto-discovery and Kubernetes-hosted microservices, installable in under an hour, built for a leading HPC manufacturer. And we modernize the engineering around HPC software: on one platform we took software test cycles from one month to one hour with CI/CD and artifact-management modernization.

If your compute is the product — research, simulation, AI training — we keep it fast, stable, and operable by your own team.

THE PROBLEM
Research stalls when the cluster does.
Nodes fail faster than anyone can diagnose them. The filesystem misbehaves under load. Scheduler queues back up, and the team is firefighting instead of improving the system.
HOW WE HELP
01
Assess — profile the system: nodes, storage, scheduling, and where time is actually lost.
02
Stabilize — automated health checks, issue resolution, and coordinated hardware remediation.
03
Automate — CI/CD, testing, and system management tooling that scales with the machine.
04
Hand off — your operations team runs it, with runbooks that match reality.
TYPICAL OUTCOMES
6 mo
TO ALL CRITICAL BENCHMARKS ON A TOP50, 10K+-NODE SUPERCOMPUTER
1 hr
SOFTWARE TEST CYCLES — DOWN FROM ONE MONTH
10k+
NODES STABILIZED ON A SINGLE HYBRID-CLOUD SYSTEM
WHAT WE TAKE ON
If this sounds like your cluster, we should talk.
01
Node failures outpace diagnosis
We co-develop automated health checks and coordinate hardware remediation at 10k+-node scale — hundreds of replacements without derailing operations.
02
Storage and scheduling misbehave under load
We resolve system issues across Lustre, job schedulers, and Kubernetes — the full path from filesystem to workload.
03
Software testing can’t keep up with the machine
We modernize CI/CD and artifact management for HPC software — on one platform, test cycles went from one month to one hour.
04
System management doesn’t scale with the system
We build cloud-native management planes — high availability, auto-discovery, Kubernetes-hosted microservices — for systems at production HPC scale.
COMMON QUESTIONS
What buyers usually ask before an engagement
Have you worked at real supercomputer scale?
Yes — we stabilized a Top50 hybrid-cloud system with more than 10,000 nodes, and we build system-management software for a leading HPC manufacturer.
Do you work with research institutions as well as vendors?
Both. We’ve embedded with operations teams running research systems and built product-grade software for HPC manufacturers.
Can our operations team run what you build?
That’s the goal — automated checks, tooling, and runbooks land with your team, matched to how the system actually behaves.
RELATED WORK · HPC · RELIABILITY
Stabilizing a Top50 hybrid-cloud supercomputer with 10k+ nodes
READ CASE STUDY →
Sound familiar?
REQUEST A CONSULTATION →
MORE FROM WHAT WE DO: 01 PLATFORM 02 DATA 03 SECURITY 05 FEDERAL
Broadwing
Engineering clarity in a world of complexity. Platform, data, security, and HPC engineering for teams that can't afford downtime.
© 2026 BROADWING LLC