IA2 MIN

Inside the Meta lab that breaks servers to build its next generation of AI

The company tests racks, cooling, networking and repair procedures before repeating a design thousands of times across its data centers.

Tom Shaw holds an open server inside Meta's infrastructure lab
Image: Meta
01

A data center before the data center

Meta has shown its Infrastructure Lab in Menlo Park, California. Engineers test servers, trays, cables, power and cooling before deploying them at scale. The lab does not train the next AI model by itself. It verifies that the machines doing that work fit together, stay cool and can be repaired.

Scale changes the cost of a mistake. A connector that is merely awkward on one prototype becomes a serious maintenance problem when the same rack is repeated thousands of times.

02

The rack becomes one machine

A rack is the metal cabinet holding multiple servers. AI clusters cannot behave as isolated computers because thousands of accelerators need to exchange data without waiting on slow links. Meta designs power, cooling and networking as one system, then tests what happens when a tray fails or must be removed.

Catalina, its NVIDIA GB200 design, uses direct liquid cooling to carry heat away near the chips. Grand Teton accepts accelerators from different vendors in a common structure. That flexibility matters in the AI chip war, where changing hardware should not require rebuilding the facility around it.

Meta Catalina rack with GB200 systems and liquid cooling
Image: Meta Engineering
03

Maintainability is performance too

A chip's peak speed is only part of useful output. If repair work disconnects too many machines, or heat forces them to slow down, the cluster completes less work over a month. The lab therefore rehearses installation, cabling, component replacement and coolant flow before a design reaches production.

Meta's tour does not publish failure rates, rack cost or power and water consumption. Those figures are necessary before comparing its claims with other companies or with the resource pressure driving policy around large AI data centers.

Open Meta Grand Teton server fitted with AMD MI300X accelerators
Image: Meta Engineering
00

The conversation starts here

Sign in with a supporter account to comment. Sign in

Nobody has commented yet. Want to go first?

KEEP READING

You may also like

FRONT PAGE