OpenAI has introduced a new networking protocol called MRC. It is developed in collaboration with major technology companies including Nvidia, Microsoft, AMD, Intel, and Broadcom.
According to the company, the protocol is built to make AI supercomputer networks faster, more reliable, and more efficient for training advanced AI models.
The company said, more than 900 million people use ChatGPT every week. Supporting AI systems at that scale requires networking technology capable of moving vast amounts of data between GPUs with minimal delays and without disrupting performance.
In a blogpost, OpenAI said, ‘Our goal was not just to build a fast network, but also to build one that delivers very predictable performance, even in the presence of failures, to keep training jobs moving.’
MRC, short for Multipath Reliable Connection, is the new protocol that improves GPU networking performance and resilience across large AI training clusters. Built into the latest 800Gb/s network interfaces, the technology is designed to help GPUs communicate more efficiently during training.
In conventional AI training systems, data usually travels along a single network path. If that path becomes congested or fails, training can slow significantly or even come to a halt. MRC addresses this by distributing data packets across hundreds of network paths simultaneously, reducing congestion and allowing the system to quickly reroute around failed connections.
MRC is already being used in OpenAI’s largest Nvidia GB200 supercomputers. OpenAI has also shared the MRC specification through the Open Compute Project so that others can use it.
See What’s Next in Tech With the Fast Forward Newsletter
Tweets From @varindiamag
Nothing to see here - yet
When they Tweet, their Tweets will show up here.
