![]() |
Wenfei WuAssistant Professor
|
Wechat ID
|
Clusters keep growing, but application performance does not grow with them. As parallelism increases, the processes that jointly perform a computation must exchange ever more intermediate state, until communication — not computation — becomes the limit. Stacking more hardware does not remove this wall.
Our goal is to build high-performance clusters that enable scientific discovery and improve the quality of service of distributed systems. The way we pursue it is to redraw the division of labour between the network and the endpoints, so that the network is no longer a transport pipe but a class of compute resource that applications in the cluster can invoke.
We build or exploit the computational capability of network devices — programmable switches, SmartNICs, DPUs — and design network protocols that compute: as data travels from source to destination, aggregation, reduction, caching and agreement are performed in place on the devices along the path, rather than after all the data has landed at the endpoints.
The benefit is structural rather than incremental. The network carries results instead of raw data, so traffic volume drops; replies can be generated mid-path, so latency shrinks; computation proceeds at packet granularity, so computation and communication overlap naturally. These gains act directly on communication and therefore do not decay as the cluster grows.
We decompose the problem into three connected and successively broader lines.
INC protocols and algorithms — accelerating distributed computation correctly. Once computation is embedded in the transmission path, a design must guarantee not only correct delivery but correct results under loss, retransmission, reordering, bounded on-chip memory and failures, while remaining compatible with the deployed transport stack. (NetReduce, ASK, Turbo, VeriNC)
INC resource management — letting concurrent INC instances share a cluster efficiently. On-chip memory is scarce, measured in megabytes, and cannot be expanded by adding servers; its placement, routing and scheduling decide both per-job speedup and cluster-wide capacity. (ATP, NetPack, INAlloc, DSA)
Realization and deployment — adapting INC to the existing network ecosystem, so vendors can build the feature and users can actually use it. We lower development and porting cost through unified abstractions, an intermediate representation with per-platform compilation, and primitives hardened into silicon. (EPIC, ClickINC, NetRPC)