On May 19, 2026, our research team, led by Professor Xingjun Wang and Researcher Haowen Shu, published a research article entitled “On-chip large-scale all-optical interconnect for ultra-low-latency deep neural network inference” online in National Science Review. The study proposes an optoelectronic distributed computing system based on on-chip all-optical supernodes. By connecting multiple computing nodes through silicon photonic transceiver chips and an optical switching chip, the researchers demonstrated low-latency pipelined parallel inference on a five-layer convolutional neural network for image denoising. The article is part of the journal’s special topic on optoelectronic integration.
As artificial intelligence models grow in scale, computing systems require not only more processing power but also more efficient data transfer. When multiple computing chips work together, inter-node bandwidth, data movement, and waiting time can limit the utilization of computing resources. Optical interconnects can provide high-speed data channels, but translating this advantage into distributed-computing performance also requires a low-loss, reconfigurable switching network and a compatible data-scheduling strategy.
Our research team combined a 400 Gb/s silicon photonic transceiver chip with a 16 × 16 non-blocking optical switching chip to establish reconfigurable all-optical connections. The optical switch achieved a total loss of no more than 5 dB near 1300 nm, reducing the need for additional optical amplification. Across the 20 switching paths tested, the system achieved error-free transmission with forward error correction enabled, providing stable links for data exchange among multiple computing nodes. Here, “all-optical” refers to the interconnection and switching between nodes; the neural-network computation itself was performed by field-programmable gate arrays (FPGAs).

Figure 1. Schematic of the optoelectronic distributed computing system based on on-chip all-optical supernodes
In the system demonstration, the researchers deployed the five layers of a convolutional neural network on five separate FPGAs and configured the optical switching network as a pipelined parallel architecture. The output of each layer was sent directly to the next computing node through the optical interconnect, allowing different layers to process consecutive inputs simultaneously. The system completed denoising inference on 1,000 images of 32 × 32 pixels, represented in 32-bit floating-point format, in a total of 105.16 μs.
Under the same task and FP32 precision used in the study, a single-GPU baseline required 15.643 ms, approximately 149 times the latency of the proposed system. The peak computing capacity estimated for the deployed FPGA resources was 1.96875 TFLOPS, or about 11.6% of the 16.96 TFLOPS reported for the GPU baseline. This comparison applies to the small convolutional network and implementation tested in the paper and shows that coordinated interconnection and pipelined computing can improve the utilization of computing resources.
The study provides an experimentally validated system approach for organizing multiple computing chips through optical interconnects. Further improvements in FPGA implementation, network scale, and optoelectronic interface speed will enable evaluation on more complex inference tasks. Zihan Tao, Yan Zhou, Weizhen Yu, and Huajin Chang are the co-first authors. Haowen Shu and Xingjun Wang are the co-corresponding authors, and Peking University is the principal institution responsible for the work.
Original article: https://doi.org/10.1093/nsr/nwag282