Tools & Libraries

fasterfly

franciscocarloserra

Event-driven Triton kernels that step the whole MaleCNS connectome, 165,000 neurons and 24.5 million synapses, at 3,455 Hz on a single RTX 3090. With a 1 ms membrane time step that is one second of fly life in 0.29 seconds of wall clock, about 3.5 times faster than real time, against 708 steps per second for a torch.sparse COO baseline. Batching sixteen or sixty-four flies trades single-fly latency for throughput and reaches 51,159 fly-steps per second. The README is explicit that the same input produces the same spikes as the torch.sparse reference.

Official site ↗

fasterfly

#triton#gpu-kernel#performance#malecns#benchmark

Sources checked

Every fact above was read from these pages on the date the entry was added.

  • github.com https://github.com/franciscocarloserra/fasterfly