Graduate Research Assistant
University of Central Florida · Zheng LabAccelerator microarchitecture, sparse GPU kernels, and efficient GNN and LLM systems. First-author work at NeurIPS and MICRO; contributor on ISCA, DAC, ICCAD, and ICCD papers.
Ph.D. Candidate, Computer Engineering
University of Central Florida · graduating Summer 2027
I make sparse and irregular AI workloads run fast — across efficient deep learning model design, hardware accelerator microarchitecture, and the GPU kernels underneath. Day to day that means hardware–software co-design, cycle-accurate simulation, CUDA, and performance modelling. My work has appeared at NeurIPS, ISCA, and MICRO.
I’m a bit-bothering researcher — “that’s just the bottleneck” has never felt like the end of the answer!
Selected projects across model architecture, accelerator design, and efficient inference — most recent first.
Gated multi-hop message passing driven by an information-theoretic objective that maximises relevant signal across aggregation hops. State-of-the-art accuracy on heterophilic benchmarks, with weight sharing and a fixed-dimensionality design that make the model dramatically cheaper to run.
Reformulates sparse–dense matrix multiplication as a graph-transformation problem: decomposition for irregular patterns, non-contiguous tiling for reuse, a runtime workload-partitioning engine, custom PE micro-architecture, and a Bidirectional Fiber Tree format that supports traversal in both directions. Verified in cycle-accurate simulation with Ramulator-based HBM modelling.
Adaptive post-training quantisation with mixed-precision allocation, guided by a layer-wise vulnerability analysis that came out of my earlier fault-injection work on single-event upsets. Power-of-Two scalers replace floating-point multiplies with bit shifts, so the arithmetic maps onto simple hardware.
Capsule-network reliability under single-event upsets · a full MIPS datapath built from gate-level primitives · parallel primitives benchmarked across CPU thread pools and CUDA · and a soccer robot whose ball-detection pipeline taught me more about latency than any course did.
Accelerator microarchitecture, sparse GPU kernels, and efficient GNN and LLM systems. First-author work at NeurIPS and MICRO; contributor on ISCA, DAC, ICCAD, and ICCD papers.
Instructor of record for the undergraduate course — lectures on MIPS/ARM ISA, datapath and control design, pipelining, and the memory hierarchy, plus assignments, exams, and processor design projects.
Multi-valued-logic arbiter PUFs in 32 nm and 14 nm CNTFET libraries, simulated in HSPICE. 99.6% (ternary) and 99.2% (quaternary) reliability across temperature and supply variation.
Quantisation and reliability of capsule networks — mixed-precision post-training quantisation, and fault-injection studies tracing single-event upsets through dynamic-routing layers.
Embedded C/C++ on ARM microcontrollers, with an OpenCV ball-detection pipeline driving the perception, control, and wireless command loop.
Efficient AI systems: hardware accelerators and GPU kernels for sparse and irregular workloads.
Ranked 1st in the programme. Thesis on capsule-network reliability under single-event upsets and quantisation.
Digital design, computer architecture, and embedded systems.
Reviewer for NeurIPS · ICML · MLSys · IEEE Trans. on Computers
Bold is where I’m strongest.
I’m open to full-time roles for 2027 and research internships before then — and always happy to talk accelerators, sparse kernels, or graph learning.