Amir Ghazizadeh

Ph.D. Candidate, Computer Engineering
University of Central Florida  ·  graduating Summer 2027

I make sparse and irregular AI workloads run fast — across efficient deep learning model design, hardware accelerator microarchitecture, and the GPU kernels underneath. Day to day that means hardware–software co-design, cycle-accurate simulation, CUDA, and performance modelling. My work has appeared at NeurIPS, ISCA, and MICRO.

I’m a bit-bothering researcher — “that’s just the bottleneck” has never felt like the end of the answer!

amir@ucf — zsh
$ whoami
amir.ghazizadeh
$ cat about.json
{
"role": "Ph.D. candidate, CompE",
"lab": "UCF · Zheng Lab",
"builds": [
"sparse accelerators",
"CUDA graph kernels",
"efficient GNN / LLM systems"
],
"venues": ["NeurIPS", "ISCA", "MICRO", "DAC"],
"reviews_for": ["NeurIPS", "ICML", "MLSys"]
}
$ echo $STATUS
// open to full-time roles & internships, 2027
$
Research

What I’ve been building

Selected projects across model architecture, accelerator design, and efficient inference — most recent first.

Diagram of the GAMMA pipeline: an input graph feeds a gated multi-hop stage that maximises relevant signal, producing a compact graph representation.

NeurIPS 2025Spring 2025

GAMMA — gated message passing for heterophilic graphs

Gated multi-hop message passing driven by an information-theoretic objective that maximises relevant signal across aggregation hops. State-of-the-art accuracy on heterophilic benchmarks, with weight sharing and a fixed-dimensionality design that make the model dramatically cheaper to run.

  • 12×Less GPU memory
  • 20×Faster inference
  • 16Benchmark datasets
Diagram of the SpMM acceleration flow: a sparse graph is transformed and tiled, then mapped onto a processing-element array and memory hierarchy.

MICRO 2025Fall 2024

Rethinking tiling and dataflow for SpMM acceleration

Reformulates sparse–dense matrix multiplication as a graph-transformation problem: decomposition for irregular patterns, non-contiguous tiling for reuse, a runtime workload-partitioning engine, custom PE micro-architecture, and a Bidirectional Fiber Tree format that supports traversal in both directions. Verified in cycle-accurate simulation with Ramulator-based HBM modelling.

  • 4.3×Speedup vs. Sextans
  • 3.2×Less off-chip access
  • 4.8×Energy efficiency
Isometric diagram of capsule-network quantisation: layer-wise vulnerability analysis feeds mixed-precision allocation and Power-of-Two scalers that map onto hardware bit shifts.

M.Sc. research2021 – 2022

Quantising capsule networks for cheap hardware

Adaptive post-training quantisation with mixed-precision allocation, guided by a layer-wise vulnerability analysis that came out of my earlier fault-injection work on single-event upsets. Power-of-Two scalers replace floating-point multiplies with bit shifts, so the arithmetic maps onto simple hardware.

  • 8.67×Weight compression
  • 4.56×Activation reduction
  • <1%Accuracy lost

Earlier work

Capsule-network reliability under single-event upsets · a full MIPS datapath built from gate-level primitives · parallel primitives benchmarked across CPU thread pools and CUDA · and a soccer robot whose ball-detection pipeline taught me more about latency than any course did.

Publications

Peer-reviewed work

Full list and citation counts on Google Scholar.

MICRO 2026

Argus: An Efficient Accelerator Design for Sparse Nonlinear Solvers

F. Ye, A. Ghazizadeh, H. Zheng

ISCA 2026

TensorPrism: Rethinking Sparse High-order Tensor Acceleration via Co-occurrence Graph

F. Ye, S. Tian, A. Ghazizadeh, H. Zheng

NeurIPS 2025

GAMMA: Gated Multi-hop Message Passing for Homophily-Agnostic Node Representation in GNNs

A. Ghazizadeh, R. Ewetz, H. Zheng

MICRO 2025

Rethinking Tiling and Dataflow for SpMM Acceleration: A Graph Transformation Framework

A. Ghazizadeh, L. Yin, S. Tian, F. Ye, F. Yao, H. Zheng

DAC 2024

EGMA: Enhancing Data Reuse and Workload Balancing in Message Passing GNN Acceleration via Gram Matrix Optimization

F. Ye, L. Yin, A. Ghazizadeh, H. Zheng

ICCKE 2024

Towards Efficient Capsule Networks Through Approximate Squash Function and Layer-Wise Quantization

M. Raji, K. Soroush, A. Ghazizadeh

ICCAD 2023

ARIES: Accelerating Distributed Training in Chiplet-Based Systems via Flexible Interconnects

L. Yin, A. Ghazizadeh, A. Louri, H. Zheng

ICCD 2023

Polyform: A Versatile Architecture for Multi-DNN Execution via Spatial and Temporal Acceleration

L. Yin, A. Ghazizadeh, S. Tian, A. Louri, H. Zheng

Background

Experience & education

Experience

2022 — present

Graduate Research Assistant

University of Central Florida · Zheng Lab

Accelerator microarchitecture, sparse GPU kernels, and efficient GNN and LLM systems. First-author work at NeurIPS and MICRO; contributor on ISCA, DAC, ICCAD, and ICCD papers.

Fall 2026

Graduate Teaching Associate

University of Central Florida · Computer Organization

Instructor of record for the undergraduate course — lectures on MIPS/ARM ISA, datapath and control design, pipelining, and the memory hierarchy, plus assignments, exams, and processor design projects.

2021 — 2022

Visiting Researcher

Graduate University of Advanced Technology, Kerman

Multi-valued-logic arbiter PUFs in 32 nm and 14 nm CNTFET libraries, simulated in HSPICE. 99.6% (ternary) and 99.2% (quaternary) reliability across temperature and supply variation.

2020 — 2022

Graduate Research Assistant

Shiraz University

Quantisation and reliability of capsule networks — mixed-precision post-training quantisation, and fault-injection studies tracing single-event upsets through dynamic-routing layers.

2018 — 2020

Robotics Team

Shahid Bahonar University of Kerman

Embedded C/C++ on ARM microcontrollers, with an OpenCV ball-detection pipeline driving the perception, control, and wireless command loop.

Education

2022 — Summer 2027 (expected)

Ph.D., Computer Engineering

University of Central Florida · Orlando, FL

Efficient AI systems: hardware accelerators and GPU kernels for sparse and irregular workloads.

2020 — 2022

M.Sc., Computer Engineering

Shiraz University · Shiraz, Iran

Ranked 1st in the programme. Thesis on capsule-network reliability under single-event upsets and quantisation.

2014 — 2020

B.Sc., Computer Engineering

Shahid Bahonar University · Kerman, Iran

Digital design, computer architecture, and embedded systems.

Reviewer for NeurIPS · ICML · MLSys · IEEE Trans. on Computers

Toolset

What I work with

Bold is where I’m strongest.

Languages
PythonC++CUDATritonBash
ML & GPU libraries
PyTorchPyGcuSPARSEcuDNNTensorRT
Profiling & optimisation
Nsight Compute / Systemskernel fusiontilingstructured & unstructured sparsityquantisation
Architecture & modelling
Cycle-accurate simulationRamulatorgem5TimeloopAccelergyCACTIRTL & synthesis
Contact

Get in touch

I’m open to full-time roles for 2027 and research internships before then — and always happy to talk accelerators, sparse kernels, or graph learning.