Mingpu is an extremely minimal GPU architecture that I wrote for fun and also to learn basic hardware design. Though, it does not have the "G" in "GPU" (yet) - it does not do any graphics, but rather focuses on pure parallel SIMD compute.
Currently, Mingpu includes:
- A control unit to dispatch instructions to compute cores.
- 80 compute cores each with:
- 2 16-bit registers.
- 16-bit word size in register and memory.
- 16-bit instruction: 7-bit opcode, 1-bit register, 8-bit address.
- 512 bytes (256 16-bit words) of local mem per core.
- A minimal 6-op ISA:
NOP.ADD:r0 <= r0 + r1(signed).MUL:r0 <= r0 * r1(signed).STORE ri, addr:mem[addr] <= ri.LOAD ri, addr:ri <= mem[addr].HALT.
- 1 8-bit program counter for all cores, which means 256 instructions (512 bytes) for maximum kernel size.
I currently use Icarus Verilog for development of this project, so have it installed and you are good to go.
make simThis runs the regression suite in ./sim/gpu_tb.sv on all cores: a register LOAD/STORE test, an 8×4 by 4×10 matrix multiply, and a restart test that re-runs the matmul with new data. Waveforms are dumped to gpu_tb.vcd.
You can configure the gpu (number of cores, local mem size, data width, etc.) in ./rtl/gpu_pkg.sv.
- Rethink better arch overall, currently this is a very naive arch and implementation from me. Though it should always stay minimal.
- Integration with real hardware, possibly an EBAZ4205, which should come with:
- Ethernet communication code.
- Driver.
- Assembler.
Copyright © 2026 Nguyen Phu Minh.
This project is licensed under the Apache 2.0 License.