Skip to content

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

mingpu

Mingpu is an extremely minimal GPU architecture that I wrote for fun and also to learn basic hardware design. Though, it does not have the "G" in "GPU" (yet) - it does not do any graphics, but rather focuses on pure parallel SIMD compute.

Architecture

Currently, Mingpu includes:

  • A control unit to dispatch instructions to compute cores.
  • 80 compute cores each with:
    • 2 16-bit registers.
    • 16-bit word size in register and memory.
    • 16-bit instruction: 7-bit opcode, 1-bit register, 8-bit address.
    • 512 bytes (256 16-bit words) of local mem per core.
    • A minimal 6-op ISA:
      • NOP.
      • ADD: r0 <= r0 + r1 (signed).
      • MUL: r0 <= r0 * r1 (signed).
      • STORE ri, addr: mem[addr] <= ri.
      • LOAD ri, addr: ri <= mem[addr].
      • HALT.
  • 1 8-bit program counter for all cores, which means 256 instructions (512 bytes) for maximum kernel size.

Setup

I currently use Icarus Verilog for development of this project, so have it installed and you are good to go.

Run testbench

make sim

This runs the regression suite in ./sim/gpu_tb.sv on all cores: a register LOAD/STORE test, an 8×4 by 4×10 matrix multiply, and a restart test that re-runs the matmul with new data. Waveforms are dumped to gpu_tb.vcd.

Configuration

You can configure the gpu (number of cores, local mem size, data width, etc.) in ./rtl/gpu_pkg.sv.

Todos

  • Rethink better arch overall, currently this is a very naive arch and implementation from me. Though it should always stay minimal.
  • Integration with real hardware, possibly an EBAZ4205, which should come with:
    • Ethernet communication code.
    • Driver.
    • Assembler.

Copyright and License

Copyright © 2026 Nguyen Phu Minh.

This project is licensed under the Apache 2.0 License.

Contributors

Languages