PROTON PLATFORMPROTON NPU / RTL REFERENCE
Prototype

One program. Scalar, vector, matrix.

An open hardware reference platform that runs scalar, RVV 1.0 vector and custom INT8 matrix instructions in one bare-metal program, verified in RTL simulation.

PUBLIC STATUS BOUNDARY

Bare-metal RTL milestone · October 2026. Functional RTL simulation with bare-metal C. These results do not establish a silicon clock, AI throughput or Linux support.

SYSTEM VIEW
CEREBRAL CHIPS
PROTON NPU / SYSTEM CONNECTIONSOne program. Three execution paths.
Instructions + completion Memory + data transfers
Proton NPU scalar, vector and matrix connectionsOne ELF runs on the 64-bit RISC-V core. The command router sends RVV instructions to the vector unit and mzero / mmacc to the matrix engine, returning completion to the CPU. Scalar and vector loads and stores use the AXI memory fabric. The CPU copies matrix operands and results through the matrix engine’s memory-mapped tile buffers. There is no direct vector-to-matrix connection or matrix DMA.instructionsissueRVVmzero / mmaccScalar loads / storesMMIOVector loads / storesONE BARE-METAL ELFScalar + vector + matrixSCALAR / RV6464-bit RISC-V coreOne core · fetch, decode, retireScalar registers + cachesINSTRUCTION DISPATCHCommand routerRequests + completionVECTOR / RVV 1.0RISC-V vector unit2 lanes · VLEN 2,048 bitsVector registers + load/storeMATRIX / INT8 → INT324 × 4 matrix engine16 PEs · 4 products per PEA / BT / C local tile buffersINT32 accumulation + outputMEMORY + DATA MOVEMENTShared AXI memory fabricShared RAM + peripheralsProton NPU scalar, vector and matrix connectionsOne ELF runs on the 64-bit RISC-V core. The command router sends RVV instructions to the vector unit and mzero / mmacc to the matrix engine, returning completion to the CPU. Scalar and vector loads and stores use the AXI memory fabric. The CPU copies matrix operands and results through the matrix engine’s memory-mapped tile buffers. There is no direct vector-to-matrix connection or matrix DMA.instructionsissue / completeRVVmzero / mmaccMMIO copiesScalar + vector loads / storesONE BARE-METAL ELFScalar + vector + matrixSCALAR / RV6464-bit RISC-V coreOne core · fetch, decode, retireScalar registers + cachesINSTRUCTION DISPATCHCommand routerRequests + completionVECTOR / RVV 1.0RISC-V vector unit2 lanes · VLEN 2,048 bitsVector registers + load/storeMATRIX / INT8 → INT324 × 4 matrix engine16 PEs · 4 products per PEA / BT / C local tile buffersINT32 accumulation + outputMEMORY + DATA MOVEMENTShared AXI memory fabricShared RAM + peripherals

Matrix data path: shared RAM → CPU copies → A / BT tile buffers → INT8 compute → INT32 C buffer → CPU copies → shared RAM. Vector code can then use the results in RAM. No direct vector-to-matrix port or matrix DMA.

01THE WORKING BASELINE

A complete instruction path, running on RTL.

One 64-bit RISC-V CPU runs a bare-metal C application. Scalar instructions handle control and checking, RVV 1.0 instructions execute on a two-lane vector unit, and custom matrix instructions launch the INT8 engine. The same program uses all three execution paths.

Verilator turns the RTL into a simulator. The application runs on the simulated CPU, with instruction retirement and matrix command traces checked against the expected results.

  • 01One RV64 scalar core
  • 02Two vector lanes, VLEN 2,048 bits and ELEN 64 bits
  • 034 × 4 matrix PEs, four signed INT8 products per PE, INT32 accumulation
02COMPUTE + DATA MOVEMENT

Small tiles, explicit transfers, inspectable results.

The C library tiles larger matrices into 4 × 4 output blocks, accumulating K in chunks of 16 and padding incomplete tiles. The CPU copies operands into AXI-mapped buffers, issues mzero and mmacc, then copies row-major INT32 results back to RAM for scalar or vector processing.

LLVM emits the scalar and RVV instructions. Inline assembler encodes the custom matrix commands; no compiler fork is needed for this milestone.

The matrix instructions are custom. This implementation has no matrix DMA or direct vector-register-to-matrix port.

03PUBLISHED VERIFICATION

Every result has a check behind it.

The combined program passes eight workloads, including dimensions that require partial tiles, in 115,650 simulated cycles. Its 52 mzero and 34 mmacc instructions reconcile with 86 issue, completion and response sequences. All 86 matrix output tiles were also checked from recorded RTL waveforms.

  • 01262,144 signed-byte/lane arithmetic combinations and 256 tile cases
  • 029/9 selected scalar/vector tests with matrix enabled and 9/9 with it disabled
  • 03Wrong-result and forced-timeout runs rejected by the verification wrapper

Cycle counts include startup, checking and console output. They are functional test evidence, not an AI benchmark, silicon frequency or TOPS claim.

04BUILD / RUN / INSPECT

Start with the same source and the same workload.

The public repository contains pinned dependencies, reviewable integration patches, bare-metal examples, build scripts, independent checks and waveform guides. The validated environment is ARM64 Ubuntu in a Lima VM on Apple Silicon.

Start with the repository README, provision the tools and run the combined scalar/vector/matrix example. The fresh-install bootstrap is documented; its full cold path still needs validation on a second clean machine.

05REFERENCE PLATFORM → RESEARCH

A working foundation for the next design decisions.

Proton NPU is the working reference platform for scalar, vector and matrix integration. The longer-term Proton LPU research explores quantization, memory movement and language-model execution beyond this baseline.

Linux, IREE and language-model runtime integration, lower-bit quantization, FPGA bring-up and physical implementation remain future work. This milestone does not establish model inference performance or a fabricated chip.

06OPEN HARDWARE / SHARED CREDIT

Build on open components. Keep their provenance visible.

Our integration combines established open-source scalar, vector and matrix compute IP with a small, explicit hardware/software contract. Cerebral Chips’ original contributions use Apache-2.0; third-party components retain their licenses and attribution.

CONTINUE THE SYSTEM

Contribute to the open build

Continue