OPEN-SOURCE EDGE AI PLATFORM

Build intelligence in the open.

Cerebral Chips is an open-source platform for people building capable, local AI. We are making the tools, evidence, and decisions visible so a community can shape intelligence that serves the machines—and people—who depend on it.

CC / 001
Proton NPU · bare-metal RTL milestone published
SYSTEM / PROTONLOCAL
INPUT / TOKEN STREAMSTATE / RESEARCH
Intelligence contained inside a compute deviceA five-node neural mesh receives a token signal and resolves it locally inside an octagonal device frame.
The Neural Mesh represents intelligence connected and contained inside the device.
SCROLL / 01
LATEST BUILD / PROTON NPURTL milestone published

One program. Scalar, vector, matrix.

An open hardware reference platform that runs scalar, RVV 1.0 vector and custom INT8 matrix instructions in one bare-metal program, verified in RTL simulation.

PROTON NPU / SYSTEM CONNECTIONSOne program. Three execution paths.
Instructions + completion Memory + data transfers
Proton NPU scalar, vector and matrix connectionsOne ELF runs on the 64-bit RISC-V core. The command router sends RVV instructions to the vector unit and mzero / mmacc to the matrix engine, returning completion to the CPU. Scalar and vector loads and stores use the AXI memory fabric. The CPU copies matrix operands and results through the matrix engine’s memory-mapped tile buffers. There is no direct vector-to-matrix connection or matrix DMA.instructionsissueRVVmzero / mmaccScalar loads / storesMMIOVector loads / storesONE BARE-METAL ELFScalar + vector + matrixSCALAR / RV6464-bit RISC-V coreOne core · fetch, decode, retireScalar registers + cachesINSTRUCTION DISPATCHCommand routerRequests + completionVECTOR / RVV 1.0RISC-V vector unit2 lanes · VLEN 2,048 bitsVector registers + load/storeMATRIX / INT8 → INT324 × 4 matrix engine16 PEs · 4 products per PEA / BT / C local tile buffersINT32 accumulation + outputMEMORY + DATA MOVEMENTShared AXI memory fabricShared RAM + peripheralsProton NPU scalar, vector and matrix connectionsOne ELF runs on the 64-bit RISC-V core. The command router sends RVV instructions to the vector unit and mzero / mmacc to the matrix engine, returning completion to the CPU. Scalar and vector loads and stores use the AXI memory fabric. The CPU copies matrix operands and results through the matrix engine’s memory-mapped tile buffers. There is no direct vector-to-matrix connection or matrix DMA.instructionsissue / completeRVVmzero / mmaccMMIO copiesScalar + vector loads / storesONE BARE-METAL ELFScalar + vector + matrixSCALAR / RV6464-bit RISC-V coreOne core · fetch, decode, retireScalar registers + cachesINSTRUCTION DISPATCHCommand routerRequests + completionVECTOR / RVV 1.0RISC-V vector unit2 lanes · VLEN 2,048 bitsVector registers + load/storeMATRIX / INT8 → INT324 × 4 matrix engine16 PEs · 4 products per PEA / BT / C local tile buffersINT32 accumulation + outputMEMORY + DATA MOVEMENTShared AXI memory fabricShared RAM + peripherals

Matrix data path: shared RAM → CPU copies → A / BT tile buffers → INT8 compute → INT32 C buffer → CPU copies → shared RAM. Vector code can then use the results in RAM. No direct vector-to-matrix port or matrix DMA.

8 / 8Combined workloads passed
86Matrix commands reconciled
115,650RTL cycles in the combined run
Inspect the verification records ↗

Functional RTL simulation with bare-metal C. These results do not establish a silicon clock, AI throughput or Linux support.

01

THE MISSION

Put the tools for local intelligence in everyone’s hands.

Language models are becoming part of products, robots, industrial systems, vehicles, personal devices, and embedded interfaces. Those systems need local intelligence with predictable latency, privacy, and operation even when the cloud is unavailable.

Cerebral Chips is building that foundation as an open-source platform—from workload analysis and software integration down to the instruction architecture, memory system, and quantized datapath—so contributors can inspect it, improve it, and carry it forward together.

01
WORKLOAD EVIDENCEModel workload

Observe operators, shapes, movement, and quantization.

02
ARCHITECTURE DECISIONSMeasured architecture

Turn measured behavior into explicit compute tradeoffs.

03
DEVICE EXECUTIONDevice execution

Keep intelligence responsive, private, and local.

02

HARDWARE × SOFTWARE CO-DESIGN

Modern edge AI needs one system, not two disconnected roadmaps.

Cerebral Chips treats models, quantization, compilers, runtimes, firmware, instruction architecture, matrix and vector execution, and memory movement as one connected design problem. Real workload behavior drives the machine; hardware constraints reshape the software path.

SOFTWARE SYSTEMMake the workload programmable
S1Models and quantization
S2Compiler and runtime
S3Kernels and profiling
ACTIVE CO-DESIGN LOOP

Start with what the model actually does.

Prefill, decode, operator mix, tensor shapes, and quantization behavior establish the first design contract.

HARDWARE SYSTEMMake the workload efficient
H1ISA and command model
H2Vector and matrix engines
H3Scratchpad, DMA, and memory
03

THE PROTON PLATFORM

A complete stack for on-device language intelligence.

The Proton platform connects model workloads, developer software, architecture simulation, agentic engineering, and accelerator hardware. Each layer exists to make the next layer measurable, programmable, and verifiable.

P-01Prototype
NPU

Proton NPU

An open hardware reference platform that runs scalar, RVV 1.0 vector and custom INT8 matrix instructions in one bare-metal program, verified in RTL simulation.

Source on GitHub ↗
P-02Research
LPU

Proton LPU

A quantization-native, memory-aware language-processing accelerator architecture.

P-03In development
SDK

Proton SDK

A planned developer path from GGUF models and llama.cpp to runtime, kernels, and profiling.

P-04Prototype
Sim

ProtonSim

A workload and architecture simulator for choosing hardware from measurable model behavior.

P-05Prototype
Forge

Proton Forge

An evidence-gated engineering system for accelerating hardware and software iteration.

04

PROTON LPU / PROPOSED ARCHITECTURE

Language-model hardware designed around real inference behavior.

Proton LPU is a proposed RISC-V-controlled accelerator architecture for quantized transformer inference. It combines vector execution, quantization-aware matrix compute, explicitly managed local memory, DMA, external-memory access, device firmware, and a practical llama.cpp software path.

SYSTEM MAP / EDITABLECandidate execution architecture
Research
01Host systemModel orchestration and command submission
02RISC-V controlFirmware, queues, and execution control
03Vector engineElementwise, reduction, and support operations
04Quantized matrix engineGEMM and GEMV execution candidates
05Local scratchpadExplicitly managed working data
06DMA + external memoryScheduled data movement and model storage

Logical architecture view. Interfaces and parameters remain subject to workload-driven evaluation.

MATRIX ENGINE / MAPPING VIEW8 × 8 outer-product candidate
Under evaluation
PARTIAL SUMS / INT32 DIRECTION
01Parallel products

Eight activation lanes interact with eight weight lanes.

02Dataflow study

GEMM, GEMV, and reduction behavior remain under evaluation.

03Complete quantization path

Scale, zero-point, accumulation, and output handling stay in scope.

Primary workloadProposed
Dense decoder-only transformer inference

The first workload contract is intentionally narrow and measurable.

Initial model envelopeUnder evaluation
Approximately 0.5B–1.5B parameters

A 3B parameter model remains a stretch target.

Arithmetic directionProposed
Four-bit weights, eight-bit activations, INT32 accumulation

Exact GGUF quantization coverage will be selected through bring-up work.

Matrix candidateUnder evaluation
8 × 8 physical MAC fabric

GEMM and GEMV dataflows are evaluated independently.

Compute scratchpadUnder evaluation
128 KB, 256 KB, and 512 KB study

256 KB is a working baseline, not a finalized choice.

05

PROTON SDK / SOFTWARE PATH

From downloadable model to device execution.

The initial software path starts with GGUF models and llama.cpp. A Proton GGML backend, userspace runtime, firmware, kernels, command queues, and profiling tools connect model execution to the accelerator. IREE remains a future frontend and compiler integration path.

INITIAL EXECUTION PATHIn development
  1. 01GGUF modelMODEL
  2. 02llama.cpp / GGMLHOST SOFTWARE
  3. 03Proton backendHOST SOFTWARE
  4. 04Proton runtimeHOST SOFTWARE
  5. 05Firmware and kernelsDEVICE PATH
  6. 06Proton LPU or ProtonSimDEVICE PATH
FUTURE FRONTEND PATH
IREEFuture
06

PROTONSIM / DESIGN SPACE

Design the hardware from workload evidence.

ProtonSim models the work, memory movement, and execution behavior of real transformer workloads before architecture choices are frozen. It is intended to connect model operators and quantization formats to decisions such as vector width, matrix shape, scratchpad capacity, DMA behavior, and external-memory bandwidth.

EXAMPLE STUDY INPUTSArchitecture explorer
Prototype
Workload phase
Quantization format
Matrix shape
NO RESULT GENERATEDConnect a validated workload model to view results.

Controls illustrate the intended design space. They do not represent benchmark results or finalized specifications.

Selected workload
Decode
Quantization study
Grouped INT4
Compute candidate
8 × 8
07

AI-NATIVE ENGINEERING

Agents accelerate the work. Evidence decides what ships.

Proton Forge is an agentic engineering system for specification analysis, architecture exploration, RTL and verification scaffolding, firmware, kernel generation, documentation, and regression triage.

Generated artifacts are never accepted on model confidence alone. They pass through deterministic builds, reference-model comparison, randomized tests, assertions, coverage, synthesis, and human review.

EVIDENCE-GATED PIPELINEPrototype
01Requirements
02Structured specification
03Workload model
04Design or code generation
05Compile and lint
06Golden-reference comparison
07Randomized testing
08Coverage and synthesis regression
09Human review
ACCEPT ONLY WITH REPRODUCIBLE EVIDENCE
08

KERNEL GENERATION / VERIFY

Target-specific kernels with reproducible verification.

The Proton toolchain is being designed to generate and tune custom kernels for vector and matrix targets from structured specifications. The accepted result is a versioned artifact backed by compilation, numerical differential testing, edge-case coverage, and performance evidence.

kernel.spec / illustrativeREAD ONLY
target: proton_vector_matrix
operation: quantized_linear
weights: grouped_int4
activations: int8_candidate
accumulation: int32
accept_if:
  - compiles
  - matches_reference
  - passes_edge_cases
01Compile
02Numerical diff
03Edge cases
04Versioned evidence
09

PROTON NPU / WORKING REFERENCE

Build useful systems before custom silicon is complete.

Proton NPU is the working reference platform for scalar, vector and matrix integration. The longer-term Proton LPU research explores quantization, memory movement and language-model execution beyond this baseline.

Proton NPU source, architecture notes, bare-metal examples and RTL verification records are public.

Explore the Proton NPU repository
OPEN SOURCE / SHARED STEWARDSHIP

A platform built in public, held in common.

Cerebral Chips is an open-source platform for everyone who wants local intelligence to be transparent, useful, and accessible. We are inviting engineers, researchers, makers, and system builders to help shape the work through ideas, review, experiments, and code.

01Architecture notes
02Workload models
03Quantization references
04RTL and verification
05Software and firmware
06FPGA bring-up
07Benchmark methodology
08Decision records
Join the community build
10

ROADMAP / STAGE-BASED

Build the evidence, then build the machine.

The roadmap follows technical evidence rather than calendar promises. Each stage reduces uncertainty for the architecture, software path, and reference system that follows.

  1. 01
    Active

    Workload contracts and RISC-V reference platforms

  2. 02
    Active

    ProtonSim and architecture performance modeling

  3. 03
    Planned

    Proton SDK, runtime, and llama.cpp backend

  4. 04
    Planned

    Proton LPU-E0 RTL and verification

  5. 05
    Future

    FPGA reference system

  6. 06
    Future

    Open ASIC exploration

  7. 07
    Future

    Proton LPU-E1 multi-tile edge architecture

  8. 08
    Long-term

    Proton LPU-S1 server research

OPEN SOURCE / COMMON CAUSE

Build the future of local intelligence with us.

This platform is for the people who need AI to be local, understandable, and useful. Bring a question, an experiment, a review, or a line of code—every contribution can move the shared foundation forward.