FPGA-Validated - TRL 4 - Pre-Tapeout

Purpose-built AI accelerator silicon, engineered in India

The Neural Computer is a custom accelerator architecture for transformer training and inference - mixed-precision compute, a scalable tiled design, and high-bandwidth memory, validated end-to-end on FPGA.

~300 TFLOPS
Projected FP16 peak for a 7nm ASIC
~1.5 TFLOPS/W
~2x A100 efficiency at equal node
50+ configs
RTL verified and parameterised
Core Specifications

One architecture, from prototype to full-scale silicon

The same parameterised RTL that runs today on FPGA scales directly to a full ASIC. The prototype uses fewer, smaller tiles due to FPGA resource limits - not architectural constraints.

Specification FPGA Prototype Full-Scale ASIC (projected)
Compute4 processing tiles (reduced)32-64 processing tiles
PrecisionFP16 compute / FP32 accumulateFP16 compute / FP32 accumulate
Peak Performance~30 GFLOPS @ 150 MHz~300 TFLOPS @ ~1 GHz
Memory InterfaceHBM2e-class - 460 GB/sHBM3-class - ~1 TB/s
TrainingValidatedSupported
InferenceValidatedSupported
Component Architecture

Cluster / tile hierarchy — Systolic compute

Hover or select any block to inspect its role. FP16 multiply, FP32 accumulate, and HBM3-class memory on a single 7nm-target ASIC.

Single ASIC · 7nm target
Cluster 1
Cluster M
FPGA Prototype Validation

Proven on synthesised hardware, not simulation

A complete transformer trains end-to-end on the accelerator running on cloud FPGA infrastructure, with cycle-accurate HBM memory timing.

Hardware proof, condensed

Validation spans the complete training path, from matrix kernels through gradient aggregation and optimizer updates.

End-to-end training. Forward pass, activations, backward pass, and SGD/Adam weight update all on-chip.

Loss convergence confirmed. Monotonically decreasing loss verifies correct gradient flow.

Multi-tile execution. Parallel execution with hardware gradient aggregation across tiles.

Configuration coverage. RTL validated across 50+ parameterised hardware configurations.

4FPGA tiles
150 MHzPrototype clock
HBM2eCycle timing
SYSTEM HEALTH: OPTIMAL4 TILES / 150 MHz / HBM2e
FP16 matrix multiplicationVALIDATED
Multi-tile parallel executionVALIDATED
HBM bandwidth saturationVALIDATED
Vector unit FP16 operationsVALIDATED
Transformer training loopVALIDATED
Multi-tile gradient aggregationVALIDATED
Performance Comparison

Datacenter-class compute at a fraction of the power

Projected full-scale ASIC (64 tiles, 1 GHz, 7nm) against leading accelerators. Figures for competitors are published specifications.

Metric Neural Computer NVIDIA A100 NVIDIA H100 Google TPUv4
Peak FP16 TFLOPS~300312989275
Memory Bandwidth~1,000 GB/s2,039 GB/s3,350 GB/s1,200 GB/s
TDP~200 W400 W700 W175 W
TFLOPS / Watt~1.50.781.411.57
Process Node7 nm7 nm4 nm7 nm

At the same 7nm node as A100, the design targets A100-class peak compute at roughly half the power, the result of a purpose-built architecture that eliminates general-purpose GPU overhead.

Supported Workloads

Optimised kernels for demanding networks

From dense GEMM to full transformer training, with a software stack that offloads activations to the vector unit automatically.

Transformer Training

Full-stack support for BERT, GPT, and LLaMA patterns: forward, backward, and optimizer steps with distributed multi-tile attention.

Attention Mechanisms

Accelerated multi-head attention with fused compute kernels and flash-attention-inspired memory access patterns.

CNNs & Diffusion

ResNet and EfficientNet convolution paths, plus latent-diffusion and U-Net operations on specialised vector units.

GEMM Multi-head Attention GELU / SiLU / ReLU Softmax Layer Norm SGD / Adam MLP Distributed Training Gradient Aggregation
Scaling Projections

From edge inference to datacenter training

The same architecture, validated at multiple scales and parameterised across tile count and clock.

ScaleTilesClockPeak FP16Memory
FPGA prototype4 (reduced)150 MHz0.03 TFLOPSHBM2e 460 GB/s
ASIC (entry)16~1 GHz~75 TFLOPSHBM3-class ~1 TB/s
ASIC (mid)32~1 GHz~150 TFLOPSHBM3-class ~1 TB/s
ASIC (full)64~1 GHz~300 TFLOPSHBM3-class ~1 TB/s
ASIC (max)64~1 GHz~300 TFLOPSHBM3-class ~1 TB/s
Roadmap

From validated RTL to fabricated silicon

01 · COMPLETE

FPGA validation

Complete transformer training loop on FPGA. More than 50 RTL configurations verified. TRL 4 achieved.

02 · IN PROGRESS

Physical design

Synthesis, floorplanning, place-and-route, timing closure and DFT for the 7nm target.

03 · NEXT

Shuttle tapeout

First silicon run to validate the RTL against physical implementation.

04 · NEXT

Post-silicon bring-up

Characterise the fabricated device: throughput, power, thermal envelope and achieved utilisation.

05 · NEXT

Reference platform

Board-level product delivered to early partners and design-in customers.

06 · NEXT

Scale-out

Progression toward the full datacenter configuration with multi-cluster boards and rack-scale deployment.

Research Papers

View research papers

Published technical documents are available here to view.

Available papers

0 files

Build with Smyx Labs

For partnerships, early access, or technical collaboration around the Neural Computer program, contact the founding team.