tomato
KiCad render of Tomato’s green circuit board, showing the two logic cells, chips, copper connections, and Penn Engineering mark

Drag to explore

Controls

Drag to orbit · scroll to zoom · 1–4 viewsDrag to orbit · pinch to zoom

Tour the board

A functional 32-bit computer
built from scratch in a dorm.

Designed from the logic level up — microarchitecture, ISA, assembler, OS, games, UI, and the tools that verify it. The ALU is designed, fabricated, and assembled. Run it in the browser or text the computer in my dorm. Built as a sophomore.

Run it in the browserText it when Tomato is awake

What is Tomato?

A functional 32-bit computer I designed from the logic level up in my dorm. The ALU is designed, fabricated, and assembled. You can run the machine in a browser, or text the one on my desk.

Its logic can change with each instruction.

See how it works

ONE ARCHITECTURE · THREE WAYS TO EXPERIENCE IT

See the hardware.
Run it. Then text it.

The complete computer runs on FPGA. The physical discrete build currently proves its 8-bit ALU slice. When a verified bridge holds a lease, Tomato is awake in Philadelphia. When it doesn’t, Tomato is sleeping — leave a message, or talk to Virtual Tomato. Virtual and physical replies are always labeled.

SHIPPED TOMATO OS

Start at Home.
Choose what to run.

This is the real operating-system menu running in the browser CPU emulator—not a mock overlay.

Loading Tomato OS…

Move through Home with the arrow controls, then press ● to open the selected app. Hold an arrow to keep moving.

Open full screen + demo chat
Current Tomato OS assembly in the browser CPU emulator · deterministic Home boot · simulated peripherals · not an FPGA result
OPEN LARGER
Nexys A7-100T · Artix-7 · HDMI

Tomato works—it boots its OS
to run Tetris, Snake, and more.

I wrote Tomato OS in the assembly language I designed for this machine. It draws the menu, reads the five-button controls, and runs Snake, Tetris, five-lane Racer, Fibonacci, Tribonacci, and Envelop.

Tomato OS · hardware recording

The discrete 8-bit ALU slice is
assembled and under bench test.

This Lot 07 board is a physical slice of Tomato’s Dual-LUT ALU. It is not the complete discrete computer; the complete working Tomato computer currently runs on the Nexys A7 FPGA.

Half-soldered Tomato 07_alu board held next to its Digital simulation
Physical Lot 07 ALU sliceDiscrete 74ACT logic on copper · assembly record, 18 Aug 2026

One circuit.
524,288 ways to configure it.

The ALU is the processor’s arithmetic and logic engine. I gave mine two programmable truth tables per bit and wired them straight into an adder.

An instruction can change those tables. So the same circuit can add numbers, combine bits, or combine bits and add—all in one trip through the ALU.

For example, the same hardware can compute A + B or A + (B AND C). Change the truth tables, and the adder receives a different calculation.

What “polymorphic” means on Tomato
Inside the ALU: two truth tables and an adder

Two rules. One sum.

Follow one calculation through the ALU.

256rules for f
256rules for g
8carry sources
524,288ALU configurations
Follow every bit.8-BIT VIEW · BINARY LIVE
Set the inputs · decimal 0–255

A + B

13
0000 11010x0D

Add A and B. In binary: 0000 1000 + 0000 0101 = 0000 1101 (13).

F OPCODE0xAA1010 1010Pass A
G OPCODE0xCC1100 1100Pass B
Trace F + G + carry in binary
f(A, B, C)0000 1000
+ g(A, B, C)0000 0101
+ carry in0
= result0000 1101

Carry out: 0 · low 8 bits shown

Read the opcodes as truth tables

Each opcode is eight answers, one for every combination of three input bits. The table index is {C, B, A}; opcode bit 0 is the answer for 000.

Input combination → output from each LUT
C B AF bitG bit
00000
00110
01001
01111
10000
10110
11001
11111

An 8-bit window into the 32-bit ALU. Choose and Majority are building blocks of SHA-256; each example shows one operation, not a complete hash.

Open the full 32-bit playground
For the engineer: what does 524,288 actually count?

Two independently programmed LUT3 planes offer 256 × 256 truth-table pairs. Eight carry-source selections give 524,288 control configurations. Some configurations produce the same function. The instruction ROM has 512 rows and selects a practical subset.

The datapath computes f(A,B,C) + g(A,B,C) + carry. One ALU evaluation is distinct from a complete CPU instruction or a whole-program speedup.

Instruction set and ROM mapping

From design to a soldered board

The ALU is built.
Here is the work behind it.

Designed, verified, fabricated, and soldered. The remaining discrete CPU boards are in progress.

  1. DoneDesign
  2. DoneSchematics
    + routing
  3. DoneSimulation
    + formal checks
  4. Done · Aug 2026Fabrication
  5. Done · Aug 2026Assembly
    + first lights

The complete computer also runs Tomato OS on FPGA. See the running system

How one instruction configures the whole datapath

Why I call it polymorphic

One datapath.
Different logic with each instruction.

I use polymorphic to describe how an instruction configures the machine’s existing paths. It starts with the Dual-LUT ALU: two programmable Boolean functions feed the same adder. Change their truth tables and carry selection, and the circuit can perform ordinary arithmetic, logic, or a compound expression such as A + (B AND C).

The flexibility extends beyond the ALU. The instruction ROM also selects operand sources, immediate interpretation, and the controls for flags, write-back, and sequencing. The immediate box is one part of that system: it constructs a value from instruction bits combinationally, without a separate conversion instruction.

Both the operation and the way its operands reach it can change. That is the idea behind the name. ISA profiles map supported operations onto these controls; a profile alone does not establish compatibility with unmodified foreign binaries.

See the instruction mappings Why I designed it this way
One instruction configures all four
9-bit opcodeInstruction ROMControl word ↓
01 / The operationTwo truth tables.
One adder.
f(A,B,C) + g(A,B,C) + cin

Arithmetic, Boolean logic, or both in the same evaluation.

02 / The inputsChoose the operands

Register values and immediate sources feed the computation.

03 / The instruction bitsInterpret the immediate

Sixteen selections construct signed, unsigned, and other forms.

04 / The resultWrite back. Set flags. Continue.

Controls select where the result goes and how execution proceeds.

Explore the complete architecture

Hardware-level compilerFind the operation once.
Reuse it in one ALU pass.

A counter FSM that acts as a primitive hardware compiler. It tests one of the ALU’s 65,536 Dual-LUT configurations per CPU clock; on the default 6.25 MHz FPGA CPU, a full sweep takes at most about 10.5 ms.

You supply fixed inputs and a target output. The FSM sweeps opcode space, freezes on a match, and latches that opcode—so I can reuse exotic single-cycle operations that would otherwise take many instructions.

A match is still a candidate to verify; sweep time scales directly with the CPU clock.

Explore the hardware compiler
Digital simulation · follow the search through the circuit.
FPGA · the same search running in hardware.

Recordings show the sweep in Digital and on the Artix-7.

  1. 01 / SPECIFYInputs + desired output
  2. 02 / SWEEPCounter tries F × G opcodes
  3. 03 / MATCHFreeze and latch the opcode
  4. 04 / VERIFYCheck more inputs and the function

If I can’t win on frequency,
I want to win on expressiveness.

A discrete machine may not chase GHz clocks—but it can pack more useful work into each cycle. Tomato’s Dual-LUT path turns compound Boolean and arithmetic into fewer ALU passes: more done per instruction, fewer steps for the same job. The examples below show where that pays off—and where base RV32I is more direct.

1 vs 1Addition · a tie

1 vs 2Signed comparison · RV32I advantage

8 vs 2Two-stage vote · Tomato advantage

RISC-V RV32I · instructionsTomato · ALU evaluations
Operation counts: ties, a RISC-V advantage, and compound-work advantages for Tomato.Vertical axis starts at zero. Each expression has one point per architecture. Exact sequences and counts appear in the expandable rows below.Operation count · lower is fewer01234567811223456681211111222AddSignedless-thanMask+ addMask− subtractChooseMajorityXOR3+ AND3TwoselectionsVote, mask+ addTwo-stagevote

Expressions run left to right. Lines connect discrete examples; this is not a scaling or timing curve.

Expression · expand for the stepsRV32I / Tomato
01Add1/1
A + B

RV32I · 1 instructionADD out, A, B

Tomato · 1 ALU evaluation1. (A, B, C) → out F=0xAA; G=0xCC; carry=0

A tie: ordinary addition needs one step on each datapath.

02Signed less-than1/2
out = signed(A) < signed(B) ? 1 : 0

RV32I · 1 instructionSLT out, A, B

Tomato · 2 ALU evaluations1. (A, B, C) → discard F=0xAA; G=0x33; carry=1; latch flags 2. (A, B, C) → out F=0x00; G=0x00; carry=LT

RV32I wins this sequence: SLT writes a Boolean directly. Tomato first subtracts and latches LT, then computes 0 + 0 + LT. This is a modeled flag-based route; it is not a claim that every possible Tomato implementation needs two steps.

03Mask + add2/1
A + (B AND C)

RV32I · 2 instructionsAND t0, B, C ADD out, A, t0

Tomato · 1 ALU evaluation1. (A, B, C) → out F=0xAA; G=0xC0; carry=0

Tomato forms the mask and sum together.

04Mask + subtract2/1
A − (B AND C)

RV32I · 2 instructionsAND t0, B, C SUB out, A, t0

Tomato · 1 ALU evaluation1. (A, B, C) → out F=0xAA; G=0x3F; carry=1

Complement the masked value and add carry-in 1 for subtraction.

05Choose3/1
Choose(A,B,C) = (A AND B) OR (NOT A AND C)

RV32I · 3 instructionsXOR t0, B, C AND t0, A, t0 XOR out, C, t0

Tomato · 1 ALU evaluation1. (A, B, C) → out F=0xD8; G=0x00; carry=0

A selects B or C independently at each bit. One Tomato truth table implements the selection.

06Majority4/1
Maj(A,B,C) = (A AND B) OR (C AND (A XOR B))

RV32I · 4 instructionsAND t1, A, B XOR t2, A, B AND t2, C, t2 OR out, t1, t2

Tomato · 1 ALU evaluation1. (A, B, C) → out F=0xE8; G=0x00; carry=0

Each output bit is the majority of the three input bits. Choose and Majority are SHA-256 building blocks, not complete hashes.

07XOR3 + AND35/1
(A XOR B XOR C) + (A AND B AND C)

RV32I · 5 instructionsXOR t1, A, B XOR t1, t1, C AND t2, A, B AND t2, t2, C ADD out, t1, t2

Tomato · 1 ALU evaluation1. (A, B, C) → out F=0x96; G=0x80; carry=0

Two different Boolean results feed the same adder.

08Two selections6/2
t = Choose(A,B,C); out = Choose(D,t,E)

RV32I · 6 instructionsXOR t0, B, C AND t0, A, t0 XOR t, C, t0 XOR t0, t, E AND t0, D, t0 XOR out, E, t0

Tomato · 2 ALU evaluations1. (A, B, C) → t F=0xD8; G=0x00; carry=0 2. (D, t, E) → out F=0xD8; G=0x00; carry=0

Two layers of bitwise selection: first A selects B/C, then D selects that result/E. The intermediate t connects the two steps.

09Vote, mask + add6/2
t = Maj(A,B,C); out = D + (t AND E)

RV32I · 6 instructionsAND t1, A, B XOR t2, A, B AND t2, C, t2 OR t, t1, t2 AND t0, t, E ADD out, D, t0

Tomato · 2 ALU evaluations1. (A, B, C) → t F=0xE8; G=0x00; carry=0 2. (D, t, E) → out F=0xAA; G=0xC0; carry=0

Compute a bitwise vote, keep selected bits, then add a bias D. Tomato fuses the mask and addition in its second operation.

10Two-stage vote8/2
t = Maj(A,B,C); out = Maj(t,D,E)

RV32I · 8 instructionsAND t1, A, B XOR t2, A, B AND t2, C, t2 OR t, t1, t2 AND t1, t, D XOR t2, t, D AND t2, E, t2 OR out, t1, t2

Tomato · 2 ALU evaluations1. (A, B, C) → t F=0xE8; G=0x00; carry=0 2. (t, D, E) → out F=0xE8; G=0x00; carry=0

A hierarchical bitwise vote. This is NOT majority-of-five: the first three values vote as one group before meeting D and E.

Counts are for the sequences shown—not proven minima, CPU cycles, measured speed, or compiled benchmarks. Inputs and temporary registers are assumed available; arithmetic wraps to 32 bits. Tomato rows model the Dual-LUT and flag latch, not guaranteed installed instruction-ROM entries. Register-bank setup and instruction mapping are excluded. Base RV32I excludes extensions; adding extensions can change these comparisons.

A small number.
Ready for a
32-bit machine.

To say “add 5,” I need a way to put the number 5 directly into an instruction. That embedded number is called an immediate.

I gave Tomato an immediate box with 16 decode choices. It can read short constants, preserve negative numbers, position upper bits, or scale a branch offset to a word address.

The result is always a 32-bit value the rest of the machine can use. A small instruction field becomes an operand, an address offset, or the upper part of a larger number.

Immediate decode and write-back

A short value, made full width.

Four examples from the immediate box.

INSTRUCTION FIELD0xFF8 bits
Fill the upper 24 bits with zeros
32-BIT OPERAND0x000000FF
00000000000000000000000011111111

Unsigned 255 stays 255. The extra bits are zero-filled.

The examples use selectors 0, 2, 3, and 11 in the instruction register.

One instruction.
A path through the machine.

I built the CPU around a simple loop: read an instruction, set up the circuits, do the work, and store the result.

  1. 01 / FETCH

    Read the instruction

    The program counter points to the next 32-bit word in memory.

  2. 02 / DECODE

    Set the controls

    A 9-bit opcode selects a row from the 512-row instruction ROM.

  3. 03 / EXECUTE

    Do the calculation

    Three register read ports supply A, B, and C. The immediate box supplies constants when needed.

  4. 04 / WRITE BACK

    Keep the result

    The write port stores the answer. A branch can change which instruction comes next.

What fits in a 32-bit instruction?

A typical register-to-register ALU instruction.

Other instruction formats reuse fields for immediates and branch controls. These stages describe the data flow, not a four-stage pipeline.Explore the instruction set

You can follow
the whole thing.

From the circuit design to the software on screen. The working prototype and the physical build each have a place in the story.

RUNNING

A complete FPGA computer

My custom CPU boots Tomato OS and drives an HDMI display. The FPGA faithfully implements the architecture, giving me a working system for developing and testing software.

Software and running system
ON THE BENCH

Discrete logic hardware

I’ve soldered the Lot 07 ALU board. The register, memory, control, and other boards continue in parallel—bringing the same architecture into discrete logic.

See the physical ALU status
WHAT’S NEXT

More software. More silicon.

The faithful FPGA gives me a working computer to build on. My next focus is broader instruction mappings, compiler support for compound operations, and the tools to make them useful. Alongside that, I’m completing the remaining discrete boards—growing the software ecosystem and the physical computer together.

Instruction set and mappings
Recorded run + reproducible harness

130 billion test vectors.
2 hours 45 minutes of simulation.

I recorded a 130-billion-vector Verilator run through the 32-bit ALU, checking each result against a reference model. That run took 2 hours 45 minutes. The repository includes the reproducible harness, while the full run log is not checked in. I also used formal verification to check the ALU’s mathematical behavior at 1, 8, and 32 bits.

These are ALU checks. They do not certify the entire computer or its physical timing. The verification page explains the scope and reproducible commands.

Inspect the evidence · Verilator + SymbiYosys

A smaller board.
A more flexible building block.

The predecessor gave me a fixed menu of operations. Tomato combines programmable Boolean logic with arithmetic—and fits its 8-bit slice into about one-seventh the board area.

PCB area7.3× smaller
Predecessor72,900 mm²
Tomato≈9,975 mm²

Same 8-bit slice width. 270 × 270 mm → 99.95 × 99.80 mm; 86.3% less board area.

Input operands1.5× as many
Predecessor2 · A, B
Tomato3 · A, B, C

A third input lets the truth tables combine three values before arithmetic.

Result flags1.75× as many
Predecessor4 flags
Tomato7 flags

Tomato exposes Z, N, C, V, LT, GT and GTE. These describe the result; they are not extra operations.

Operation control spaceFixed → configurable
Predecessor19 fixed operations
Tomato524,288 settings

Different counts: a fixed operation menu versus configurable controls. Multiple settings can implement the same function; this is not a speed comparison.

Board area per bit7.3× denser layout
Predecessor≈9,112 mm² / bit
Tomato≈1,247 mm² / bit

Area divided by the same 8-bit width. This measures PCB packing, not transistor density.

Simulation cases130 billion checked
Predecessor1.24M+ vectors
Tomato130B cases

About 105,000× the rounded predecessor count. Different test harnesses and coverage; more cases alone do not establish correctness.

Compare the implementation, timing, power, and verification
MeasurePredecessorTomato
Building blockMonolithic 8-bit ALU8-bit Dual-LUT slice; two slices make 16 bits, four make 32
Logic implementation3,488 transistors: 624 discrete MOSFETs + 2,864 inside 74HC ICs25 ICs per 8-bit slice; 74ACT151 truth tables and 74ACT283 adders
Arithmetic pathSeparate arithmetic and logic paths; output selects the resultTwo Boolean planes feed the same adder; no final arithmetic/logic selection mux
Control space19 fixed operations256 × 256 × 8 = 524,288 configurations, ≈27,600× the menu count; not that many distinct functions
Timing record≈445 ns worst-case path; ≈16.5 MHz reported for a separate carry-bypass pathCombinational 74ACT logic and carry-select; no like-for-like measured speedup claimed
Power5 V; documented 0.5–1 A range5 V ACT/HCT logic; comparable measured current is not recorded here
Verification1.24M+ automated vectors10B cases at 8 bits; 130B at 32 bits in 2 h 45 min; 476 directed vectors and 5/5 separate formal jobs
Build recordPredecessor design, set aside before assemblyFabricated and soldered in August 2026; DRC-clean board design

Transistors and packaged ICs are different units. Delay paths, widths, and verification scopes differ; these figures do not establish a power or performance ratio.

Tyrone Marhguy, the designer of Tomato
TYRONE MARHGUY · GHANA → PHILADELPHIA

A little about me

First principles.
Then better primitives.

I’m Tyrone Marhguy, a Computer Engineering student at the University of Pennsylvania, Class of 2028.

I start from first principles, reduce the problem to the right primitive, and iterate until the design is both correct and efficient. Tomato grew from that process: rethink the ALU, build a processor around it, then write the software that puts it to work.

The 8-bit ALU was my starting point. The Dual-LUT is the next iteration—less board area, a more flexible primitive, and a path to a complete computer.

Still wondering?

The questions people ask first—answered briefly here.

Why two LUTs instead of one?

A second LUT costs two extra packages per nibble, but gives the adder an independent Boolean plane. That is the trade for compound work like XOR, NAND, and arithmetic in one trip. More

Why three inputs—not two or four?

Two felt too close to a 74181. Four blows up the lookup hardware. Three keeps each truth table at eight entries while still fitting A + (B AND C) in one evaluation. More

What actually happens in the ALU?

Inputs feed two lookup planes; their outputs go straight into the adder: f(A,B,C) + g(A,B,C) + cin. There is no final mux choosing between separate logic and arithmetic results. More

Are there really 524,288 different operations?

256 × 256 × 8 carry sources = 524,288 control settings. Many implement the same function. That is control space—not 524,288 unique instructions in the ROM. More

What is the hardware compiler?

A counter FSM tests one Dual-LUT configuration per CPU clock. On the default 6.25 MHz FPGA CPU, a 65,536-configuration sweep takes at most about 10.5 ms before it freezes on a match or finishes. A match is still a candidate to verify. More

Has the entire CPU been formally verified?

Not yet. The ALU has simulation and formal checks. The whole CPU has Verilog simulation and software running on FPGA—different evidence. ALU proofs are not a proof of the complete computer. More

All questions on the FAQ page

Go as deep as you like.