Research report ·

TLDR

    Key insights

    The spec

    THE DELIVERABLE

    A dedicated batch-1 decode engine. Roughly 15–30 TOPS INT4-equivalent, a three-tier memory hierarchy with no off-package HBM, and a 5W envelope — targeting 8–15 tok/s on a 100–500B-total heavily-sparse model.

    Memory hierarchy

    TIER 0
    SRAM on-die32–64 MB · >10 TB/s Shared experts, router, draft model. Never leaves the die.
    TIER 1
    LPDDR6 on-package24–48 GB · 150–250 GB/s · 4–6 pJ/bit Active working set + KV / SSM state. This tier sets the token rate.
    TIER 2
    NAND cold store256 GB – 1 TB · ~4 GB/s Model residency only. Disqualified from the decode loop — 40–80× too slow.

    Power budget — 5W

    The arithmetic

    Hover a row.
    ScenarioGB / token@100 GB/s@273 GB/s @614 GB/s@3 TB/s@5W energy-limited
    CROSSOVER
    P* = Bandwidth × J/byte
    Independent of model size. At 100 GB/s phone memory with a system-level 0.16 nJ/byte, P* ≈ 16W — below which the device is energy-bound, above which it is bandwidth-bound. A 5W handheld sits firmly on the energy side. The 5W column is the binding one.

    Citation graph

    Drag to pan · scroll to zoom · double-click a node to focus

    Timeline

    Drag to orbit · X time · Y citations · Z angle lane

    Next steps

      Idea deck

      Report quality

      Citations