← Research ReadingBatch 1 is a different machine
Batch 1 is a different machine
Model inference but for a single user, on a local machine. You lose batching at scale and the property that all MoE experts activate at large batch size. For privacy. Find the exact nature of computation for a single thread of inference and spec out what the ideal dedicated hardware would look like. Goal: run 100B-500B models on a handheld.