We are three physicists. Inference optimization turned out to be mostly back-of-the-envelope estimates and error bars, so we felt right at home.
Team
How we got here
We come from places where efficient code was not optional (think academia and broke students in dorm rooms). We started with a few ThinkPads and zero GPUs, which teaches you fast what parts of a program are completely useless. Later we got lucky and got our hands on a 5090. Then we physically built and optimized a 4x RTX PRO 6000 cluster for our first client.
Built along the way
Image detection software that had to run in your browser (in WebAssembly).
Quantum physics informed neural nets, simulated on laptops, solving PDEs like Navier-Stokes.
Local LLM inference at 350 tok/s with high retention on coding/SWE tasks.
Local finetuning for autoformalization (turning written math into Lean) with adapters and RL.
Generative LEGO pipelines based on LLMs and diffusion models.