Skip to content
View alesha-pro's full-sized avatar
❤️
I love my life
❤️
I love my life

Highlights

  • Pro

Block or report alesha-pro

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
alesha-pro/README.md
alexey@rig

X / @superalesha · alesha.pro · t.me/fuckup_files

About

By day I run infrastructure. The rest of the time I bench LLMs on a hand-built 4x RTX 3090 rig with 96 GB of VRAM, and I publish every measurement with raw traces and per-run provenance attached. People argue about quants and reasoning effort from vibes. I run the thing and count.

Featured work

Project Stars What it is
atlas stars Interactive canvas that takes an LLM apart tensor by tensor. It measures INT8/INT4/FP8 error, distributions, spectra and outlier channels on every weight tensor
evalatro stars The open benchmark where language models play real Balatro
4x3090-llm-benchmarks stars Public archive of local LLM speed measurements from the rig, with exact launch commands and per-run provenance
vllm-cross-model-kv stars Cross-model KV cache transfer in vLLM: a smaller model prefills, a bigger one answers from its cache
qwen38-27b-bench-4x3090 stars Raw traces from the 67-hour Qwen3.8-27B quant grinder, GGUF sweeps plus a 288-point vLLM vs SGLang matrix
tools stars ComfyUI workflows and benchmark configs from the rig

The rig

The rig is 4x RTX 3090 with 96 GB of VRAM in total. P2P is verified on all 12 pairs, the cards run at 220 W each on PCIe 3.0 x16. Engines on the box: vLLM, llama.cpp, exllamav3 + tabbyAPI, plus SGLang when it behaves. Ampere sm_86 has no native FP8, and most quant advice out there assumes everyone owns a Hopper. I enjoy measuring what actually wins on these cards.

Highlights from X

Numbers

GitHub streak GitHub stats


Every number on this profile comes from a real run. If it is not in a repo with raw data, I did not say it.

Pinned Loading

  1. atlas atlas Public

    Interactive canvas for taking an LLM apart tensor by tensor: measured INT8/INT4/FP8 error, distributions, spectra and outlier channels for every weight tensor

    TypeScript 66 7

  2. 4x3090-llm-benchmarks 4x3090-llm-benchmarks Public

    Local LLM speed measurements from a 4x RTX 3090 rig, with exact launch commands and per-run provenance.

    Python 17 2

  3. tools tools Public

    Tools, ComfyUI workflows and benchmark configs from a 4x RTX 3090 local-inference rig

    Python 22 1

  4. evalatro evalatro Public

    🃏 The open benchmark where language models play real Balatro.

    TypeScript 26 4

  5. vllm-cross-model-kv vllm-cross-model-kv Public

    Cross-model KV cache transfer in vLLM: a smaller model prefills, a bigger one answers from its cache (independent implementation of arXiv:2608.03893, measured on 2x RTX 3090)

    Python 14 2

  6. qwen38-27b-bench-4x3090 qwen38-27b-bench-4x3090 Public

    Raw speed benchmark traces: Qwen3.8-27B on 4x RTX 3090. llama.cpp GGUF sweeps + 288-point vLLM vs SGLang, NVFP4 vs FP8 matrix.

    Python 11 1