Hi, I'm Jędrzej Maczan
I do artificial intelligence research
GitHub
|
Google Scholar
|
X
|
ORCID
publications
Paper
2026-07-31
Measuring and Reducing WebGPU Dispatch Overhead for LLM Inference
The Fourth UK AI Conference 2026
Nottingham, UK
OpenReview
Also presented in
COLM 2026 MOSS
and
COLM 2026 Efficient Reasoning
Paper
Poster
2026-07-11
Memory-Efficient Activation Checkpointing with Sliding Window and Hirschberg's Algorithm for 0/1 Knapsack Solving in PyTorch
Fast Machine Learning for Science Conference 2026
University of California, San Diego, USA
Indico CERN
·
OpenReview
Also presented in
COLM 2026 Efficient Reasoning
Paper
2026-02-09
Characterizing WebGPU Dispatch Overhead for LLM Inference Across Four GPU Vendors, Three Backends, and Three Browsers
arXiv
Open source
workshops / talks
Paper
2026-10-06
Memory-Efficient Activation Checkpointing with Sliding Window and Hirschberg's Algorithm for 0/1 Knapsack Solving in PyTorch
COLM 2026 Workshop on Efficient Reasoning
San Francisco, USA
OpenReview
Also presented in
Fast ML 2026
Paper
2026-10-06
"As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It
COLM 2026 Workshop on Efficient Reasoning
San Francisco, USA
OpenReview
Also presented in
KONVENS 2026 Eval4SD
Paper
2026-10-06
Dispatch Overhead, Not Kernel Quality: The Performance Bottleneck in Single-Stream WebGPU LLM Inference
COLM 2026 Workshop on Methods and Opportunities at Small Scale (MOSS)
San Francisco, USA
OpenReview
Also presented in
COLM 2026 Efficient Reasoning
and
UK AI 2026
Paper
2026-10-06
Dispatch Overhead, Not Kernel Quality: The Performance Bottleneck in Single-Stream WebGPU LLM Inference
COLM 2026 Workshop on Efficient Reasoning
San Francisco, USA
OpenReview
Also presented in
COLM 2026 MOSS
and
UK AI 2026
Poster
2026-09-14
"As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It
KONVENS 2026 First Workshop on Evaluating LLMs for Specialized Domains (Eval4SD)
University of Hamburg, Germany
OpenReview
Also presented in
COLM 2026 Efficient Reasoning
Talk
2026-06-05
Online Softmax
Cohere Labs - Open Science ML Math Community
articles
2026-07-27
The cuBLAS transposition trick
Paged Out!, Issue #9, page 14
Open source
2026-05-29
From Boltzmann-Gibbs distribution to numerically stable online softmax on GPU with proofs
2026-02-19
Solving 0/1 Knapsack problem with sliding window and Hirschberg algorithm
Paged Out!, Issue #8, page 8
Open source
2025-11-21
Use 20x less peak RAM with dp_knapsack_sliding_hirschberg, a new activation memory budget solver for PyTorch
Open source
2025-10-01
Self-contained handwritten digit recognizer
Paged Out!, Issue #7, page 11
Open source
2025-03-01
A primer on Differentiable Architecture Search
Paged Out!, Issue #6, page 5
Open source
2024-12-01
GPT in PyTorch
Paged Out!, Issue #5, page 6
Open source
2024-06-01
Building automated machine learning with type inference
Paged Out!, Issue #4, page 4
Open source
[In progress]
tiny-vllm: Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
GitHub
Open source
about me
I like computers!
I want to be a full time research scientist and work on the intersection of math, AI and low-level systems
I worked as software engineer for 7 years and then as AI engineer for 2 years
My name is spelled as yenjay in English and 延杰 in Chinese
Email: jedrzej@maczan.pl
I'm doing Foundations of Probability and Statistics Specialization at University of Colorado Boulder, to get admitted to Master's program in AI
I got bachelor degree in computer science from Wrocław University of Science and Technology
Art of Anita:
anitamaczan.pl
, instagram:
@anita_maczan
Other links (not very active):
Mastodon:
mathstodon.xyz/@jmaczan
LinkedIn:
click
X/Twitter:
@jedmaczan
HackTheBox:
@7563687575
(inactive,
had
fun
some time ago)
. . ...