Tylo
Tylo is an experimental library for writing NVIDIA GPU kernels in Julia. It provides logical memory views, distributed register values, and specific copy and matrix instructions that connect them. PTX.jl supplies instruction bindings; the kernel author controls the launch, work assignment, allocation and pipeline.
The central question is which values live where, who holds them, and when may the next operation use them? Tylo makes these different facts explicit. The implementation has complete GEMM and streaming-attention examples, but a much smaller operation set than CuTe or ThunderKittens.
A reading path
- Design and boundaries: the model, the decisions behind it, and the tradeoffs that remain open.
- Layouts and register fragments: follow a logical coordinate into memory or into a thread's registers. These pages include examples you can run without a GPU.
- Walk through GEMM: see the pieces in one complete kernel, including who waits and who releases a shared buffer.
- Read the implementation: source files, dispatch paths, generated code, and the tests that establish their contracts.
- Current status: distinguish implemented behavior, assembly evidence, runtime coverage, and unfinished work.
Then choose a hardware path: TMA/WGMMA, TMEM, streaming attention, or boundary tiles. API reference collects docstrings separately from the explanation.
A fragment is one thread's share
The register distribution of a warp matrix-multiply accumulator can be inspected on the CPU. Here 32 threads would jointly own a 16×8 logical tile, with four FP32 values per thread:
julia> using Tylo
julia> atom = MMAAtom((16,8,16),Tylo.BFloat16);
julia> f = Fragment(zero_accumulator(atom));
julia> size(Tylo.Layouts.layout(f)), length(f.data)
((16, 8), 4)
julia> [Tylo.Layouts.coordinate(Tylo.Layouts.layout(f), 0, Val(i)) for i in 0:3]
4-element Vector{Tuple{Int64, Int64}}:
(0, 0)
(0, 1)
(8, 0)
(8, 1)
julia> (2f0 .* (f .+ 1f0)).data
(2.0f0, 2.0f0, 2.0f0, 2.0f0)The tuple is local payload; its four entries are not a four-element logical matrix. Elementwise arithmetic works locally. A logical reduction may need values from other lanes and therefore a GPU collective. A memory layout is a separate map: these logical coordinates do not say where a global or shared matrix is stored.
What is usable now?
The warp-MMA GEMM, fragment arithmetic, TMA copies and complete BF16 streaming attention have GB10 runtime coverage. Hopper WGMMA and datacenter Blackwell TMEM paths assemble, with hardware tests prepared but not yet executed on their required devices. Tylo does not yet provide tcgen05 MMA, arbitrary register redistribution, automatic allocation or a uniform array interface for every representation. See the capability table.
Run and inspect
From this checkout, host tests need no GPU:
julia --project=. -e 'using Pkg; Pkg.instantiate(); Pkg.test()'The GPU environment expects the jool registry (which provides PTX) and Julia 1.10 or later:
julia --project=. -e 'using Pkg; Pkg.instantiate(; workspace=true)'
julia --project=test test/runtests.jl --jobs=4
julia --project=test examples/gemm/run.jl 65 97 73Current local checks use Julia 1.13. Compatibility is declared from Julia 1.10; the full suite also passes on 1.10.12 and 1.11.9 (Julia 1.10 unrolls the FlashAttention K loop only partially, which the comparison accounts for). For the datacenter attention comparison, assembly artifacts, and documentation build instructions, see Current status and validation.