01 · Exact surrogate adjoint
One reference tile
For the reported two-step refinement path, the four staircase plan factors can be reconstructed from one resident reference tile and explicit row and column modifiers.
Optimal transport / efficient attention / ML systems
I work on learning systems that have to survive both theoretical scrutiny and real machines: differentiable transport, long-context attention, sparse derivatives, and implementations designed around the memory hierarchy.
Featured paper · arXiv:2605.08123 · April 2026
Tail-refinement gradients with a gap-aware dustbin bridge: a memory-efficient route to differentiating balanced entropic optimal-transport attention at long context.
Block-wise transport attention
The active tile moves through banded support while the rest of the plan stays out of memory.
01 · Exact surrogate adjoint
For the reported two-step refinement path, the four staircase plan factors can be reconstructed from one resident reference tile and explicit row and column modifiers.
02 · Hardware-aware schedule
Fixed-width block-wise execution yields O((T+R)LW) work, O(Ld) input storage, and O(L) additional HBM for fixed head dimension and band width.
03 · Structured gaps
The implemented gap-aware path is formalized as the same balanced surrogate on an augmented support, allowing the adjoint schedule to lift without claiming a general unbalanced model.
Publications
This index is designed to grow as new papers are released, with each paper receiving its own explanation, artifacts, and evidence boundary.
Earlier research threads
Reactive exploration methods for measuring domain-shift magnitude across controlled reinforcement-learning environments.
Research across biomarker discovery, protein representations, regulatory networks, and learning from heterogeneous biological signals.
Patch tokenization, self-attention, and inter-slice aggregation for volume-aware diagnostic classification.