Posts by Category


🚀 I’m always open to collaborate, exchange ideas or just talk about anything!

👨🏻‍💻 I’m eager to work with anyone who has great ideas, wants to learn more and more and also share their experience to others. Don’t hesitate to write me if you’d like to propose your help or ask for mine on a project, research, paper-idea, or a moonshot you’re cooking up.

👉 Email Me ✉️


basics

Output and Gated Activations: Softmax, Sparsemax, GLU, and SIREN

6 minute read

Published:

The last activation in your network is not a modelling preference, it is a contract with your loss function. Break it and training stops meaning anything. Here is the contract, and what changes when the activation itself becomes learned.

Activation Functions in Neural Networks: Why Non-Linearity Matters

7 minute read

Published:

Activation functions are the reason neural networks can model curved decision boundaries instead of collapsing into one giant linear map. This chapter builds the intuition first, then walks through the classical functions that shaped deep learning.

RNNs, LSTMs, and GRUs: Sequence Models Before Attention

6 minute read

Published:

A recurrent network shares weights across time exactly as a convolution shares them across space. The trouble is that gradients then travel through a product of Jacobians, and a product of a hundred numbers slightly below one is zero.

Convolutions and CNNs: Weight Sharing as a Prior

22 minute read

Published:

A dense layer from a 224-by-224 colour image to 1000 units holds 150.5 million weights; a 3-by-3, 64-filter convolution holds 1,792, and the two constraints that buy that factor of 84,000 are exactly the prior that makes it work on images.

PCA: Maximum Variance and Minimum Reconstruction Error

6 minute read

Published:

PCA can be derived by asking for the directions of greatest spread, or by asking for the subspace that loses the least when you project onto it. The two questions look unrelated and have the same answer, which is the most useful thing to understand about it.

Clustering: What Each Algorithm Assumes a Cluster Is

29 minute read

Published:

k-means says a cluster is a ball around a centroid, DBSCAN says it is a connected dense region, and a Gaussian mixture says it is a bump in a density, pick the algorithm and you have already picked the answer.

Support Vector Machines: Margins and the Kernel Trick

8 minute read

Published:

Among all the hyperplanes that separate two classes, one sits furthest from both. Finding it turns out to depend on the data only through inner products, and that single fact is what lets you work in a space you never build.

Trees, Forests, and Boosting: Axis-Aligned Everything

15 minute read

Published:

A decision tree chops feature space into axis-aligned boxes and predicts one number per box, which explains why it needs no feature scaling, why it approximates a diagonal boundary as a staircase, and why it cannot extrapolate a single step beyond the training range.

Logistic Regression: Linear in the Log-Odds

19 minute read

Published:

Logistic regression is not a squashed linear regression, it is a straight line drawn in log-odds space, which is why one coefficient means one multiplication of the odds, and why perfectly separable data drives that coefficient to infinity.

Linear Regression: Least Squares as Projection

20 minute read

Published:

Fitting a line by least squares is not an optimisation trick, it is the orthogonal projection of the observation vector onto the column space of the design matrix, and once you see that, the normal equations, the failure modes, and the reason we solve by QR instead of inverting all follow from one picture.

Gradient Descent and Backpropagation: How a Model Learns

27 minute read

Published:

Training is one loop: measure the loss, ask backpropagation which way is downhill, take a small step. This chapter derives exactly how small that step has to be, why the answer is 2/a for a quadratic, and why reverse-mode differentiation gets you every gradient for roughly the price of one forward pass.

Machine Learning Before Transformers: A Working Foundation

4 minute read

Published:

Every later book on this site assumes you already know what a loss is, why gradient descent works, and what a convolution buys you. This book supplies that, and follows one thread through it: how much structure you build in versus how much you let the data decide.

cs-basics

Recursion and Dynamic Programming: Two Conditions, One Filled Table

7 minute read

Published:

Dynamic programming is not a trick, it is a diagnosis: if a problem has optimal substructure and overlapping subproblems, exhaustive recursion is doing the same work exponentially often and a table fixes it. Here is the diagnosis, and edit distance worked out cell by cell.

diffusion

Diffusion vs Flow Matching: Two Names for One Family

5 minute read

Published:

Flow matching is often presented as the successor to diffusion. It is more accurate, and more useful, to say that diffusion is one particular probability path inside the flow-matching framework, and not the straightest one available.

Flow Matching: Training a Velocity Field Without Ever Solving an ODE

7 minute read

Published:

Continuous normalising flows were elegant and nearly untrainable, every gradient step needed an ODE solve and a divergence estimate. Flow matching removes both by regressing a velocity field against a target you can write down in closed form, one example at a time.

Distillation and Consistency Models: Getting to Four Steps

6 minute read

Published:

Better ODE solvers bottom out around ten evaluations because the trajectory is genuinely curved. To go lower you have to change the model, either teach a student to take two teacher steps at once, or train a network that jumps to the end of the trajectory from anywhere on it.

Latent Diffusion: Denoise Where the Information Is

6 minute read

Published:

Most of the bits in a photograph encode texture no one can see. Latent diffusion throws them away first with an autoencoder, then runs the entire diffusion process in a space roughly forty-eight times smaller, which is how Stable Diffusion fits on a consumer GPU.

DDIM: Same Marginals, Fewer Steps, and a Latent Space Worth Having

5 minute read

Published:

DDPM’s objective never actually required the forward process to be Markov, only that its marginals be Gaussian. Dropping the Markov assumption exposes a whole family of samplers a trained model already supports, including a deterministic one that runs in 20 steps and gives an invertible latent space.

Score Matching and the SDE View: DDPM as One Discretisation Among Many

5 minute read

Published:

Noise prediction and score estimation are the same network in different units. Taking the step size to zero turns the whole method into a stochastic differential equation, and reveals a deterministic ODE with identical marginals hiding inside it.

DDPM Training: From a Variational Bound to Four Lines of PyTorch

5 minute read

Published:

The DDPM objective starts as a variational bound with T+1 KL terms and ends as a plain mean-squared error on noise. Following the collapse shows why the discarded weighting term is not an approximation you tolerate but a reweighting that improves samples.

The Forward Process: A Corruption Engineered to Be Jumped Into

6 minute read

Published:

The forward process looks like the trivial half of diffusion, but every term in it is load-bearing. Drop the shrink factor and the variance diverges; pick the wrong schedule and a third of your timesteps are spent denoising static.

Diffusion Models: Learning to Undo Noise

4 minute read

Published:

Destroying an image is easy and needs no learning at all. Diffusion models exploit that asymmetry: they define a trivial forward corruption, then train a network to walk it backwards one small step at a time.

gdl

Gauges: When There Is No Shared Frame

7 minute read

Published:

The previous chapter ended on an arbitrary choice that could not be eliminated. Gauge theory’s answer is to stop trying: keep every local frame, transport between them explicitly, and require the model to be indifferent to which frames you picked.

Geodesics: Learning on Curved Domains

6 minute read

Published:

On a surface there is no global grid to slide a filter along, and no canonical direction to call ‘up’. What survives is distance, and building convolution out of distance alone exposes exactly one ambiguity, which is where the next chapter starts.

Groups: Equivariance Beyond Translation

6 minute read

Published:

Translation is one group. Swap it for rotations, reflections, or the rigid motions of 3-D space and the same construction produces a different architecture, with the parameter count cut by exactly the size of the orbit.

Grids: Why Translation Equivariance Forces Convolution

5 minute read

Published:

Convolution is not a clever idea someone had about images. It is the only linear map that commutes with translation, a theorem, not a design choice, and one you can verify by exhaustion on a small enough case.

The Hardware Lottery: Why the Dense Version Won

6 minute read

Published:

Message passing on a sparse graph does asymptotically less work than attention over every pair. It is still the slower one to train. This chapter is about why the architecture that wins is the one your hardware happens to like.

Transformers Are GNNs on Fully Connected Graphs

6 minute read

Published:

Write self-attention in the message-passing template and you get the Graph Attention Network equations with one substitution: the neighbourhood becomes the whole input. The two architectures are not analogous, they are the same operator on different graphs.

Permutation Symmetry and the Message-Passing Blueprint

7 minute read

Published:

A graph’s nodes have no canonical order, so any model that reads one must give the same answer under relabelling. That single requirement forces the three-step message, aggregate, update template, it is not a design choice.

Geometric Deep Learning: One Blueprint Behind Every Architecture

6 minute read

Published:

CNNs, GNNs, Transformers and sheaf models look like separate inventions. They are the same recipe applied to different domains: identify the symmetry of your data, then build layers that respect it. This book is that recipe, and the arguments that follow from it.

geometry-basics

Symmetry and Groups: Invariance, Equivariance, and Why You Build It In

6 minute read

Published:

If rotating a molecule cannot change its energy, that is a fact about the target function you know before training starts. Encoding it in the architecture makes it true everywhere; learning it from augmented data makes it approximately true where you happened to have samples.

Curvature: Why a Flat Map of the Earth Must Lie

6 minute read

Published:

Curvature at a point is one over the radius of the circle that best hugs the curve there. Push that idea up to surfaces and Gauss’s Theorema Egregium falls out: some curvature is visible from inside the surface, which is why no map projection can ever get distances right.

Linear Maps as Geometry: Rotate, Scale, Rotate

6 minute read

Published:

A matrix is not a table of numbers, it is a deformation of space. The SVD says every deformation is the same three moves in sequence: rotate, stretch along axes, rotate again.

Euclidean Space: Inner Products, Projections and the Margin

6 minute read

Published:

One bilinear form generates the whole of flat geometry: lengths, angles, orthogonality, projections and the distance from a point to a hyperplane. Get the projection formula and you get the SVM margin for free.

Geometry for Machine Learning: Why Shape Keeps Coming Back

5 minute read

Published:

Every embedding you have ever trained lives in a metric space, every dataset you have ever fitted sits near a surface far thinner than its ambient dimension, and every architecture you trust encodes a symmetry. Geometry is not decoration on top of ML, it is what makes the problems tractable.

gnn

GNNs for Computer Vision: Scene Graphs and Beyond

8 minute read

Published:

Computer vision tasks increasingly require relational reasoning, understanding how objects relate to each other, not just what they are. Scene graph generation, visual question answering, action recognition from skeletons, and 3D point cloud processing all benefit from GNN-based relational modelling.

GNNs for Robotics: Planning, Manipulation, and Multi-Agent Systems

8 minute read

Published:

Robots interact with structured environments: objects have relationships, joints form kinematic chains, agents communicate through interaction graphs. GNNs encode these relational structures, enabling generalisation across object configurations, robot morphologies, and multi-agent scenarios.

GNNs for Knowledge Graphs: Reasoning and Completion

8 minute read

Published:

Knowledge graphs encode human knowledge as typed entity-relation triples. GNNs enable structure-aware entity representation, multi-hop reasoning, knowledge base completion, and entity alignment, tasks that shallow embedding methods cannot fully solve.

GNNs for Traffic Forecasting

8 minute read

Published:

Traffic prediction is a canonical spatio-temporal graph task: sensors on roads form a fixed graph, and speed/volume measurements evolve over time. GNNs capture spatial correlations between sensors; RNNs or convolutions capture temporal patterns. Together they achieve state-of-the-art traffic forecasting.

GNNs for Social Networks: Influence, Communities, and Misinformation

7 minute read

Published:

Social networks are large sparse graphs with rich node features (user profiles) and heterogeneous edges (friendship, follow, retweet). GNNs predict user behaviour, detect communities, identify influential spreaders, and flag misinformation, tasks with significant real-world impact.

GNNs for Recommender Systems

6 minute read

Published:

Recommendation is naturally a graph problem: users and items are nodes, interactions are edges. GNNs on bipartite user-item graphs capture higher-order collaborative filtering signals, friends of friends liked this, that matrix factorisation cannot represent.

GNNs for Molecules: Drug Discovery and Material Design

7 minute read

Published:

Graph neural networks are transforming computational drug discovery. Molecules are natural graphs, and GNNs learn molecular representations that predict toxicity, solubility, binding affinity, and synthesis feasibility, tasks that previously required expensive laboratory experiments.

Polynomial Neural Sheaf Diffusion

8 minute read

Published:

Polynomial Neural Sheaf Diffusion (PNSD) replaces the fixed diffusion operator (I - Δ_F) with a learnable polynomial of the Sheaf Laplacian. This gives the model spectral flexibility, it can learn to amplify or suppress different frequency components of the sheaf signal.

Equivariant Sheaf Neural Networks

9 minute read

Published:

Sheaves with orthogonal restriction maps define a connection on the graph, a parallel transport structure over edges. This connects sheaf GNNs to differential geometry and enables equivariant processing of data with local coordinate frames at each node.

Sheaf Neural Networks and Heterophily

9 minute read

Published:

Sheaf GNNs are the principled solution to heterophily: by learning per-edge maps that transform features before comparison, they can perform diffusion that converges within classes and diverges across classes, the exact opposite of standard GCN’s collapse.

Diagonal, Orthogonal, and General Sheaf Maps

8 minute read

Published:

The restriction maps in a cellular sheaf can be constrained to different matrix classes: scalars, diagonal matrices, orthogonal matrices, or general matrices. Each class offers a different trade-off between expressivity and computational cost.

Neural Sheaf Diffusion: Learning Sheaves End-to-End

10 minute read

Published:

Neural Sheaf Diffusion (Bodnar et al., 2022) learns the sheaf restriction maps from data using a neural network, then performs diffusion with the learned Sheaf Laplacian. This gives a principled, topology-grounded GNN that handles heterophily without heuristic fixes.

The Sheaf Laplacian: Spectral Theory for Sheaves

7 minute read

Published:

The Sheaf Laplacian generalises the graph Laplacian by incorporating per-edge restriction maps. Its spectrum reveals how consistent data is under the sheaf. Sheaf diffusion with this Laplacian generalises GCN to handle heterophilic graphs.

What Is a Sheaf? From Topology to Graph Learning

8 minute read

Published:

A sheaf is a mathematical object from algebraic topology that assigns vector spaces to cells and linear maps between them. On graphs, sheaves assign feature spaces to nodes and edges, with restriction maps encoding how node features relate across edges.

Why Message Passing Is Not Enough: The Case for Sheaves

7 minute read

Published:

Standard message passing aggregates neighbour features and averages. On heterophilic graphs (where neighbours often disagree), this is harmful. Cellular sheaves provide a mathematically principled framework to model per-edge relationships between node features, going beyond mere averaging.

Molecular GNNs: Learning on Atoms and Bonds

9 minute read

Published:

Molecules are graphs. Molecular GNNs predict chemical properties from structure. The best models use 3D coordinates and bond angles, not just connectivity.

Tensor Field Networks and Geometric Deep Learning

8 minute read

Published:

Tensor Field Networks (TFN) were the first architecture to achieve SE(3) equivariance using spherical harmonics and Clebsch-Gordan tensor products. They laid the theoretical foundation for NequIP and MACE, the current state-of-the-art in equivariant molecular force fields.

SE(3)-Transformers: Attention with 3D Symmetry

9 minute read

Published:

SE(3)-Transformers extend self-attention to 3D point clouds and molecular graphs while maintaining SE(3) equivariance. Attention weights are learned between node pairs; values are equivariant features built from spherical harmonics.

EGNN: E(n)-Equivariant Graph Neural Networks

9 minute read

Published:

EGNN achieves E(n)-equivariance with a simple update rule: positions updated via weighted sums of relative position vectors, features updated via invariant distances. No spherical harmonics needed.

Equivariance: What It Means and Why It Matters

9 minute read

Published:

Equivariance formalises the idea that a function should ‘commute with symmetry transformations.’ A rotation-equivariant model applied to rotated input gives the rotated output, no extra training needed. This is the foundation for geometric deep learning.

Why Geometry Matters in Graph Neural Networks

7 minute read

Published:

Many real-world graphs are embedded in 3D space, molecules, proteins, point clouds, crystal structures. Standard GNNs ignore coordinates and only use connectivity. Geometric GNNs incorporate spatial positions and must respect physical symmetries.

Spatio-Temporal GNNs: Learning on Graphs Through Time

8 minute read

Published:

Spatio-temporal GNNs combine spatial message passing with temporal sequence modelling. They are the dominant approach for traffic forecasting, weather prediction, and any task where measurements at sensor nodes evolve over time on a fixed graph.

Graph Neural ODEs: Continuous-Time Graph Dynamics

9 minute read

Published:

Neural ODEs replace discrete layer-by-layer computation with continuous dynamics governed by a differential equation. Graph Neural ODEs apply this to graph data, treating node embeddings as a dynamical system evolving in continuous time.

Temporal Graph Networks: Learning from Events

8 minute read

Published:

TGN (Temporal Graph Network) is the leading framework for continuous-time dynamic graphs. It maintains a per-node memory that is updated upon each interaction, enabling efficient inductive link prediction on event streams.

Static vs Dynamic Graphs: When Structure Changes Over Time

6 minute read

Published:

Most GNN research assumes a fixed graph. Real graphs evolve: edges appear and disappear, node features drift, new nodes arrive. Dynamic graph learning addresses how to model and predict on graphs whose structure changes over time.

Temporal Knowledge Graphs: Facts That Change Over Time

7 minute read

Published:

Most knowledge graphs treat facts as timeless, but facts change. Barack Obama was president from 2009 to 2017. Temporal Knowledge Graphs add timestamps to triples, requiring models to reason about what was true when.

Knowledge Graph Embeddings vs GNNs

11 minute read

Published:

Knowledge graph completion can be solved with shallow KG embeddings (TransE, DistMult, ComplEx) or with structural GNNs (R-GCN, CompGCN). Each approach has different inductive biases and failure modes. Understanding when to use each is the central design decision for KG tasks.

HAN: Heterogeneous Graph Attention Networks

6 minute read

Published:

HAN combines meta-path decomposition with two levels of attention: node-level attention weights neighbours along a meta-path, and semantic-level attention weights different meta-paths. This lets the model learn which relationships matter most for a given task.

R-GCN: Relational Graph Convolutional Networks

7 minute read

Published:

R-GCN extends GCN to multi-relational graphs by learning a separate weight matrix for each relation type. It handles knowledge graphs with typed edges and powers both entity classification and link prediction tasks.

Heterogeneous Graphs: When Nodes and Edges Have Types

5 minute read

Published:

Most real-world graphs are heterogeneous, they contain multiple node types (users, items, tags) and edge types (clicks, rates, authors). Standard GNNs treat all nodes and edges identically, making them blind to this type structure.

Graph Classification: From Node Embeddings to Graph Embeddings

7 minute read

Published:

Graph classification is the task of predicting a label for an entire graph. It requires composing message passing (node embeddings), readout (graph embedding), and a classifier, and all three choices interact to determine model expressiveness.

Set2Set and Attention Readout: Order-Invariant Graph Summaries

8 minute read

Published:

Mean and sum readout treat all nodes equally. Attention readout learns which nodes matter most for a given task. Set2Set goes further, it uses an LSTM to iteratively query the node set, producing richer graph representations than single-pass pooling.

TopKPool and SAGPool: Sparse Graph Pooling

8 minute read

Published:

Instead of soft cluster assignment (DiffPool), TopKPool and SAGPool select a subset of the most important nodes, producing a smaller but sparser graph at each level. Hard selection is scalable but requires careful score learning.

DiffPool: Learning Hierarchical Graph Pooling

8 minute read

Published:

DiffPool learns to hierarchically cluster nodes into super-nodes across layers, like a convolutional pyramid for graphs. Unlike flat global pooling, it captures multi-scale graph structure by differentiably assigning nodes to clusters.

Global Pooling in GNNs: Mean, Sum, and Max

6 minute read

Published:

To predict a property of an entire graph, node embeddings must be aggregated into a single vector. The choice of global pooling, mean, sum, or max, is not arbitrary: each has distinct expressive power and fits different tasks.

Sign Ambiguity in Laplacian Eigenvectors

9 minute read

Published:

Laplacian eigenvectors are only defined up to sign: if u is an eigenvector, so is -u. This seemingly minor issue creates a fundamental problem for learning with LapPE. Here is the problem, its consequences, and how SignNet solves it.

Structural vs Positional Encodings in Graphs

7 minute read

Published:

Positional encodings say where a node is in the graph. Structural encodings say what role it plays. They are complementary, and confusing them leads to poor design choices.

Shortest-Path Encodings for Graph Transformers

6 minute read

Published:

Shortest-path distances between nodes can be encoded as attention biases or node features, directly informing the model about graph proximity without requiring message passing.

Random Walk Positional Encodings

7 minute read

Published:

Random walk positional encodings encode each node’s structural context by computing the probability of returning to it from itself in k steps, a computationally efficient alternative to Laplacian eigenvectors with no sign ambiguity.

Laplacian Eigenvectors as Graph Positional Encodings

11 minute read

Published:

The k smallest eigenvectors of the graph Laplacian form a natural positional embedding space, the graph’s own coordinate system. They capture global structure, symmetry, and community membership.

Why GNNs Need Positional Encodings

7 minute read

Published:

Message-passing GNNs are permutation-equivariant by design, they cannot assign unique positions to nodes. Without positional encodings, symmetric nodes are indistinguishable. Here is why that matters and how to fix it.

Depth in GNNs: Why Deeper Is Not Always Better

8 minute read

Published:

In Transformers, depth = expressiveness. In GNNs, depth = both expressiveness AND over-smoothing. The optimal GNN depth is rarely more than 3-4 layers, fundamentally different from the hundreds of layers in modern LLMs.

Over-smoothing vs Over-squashing: The Difference

8 minute read

Published:

Oversmoothing and oversquashing are both problems with deep GNNs, but they affect different nodes, have different causes, and require different fixes. Confusing them leads to applying the wrong solution.

Oversmoothing: When All Node Embeddings Become the Same

9 minute read

Published:

Stack enough GNN layers and all node embeddings converge to the same vector, making the model useless. Oversmoothing is not a training problem; it is a mathematical inevitability of iterated averaging.

The Weisfeiler-Lehman Test: How Powerful Are GNNs?

11 minute read

Published:

The 1-WL graph isomorphism test provides the exact upper bound on message-passing GNN expressivity. GIN achieves this bound. Any pair of graphs that 1-WL cannot distinguish cannot be distinguished by any MPNN.

MPNN: The General Message Passing Neural Network Framework

6 minute read

Published:

The MPNN framework (Gilmer et al., 2017) unifies GCN, GAT, GIN, GraphSAGE, and almost all spatial GNNs under one abstraction: message functions, aggregation, and update. Understanding MPNN means understanding the whole GNN family.

Graphormer: Transformers with Structural Biases for Graphs

9 minute read

Published:

Graphormer encodes graph structure directly into Transformer attention via three biases: node centrality, spatial encoding (shortest paths), and edge encoding. It won the OGB-LSC 2021 competition on molecular property prediction.

Graph Transformers: Bringing Attention to Graphs

6 minute read

Published:

Graph Transformers replace or augment local message passing with full pairwise attention, every node attends to every other node. This solves long-range dependencies and over-squashing at the cost of O(N²) computation.

APPNP: Personalized PageRank Meets Graph Neural Networks

7 minute read

Published:

APPNP decouples feature transformation from propagation. A neural network transforms features first; then Personalized PageRank propagates the result. This enables deep propagation without over-smoothing.

SGC: Simple Graph Convolution

7 minute read

Published:

SGC removes all nonlinearities between GCN layers and collapses the entire propagation into a single pre-computed matrix power. Surprisingly, it matches GCN on most benchmarks, revealing that nonlinearities between layers may be unnecessary.

Graph Fourier Transform: The Spectral View of Graphs

8 minute read

Published:

The Graph Fourier Transform decomposes a signal on a graph into frequency components using the Laplacian’s eigenvectors. This spectral view is the mathematical foundation behind spectral GNNs like ChebNet and GCN.

Graph Tasks: Node, Edge, and Graph-Level Prediction

7 minute read

Published:

GNNs can predict at three levels: properties of individual nodes, existence or type of edges, or properties of entire graphs. Each level requires a different output head and training setup.

GIN: Graph Isomorphism Network, The Most Expressive GNN

5 minute read

Published:

How powerful can a GNN be? Xu et al. (2019) answered with a theoretical bound, and GIN is the architecture that achieves it. The secret: use sum aggregation and an MLP, not mean or max.

GraphSAGE: Inductive Learning on Large Graphs

4 minute read

Published:

GCN and GAT learn embeddings for fixed graphs, add a new node and you’re stuck. GraphSAGE (Hamilton et al., 2017) learns an aggregation function instead, so it can generate embeddings for entirely new nodes at inference time.

GAT: Graph Attention Networks

5 minute read

Published:

GCN assigns the same (degree-based) weight to every neighbour. GAT learns which neighbours actually matter, using attention coefficients on edges. More expressive, more interpretable.

GCN: Graph Convolutional Networks

5 minute read

Published:

GCN (Kipf & Welling, 2016) is the ‘hello world’ of GNNs. It simplifies spectral graph convolution into a single elegant layer: normalised neighbourhood averaging with a learned linear transformation.

Message Passing: The Universal GNN Framework

4 minute read

Published:

Every GNN, GCN, GAT, GraphSAGE, GIN, is a special case of message passing. Learn the three-step loop that defines them all: compute messages, aggregate, update.

The Graph Laplacian: Spectral Graph Theory Explained Simply

6 minute read

Published:

The Graph Laplacian is L = D - A. Its eigenvectors reveal the graph’s community structure; its eigenvalues tell you how well-connected the graph is. It’s also the mathematical bridge from spectral theory to GNNs like GCN.

The Graph Adjacency Matrix: A Graph in Matrix Form

4 minute read

Published:

Before understanding GNNs, you need to understand how graphs are represented mathematically. The adjacency matrix is the foundation, a simple grid that tells you which nodes are connected.

Graph Neural Networks: Learning on Graphs

5 minute read

Published:

Graphs are everywhere, molecules, social networks, road maps, knowledge bases. Graph Neural Networks learn from this relational structure by propagating information between connected nodes. Here’s the complete picture.

math-basics

Convexity: What It Guarantees, and Why Deep Learning Works Without It

8 minute read

Published:

Convexity buys one enormous guarantee, every local minimum is global, and it says nothing at all about speed. Deep learning throws the guarantee away and still works, and it is worth being precise about how much of that we actually understand.

Norms, Inner Products, and the Geometry Behind L1 Sparsity

7 minute read

Published:

L1 regularisation produces exact zeros and L2 does not. The reason is not statistical, it is geometric: the L1 unit ball has corners on the axes, and corners are what optimisation solutions stick to.

Jacobians, Hessians, and Why Newton’s Method Loses at Scale

7 minute read

Published:

The Jacobian tells you how a map distorts volume, which is exactly the term normalising flows have to pay. The Hessian tells you the shape of the valley you are descending. Both are indispensable to reason with and, at a billion parameters, hopeless to form.

The Derivative Is a Linear Approximation, and the Gradient Is a Covector

6 minute read

Published:

Treating the derivative as a slope stops working the moment there is more than one input. Treating it as the best linear approximation keeps working forever, and makes backpropagation an obvious consequence of the chain rule rather than an algorithm to memorise.

Eigenvectors, the Spectral Theorem, and Why the SVD Always Exists

7 minute read

Published:

Eigenvectors are the directions a matrix does not rotate, when they exist. The SVD asks a weaker question that always has an answer, and that is exactly why it, not the eigendecomposition, is the workhorse of applied linear algebra.

Matrices as Linear Maps: Span, Rank, and the Subspaces They Create

7 minute read

Published:

A matrix is not a grid of numbers, it is a map. Once you read it that way, rank, column space, null space and rank–nullity stop being definitions to memorise and become one geometric statement about what the map keeps and what it destroys.

What Mathematics an ML Engineer Actually Needs

5 minute read

Published:

Almost every mathematical question asked in an ML interview reduces to two things: what a matrix does to space, and how to differentiate a composition. This book covers those two things properly and is honest about what you can safely forget.

physics-basics

Noether’s Theorem: Symmetry, Conservation, and Equivariant Networks

6 minute read

Published:

Energy is conserved because the laws of physics do not care what time it is. That single sentence is Noether’s theorem, and its machine learning descendant is the reason an equivariant network needs less data than one that must learn the symmetry from examples.

Entropy and Free Energy: From the Second Law to the ELBO

6 minute read

Published:

Thermodynamic entropy and Shannon entropy differ by a constant with units. Once you accept that, the ELBO stops being an inference trick and becomes a free energy, and diffusion models stop being a clever architecture and become a driven non-equilibrium process.

Statistical Mechanics: The Boltzmann Distribution and the Cost of Z

6 minute read

Published:

Counting microstates gives you the Boltzmann distribution, and the Boltzmann distribution gives you softmax, simulated annealing and energy-based models. The partition function is not a bookkeeping constant, it is the object that contains every thermodynamic quantity, and it is intractable for exactly that reason.

Hamiltonian Dynamics: Phase Space, Liouville, and Why HMC Works

6 minute read

Published:

Hamiltonian Monte Carlo is not a heuristic that happens to move well. It is exact because Hamiltonian flow preserves phase-space volume and is reversible, which is what makes the Metropolis acceptance ratio collapse to a difference of energies.

Lagrangian Mechanics: Why Nature Optimises a Functional

6 minute read

Published:

Newton says a particle moves because a force pushes it. Lagrange says it moves along the path that makes the action stationary. The second statement is harder to believe and far easier to use, and it is the one machine learning inherited.

Physics for Machine Learning: Why the Same Equations Keep Coming Back

5 minute read

Published:

Diffusion models, energy-based models and Hamiltonian Monte Carlo were not inspired by physics, they are physics, rewritten with a neural network in place of an analytic potential. Knowing which physics saves you from re-deriving it badly.

prob-basics

Expectation and Variance: Linearity Is Free, Additivity Is Not

5 minute read

Published:

Expectation adds up no matter how tangled the dependence. Variance does not, and the correction term, covariance, is where most of the interesting behaviour of ensembles, portfolios and minibatch gradients lives.

Random Variables: A Density Is Not a Probability

5 minute read

Published:

A probability density can be 2, or 200, and nothing is wrong. Getting clear on what a PDF actually is fixes half the confusion about continuous distributions, and explains the Jacobian term that makes normalising flows work.

Probability for ML Interviews: The Eight Ideas Worth Re-deriving

5 minute read

Published:

Almost every loss function in machine learning is a negative log-likelihood in disguise, and almost every model output is a distribution. This book rebuilds the probability you need to read those objects fluently, and flags the eight places interviewers know people slip.

python-primer

Files and Context Managers: Why with Is Not Optional

7 minute read

Published:

A file object that falls out of scope without with does get closed eventually, but ‘eventually’ means whenever the garbage collector gets around to it, which is not a promise any program handling more than a handful of files can live with.

Errors and Exceptions: EAFP, the Hierarchy, and Reading a Traceback

7 minute read

Published:

Exceptions in Python are not exceptional. They are a normal control-flow mechanism that the language leans on so heavily that the idiomatic style is to try the operation and handle the failure, rather than check first, which is faster, shorter, and free of race conditions.

Modules, Packages, and How Python Actually Finds Your Code

7 minute read

Published:

An import is not a textual include, it executes a file once, caches the result, and binds a name. Almost every confusing import error, from circular imports to ‘attempted relative import with no known parent package’, follows directly from that one sentence.

Dunder Methods: How Python’s Protocols Replace Interfaces

6 minute read

Published:

Python has almost no interfaces to implement and no operators to declare. Instead, every piece of syntax, len(x), x[i], for y in x, a + b, with r as f, is a documented call to a method with a double-underscore name. Learn the mapping and your own types stop being second-class citizens.

Classes and Objects: self, Attributes, and When Inheritance Is the Wrong Tool

6 minute read

Published:

A Python class is a factory for namespaces, not a sealed blueprint. Once you see where an attribute actually lives, on the instance or on the class, the mutable-default trap, the point of self, and the reason composition usually beats inheritance all fall out of the same rule.

Dicts and Sets: Hashing, Defaults, and What Makes a Key Legal

7 minute read

Published:

A dict trades memory for the ability to skip the search entirely. Everything that follows, why keys must be hashable, why lists cannot be keys, and why 1, 1.0 and True collide, comes from that single trade.

Names, Not Boxes: Python Syntax, Variables and the Built-in Types

7 minute read

Published:

A Python variable is not a container that holds a value, it is a label stuck onto an object that lives somewhere else. Almost every early surprise, from shared lists to 0.1 + 0.2, follows from taking that sentence literally.

research

BrainDyn: Sheaves Meet Neural ODEs for Brain Dynamics

11 minute read

Published:

Brain regions do not encode information in a shared feature space, which is exactly the assumption scalar message passing makes. BrainDyn puts learnable restriction maps between regions and integrates the result as a continuous-time system, the first pairing of cellular sheaves with neural ODEs.

DNSD: Making Sheaf Diffusion Work at Depth

12 minute read

Published:

Neural Sheaf Diffusion has a theoretical guarantee against representation collapse that does not survive contact with depth. DNSD diagnoses why, the Laplacian’s disagreement signal vanishes as diffusion succeeds, and replaces the operator rather than patching around it.

Beyond Simplices: Cell and Combinatorial Complexes

20 minute read

Published:

A simplicial complex cannot hold a benzene ring as a single cell, filling the hexagon costs three edges between atoms that share no bond, and this one constraint is what cell and combinatorial complexes exist to remove.

Message Passing on Simplicial Complexes

20 minute read

Published:

A simplex has four kinds of neighbour rather than one, and separating them lets a network see the difference between a filled triangle and an empty one, a distinction no graph neural network can make.

Topological Deep Learning Is Not Topological Data Analysis

17 minute read

Published:

TDA computes a topological descriptor and hands it to a model; TDL makes the topological object the domain the model runs on. Telling the two apart is the difference between a preprocessing step and an architecture.

GAPE: Remember to Forget, Gated Adaptive Positional Encoding

9 minute read

Published:

GAPE is a drop-in RoPE augmentation that adds content-aware attention logit biases: a query-gate suppresses irrelevant distant context while a key-gate preserves salient distant tokens. Provably sharper attention and improved long-context robustness, no architecture changes needed.

PolyNSD: Polynomial Neural Sheaf Diffusion

10 minute read

Published:

PolyNSD replaces the NSD propagation operator with a degree-K Chebyshev polynomial in the normalised sheaf Laplacian, achieving SOTA on homo- and heterophilic benchmarks with only diagonal restriction maps and dramatically lower memory usage.

Z-SASLM: Zero-Shot Style Blending via Spherical Interpolation

9 minute read

Published:

Z-SASLM is a zero-shot, fine-tuning-free style blending pipeline that replaces linear latent interpolation with SLERP along the geodesic of the hypersphere, preserving latent manifold structure when blending multiple styles. Published at CVPR 2025 Workshop.

HetSheaf: Heterogeneous Graphs Meet Cellular Sheaves

9 minute read

Published:

HetSheaf encodes graph heterogeneity directly in the sheaf data structure, type-aware stalks and restriction maps conditioned on node and edge types, instead of specialised architectural components, achieving +2pp on HGB with 10× fewer parameters.

sheaf

Equivariant Sheaf Neural Networks: Learning Geometric Transport on Graphs

33 minute read

Published:

Sheaf networks move features through matrix-valued maps but ignore the symmetries of physical space; equivariant GNNs respect those symmetries but move vectors with scalars. ESNN does both, learned, directed, matrix-valued transport that is exactly E(n)-equivariant, and proves that when displacement is the only geometric input, the radial–tangential family is all the linear transport there is.

Sheaf4Rec: What a Recommender Gains from a Vector Space per Node

13 minute read

Published:

Collaborative filtering represents every user and item as one static vector. Sheaf4Rec replaces each with a vector space, and reports consistent gains on ranking metrics, though the wins come from recall rather than precision, and the headline efficiency claim is hard to reconcile with the timing table.

Cooperative Sheaf Neural Networks: Listening Without Speaking

14 minute read

Published:

A sheaf gives every node a matrix-valued say in how its neighbours reach it, but not in whether they do. Set a node’s restriction maps to zero to stop it listening and you also stop it speaking. Fixing that needs sheaves on directed graphs, and the fix costs the Laplacian its positive semi-definiteness.

Sheaf Hypergraph Networks: Apparent Consensus in Higher-Order Relations

13 minute read

Published:

A graph edge relates two things. A hyperedge relates any number of them, and hypergraph networks aggregate over it uniformly, every member contributes the same way. Attaching a sheaf gives each member its own linear map into the group, and turns forced consensus into apparent consensus.

Bayesian Sheaf Neural Networks: Putting a Distribution on the Geometry

13 minute read

Published:

If a sheaf neural network learns its geometry from data, it can learn the wrong geometry and have no way of knowing. Treating the sheaf Laplacian as a latent random variable fixes that, but requires a reparameterisable distribution on SO(n) with a tractable density, which did not exist.

Joint Diffusion and Rotation Invariance: Sheaves That Learn to Lie

12 minute read

Published:

Every sheaf network so far predicts restriction maps with an MLP on concatenated features, a universal approximator with no inductive bias for heterophily at all. Two alternatives drawn from opinion dynamics get the bias for free, and stop the parameter count scaling with the feature dimension.

Sheaf-Based Positional Encodings: Letting Node Features Into the Spectrum

12 minute read

Published:

Laplacian eigenvector positional encodings tell a node where it sits in the graph, but the graph Laplacian only knows adjacency, so two structurally identical nodes get identical encodings no matter how different their features are. Swap in the sheaf Laplacian and the features enter the spectrum.

Surfing on the Neural Sheaf: What Happens If You Use the Wave Equation

9 minute read

Published:

Every sheaf model so far discretises the heat equation, which dissipates energy. Suk et al. try the wave equation instead, which conserves it, a one-line change of PDE with a clean theoretical motivation and a genuinely mixed empirical result.

Conn-NSD: Computing the Sheaf Instead of Learning It

11 minute read

Published:

Neural Sheaf Diffusion learns the restriction maps by gradient descent. Conn-NSD computes them once, before training, by assuming the data lies on a manifold and optimally aligning neighbouring tangent spaces, matching the learned models on small graphs at roughly half the cost per epoch.

SheafPool: Basis-Invariant Graph Readout for Sheaf Neural Networks

8 minute read

Published:

SheafPool solves a key missing piece in sheaf GNNs: graph-level pooling. Instead of averaging stalk vectors in arbitrary local bases, it aligns them into a shared canonical frame and builds a readout that is invariant to local basis changes.

Sheaf Attention Networks: GAT with Matrices Instead of Scalars

9 minute read

Published:

GAT weights a neighbour by a scalar. SheafAN keeps the scalar and adds a learned orthogonal transport matrix alongside it, recovering GAT exactly at d = 1, and turning a model that goes numerically unstable past eight layers into one that runs to sixty-four.

Sheaf Neural Networks: A Complete Research Guide

9 minute read

Published:

Standard GNNs assume neighbouring nodes should agree. Sheaf Neural Networks replace that assumption with a learned linear map on every edge, which turns heterophily, oversmoothing, and directional structure into one operator: the sheaf Laplacian.

stats-basics

Bayesian vs Frequentist: Two Meanings of the Word Probability

6 minute read

Published:

One school says probability is a long-run frequency, so parameters cannot have probabilities. The other says probability is a degree of belief, so they can. Everything else, priors, credible intervals, the whole argument, follows from that one disagreement.

The Bootstrap: Uncertainty When the Algebra Runs Out

6 minute read

Published:

If you cannot resample from the population, resample from your sample instead. That one substitution gives standard errors and intervals for statistics whose sampling distributions nobody can write down, and it fails in ways worth memorising.

Hypothesis Testing: What a p-Value Actually Measures

7 minute read

Published:

A p-value answers one narrow question: if nothing were going on, how often would data look at least this extreme? It says nothing about whether the effect is real, large, or worth shipping, and running twenty of them changes the meaning of all twenty.

Confidence Intervals: The Interval Is Random, the Parameter Is Not

5 minute read

Published:

A 95% confidence interval does not say the parameter is 95% likely to be inside it. It says the recipe that produced the interval succeeds 95% of the time. That distinction is the single most-failed question in statistics interviews.

Estimators, Bias and Variance: Why Unbiased Is Not the Same as Good

6 minute read

Published:

An estimator is a random variable, so it has a mean and a spread. Squared error splits exactly into those two pieces, and once you see the split, it becomes obvious that deliberately biasing an estimator can make it strictly better.

Statistics Basics: Reasoning Backwards From Data to Model

5 minute read

Published:

Probability runs forwards: pick a model, predict the data. Statistics runs backwards, and backwards is harder, many models could have produced what you saw. Everything else in this book is machinery for handling that ambiguity honestly.

tdl

transformers

FoPE: Fourier Position Embedding for Length Generalization

5 minute read

Published:

FoPE rethinks long-context positional encoding from a frequency-domain perspective. Instead of only stretching RoPE heuristically, it explicitly improves attention’s periodic extension so Transformers generalize more gracefully to longer sequences.

Position Interpolation: Extending RoPE with Minimal Fine-Tuning

5 minute read

Published:

Position Interpolation rescales positions before applying RoPE so a model trained on short contexts can be adapted to longer ones with surprisingly little fine-tuning. It became the reference baseline for long-context RoPE extension.

XPos: Length-Extrapolatable Rotary Embeddings

4 minute read

Published:

XPos modifies RoPE with a multiplicative decay that keeps relative rotations while stabilising magnitude at long distance. It is one of the cleanest attempts to make rotary embeddings extrapolate better.

p-RoPE: What Makes Rotary Positional Encodings Useful?

6 minute read

Published:

This paper does two things at once: it explains what RoPE is really doing inside a trained LLM, and it proposes p-RoPE, a partial rotary variant that drops the lowest frequencies to preserve stronger semantic channels.

LongRoPE: Extending Context to 2 Million Tokens

7 minute read

Published:

LongRoPE (Microsoft, 2024) pushes RoPE-based context to 2M tokens by searching for optimal per-dimension rescaling factors, far outperforming NTK or YaRN at extreme lengths.

YaRN: Yet Another RoPE Extensionn Method

6 minute read

Published:

YaRN combines NTK scaling for high-frequency dimensions with linear interpolation for low-frequency ones, plus a temperature correction, achieving better long-context performance with minimal fine-tuning.

The Transformer Block: Putting It All Together

6 minute read

Published:

A single Transformer block combines attention, residuals, layer norm, and an FFN into one reusable unit. Understanding this block is understanding the Transformer.

Residual Connections: Why Transformers Can Be Deep

7 minute read

Published:

Without residual connections, training a 96-layer Transformer would be practically impossible. The skip connection is a simple addition that solves the vanishing gradient problem and enables arbitrary depth.

Layer Normalization in Transformers

6 minute read

Published:

Layer norm is not optional plumbing. It determines training stability, gradient flow, and whether deep Transformers converge at all. Pre-LN vs Post-LN is not a detail, it changes training dynamics fundamentally.

Query, Key, Value: The Intuition Behind QKV

6 minute read

Published:

Q, K, and V are not arbitrary labels. They map precisely onto search queries, database labels, and retrieved content, a framework you already understand.

ALiBi: Attention with Linear Biases

4 minute read

Published:

ALiBi skips traditional positional embeddings entirely and just subtracts a distance penalty from attention scores. Zero extra parameters, excellent extrapolation. Press et al., 2022.

RoPE: Rotary Position Embeddings

4 minute read

Published:

RoPE encodes position by rotating query and key vectors by an angle proportional to position. The clever result: absolute encoding produces relative attention for free, and it’s now the dominant PE for large language models.

Relative Positional Encodings: It’s All About Distance

4 minute read

Published:

Instead of asking ‘where am I?’, relative PEs ask ‘how far are these two tokens apart?’ Shaw et al. and T5 both use this idea to build models that generalise better to variable-length inputs.

Learned Positional Encodings: Data-Driven Position

3 minute read

Published:

Instead of a fixed formula, why not just train position embeddings from scratch, like word embeddings? That’s exactly what BERT and GPT-1 do. Here’s how and when it works.

Sinusoidal Positional Encodings: The Original Solution

3 minute read

Published:

The PE method from the 2017 ‘Attention Is All You Need’ paper uses sine and cosine waves at different frequencies. Learn why this elegant choice encodes position without any training.

Positional Encodings: Why Position Matters

3 minute read

Published:

Transformers see all tokens at once, which means without help they’d treat ‘cat ate mouse’ and ‘mouse ate cat’ the same. Positional encodings fix this. Here’s the full landscape.

Multi-Head Attention: Many Eyes on the Data

4 minute read

Published:

One attention head sees one relationship. Multiple heads running in parallel let the model capture syntax, semantics, and coreference simultaneously, here’s how.

Self-Attention: Teaching Machines to Focus

5 minute read

Published:

Self-attention is the core of every Transformer. Learn how Query, Key, and Value vectors let every token directly attend to every other, and why that matters.

Transformers: The Architecture That Changed AI

7 minute read

Published:

A self-contained guide to the Transformer, the engine behind GPT, BERT, and modern AI. Learn how attention replaces recurrence and why every major AI system uses it.