Sitemap
A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.
Pages
Posts
Idiomatic and Performant Python: Writing It Well, Then Measuring Before You Optimise
Published:
Pythonic code and fast code are usually the same code, iterating over objects instead of indices is both clearer and quicker, but the two goals diverge exactly at the point where people start guessing instead of profiling.
The Standard Library Modules Worth Knowing Before You pip install
Published:
Python ships with a counter, a queue, a memoiser, a CLI parser, a test framework’s smaller cousin, and a JSON-safe date library already installed, most of what people reach for a package to solve is one import away.
Files and Context Managers: Why with Is Not Optional
Published:
A file object that falls out of scope without with does get closed eventually, but ‘eventually’ means whenever the garbage collector gets around to it, which is not a promise any program handling more than a handful of files can live with.
Errors and Exceptions: EAFP, the Hierarchy, and Reading a Traceback
Published:
Exceptions in Python are not exceptional. They are a normal control-flow mechanism that the language leans on so heavily that the idiomatic style is to try the operation and handle the failure, rather than check first, which is faster, shorter, and free of race conditions.
Modules, Packages, and How Python Actually Finds Your Code
Published:
An import is not a textual include, it executes a file once, caches the result, and binds a name. Almost every confusing import error, from circular imports to ‘attempted relative import with no known parent package’, follows directly from that one sentence.
Dunder Methods: How Python’s Protocols Replace Interfaces
Published:
Python has almost no interfaces to implement and no operators to declare. Instead, every piece of syntax, len(x), x[i], for y in x, a + b, with r as f, is a documented call to a method with a double-underscore name. Learn the mapping and your own types stop being second-class citizens.
Classes and Objects: self, Attributes, and When Inheritance Is the Wrong Tool
Published:
A Python class is a factory for namespaces, not a sealed blueprint. Once you see where an attribute actually lives, on the instance or on the class, the mutable-default trap, the point of self, and the reason composition usually beats inheritance all fall out of the same rule.
Comprehensions and Generators: Building Sequences Without Building Them
Published:
A comprehension builds the whole result before you touch any of it. Change one pair of brackets to parentheses and nothing is built at all until you ask, the difference, on a million elements, is 40 MB against 432 bytes.
Functions: Arguments, Scope, Closures, and the Default That Remembers
Published:
A def is executed, not declared: it builds a function object once, evaluating the default arguments there and then. Everything odd about mutable defaults and late-binding closures follows from taking that seriously.
Strings and f-strings: Immutability, Formatting, and the Quadratic Loop
Published:
Strings cannot be modified, so every method that looks like it edits one actually builds another. That single fact explains the whole str API, and why growing a string inside a loop can quietly become quadratic.
Dicts and Sets: Hashing, Defaults, and What Makes a Key Legal
Published:
A dict trades memory for the ability to skip the search entirely. Everything that follows, why keys must be hashable, why lists cannot be keys, and why 1, 1.0 and True collide, comes from that single trade.
Lists and Tuples: Slicing, Mutation, and the Cost of Every Method
Published:
A Python list is a growable array of pointers, and that one implementation fact predicts every complexity in its API: appending is cheap, inserting at the front is not, and in costs a full scan.
Operators and Control Flow: is Is Not ==, and for Has an else
Published:
== asks whether two objects are equal; is asks whether they are the same object. CPython’s caching of small integers makes the two agree just often enough to teach you the wrong lesson.
Names, Not Boxes: Python Syntax, Variables and the Built-in Types
Published:
A Python variable is not a container that holds a value, it is a label stuck onto an object that lives somewhere else. Almost every early surprise, from shared lists to 0.1 + 0.2, follows from taking that sentence literally.
Python from the Ground Up: What Actually Happens When You Run a Script
Published:
Python is not interpreted line by line, and python is not the language. Knowing what CPython compiles your file into, and where it puts the packages you install, explains most of the confusion beginners hit in their first week.
Memory and Concurrency in Python: References, the GIL, and Why Data Loaders Fork
Published:
Python names are references, so passing a list to a function hands over the original. Python threads hold one interpreter lock, so two CPU-bound threads take as long as one after the other. Both facts explain bugs you have already had.
Recursion and Dynamic Programming: Two Conditions, One Filled Table
Published:
Dynamic programming is not a trick, it is a diagnosis: if a problem has optimal substructure and overlapping subproblems, exhaustive recursion is doing the same work exponentially often and a table fixes it. Here is the diagnosis, and edit distance worked out cell by cell.
Sorting and Searching: The n log n Lower Bound, and How to Beat It Legally
Published:
No comparison sort can beat n log n, that is a theorem with a two-line counting proof, not a limit of current cleverness. Counting sort beats it anyway, because it never compares two elements. Also: how to write a binary search whose boundaries are right.
Graphs: Representations, BFS vs DFS, and the Assumption Dijkstra Makes Silently
Published:
Dijkstra is greedy: it finalises the closest unvisited vertex and never revisits it. That is valid only because every edge adds weight, put one negative edge in the graph and the algorithm returns a wrong answer without complaining.
Trees, Traversals and BSTs: Why Balance Is the Only Thing Standing Between You and a Linked List
Published:
A binary search tree gives O(log n) search only if it is balanced, and inserting sorted keys, the most natural thing a caller ever does, makes it a linked list with extra pointers. Here is what each traversal is for, and what AVL and red-black trees actually guarantee.
Stacks, Queues and Heaps: Three Restrictions That Buy You Speed
Published:
A heap is a tree with no pointers, just an array and two index formulas, and building one from n items takes O(n), not O(n log n). The sum that proves it is the most quotable derivation in this book.
Arrays, Dynamic Arrays and Hash Tables: Where O(1) Comes From and What It Costs
Published:
A linked list has O(1) insertion and loses to an array anyway. Appending is O(1) but a single append can copy a million elements. Hash lookup is O(1) until every key lands in one bucket. All three claims are true, with the qualifier that got dropped.
Big-O Precisely: What O, Ω and Θ Each Claim, and When the Faster Algorithm Loses
Published:
Big-O is an upper bound on growth, not a running time, and saying an algorithm is O(n²) is compatible with it being O(n). Here is what each symbol actually claims, how to read a cost off loops and recursions, and why the asymptotically better algorithm sometimes loses on real hardware.
Computer Science Basics for ML Engineers: What the Interview Actually Tests
Published:
An ML interview that asks you to invert a binary tree is not testing your tree knowledge, it is testing whether you can state a cost, defend it, and write a loop whose boundaries are right the first time. Here is the eight-post map of what that requires.
Noether’s Theorem: Symmetry, Conservation, and Equivariant Networks
Published:
Energy is conserved because the laws of physics do not care what time it is. That single sentence is Noether’s theorem, and its machine learning descendant is the reason an equivariant network needs less data than one that must learn the symmetry from examples.
Brownian Motion, Langevin and Fokker–Planck: The Physics of a Sampler
Published:
A Langevin step is a gradient step plus noise. Change the potential to the negative log-density and it samples from that density instead of minimising it, which is the entire sampler behind score-based generative models.
Entropy and Free Energy: From the Second Law to the ELBO
Published:
Thermodynamic entropy and Shannon entropy differ by a constant with units. Once you accept that, the ELBO stops being an inference trick and becomes a free energy, and diffusion models stop being a clever architecture and become a driven non-equilibrium process.
Statistical Mechanics: The Boltzmann Distribution and the Cost of Z
Published:
Counting microstates gives you the Boltzmann distribution, and the Boltzmann distribution gives you softmax, simulated annealing and energy-based models. The partition function is not a bookkeeping constant, it is the object that contains every thermodynamic quantity, and it is intractable for exactly that reason.
Hamiltonian Dynamics: Phase Space, Liouville, and Why HMC Works
Published:
Hamiltonian Monte Carlo is not a heuristic that happens to move well. It is exact because Hamiltonian flow preserves phase-space volume and is reversible, which is what makes the Metropolis acceptance ratio collapse to a difference of energies.
Lagrangian Mechanics: Why Nature Optimises a Functional
Published:
Newton says a particle moves because a force pushes it. Lagrange says it moves along the path that makes the action stationary. The second statement is harder to believe and far easier to use, and it is the one machine learning inherited.
Physics for Machine Learning: Why the Same Equations Keep Coming Back
Published:
Diffusion models, energy-based models and Hamiltonian Monte Carlo were not inspired by physics, they are physics, rewritten with a neural network in place of an analytic potential. Knowing which physics saves you from re-deriving it badly.
Symmetry and Groups: Invariance, Equivariance, and Why You Build It In
Published:
If rotating a molecule cannot change its energy, that is a fact about the target function you know before training starts. Encoding it in the architecture makes it true everywhere; learning it from augmented data makes it approximately true where you happened to have samples.
Riemannian Geometry: Geodesics, Exponential Maps and Why Trees Prefer Hyperbolic Space
Published:
Put an inner product on every tangent space and let it vary smoothly: that single object determines lengths, angles, distances, straight lines and volume. It also explains why a hierarchy embeds badly in Euclidean space and beautifully in hyperbolic space, the volume is in the wrong place.
Manifolds and Tangent Spaces: What ‘The Data Lies on a Manifold’ Actually Claims
Published:
A manifold is a space that looks flat close up and can be curved or knotted globally. That gap between local and global is why charts exist, why a single autoencoder latent space sometimes cannot fit the data, and what dimensionality reduction is really exploiting.
Curvature: Why a Flat Map of the Earth Must Lie
Published:
Curvature at a point is one over the radius of the circle that best hugs the curve there. Push that idea up to surfaces and Gauss’s Theorema Egregium falls out: some curvature is visible from inside the surface, which is why no map projection can ever get distances right.
Linear Maps as Geometry: Rotate, Scale, Rotate
Published:
A matrix is not a table of numbers, it is a deformation of space. The SVD says every deformation is the same three moves in sequence: rotate, stretch along axes, rotate again.
Euclidean Space: Inner Products, Projections and the Margin
Published:
One bilinear form generates the whole of flat geometry: lengths, angles, orthogonality, projections and the distance from a point to a hyperplane. Get the projection formula and you get the SVM margin for free.
Geometry for Machine Learning: Why Shape Keeps Coming Back
Published:
Every embedding you have ever trained lives in a metric space, every dataset you have ever fitted sits near a surface far thinner than its ambient dimension, and every architecture you trust encodes a symmetry. Geometry is not decoration on top of ML, it is what makes the problems tractable.
Equivariant Sheaf Neural Networks: Learning Geometric Transport on Graphs
Published:
Sheaf networks move features through matrix-valued maps but ignore the symmetries of physical space; equivariant GNNs respect those symmetries but move vectors with scalars. ESNN does both, learned, directed, matrix-valued transport that is exactly E(n)-equivariant, and proves that when displacement is the only geometric input, the radial–tangential family is all the linear transport there is.
Entropy, Cross-Entropy and KL: Why Your Classification Loss Is a Divergence
Published:
Minimising cross-entropy is minimising a KL divergence plus a constant you cannot change. Once that clicks, maximum likelihood, variational inference and the mode-collapse of a VAE all turn out to be the same argument run in different directions.
Limit Theorems: What the CLT Promises and What It Refuses to Promise
Published:
The central limit theorem is an asymptotic statement about the centre of a distribution. It says nothing about your sample size, and nothing about the tails you actually care about, which is why concentration inequalities exist.
The Ten Distributions You Need, and the Geometry of the Multivariate Gaussian
Published:
Ten families cover almost everything you will meet in a modelling interview. Nine of them fit in a table; the multivariate Gaussian deserves a page, because its covariance matrix is a statement about geometry.
Expectation and Variance: Linearity Is Free, Additivity Is Not
Published:
Expectation adds up no matter how tangled the dependence. Variance does not, and the correction term, covariance, is where most of the interesting behaviour of ensembles, portfolios and minibatch gradients lives.
Random Variables: A Density Is Not a Probability
Published:
A probability density can be 2, or 200, and nothing is wrong. Getting clear on what a PDF actually is fixes half the confusion about continuous distributions, and explains the Jacobian term that makes normalising flows work.
Conditional Probability and Bayes: Why a 99% Accurate Test Means Almost Nothing
Published:
A test catches 99% of a disease and false-alarms only 5% of the time. You test positive. The chance you are ill is one in six. Here is the computation, and why almost everyone guesses ten times too high.
Sample Spaces and the Kolmogorov Axioms: What a Probability Actually Is
Published:
Probability is not defined as a long-run frequency or a degree of belief. It is defined as a function on sets that obeys three rules, and the third rule, countable additivity, is the one doing all the work.
Probability for ML Interviews: The Eight Ideas Worth Re-deriving
Published:
Almost every loss function in machine learning is a negative log-likelihood in disguise, and almost every model output is a distribution. This book rebuilds the probability you need to read those objects fluently, and flags the eight places interviewers know people slip.
Bayesian vs Frequentist: Two Meanings of the Word Probability
Published:
One school says probability is a long-run frequency, so parameters cannot have probabilities. The other says probability is a degree of belief, so they can. Everything else, priors, credible intervals, the whole argument, follows from that one disagreement.
Sheaf4Rec: What a Recommender Gains from a Vector Space per Node
Published:
Collaborative filtering represents every user and item as one static vector. Sheaf4Rec replaces each with a vector space, and reports consistent gains on ranking metrics, though the wins come from recall rather than precision, and the headline efficiency claim is hard to reconcile with the timing table.
The Bootstrap: Uncertainty When the Algebra Runs Out
Published:
If you cannot resample from the population, resample from your sample instead. That one substitution gives standard errors and intervals for statistics whose sampling distributions nobody can write down, and it fails in ways worth memorising.
Hypothesis Testing: What a p-Value Actually Measures
Published:
A p-value answers one narrow question: if nothing were going on, how often would data look at least this extreme? It says nothing about whether the effect is real, large, or worth shipping, and running twenty of them changes the meaning of all twenty.
Cooperative Sheaf Neural Networks: Listening Without Speaking
Published:
A sheaf gives every node a matrix-valued say in how its neighbours reach it, but not in whether they do. Set a node’s restriction maps to zero to stop it listening and you also stop it speaking. Fixing that needs sheaves on directed graphs, and the fix costs the Laplacian its positive semi-definiteness.
Toward a Spectral Theory of Cellular Sheaves: The Paper Everything Else Cites
Published:
Every sheaf neural network is built on an operator defined in this 2019 paper. It is not a machine learning paper but a programme for lifting spectral graph theory to cellular sheaves, and it is unusually candid about which parts of the lift fail.
Confidence Intervals: The Interval Is Random, the Parameter Is Not
Published:
A 95% confidence interval does not say the parameter is 95% likely to be inside it. It says the recipe that produced the interval succeeds 95% of the time. That distinction is the single most-failed question in statistics interviews.
Sheaf Hypergraph Networks: Apparent Consensus in Higher-Order Relations
Published:
A graph edge relates two things. A hyperedge relates any number of them, and hypergraph networks aggregate over it uniformly, every member contributes the same way. Attaching a sheaf gives each member its own linear map into the group, and turns forced consensus into apparent consensus.
Maximum Likelihood: The Default Recipe for Turning Data Into Parameters
Published:
Pick the parameter under which the data you actually observed would have been least surprising. That single sentence generates the Gaussian mean, cross-entropy loss, and, once you add a prior, L2 regularisation.
Estimators, Bias and Variance: Why Unbiased Is Not the Same as Good
Published:
An estimator is a random variable, so it has a mean and a spread. Squared error splits exactly into those two pieces, and once you see the split, it becomes obvious that deliberately biasing an estimator can make it strictly better.
Bayesian Sheaf Neural Networks: Putting a Distribution on the Geometry
Published:
If a sheaf neural network learns its geometry from data, it can learn the wrong geometry and have no way of knowing. Treating the sheaf Laplacian as a latent random variable fixes that, but requires a reparameterisable distribution on SO(n) with a tractable density, which did not exist.
Joint Diffusion and Rotation Invariance: Sheaves That Learn to Lie
Published:
Every sheaf network so far predicts restriction maps with an MLP on concatenated features, a universal approximator with no inductive bias for heterophily at all. Two alternatives drawn from opinion dynamics get the bias for free, and stop the parameter count scaling with the feature dimension.
Descriptive Statistics: What a Summary Keeps and What It Discards
Published:
A summary statistic is a lossy compression of a sample. Knowing which information each one discards, and why sample variance divides by n minus 1 rather than n, is most of what descriptive statistics has to teach.
Statistics Basics: Reasoning Backwards From Data to Model
Published:
Probability runs forwards: pick a model, predict the data. Statistics runs backwards, and backwards is harder, many models could have produced what you saw. Everything else in this book is machinery for handling that ambiguity honestly.
Sheaf-Based Positional Encodings: Letting Node Features Into the Spectrum
Published:
Laplacian eigenvector positional encodings tell a node where it sits in the graph, but the graph Laplacian only knows adjacency, so two structurally identical nodes get identical encodings no matter how different their features are. Swap in the sheaf Laplacian and the features enter the spectrum.
Surfing on the Neural Sheaf: What Happens If You Use the Wave Equation
Published:
Every sheaf model so far discretises the heat equation, which dissipates energy. Suk et al. try the wave equation instead, which conserves it, a one-line change of PDE with a clean theoretical motivation and a genuinely mixed empirical result.
Conn-NSD: Computing the Sheaf Instead of Learning It
Published:
Neural Sheaf Diffusion learns the restriction maps by gradient descent. Conn-NSD computes them once, before training, by assuming the data lies on a manifold and optimally aligning neighbouring tangent spaces, matching the learned models on small graphs at roughly half the cost per epoch.
Gauges: When There Is No Shared Frame
Published:
The previous chapter ended on an arbitrary choice that could not be eliminated. Gauge theory’s answer is to stop trying: keep every local frame, transport between them explicitly, and require the model to be indifferent to which frames you picked.
Convexity: What It Guarantees, and Why Deep Learning Works Without It
Published:
Convexity buys one enormous guarantee, every local minimum is global, and it says nothing at all about speed. Deep learning throws the guarantee away and still works, and it is worth being precise about how much of that we actually understand.
Norms, Inner Products, and the Geometry Behind L1 Sparsity
Published:
L1 regularisation produces exact zeros and L2 does not. The reason is not statistical, it is geometric: the L1 unit ball has corners on the axes, and corners are what optimisation solutions stick to.
Geodesics: Learning on Curved Domains
Published:
On a surface there is no global grid to slide a filter along, and no canonical direction to call ‘up’. What survives is distance, and building convolution out of distance alone exposes exactly one ambiguity, which is where the next chapter starts.
Matrix Calculus Without Sign Errors: The Identities and the Layout Trap
Published:
Most matrix-calculus mistakes are not mistakes about calculus. They are mistakes about which convention you were using, and they produce an answer that is correct up to a transpose, which is to say, wrong.
Groups: Equivariance Beyond Translation
Published:
Translation is one group. Swap it for rotations, reflections, or the rigid motions of 3-D space and the same construction produces a different architecture, with the parameter count cut by exactly the size of the orbit.
Jacobians, Hessians, and Why Newton’s Method Loses at Scale
Published:
The Jacobian tells you how a map distorts volume, which is exactly the term normalising flows have to pay. The Hessian tells you the shape of the valley you are descending. Both are indispensable to reason with and, at a billion parameters, hopeless to form.
Grids: Why Translation Equivariance Forces Convolution
Published:
Convolution is not a clever idea someone had about images. It is the only linear map that commutes with translation, a theorem, not a design choice, and one you can verify by exhaustion on a small enough case.
The Hardware Lottery: Why the Dense Version Won
Published:
Message passing on a sparse graph does asymptotically less work than attention over every pair. It is still the slower one to train. This chapter is about why the architecture that wins is the one your hardware happens to like.
The Derivative Is a Linear Approximation, and the Gradient Is a Covector
Published:
Treating the derivative as a slope stops working the moment there is more than one input. Treating it as the best linear approximation keeps working forever, and makes backpropagation an obvious consequence of the chain rule rather than an algorithm to memorise.
Transformers Are GNNs on Fully Connected Graphs
Published:
Write self-attention in the message-passing template and you get the Graph Attention Network equations with one substitution: the neighbourhood becomes the whole input. The two architectures are not analogous, they are the same operator on different graphs.
Eigenvectors, the Spectral Theorem, and Why the SVD Always Exists
Published:
Eigenvectors are the directions a matrix does not rotate, when they exist. The SVD asks a weaker question that always has an answer, and that is exactly why it, not the eigendecomposition, is the workhorse of applied linear algebra.
Matrices as Linear Maps: Span, Rank, and the Subspaces They Create
Published:
A matrix is not a grid of numbers, it is a map. Once you read it that way, rank, column space, null space and rank–nullity stop being definitions to memorise and become one geometric statement about what the map keeps and what it destroys.
Permutation Symmetry and the Message-Passing Blueprint
Published:
A graph’s nodes have no canonical order, so any model that reads one must give the same answer under relabelling. That single requirement forces the three-step message, aggregate, update template, it is not a design choice.
What Mathematics an ML Engineer Actually Needs
Published:
Almost every mathematical question asked in an ML interview reduces to two things: what a matrix does to space, and how to differentiate a composition. This book covers those two things properly and is honest about what you can safely forget.
Geometric Deep Learning: One Blueprint Behind Every Architecture
Published:
CNNs, GNNs, Transformers and sheaf models look like separate inventions. They are the same recipe applied to different domains: identify the symmetry of your data, then build layers that respect it. This book is that recipe, and the arguments that follow from it.
BrainDyn: Sheaves Meet Neural ODEs for Brain Dynamics
Published:
Brain regions do not encode information in a shared feature space, which is exactly the assumption scalar message passing makes. BrainDyn puts learnable restriction maps between regions and integrates the result as a continuous-time system, the first pairing of cellular sheaves with neural ODEs.
DNSD: Making Sheaf Diffusion Work at Depth
Published:
Neural Sheaf Diffusion has a theoretical guarantee against representation collapse that does not survive contact with depth. DNSD diagnoses why, the Laplacian’s disagreement signal vanishes as diffusion succeeds, and replaces the operator rather than patching around it.
Beyond Simplices: Cell and Combinatorial Complexes
Published:
A simplicial complex cannot hold a benzene ring as a single cell, filling the hexagon costs three edges between atoms that share no bond, and this one constraint is what cell and combinatorial complexes exist to remove.
Message Passing on Simplicial Complexes
Published:
A simplex has four kinds of neighbour rather than one, and separating them lets a network see the difference between a filled triangle and an empty one, a distinction no graph neural network can make.
Topological Deep Learning Is Not Topological Data Analysis
Published:
TDA computes a topological descriptor and hands it to a model; TDL makes the topological object the domain the model runs on. Telling the two apart is the difference between a preprocessing step and an architecture.
Diffusion Beyond Images: Inpainting, Audio, Video, Molecules, Robots
Published:
Diffusion travels well between modalities, but not because it is generic. It fits wherever the answer is genuinely one-to-many, and the specific reason differs for a spectrogram, a protein backbone and a robot’s next ten actions.
Diffusion vs Flow Matching: Two Names for One Family
Published:
Flow matching is often presented as the successor to diffusion. It is more accurate, and more useful, to say that diffusion is one particular probability path inside the flow-matching framework, and not the straightest one available.
Rectified Flow: Straightening Trajectories to Buy Few-Step Sampling
Published:
A perfectly straight generative trajectory is exactly solvable in a single Euler step. Every extra sampling step you pay for is buying back curvature, so rectified flow attacks the curvature itself rather than the solver.
Flow Matching: Training a Velocity Field Without Ever Solving an ODE
Published:
Continuous normalising flows were elegant and nearly untrainable, every gradient step needed an ODE solve and a divergence estimate. Flow matching removes both by regressing a velocity field against a target you can write down in closed form, one example at a time.
Distillation and Consistency Models: Getting to Four Steps
Published:
Better ODE solvers bottom out around ten evaluations because the trajectory is genuinely curved. To go lower you have to change the model, either teach a student to take two teacher steps at once, or train a network that jumps to the end of the trajectory from anywhere on it.
Samplers and Schedulers: Diffusion Sampling Is Numerical Integration
Published:
Once you see the reverse process as an ODE, the thousand-step sampling loop stops being a fact of life and becomes what it really is, a first-order solver with a terrible step size. Everything after that is numerical analysis.
Latent Diffusion: Denoise Where the Information Is
Published:
Most of the bits in a photograph encode texture no one can see. Latent diffusion throws them away first with an autoencoder, then runs the entire diffusion process in a space roughly forty-eight times smaller, which is how Stable Diffusion fits on a consumer GPU.
Classifier-Free Guidance: Buying Prompt Adherence with a Second Forward Pass
Published:
Conditioning a diffusion model on a prompt is easy; making it actually obey the prompt is not. Classifier-free guidance fixes that by training one network to do two jobs and then extrapolating between its own answers.
DDIM: Same Marginals, Fewer Steps, and a Latent Space Worth Having
Published:
DDPM’s objective never actually required the forward process to be Markov, only that its marginals be Gaussian. Dropping the Markov assumption exposes a whole family of samplers a trained model already supports, including a deterministic one that runs in 20 steps and gives an invertible latent space.
Score Matching and the SDE View: DDPM as One Discretisation Among Many
Published:
Noise prediction and score estimation are the same network in different units. Taking the step size to zero turns the whole method into a stochastic differential equation, and reveals a deterministic ODE with identical marginals hiding inside it.
DDPM Training: From a Variational Bound to Four Lines of PyTorch
Published:
The DDPM objective starts as a variational bound with T+1 KL terms and ends as a plain mean-squared error on noise. Following the collapse shows why the discarded weighting term is not an approximation you tolerate but a reweighting that improves samples.
The Reverse Process: Why a Gaussian Is Enough, and What the Network Should Predict
Published:
Reversing a corruption is generally intractable. Diffusion escapes by taking steps small enough that the reverse conditional is itself Gaussian, so the network only has to output a mean. Which mean it outputs, though, turns out to matter a great deal.
The Forward Process: A Corruption Engineered to Be Jumped Into
Published:
The forward process looks like the trivial half of diffusion, but every term in it is load-bearing. Drop the shrink factor and the variance diverges; pick the wrong schedule and a third of your timesteps are spent denoising static.
Diffusion Models: Learning to Undo Noise
Published:
Destroying an image is easy and needs no learning at all. Diffusion models exploit that asymmetry: they define a trivial forward corruption, then train a network to walk it backwards one small step at a time.
Output and Gated Activations: Softmax, Sparsemax, GLU, and SIREN
Published:
The last activation in your network is not a modelling preference, it is a contract with your loss function. Break it and training stops meaning anything. Here is the contract, and what changes when the activation itself becomes learned.
Modern Activation Functions: GELU, SiLU, Mish, and Smooth Gating
Published:
Once ReLU became the default, researchers started asking a better question: can we keep the easy optimization while making the activation smoother, softer, and more expressive? This chapter covers the modern answers.
Activation Functions in Neural Networks: Why Non-Linearity Matters
Published:
Activation functions are the reason neural networks can model curved decision boundaries instead of collapsing into one giant linear map. This chapter builds the intuition first, then walks through the classical functions that shaped deep learning.
FoPE: Fourier Position Embedding for Length Generalization
Published:
FoPE rethinks long-context positional encoding from a frequency-domain perspective. Instead of only stretching RoPE heuristically, it explicitly improves attention’s periodic extension so Transformers generalize more gracefully to longer sequences.
Position Interpolation: Extending RoPE with Minimal Fine-Tuning
Published:
Position Interpolation rescales positions before applying RoPE so a model trained on short contexts can be adapted to longer ones with surprisingly little fine-tuning. It became the reference baseline for long-context RoPE extension.
XPos: Length-Extrapolatable Rotary Embeddings
Published:
XPos modifies RoPE with a multiplicative decay that keeps relative rotations while stabilising magnitude at long distance. It is one of the cleanest attempts to make rotary embeddings extrapolate better.
p-RoPE: What Makes Rotary Positional Encodings Useful?
Published:
This paper does two things at once: it explains what RoPE is really doing inside a trained LLM, and it proposes p-RoPE, a partial rotary variant that drops the lowest frequencies to preserve stronger semantic channels.
SheafPool: Basis-Invariant Graph Readout for Sheaf Neural Networks
Published:
SheafPool solves a key missing piece in sheaf GNNs: graph-level pooling. Instead of averaging stalk vectors in arbitrary local bases, it aligns them into a shared canonical frame and builds a readout that is invariant to local basis changes.
GAPE: Remember to Forget, Gated Adaptive Positional Encoding
Published:
GAPE is a drop-in RoPE augmentation that adds content-aware attention logit biases: a query-gate suppresses irrelevant distant context while a key-gate preserves salient distant tokens. Provably sharper attention and improved long-context robustness, no architecture changes needed.
PolyNSD: Polynomial Neural Sheaf Diffusion
Published:
PolyNSD replaces the NSD propagation operator with a degree-K Chebyshev polynomial in the normalised sheaf Laplacian, achieving SOTA on homo- and heterophilic benchmarks with only diagonal restriction maps and dramatically lower memory usage.
Z-SASLM: Zero-Shot Style Blending via Spherical Interpolation
Published:
Z-SASLM is a zero-shot, fine-tuning-free style blending pipeline that replaces linear latent interpolation with SLERP along the geodesic of the hypersphere, preserving latent manifold structure when blending multiple styles. Published at CVPR 2025 Workshop.
HetSheaf: Heterogeneous Graphs Meet Cellular Sheaves
Published:
HetSheaf encodes graph heterogeneity directly in the sheaf data structure, type-aware stalks and restriction maps conditioned on node and edge types, instead of specialised architectural components, achieving +2pp on HGB with 10× fewer parameters.
LongRoPE: Extending Context to 2 Million Tokens
Published:
LongRoPE (Microsoft, 2024) pushes RoPE-based context to 2M tokens by searching for optimal per-dimension rescaling factors, far outperforming NTK or YaRN at extreme lengths.
YaRN: Yet Another RoPE Extensionn Method
Published:
YaRN combines NTK scaling for high-frequency dimensions with linear interpolation for low-frequency ones, plus a temperature correction, achieving better long-context performance with minimal fine-tuning.
NTK-Aware Scaling: Extending Context Without Fine-Tuning
Published:
NTK-Aware Scaling extends the context window of RoPE-based models by rescaling frequencies using Neural Tangent Kernel theory, with no fine-tuning required.
The Transformer Block: Putting It All Together
Published:
A single Transformer block combines attention, residuals, layer norm, and an FFN into one reusable unit. Understanding this block is understanding the Transformer.
Feed-Forward Networks: The Forgotten Half of Transformers
Published:
The FFN block holds two-thirds of a Transformer’s parameters and does most of its factual recall. Yet it is almost always overlooked in introductions to attention.
Residual Connections: Why Transformers Can Be Deep
Published:
Without residual connections, training a 96-layer Transformer would be practically impossible. The skip connection is a simple addition that solves the vanishing gradient problem and enables arbitrary depth.
Layer Normalization in Transformers
Published:
Layer norm is not optional plumbing. It determines training stability, gradient flow, and whether deep Transformers converge at all. Pre-LN vs Post-LN is not a detail, it changes training dynamics fundamentally.
Encoder vs Decoder vs Encoder-Decoder Transformers
Published:
BERT, GPT, and T5 are all Transformers, but their architectures are fundamentally different. One comparison table clarifies the entire landscape.
Cross-Attention: How Models Attend to Another Sequence
Published:
Cross-attention lets one sequence query information from a completely different sequence. It is the bridge between encoder and decoder, and the core of multimodal AI.
Attention Masks: Causal, Padding, and Bidirectional
Published:
The difference between GPT, BERT, and T5 is largely a masking decision. Learn how causal, padding, and bidirectional masks shape what each token is allowed to see.
Query, Key, Value: The Intuition Behind QKV
Published:
Q, K, and V are not arbitrary labels. They map precisely onto search queries, database labels, and retrieved content, a framework you already understand.
Scaled Dot-Product Attention: Why the √d Matters
Published:
Dividing by √d_k is not just a trick, it prevents softmax from saturating and dying in high-dimensional spaces. Here’s the math and the intuition.
ALiBi: Attention with Linear Biases
Published:
ALiBi skips traditional positional embeddings entirely and just subtracts a distance penalty from attention scores. Zero extra parameters, excellent extrapolation. Press et al., 2022.
RoPE: Rotary Position Embeddings
Published:
RoPE encodes position by rotating query and key vectors by an angle proportional to position. The clever result: absolute encoding produces relative attention for free, and it’s now the dominant PE for large language models.
Relative Positional Encodings: It’s All About Distance
Published:
Instead of asking ‘where am I?’, relative PEs ask ‘how far are these two tokens apart?’ Shaw et al. and T5 both use this idea to build models that generalise better to variable-length inputs.
Learned Positional Encodings: Data-Driven Position
Published:
Instead of a fixed formula, why not just train position embeddings from scratch, like word embeddings? That’s exactly what BERT and GPT-1 do. Here’s how and when it works.
Sinusoidal Positional Encodings: The Original Solution
Published:
The PE method from the 2017 ‘Attention Is All You Need’ paper uses sine and cosine waves at different frequencies. Learn why this elegant choice encodes position without any training.
Positional Encodings: Why Position Matters
Published:
Transformers see all tokens at once, which means without help they’d treat ‘cat ate mouse’ and ‘mouse ate cat’ the same. Positional encodings fix this. Here’s the full landscape.
Multi-Head Attention: Many Eyes on the Data
Published:
One attention head sees one relationship. Multiple heads running in parallel let the model capture syntax, semantics, and coreference simultaneously, here’s how.
Self-Attention: Teaching Machines to Focus
Published:
Self-attention is the core of every Transformer. Learn how Query, Key, and Value vectors let every token directly attend to every other, and why that matters.
Transformers: The Architecture That Changed AI
Published:
A self-contained guide to the Transformer, the engine behind GPT, BERT, and modern AI. Learn how attention replaces recurrence and why every major AI system uses it.
RNNs, LSTMs, and GRUs: Sequence Models Before Attention
Published:
A recurrent network shares weights across time exactly as a convolution shares them across space. The trouble is that gradients then travel through a product of Jacobians, and a product of a hundred numbers slightly below one is zero.
Convolutions and CNNs: Weight Sharing as a Prior
Published:
A dense layer from a 224-by-224 colour image to 1000 units holds 150.5 million weights; a 3-by-3, 64-filter convolution holds 1,792, and the two constraints that buy that factor of 84,000 are exactly the prior that makes it work on images.
PCA: Maximum Variance and Minimum Reconstruction Error
Published:
PCA can be derived by asking for the directions of greatest spread, or by asking for the subspace that loses the least when you project onto it. The two questions look unrelated and have the same answer, which is the most useful thing to understand about it.
Clustering: What Each Algorithm Assumes a Cluster Is
Published:
k-means says a cluster is a ball around a centroid, DBSCAN says it is a connected dense region, and a Gaussian mixture says it is a bump in a density, pick the algorithm and you have already picked the answer.
Support Vector Machines: Margins and the Kernel Trick
Published:
Among all the hyperplanes that separate two classes, one sits furthest from both. Finding it turns out to depend on the data only through inner products, and that single fact is what lets you work in a space you never build.
Trees, Forests, and Boosting: Axis-Aligned Everything
Published:
A decision tree chops feature space into axis-aligned boxes and predicts one number per box, which explains why it needs no feature scaling, why it approximates a diagonal boundary as a staircase, and why it cannot extrapolate a single step beyond the training range.
Logistic Regression: Linear in the Log-Odds
Published:
Logistic regression is not a squashed linear regression, it is a straight line drawn in log-odds space, which is why one coefficient means one multiplication of the odds, and why perfectly separable data drives that coefficient to infinity.
Linear Regression: Least Squares as Projection
Published:
Fitting a line by least squares is not an optimisation trick, it is the orthogonal projection of the observation vector onto the column space of the design matrix, and once you see that, the normal equations, the failure modes, and the reason we solve by QR instead of inverting all follow from one picture.
Gradient Descent and Backpropagation: How a Model Learns
Published:
Training is one loop: measure the loss, ask backpropagation which way is downhill, take a small step. This chapter derives exactly how small that step has to be, why the answer is 2/a for a quadratic, and why reverse-mode differentiation gets you every gradient for roughly the price of one forward pass.
Bias, Variance, and Regularisation: Why Models Fail to Generalise
Published:
A model can fail for two opposite reasons: it is too rigid to represent the truth, or so flexible that it chases the noise. Squared error splits cleanly into exactly those two terms plus a floor you cannot beat.
Machine Learning Before Transformers: A Working Foundation
Published:
Every later book on this site assumes you already know what a loss is, why gradient descent works, and what a convolution buys you. This book supplies that, and follows one thread through it: how much structure you build in versus how much you let the data decide.
Topological Deep Learning: From Persistent Homology to Higher-Order Message Passing
Published:
This book has two halves. One computes topological summaries of data and feeds them to a model. The other makes the topology itself the domain the network runs on. This guide sets out both, and the pipeline that connects them.
Sheaf Attention Networks: GAT with Matrices Instead of Scalars
Published:
GAT weights a neighbour by a scalar. SheafAN keeps the scalar and adds a learned orthogonal transport matrix alongside it, recovering GAT exactly at d = 1, and turning a model that goes numerically unstable past eight layers into one that runs to sixty-four.
Neural Sheaf Diffusion: Heterophily and Oversmoothing Are the Same Problem
Published:
GNNs fail on heterophilic graphs and they oversmooth with depth. Bodnar et al. show these are one failure with one cause, the graph is implicitly equipped with a trivial sheaf, and prove a hierarchy of exactly what richer sheaves buy you.
Sheaf Neural Networks: A Complete Research Guide
Published:
Standard GNNs assume neighbouring nodes should agree. Sheaf Neural Networks replace that assumption with a learned linear map on every edge, which turns heterophily, oversmoothing, and directional structure into one operator: the sheaf Laplacian.
GNNs for Computer Vision: Scene Graphs and Beyond
Published:
Computer vision tasks increasingly require relational reasoning, understanding how objects relate to each other, not just what they are. Scene graph generation, visual question answering, action recognition from skeletons, and 3D point cloud processing all benefit from GNN-based relational modelling.
GNNs for Robotics: Planning, Manipulation, and Multi-Agent Systems
Published:
Robots interact with structured environments: objects have relationships, joints form kinematic chains, agents communicate through interaction graphs. GNNs encode these relational structures, enabling generalisation across object configurations, robot morphologies, and multi-agent scenarios.
GNNs for Knowledge Graphs: Reasoning and Completion
Published:
Knowledge graphs encode human knowledge as typed entity-relation triples. GNNs enable structure-aware entity representation, multi-hop reasoning, knowledge base completion, and entity alignment, tasks that shallow embedding methods cannot fully solve.
GNNs for Traffic Forecasting
Published:
Traffic prediction is a canonical spatio-temporal graph task: sensors on roads form a fixed graph, and speed/volume measurements evolve over time. GNNs capture spatial correlations between sensors; RNNs or convolutions capture temporal patterns. Together they achieve state-of-the-art traffic forecasting.
GNNs for Social Networks: Influence, Communities, and Misinformation
Published:
Social networks are large sparse graphs with rich node features (user profiles) and heterogeneous edges (friendship, follow, retweet). GNNs predict user behaviour, detect communities, identify influential spreaders, and flag misinformation, tasks with significant real-world impact.
GNNs for Recommender Systems
Published:
Recommendation is naturally a graph problem: users and items are nodes, interactions are edges. GNNs on bipartite user-item graphs capture higher-order collaborative filtering signals, friends of friends liked this, that matrix factorisation cannot represent.
GNNs for Molecules: Drug Discovery and Material Design
Published:
Graph neural networks are transforming computational drug discovery. Molecules are natural graphs, and GNNs learn molecular representations that predict toxicity, solubility, binding affinity, and synthesis feasibility, tasks that previously required expensive laboratory experiments.
Polynomial Neural Sheaf Diffusion
Published:
Polynomial Neural Sheaf Diffusion (PNSD) replaces the fixed diffusion operator (I - Δ_F) with a learnable polynomial of the Sheaf Laplacian. This gives the model spectral flexibility, it can learn to amplify or suppress different frequency components of the sheaf signal.
Equivariant Sheaf Neural Networks
Published:
Sheaves with orthogonal restriction maps define a connection on the graph, a parallel transport structure over edges. This connects sheaf GNNs to differential geometry and enables equivariant processing of data with local coordinate frames at each node.
Sheaf Neural Networks and Heterophily
Published:
Sheaf GNNs are the principled solution to heterophily: by learning per-edge maps that transform features before comparison, they can perform diffusion that converges within classes and diverges across classes, the exact opposite of standard GCN’s collapse.
Diagonal, Orthogonal, and General Sheaf Maps
Published:
The restriction maps in a cellular sheaf can be constrained to different matrix classes: scalars, diagonal matrices, orthogonal matrices, or general matrices. Each class offers a different trade-off between expressivity and computational cost.
Neural Sheaf Diffusion: Learning Sheaves End-to-End
Published:
Neural Sheaf Diffusion (Bodnar et al., 2022) learns the sheaf restriction maps from data using a neural network, then performs diffusion with the learned Sheaf Laplacian. This gives a principled, topology-grounded GNN that handles heterophily without heuristic fixes.
The Sheaf Laplacian: Spectral Theory for Sheaves
Published:
The Sheaf Laplacian generalises the graph Laplacian by incorporating per-edge restriction maps. Its spectrum reveals how consistent data is under the sheaf. Sheaf diffusion with this Laplacian generalises GCN to handle heterophilic graphs.
What Is a Sheaf? From Topology to Graph Learning
Published:
A sheaf is a mathematical object from algebraic topology that assigns vector spaces to cells and linear maps between them. On graphs, sheaves assign feature spaces to nodes and edges, with restriction maps encoding how node features relate across edges.
Why Message Passing Is Not Enough: The Case for Sheaves
Published:
Standard message passing aggregates neighbour features and averages. On heterophilic graphs (where neighbours often disagree), this is harmful. Cellular sheaves provide a mathematically principled framework to model per-edge relationships between node features, going beyond mere averaging.
Molecular GNNs: Learning on Atoms and Bonds
Published:
Molecules are graphs. Molecular GNNs predict chemical properties from structure. The best models use 3D coordinates and bond angles, not just connectivity.
Tensor Field Networks and Geometric Deep Learning
Published:
Tensor Field Networks (TFN) were the first architecture to achieve SE(3) equivariance using spherical harmonics and Clebsch-Gordan tensor products. They laid the theoretical foundation for NequIP and MACE, the current state-of-the-art in equivariant molecular force fields.
SE(3)-Transformers: Attention with 3D Symmetry
Published:
SE(3)-Transformers extend self-attention to 3D point clouds and molecular graphs while maintaining SE(3) equivariance. Attention weights are learned between node pairs; values are equivariant features built from spherical harmonics.
EGNN: E(n)-Equivariant Graph Neural Networks
Published:
EGNN achieves E(n)-equivariance with a simple update rule: positions updated via weighted sums of relative position vectors, features updated via invariant distances. No spherical harmonics needed.
Equivariance: What It Means and Why It Matters
Published:
Equivariance formalises the idea that a function should ‘commute with symmetry transformations.’ A rotation-equivariant model applied to rotated input gives the rotated output, no extra training needed. This is the foundation for geometric deep learning.
Why Geometry Matters in Graph Neural Networks
Published:
Many real-world graphs are embedded in 3D space, molecules, proteins, point clouds, crystal structures. Standard GNNs ignore coordinates and only use connectivity. Geometric GNNs incorporate spatial positions and must respect physical symmetries.
Spatio-Temporal GNNs: Learning on Graphs Through Time
Published:
Spatio-temporal GNNs combine spatial message passing with temporal sequence modelling. They are the dominant approach for traffic forecasting, weather prediction, and any task where measurements at sensor nodes evolve over time on a fixed graph.
Graph Neural ODEs: Continuous-Time Graph Dynamics
Published:
Neural ODEs replace discrete layer-by-layer computation with continuous dynamics governed by a differential equation. Graph Neural ODEs apply this to graph data, treating node embeddings as a dynamical system evolving in continuous time.
Temporal Graph Networks: Learning from Events
Published:
TGN (Temporal Graph Network) is the leading framework for continuous-time dynamic graphs. It maintains a per-node memory that is updated upon each interaction, enabling efficient inductive link prediction on event streams.
Static vs Dynamic Graphs: When Structure Changes Over Time
Published:
Most GNN research assumes a fixed graph. Real graphs evolve: edges appear and disappear, node features drift, new nodes arrive. Dynamic graph learning addresses how to model and predict on graphs whose structure changes over time.
Temporal Knowledge Graphs: Facts That Change Over Time
Published:
Most knowledge graphs treat facts as timeless, but facts change. Barack Obama was president from 2009 to 2017. Temporal Knowledge Graphs add timestamps to triples, requiring models to reason about what was true when.
Knowledge Graph Embeddings vs GNNs
Published:
Knowledge graph completion can be solved with shallow KG embeddings (TransE, DistMult, ComplEx) or with structural GNNs (R-GCN, CompGCN). Each approach has different inductive biases and failure modes. Understanding when to use each is the central design decision for KG tasks.
HAN: Heterogeneous Graph Attention Networks
Published:
HAN combines meta-path decomposition with two levels of attention: node-level attention weights neighbours along a meta-path, and semantic-level attention weights different meta-paths. This lets the model learn which relationships matter most for a given task.
R-GCN: Relational Graph Convolutional Networks
Published:
R-GCN extends GCN to multi-relational graphs by learning a separate weight matrix for each relation type. It handles knowledge graphs with typed edges and powers both entity classification and link prediction tasks.
Heterogeneous Graphs: When Nodes and Edges Have Types
Published:
Most real-world graphs are heterogeneous, they contain multiple node types (users, items, tags) and edge types (clicks, rates, authors). Standard GNNs treat all nodes and edges identically, making them blind to this type structure.
Graph Classification: From Node Embeddings to Graph Embeddings
Published:
Graph classification is the task of predicting a label for an entire graph. It requires composing message passing (node embeddings), readout (graph embedding), and a classifier, and all three choices interact to determine model expressiveness.
Set2Set and Attention Readout: Order-Invariant Graph Summaries
Published:
Mean and sum readout treat all nodes equally. Attention readout learns which nodes matter most for a given task. Set2Set goes further, it uses an LSTM to iteratively query the node set, producing richer graph representations than single-pass pooling.
TopKPool and SAGPool: Sparse Graph Pooling
Published:
Instead of soft cluster assignment (DiffPool), TopKPool and SAGPool select a subset of the most important nodes, producing a smaller but sparser graph at each level. Hard selection is scalable but requires careful score learning.
DiffPool: Learning Hierarchical Graph Pooling
Published:
DiffPool learns to hierarchically cluster nodes into super-nodes across layers, like a convolutional pyramid for graphs. Unlike flat global pooling, it captures multi-scale graph structure by differentiably assigning nodes to clusters.
Global Pooling in GNNs: Mean, Sum, and Max
Published:
To predict a property of an entire graph, node embeddings must be aggregated into a single vector. The choice of global pooling, mean, sum, or max, is not arbitrary: each has distinct expressive power and fits different tasks.
Sign Ambiguity in Laplacian Eigenvectors
Published:
Laplacian eigenvectors are only defined up to sign: if u is an eigenvector, so is -u. This seemingly minor issue creates a fundamental problem for learning with LapPE. Here is the problem, its consequences, and how SignNet solves it.
Structural vs Positional Encodings in Graphs
Published:
Positional encodings say where a node is in the graph. Structural encodings say what role it plays. They are complementary, and confusing them leads to poor design choices.
Shortest-Path Encodings for Graph Transformers
Published:
Shortest-path distances between nodes can be encoded as attention biases or node features, directly informing the model about graph proximity without requiring message passing.
Random Walk Positional Encodings
Published:
Random walk positional encodings encode each node’s structural context by computing the probability of returning to it from itself in k steps, a computationally efficient alternative to Laplacian eigenvectors with no sign ambiguity.
Laplacian Eigenvectors as Graph Positional Encodings
Published:
The k smallest eigenvectors of the graph Laplacian form a natural positional embedding space, the graph’s own coordinate system. They capture global structure, symmetry, and community membership.
Why GNNs Need Positional Encodings
Published:
Message-passing GNNs are permutation-equivariant by design, they cannot assign unique positions to nodes. Without positional encodings, symmetric nodes are indistinguishable. Here is why that matters and how to fix it.
Why Some Graphs Fool GNNs: The Structural Indistinguishability Problem
Published:
Certain graph structures are invisible to message-passing GNNs, not because of bad training, but because of fundamental mathematical limits. Two structurally distinct graphs can produce identical embeddings in any MPNN.
Depth in GNNs: Why Deeper Is Not Always Better
Published:
In Transformers, depth = expressiveness. In GNNs, depth = both expressiveness AND over-smoothing. The optimal GNN depth is rarely more than 3-4 layers, fundamentally different from the hundreds of layers in modern LLMs.
Over-smoothing vs Over-squashing: The Difference
Published:
Oversmoothing and oversquashing are both problems with deep GNNs, but they affect different nodes, have different causes, and require different fixes. Confusing them leads to applying the wrong solution.
Oversquashing: When Too Much Information Passes Through Bottlenecks
Published:
Oversquashing occurs when exponentially many node features must be compressed into a fixed-size embedding through a bottleneck edge. It is the reason GNNs struggle with long-range dependencies, not just oversmoothing.
Oversmoothing: When All Node Embeddings Become the Same
Published:
Stack enough GNN layers and all node embeddings converge to the same vector, making the model useless. Oversmoothing is not a training problem; it is a mathematical inevitability of iterated averaging.
The Weisfeiler-Lehman Test: How Powerful Are GNNs?
Published:
The 1-WL graph isomorphism test provides the exact upper bound on message-passing GNN expressivity. GIN achieves this bound. Any pair of graphs that 1-WL cannot distinguish cannot be distinguished by any MPNN.
MPNN: The General Message Passing Neural Network Framework
Published:
The MPNN framework (Gilmer et al., 2017) unifies GCN, GAT, GIN, GraphSAGE, and almost all spatial GNNs under one abstraction: message functions, aggregation, and update. Understanding MPNN means understanding the whole GNN family.
Graphormer: Transformers with Structural Biases for Graphs
Published:
Graphormer encodes graph structure directly into Transformer attention via three biases: node centrality, spatial encoding (shortest paths), and edge encoding. It won the OGB-LSC 2021 competition on molecular property prediction.
Graph Transformers: Bringing Attention to Graphs
Published:
Graph Transformers replace or augment local message passing with full pairwise attention, every node attends to every other node. This solves long-range dependencies and over-squashing at the cost of O(N²) computation.
APPNP: Personalized PageRank Meets Graph Neural Networks
Published:
APPNP decouples feature transformation from propagation. A neural network transforms features first; then Personalized PageRank propagates the result. This enables deep propagation without over-smoothing.
SGC: Simple Graph Convolution
Published:
SGC removes all nonlinearities between GCN layers and collapses the entire propagation into a single pre-computed matrix power. Surprisingly, it matches GCN on most benchmarks, revealing that nonlinearities between layers may be unnecessary.
ChebNet: Spectral Graph Convolutions via Chebyshev Polynomials
Published:
| ChebNet avoids the expensive full eigendecomposition by approximating spectral filters with Chebyshev polynomials, achieving O( | E | ) computation and spatial locality without sacrificing expressiveness. |
Graph Fourier Transform: The Spectral View of Graphs
Published:
The Graph Fourier Transform decomposes a signal on a graph into frequency components using the Laplacian’s eigenvectors. This spectral view is the mathematical foundation behind spectral GNNs like ChebNet and GCN.
Graph Tasks: Node, Edge, and Graph-Level Prediction
Published:
GNNs can predict at three levels: properties of individual nodes, existence or type of edges, or properties of entire graphs. Each level requires a different output head and training setup.
Homophily vs Heterophily: When Neighbours Are Similar or Different
Published:
Most GNNs assume nearby nodes are similar, the homophily assumption. When this breaks (heterophilic graphs), standard message passing hurts performance. Understanding this distinction is essential for modern GNN design.
Directed, Undirected, Weighted, and Heterogeneous Graphs
Published:
Not all graphs are equal. Directed edges, edge weights, multiple node/edge types, each variant requires different GNN design choices.
What Is a Graph? Nodes, Edges, Features, and Labels
Published:
A graph is a set of nodes connected by edges, but the power of GNNs comes from the features attached to nodes and edges, and the labels we want to predict.
GIN: Graph Isomorphism Network, The Most Expressive GNN
Published:
How powerful can a GNN be? Xu et al. (2019) answered with a theoretical bound, and GIN is the architecture that achieves it. The secret: use sum aggregation and an MLP, not mean or max.
GraphSAGE: Inductive Learning on Large Graphs
Published:
GCN and GAT learn embeddings for fixed graphs, add a new node and you’re stuck. GraphSAGE (Hamilton et al., 2017) learns an aggregation function instead, so it can generate embeddings for entirely new nodes at inference time.
GAT: Graph Attention Networks
Published:
GCN assigns the same (degree-based) weight to every neighbour. GAT learns which neighbours actually matter, using attention coefficients on edges. More expressive, more interpretable.
GCN: Graph Convolutional Networks
Published:
GCN (Kipf & Welling, 2016) is the ‘hello world’ of GNNs. It simplifies spectral graph convolution into a single elegant layer: normalised neighbourhood averaging with a learned linear transformation.
Message Passing: The Universal GNN Framework
Published:
Every GNN, GCN, GAT, GraphSAGE, GIN, is a special case of message passing. Learn the three-step loop that defines them all: compute messages, aggregate, update.
The Graph Laplacian: Spectral Graph Theory Explained Simply
Published:
The Graph Laplacian is L = D - A. Its eigenvectors reveal the graph’s community structure; its eigenvalues tell you how well-connected the graph is. It’s also the mathematical bridge from spectral theory to GNNs like GCN.
The Graph Adjacency Matrix: A Graph in Matrix Form
Published:
Before understanding GNNs, you need to understand how graphs are represented mathematically. The adjacency matrix is the foundation, a simple grid that tells you which nodes are connected.
Graph Neural Networks: Learning on Graphs
Published:
Graphs are everywhere, molecules, social networks, road maps, knowledge bases. Graph Neural Networks learn from this relational structure by propagating information between connected nodes. Here’s the complete picture.
PartecipationsAndTalks
portfolio
Portfolio item number 1
Short description of portfolio item number 1
Portfolio item number 2
Short description of portfolio item number 2 
projects
AdaViT (Adaptive Vision Transformers)
Adaptive Vision Transformer with dynamic token sparsification and halting for efficient image classification.
ALPR: Automatic License Plate Recognition System
End-to-end real-time license plate detection and OCR pipeline with dual GUIs, one for security managers, one for drivers, built with PyTorch and Streamlit.
(AMR) Autonomous Mobile Robotics Cleaning Robot
Autonomous mobile robot for indoor cleaning with navigation, obstacle avoidance, and task orchestration.
Archaic Italian Modernization: Historical Italian Rewriting with Transformers and LLM Judges
Automatic modernization of 13th-15th century Italian into modern Italian using multilingual transformers, prompted LLMs, and LLM-as-a-judge evaluation.
AutoDriveCarSimulator: Autonomous Driving with CNNs
A simulation platform for developing and testing autonomous driving algorithms, using CNNs to map raw camera frames to steering and throttle commands.
BioHeat PINNs: Temperature Estimation with Bio-Heat Equation using Physics-Informed Neural Networks
Physics-Informed Neural Networks for real-time temperature estimation via the Pennes Bio-Heat Equation, supporting hyperthermia therapy control.
CareConnect: AI-Driven Hospital Environment Monitoring System
An AI system for querying hospital environmental sensor data via natural language chat, generating real-time graphs, and triggering automated actions via LangChain and MQTT.
Clustering-Deepening: Clustering Algorithms for Object Tracking & Image Segmentation
An in-depth study of clustering algorithms, from k-Means to DBSCAN and GMMs, applied to object tracking and image segmentation.
ElectricCompany-TicketingSystem: IT Infrastructure & Ticketing for an Electric Consultancy
End-to-end analysis and implementation of IT infrastructure (disaster recovery, smart working, fleet management) and a full ticketing system for an electric consultancy firm.
EmailSpamDetector: Spam Detection with Bidirectional LSTMs
Classifies spam and ham emails using a Bidirectional LSTM, capturing both forward and backward temporal context in email text for high-accuracy filtering.
HelpDeskSystem: Web-Based Customer Support Platform
A full-stack web help desk for issue tracking and customer support, with ticket management, user authentication, and real-time status updates.
Home-Automation: Smart Home with IoT and Arduino
End-to-end smart home system, from a physical miniature house build to Arduino-powered sensors, automated routines, and a companion mobile app.
InstaSocial: Photo-Sharing Social Platform
A full-stack Instagram-like photo sharing app, upload, explore, like, and comment, built with Vue.js frontend, Go REST API, and Docker deployment.
Java-CategoryTheory: A Category Theory Library in Java
A Java library that models core Category Theory constructs, categories, functors, natural transformations, and demonstrates their practical role in software design.
MLPipelineOptimizationStudy: End-to-End ML Pipeline Exploration
A systematic exploration of ML pipeline optimisation, covering preprocessing, feature engineering, model selection, and hyperparameter tuning across multiple algorithms.
MoonBot Navigation
Autonomous lunar rover navigation and interaction, winner of the TESP 2025 Competition.
NSIO: Neural Search Indexing Optimization
Optimising the Differentiable Search Index (DSI) with data augmentation and parameter-efficient fine-tuning (LoRA, QLoRA, AdaLoRA), evaluated on MS MARCO.
PC-Performance-Monitoring: Statistical Analysis & ML for System Metrics
Collects, analyses, and visualises PC performance metrics, then applies ML clustering to detect anomalies and performance degradation patterns.
QRCodeGenerator: Custom Static QR Code Generator
Generate static, unlimited-use QR codes with custom styles, embedded icons, and optional captions, entirely in Python.
RealTime-VLM: Real-Time Vision-Language Model Inference in the Browser
Browser-based real-time VLM inference, continuously captures webcam frames and feeds them to any OpenAI-compatible vision API with sub-second latency.
RoboMAT
MATLAB library for robotics simulations, kinematics, dynamics, control, and path planning.
RTAD5G: Real-Time Anomaly Detection in 5G Networks
A real-time anomaly detection pipeline for 5G network telemetry, developed in collaboration with Hewlett Packard Enterprise (HPE).
SkinMe: Deep Learning for Skin Disease Detection
A deep learning application that classifies skin conditions from dermoscopic images using CNNs and LSTMs, supporting early diagnosis assistance.
StyleAligned: Zero-Shot Style Alignment in Text-to-Image Generation
A zero-shot framework for consistent style transfer in text-to-image generation, using minimal shared attention to propagate a reference style without fine-tuning.
UniDrive: University Carpooling App
A Flutter/Dart mobile app that connects university students for ride-sharing, schedule, match, and split commutes within the campus community.
XGNNs: Model-level Explanation of Graph Neural Networks with RL through Graph Generation
Model-level explanations for GNNs via reinforcement-learned graph generation on MUTAG.
Z-SASLM: Zero-Shot Multi-Style Image Synthesis via Spherical Linear Interpolation
CVPR 2025 workshop paper, a zero-shot framework for smooth multi-style image synthesis using Spherical Linear Interpolation in the latent space of diffusion models.
publications
Heterogeneous Sheaf Neural Networks
Published in arXiv preprint arXiv:2409.08036, 2024
HetSheaf is a cellular-sheaf framework for heterogeneous graphs that encodes node and edge types through type-aware local feature spaces and learned restriction maps, without specialised architectural components. The companion SheafPool readout is invariant to basis changes and enables graph-level prediction. Gains of up to +2 pp on the Heterogeneous Graph Benchmark with up to 10× fewer parameters.
Recommended citation: Braithwaite, L.; Borgi, A.; Onorato, G.; Tarantelli, K.; Restuccia, F.; Silvestri, F.; Liò, P. (2024). "Heterogeneous Sheaf Neural Networks." arXiv:2409.08036.
Go to the Webpage | Download Paper | Download Bibtex
Z-SASLM: Zero-Shot Style-Aligned SLI Blending for Latent Manipulation
Published in CVPR (Computer Vision and Pattern Recognition) 2025 Workshops (Nashville, USA 🇺🇸), 2025
Z-SASLM introduces a zero-shot, fine-tuning-free approach to style alignment in diffusion models by blending multiple reference styles directly in latent space using spherical linear interpolation (SLI) with learned, context-aware weights. The method avoids model retraining, preserves content semantics, and yields consistent style transfer across prompts and seeds.
Recommended citation: Borgi, A.; Maiano, L.; Amerini, I. (2025). "Z-SASLM: Zero-Shot Style-Aligned SLI Blending for Latent Manipulation." CVPR 2025 Workshops.
Go to the Webpage | Download Paper | Download Poster | Download Bibtex | GitHub Code
Polynomial Neural Sheaf Diffusion: A Spectral Filtering Approach on Cellular Sheaves
Published in arXiv preprint arXiv:2512.00242, 2025
ArXiv preprint on Polynomial Neural Sheaf Diffusion: A Spectral Filtering Approach on Cellular Sheaves.
Recommended citation: Borgi, A.; Silvestri F.; Liò P. (2025). "Polynomial Neural Sheaf Diffusion: A Spectral Filtering Approach on Cellular Sheaves.
Go to the Webpage | Download Paper | Download Bibtex
Remember to Forget: Gated Adaptive Positional Encoding
Published in arXiv preprint arXiv:2605.10414, 2026
GAPE (Gated Adaptive Positional Encoding) addresses core limitations of RoPE in long-context language models. A content-aware bias is injected directly into attention logits while preserving rotary geometry: query-dependent and key-dependent gates suppress irrelevant distant tokens while protecting salient context, improving attention sharpness and long-context performance on retrieval and standard benchmarks.
Recommended citation: Ali, R.; Borgi, A.; Irwin, C.; Severino, M.; Liò, P. (2026). "Remember to Forget: Gated Adaptive Positional Encoding." arXiv:2605.10414.
Go to the Webpage | Download Paper | Download Bibtex
Equivariant Sheaf Neural Networks: Learning Geometric Transport on Graphs
Published in arXiv preprint arXiv:2608.28853, 2026
ESNN learns directed, matrix-valued transport between neighbouring vector features while preserving exact Euclidean equivariance. It captures radial and tangential geometric interactions, supports controlled symmetry relaxation, and improves performance across dynamics, mesh simulation, point clouds, and molecular-property prediction.
Recommended citation: Borgi, A.; Severino, M.; Silvestri, F.; Liò, P. (2026). "Equivariant Sheaf Neural Networks: Learning Geometric Transport on Graphs." arXiv:2608.28853.
Go to the Webpage | Download Paper | Download Bibtex
talks
Talk 1 on Relevant Topic in Your Field
Published:
This is a description of your talk, which is a markdown file that can be all markdown-ified like any other post. Yay markdown!
Conference Proceeding talk 3 on Relevant Topic in Your Field
Published:
This is a description of your conference proceedings talk, note the different field in type. You can put anything in this field.
teaching
Teaching experience 1
Undergraduate course, University 1, Department, 2014
This is a description of a teaching experience. You can use markdown like any other post.
Teaching experience 2
Workshop, University 1, Department, 2015
This is a description of a teaching experience. You can use markdown like any other post.
