The Standard Library Modules Worth Knowing Before You pip install

5 minute read

Published:

TL;DR: Before adding a dependency, check whether the standard library already does it. collections gives you a counting dict, a two-ended queue and a dict with defaults built in; itertools and functools give you lazy iteration and memoisation without hand-rolled loops; datetime, re, random, math/statistics cover the everyday cases directly; argparse and logging replace hand-parsed sys.argv and scattered print calls; pytest (a near-universal third-party addition, not stdlib, but assume it) and typing round out the tools that make code checkable before it runs.

Counting, queues, and dicts with defaults, collections

from collections import Counter, defaultdict, deque

words = "the quick brown fox the lazy dog the fox".split()
counts = Counter(words)
print(counts.most_common(2))   # -> [('the', 3), ('fox', 2)]

groups = defaultdict(list)
for w in words:
    groups[len(w)].append(w)
print(groups[3])   # -> ['the', 'fox', 'the', 'dog', 'the', 'fox']

dq = deque(maxlen=3)          # a bounded ring buffer
for x in range(5):
    dq.append(x)
print(dq)   # -> deque([2, 3, 4], maxlen=3)

Counter is a dict subclass that never raises KeyError for a missing key, it returns 0, and adds most_common(). defaultdict calls its factory (list, int, set, or any zero-argument callable) the first time a key is missing, which removes the if key not in d: boilerplate covered in dicts and sets. deque is a doubly linked queue with O(1) appends and pops from either end, a plain list is O(n) at the front, and maxlen turns it into a fixed-size sliding window for free.

Iterating without writing the loop, itertools and functools

from itertools import chain, groupby, islice
from functools import lru_cache, partial, reduce

print(list(chain([1, 2], [3, 4])))          # -> [1, 2, 3, 4]
print(list(islice(range(1_000_000), 3)))    # -> [0, 1, 2], stops early, no full list built

data = sorted([1, 1, 2, 2, 2, 3])
for key, group in groupby(data):
    print(key, list(group))
# -> 1 [1, 1]
# -> 2 [2, 2, 2]
# -> 3 [3]

@lru_cache(maxsize=None)
def fib(n):
    return n if n < 2 else fib(n - 1) + fib(n - 2)
print(fib(35))   # -> 9227465, instant on the second call, exponential-time recursion made linear

add_tax = partial(lambda price, rate: price * (1 + rate), rate=0.2)
print(add_tax(100))   # -> 120.0

print(reduce(lambda acc, x: acc * x, [1, 2, 3, 4], 1))   # -> 24

itertools functions return iterators, not lists, islice on a million-element range never materialises the million elements, echoing the laziness from comprehensions and generators. groupby only groups consecutive runs, so sort first if the groups are not already adjacent. lru_cache memoises a pure function’s return value by its arguments, turning naive recursive Fibonacci from exponential into linear time by never recomputing a call it has already seen. partial fixes some arguments of a callable ahead of time; reduce folds a sequence down to one value, useful, but a plain loop or sum/math.prod is usually more readable for the common cases.

Dates, text patterns, randomness, and numbers

from datetime import date, datetime, timedelta

today = date(2026, 11, 14)
print(today + timedelta(days=30))   # -> 2026-12-14
print(datetime.now().isoformat())   # -> e.g. 2026-11-14T09:12:03.481920
import re

m = re.search(r"(\d{4})-(\d{2})-(\d{2})", "seen on 2026-11-14")
print(m.group(0), m.group(1))   # -> 2026-11-14 2026
print(re.sub(r"\s+", " ", "too   many    spaces"))  # -> too many spaces
import random, math, statistics

random.seed(0)
print(random.choice(["a", "b", "c"]))   # -> 'b' (fixed by the seed)
print(math.sqrt(2), math.gcd(48, 18))   # -> 1.4142135623730951 6
print(statistics.mean([1, 2, 3, 4]), statistics.median([1, 2, 3, 4]))  # -> 2.5 2.5

Reach for re for patterns, not for structured formats, an email address or a URL has enough edge cases that a dedicated parser beats a regex; re is right for “does this line start with a timestamp”, not “is this a valid email”. random is not cryptographically secure, use the secrets module for tokens or passwords. statistics is exact and pure Python, appropriate for small datasets; anything large belongs in NumPy.

Talking to the system, os and sys

import os, sys

print(os.environ.get("HOME"))       # -> /Users/you  (or None if unset)
print(sys.argv)                     # -> ['script.py', 'arg1'] when run as `python script.py arg1`
print(sys.version_info[:2])         # -> (3, 13)
sys.exit(1)                         # exits the process with status 1

os covers the environment and process-level operations that pathlib does not (environment variables, process IDs, os.cpu_count()); sys covers the interpreter itself, arguments, the module search path, and exit codes that shells and CI systems check.

Parsing arguments and logging properly

import argparse

parser = argparse.ArgumentParser(description="Greet someone")
parser.add_argument("name")
parser.add_argument("--shout", action="store_true")
args = parser.parse_args(["ada", "--shout"])
greeting = f"hello, {args.name}"
print(greeting.upper() if args.shout else greeting)   # -> HELLO, ADA

Hand-parsing sys.argv with slicing breaks the moment a flag is optional or reordered; argparse generates --help, type-checks (type=int), and reports usage errors with the right exit code, all from a declarative spec.

import logging

logging.basicConfig(level=logging.INFO, format="%(levelname)s: %(message)s")
log = logging.getLogger(__name__)
log.info("starting job")
log.warning("retrying, attempt %d", 2)   # lazy formatting: only built if WARNING is enabled

logging beats print for anything beyond a throwaway script: it has severity levels you can filter by without editing call sites, timestamps and module names for free, and routes to a file or a log aggregator with the same call sites that write to the console today. print debugging has to be found and deleted; a logging.debug call can just be left there, silent, until the level is turned up.

Testing and types

# test_math_utils.py
def add(a, b):
    return a + b

def test_add():
    assert add(2, 3) == 5

def test_add_negative():
    assert add(-1, 1) == 0

pytest test_math_utils.py discovers any function named test_* with no boilerplate class or registration, a plain assert is the whole API, and a failure prints the actual values on both sides of the comparison automatically.

from typing import Optional

def greet(name: str, times: int = 1) -> list[str]:
    return [f"hello, {name}"] * times

def find(items: list[int], target: int) -> Optional[int]:
    return items.index(target) if target in items else None

Type hints are not enforced at runtime, greet(5) runs without complaint, but a checker like mypy or pyright catches the mismatch before the code ships, and the annotations double as documentation that cannot silently drift out of date the way a comment can.

Classic trap, reaching for a package before checking the standard library. A counting dict, a memoised function, and a CLI parser are each one import away, with no dependency to pin, audit, or update. The reverse trap also happens: treating re or hand-rolled datetime maths as sufficient for genuinely hard problems like timezone-aware scheduling or RFC-822 email parsing, where a battle-tested third-party library earns its place. The standard library is the first stop, not the only one.
Key Insight, most of these modules exist to replace a loop with a name. Counter replaces a manual for loop plus a defaultdict(int); groupby replaces a hand-written accumulator; lru_cache replaces a manual memoisation dict; reduce replaces an accumulator loop. None of them compute anything a loop could not, they just name the pattern, which makes the intent visible at the call site instead of buried in five lines of bookkeeping.

Recap

  • collections: Counter for tallies, defaultdict for grouping, deque for O(1) operations at either end.
  • itertools and functools replace loops with named, often lazy, patterns, lru_cache alone turns exponential recursion into linear time.
  • datetime, re, random, math/statistics cover everyday needs; use secrets instead of random for anything security-sensitive.
  • argparse beats hand-parsing sys.argv; logging beats print because it has levels, timestamps, and routable output.
  • Type hints are unchecked at runtime, they only help through an external checker like mypy, but they never silently rot the way a stale comment does.

Next: idiomatic and performant Python, where these building blocks get put to work at speed.

References

  1. Python documentation. collections, Container datatypes.
  2. Python documentation. itertools, Functions creating iterators for efficient looping and functools, Higher-order functions.
  3. Python documentation. datetime, re, random, statistics.
  4. Python documentation. argparse, Parser for command-line options and logging, Logging facility for Python.
  5. Python documentation. typing, Support for type hints.
  6. pytest documentation. Get Started.
  7. van Rossum, G., Lehtosalo, J., & Langa, Ł. PEP 484, Type Hints, 2014.