The Standard Library Modules Worth Knowing Before You pip install

5 minute read

Published:

TL;DR: Before adding a dependency, check whether the standard library already does it. collections gives you a counting dict, a two-ended queue and a dict with defaults built in; itertools and functools give you lazy iteration and memoisation without hand-rolled loops; datetime, re, random, math/statistics cover the everyday cases directly; argparse and logging replace hand-parsed sys.argv and scattered print calls; pytest (a near-universal third-party addition, not stdlib, but assume it) and typing round out the tools that make code checkable before it runs.

Counting, queues, and dicts with defaults — collections

from collections import Counter, defaultdict, deque

words = "the quick brown fox the lazy dog the fox".split()
counts = Counter(words)
print(counts.most_common(2))   # -> [('the', 3), ('fox', 2)]

groups = defaultdict(list)
for w in words:
    groups[len(w)].append(w)
print(groups[3])   # -> ['the', 'fox', 'the', 'dog', 'the', 'fox']

dq = deque(maxlen=3)          # a bounded ring buffer
for x in range(5):
    dq.append(x)
print(dq)   # -> deque([2, 3, 4], maxlen=3)

Counter is a dict subclass that never raises KeyError for a missing key — it returns 0 — and adds most_common(). defaultdict calls its factory (list, int, set, or any zero-argument callable) the first time a key is missing, which removes the if key not in d: boilerplate covered in dicts and sets. deque is a doubly linked queue with O(1) appends and pops from either end — a plain list is O(n) at the front — and maxlen turns it into a fixed-size sliding window for free.

Iterating without writing the loop — itertools and functools

from itertools import chain, groupby, islice
from functools import lru_cache, partial, reduce

print(list(chain([1, 2], [3, 4])))          # -> [1, 2, 3, 4]
print(list(islice(range(1_000_000), 3)))    # -> [0, 1, 2] — stops early, no full list built

data = sorted([1, 1, 2, 2, 2, 3])
for key, group in groupby(data):
    print(key, list(group))
# -> 1 [1, 1]
# -> 2 [2, 2, 2]
# -> 3 [3]

@lru_cache(maxsize=None)
def fib(n):
    return n if n < 2 else fib(n - 1) + fib(n - 2)
print(fib(35))   # -> 9227465, instant on the second call, exponential-time recursion made linear

add_tax = partial(lambda price, rate: price * (1 + rate), rate=0.2)
print(add_tax(100))   # -> 120.0

print(reduce(lambda acc, x: acc * x, [1, 2, 3, 4], 1))   # -> 24

itertools functions return iterators, not lists — islice on a million-element range never materialises the million elements, echoing the laziness from comprehensions and generators. groupby only groups consecutive runs, so sort first if the groups are not already adjacent. lru_cache memoises a pure function’s return value by its arguments, turning naive recursive Fibonacci from exponential into linear time by never recomputing a call it has already seen. partial fixes some arguments of a callable ahead of time; reduce folds a sequence down to one value — useful, but a plain loop or sum/math.prod is usually more readable for the common cases.

Dates, text patterns, randomness, and numbers

from datetime import date, datetime, timedelta

today = date(2026, 11, 14)
print(today + timedelta(days=30))   # -> 2026-12-14
print(datetime.now().isoformat())   # -> e.g. 2026-11-14T09:12:03.481920
import re

m = re.search(r"(\d{4})-(\d{2})-(\d{2})", "seen on 2026-11-14")
print(m.group(0), m.group(1))   # -> 2026-11-14 2026
print(re.sub(r"\s+", " ", "too   many    spaces"))  # -> too many spaces
import random, math, statistics

random.seed(0)
print(random.choice(["a", "b", "c"]))   # -> 'b' (fixed by the seed)
print(math.sqrt(2), math.gcd(48, 18))   # -> 1.4142135623730951 6
print(statistics.mean([1, 2, 3, 4]), statistics.median([1, 2, 3, 4]))  # -> 2.5 2.5

Reach for re for patterns, not for structured formats — an email address or a URL has enough edge cases that a dedicated parser beats a regex; re is right for “does this line start with a timestamp”, not “is this a valid email”. random is not cryptographically secure — use the secrets module for tokens or passwords. statistics is exact and pure Python, appropriate for small datasets; anything large belongs in NumPy.

Talking to the system — os and sys

import os, sys

print(os.environ.get("HOME"))       # -> /Users/you  (or None if unset)
print(sys.argv)                     # -> ['script.py', 'arg1'] when run as `python script.py arg1`
print(sys.version_info[:2])         # -> (3, 13)
sys.exit(1)                         # exits the process with status 1

os covers the environment and process-level operations that pathlib does not (environment variables, process IDs, os.cpu_count()); sys covers the interpreter itself — arguments, the module search path, and exit codes that shells and CI systems check.

Parsing arguments and logging properly

import argparse

parser = argparse.ArgumentParser(description="Greet someone")
parser.add_argument("name")
parser.add_argument("--shout", action="store_true")
args = parser.parse_args(["ada", "--shout"])
greeting = f"hello, {args.name}"
print(greeting.upper() if args.shout else greeting)   # -> HELLO, ADA

Hand-parsing sys.argv with slicing breaks the moment a flag is optional or reordered; argparse generates --help, type-checks (type=int), and reports usage errors with the right exit code — all from a declarative spec.

import logging

logging.basicConfig(level=logging.INFO, format="%(levelname)s: %(message)s")
log = logging.getLogger(__name__)
log.info("starting job")
log.warning("retrying, attempt %d", 2)   # lazy formatting: only built if WARNING is enabled

logging beats print for anything beyond a throwaway script: it has severity levels you can filter by without editing call sites, timestamps and module names for free, and routes to a file or a log aggregator with the same call sites that write to the console today. print debugging has to be found and deleted; a logging.debug call can just be left there, silent, until the level is turned up.

Testing and types

# test_math_utils.py
def add(a, b):
    return a + b

def test_add():
    assert add(2, 3) == 5

def test_add_negative():
    assert add(-1, 1) == 0

pytest test_math_utils.py discovers any function named test_* with no boilerplate class or registration — a plain assert is the whole API, and a failure prints the actual values on both sides of the comparison automatically.

from typing import Optional

def greet(name: str, times: int = 1) -> list[str]:
    return [f"hello, {name}"] * times

def find(items: list[int], target: int) -> Optional[int]:
    return items.index(target) if target in items else None

Type hints are not enforced at runtime — greet(5) runs without complaint — but a checker like mypy or pyright catches the mismatch before the code ships, and the annotations double as documentation that cannot silently drift out of date the way a comment can.

Classic trap — reaching for a package before checking the standard library. A counting dict, a memoised function, and a CLI parser are each one import away, with no dependency to pin, audit, or update. The reverse trap also happens: treating re or hand-rolled datetime maths as sufficient for genuinely hard problems like timezone-aware scheduling or RFC-822 email parsing, where a battle-tested third-party library earns its place. The standard library is the first stop, not the only one.
Key Insight — most of these modules exist to replace a loop with a name. Counter replaces a manual for loop plus a defaultdict(int); groupby replaces a hand-written accumulator; lru_cache replaces a manual memoisation dict; reduce replaces an accumulator loop. None of them compute anything a loop could not — they just name the pattern, which makes the intent visible at the call site instead of buried in five lines of bookkeeping.

Recap

  • collections: Counter for tallies, defaultdict for grouping, deque for O(1) operations at either end.
  • itertools and functools replace loops with named, often lazy, patterns — lru_cache alone turns exponential recursion into linear time.
  • datetime, re, random, math/statistics cover everyday needs; use secrets instead of random for anything security-sensitive.
  • argparse beats hand-parsing sys.argv; logging beats print because it has levels, timestamps, and routable output.
  • Type hints are unchecked at runtime — they only help through an external checker like mypy, but they never silently rot the way a stale comment does.

Next: idiomatic and performant Python, where these building blocks get put to work at speed.

References

  1. Python documentation. collections — Container datatypes.
  2. Python documentation. itertools — Functions creating iterators for efficient looping and functools — Higher-order functions.
  3. Python documentation. datetime, re, random, statistics.
  4. Python documentation. argparse — Parser for command-line options and logging — Logging facility for Python.
  5. Python documentation. typing — Support for type hints.
  6. pytest documentation. Get Started.
  7. van Rossum, G., Lehtosalo, J., & Langa, Ł. PEP 484 — Type Hints, 2014.