Why start with StatsBomb open data
StatsBomb publishes a set of free competitions with full event data: every pass, carry, pressure and shot, with coordinates and context. It is the best sandbox in football analytics — no API key, no terms negotiation, no scraping.
The full catalogue lives in the open-data repository on GitHub.
Step 1 — Pick a competition, load the events
The statsbombpy package wraps the raw JSON:
from statsbombpy import sb
competitions = sb.competitions()
matches = sb.matches(competition_id=11, season_id=90) # La Liga 2020/21
events = sb.events(match_id=matches.iloc[0].match_id)
If you prefer raw files — recommended for learning, because you see the schema — download events/{match_id}.json from the repository and parse it yourself with pandas.json_normalize.
Step 2 — Turn events into shots
shots = (
events[events.type == "Shot"]
.assign(
x=lambda df: df.location.apply(lambda loc: loc[0]),
y=lambda df: df.location.apply(lambda loc: loc[1]),
goal=lambda df: df.outcome == "Goal",
)
[["x", "y", "goal", "player", "shot_statsbomb_xg"]]
)
print(shots.groupby("player").agg(shots=("goal", "size"), goals=("goal", "sum")))
Two things to notice: coordinates are in a 120×80 pitch, and the events already carry shot_statsbomb_xg. Use that column to validate your own model later.
Step 3 — Draw the shot map
import mplsoccer
import matplotlib.pyplot as plt
pitch = mplsoccer.Pitch(pitch_type="statsbomb", pitch_color="#0d1117")
fig, ax = pitch.draw()
sizes = shots["shot_statsbomb_xg"] * 900 + 60
colors = shots["goal"].map({True: "#00e08a", False: "#4b5563"})
pitch.scatter(shots.x, shots.y, s=sizes, c=colors, alpha=0.85, ax=ax)
ax.set_title("Shots — size = xG, colour = goal", color="white")
plt.show()
Step 4 — Make it repeatable
A pipeline that downloads, transforms and plots in one script becomes unmanageable by the third match. Split it early:
ingest/— download raw JSON, store it unchanged (cache everything; the API is slow and rate-limited).transform/— parse events into tidy frames, one function per entity (shots,passes,pressures).analysis/— pure functions over the tidy frames: aggregations, rolling windows, comparisons.notebooks/— only exploration; the moment a transform is reused, it moves totransform/.
This is the same pattern every serious football-data project uses, and it is the one that survives contact with a second dataset.
Where to go next
- Features for xG — add pressure, body part, first-time shots; compare your AUC against
shot_statsbomb_xg. - Pass networks — aggregate passes by player pair and position.
- Pressing metrics — count pressures by zone and see which teams break structure first.
Pick one question, build the smallest pipeline that answers it, and publish what you learn. That loop — question, data, answer — is the actual job.