Open, hands-on, built to scale
College of Computing, Data Science, and Society, UC Berkeley
2026-10-09
Open, hands-on, built to scale
Eric Van Dusen
College of Computing, Data Science, and Society, UC Berkeley
October 9, 2026
github.com/ericvd-ucb
1. What do we know works?
A first course in computing and statistics, for any major, in the browser, with real data from day one.
Largest course at UC Berkeley - up to 50% of undergraduates take it, and the gateway to Berkeley’s biggest major.
inferentialthinking.com
We didn’t make students learn all the machinery before we let them do data science.
1. What do we know works?
1. What do we know works?
JupyterHub means no installs and no laptop gap. Students meet ideas before they fight environments.
Connector courses (like Data 88E) provide domain context that makes methods stick.
Github Deployment, Open Stack of tools, Open textbook, notebooks, and autograding are what let it travel.
1. What do we know works? How to get 25 community colleges to adopt it?
How much of this can we do for AI education?
2. Can we adapt what we teach?
The notebook can become a laboratory for AI.
2. Can we adapt what we teach?
Student learns to consume AI
Student can inspect, modify, evaluate and build
2. Can we adapt what we teach?
You don’t need a frontier model for every experiment. A small open model is just a file students can count, inspect, and test.
from transformers import AutoTokenizer, AutoModelForCausalLM
name = "HuggingFaceTB/SmolLM2-135M" # a 270 MB file
tok = AutoTokenizer.from_pretrained(name)
lm = AutoModelForCausalLM.from_pretrained(name)
ids = tok("Amsterdam is famous for its", return_tensors="pt")
probs = lm(**ids).logits[0, -1].softmax(-1)
for p, i in zip(*probs.topk(5)):
print(tok.decode(i), f"{p:.0%}")2. Can we adapt what we teach?
The model downloads into a shared location on JupyterHub and runs in browser. Students can inspect the model, change parameters, and see what happens, e.g. with no setup.
Admin sets up environment, prof sets up curriculum, student explores.
Cloudbank Classroom: cloudbank.org/training/access-cloudbank-classroom
2. Can we adapt what we teach?
The whole notebook, and the model, runs on your own laptop. No account. No server. No cost.
Thanks to Jeremy Tuloup!
2. Can we adapt what we teach?
Traditional computing / ML
Modern AI workflow
2. Can we adapt what we teach?
A student can understand gradient descent but have no idea how an AI application actually works.
A student can now build an AI application without understanding gradient descent.
Both are educational problems.
Teach enough engineering that the system stops looking like magic.
3. What investments do we need?
1
JupyterLite
Small models run on the student’s own laptop. Free, private, zero setup.
2
JupyterHub
Shared model folders, curated environments, GPUs when a course needs them.
3
Jupyter as the front door
A small allocation for every student who needs one. Powerful but hard to use today, so invest in making it easy.
Teach the stack along the way: environments, containers, git, job schedulers, cloud costs. DevOps is part of AI literacy.
3. What investments do we need?
3. What investments do we need?
Data centres used about 485 TWh last year, roughly as much as Germany, and the IEA expects that to double by 2030. Communities are pushing back over power and water.
Laptops and campus servers use little power when idle and need little cooling. Open-weight models trail the frontier by only about four months.
Models on institutional servers keep control of the tools, the data and privacy. Subsidized AI-for-rent can’t last indefinitely.
“Educators can train students not only how to use AI models, but also how to determine when local models are preferable to cloud-based systems.”
Buhler, Pérez & Boettiger, “Scientists should shift away from AI data centres”, Nature 656, 296–298 (2026)
3. What investments do we need?
Europe doesn’t have to win the race for the biggest model.
It can compete differently, and it already has pieces at every rung.
Plus a mandate: AI literacy under Article 4 of the EU AI Act
| Rung | European pieces |
|---|---|
| Browser | QuantStack, JupyterLite |
| Campus hub | SURF and national research infrastructure |
| Public compute | EuroHPC |
| Models | GPT-NL, OpenEuroLLM |
Can Europe build a lighter-weight, open AI ecosystem rather than simply reproducing the hyperscalers?
3. What investments do we need?
Students across the disciplines
Open curriculum
Interactive notebooks
Open-source tooling
Open models
Public / shared compute
Goal 1: shared open curriculum that runs at multiple levels and can move between universities.
Goal 2: a small, fully open, multilingual teaching model with documented data.
We know how to make technical ideas accessible through exploration.
Students need to understand workflows, not only algorithms.
Universities should help students become builders, not simply customers.
We are still scaling data science. Next: give everyone the ability to look inside AI.
For the panel: who should own the infrastructure that teaches the next generation what AI is?
ericvd@berkeley.edu · github.com/ericvd-ucb
AI at Sciencepark | LAB42 | Eric Van Dusen, UC Berkeley