LLMs Unplugged

Understand AI by building it yourself

Acknowledgement of Country

EddieBenColeThomasHannahMaiaSungyeonSafiyaSuiLorennUshiniRob

activity

everyone stand up

sit down if you have

never used ChatGPT/Claude

sit down if you haven’t

used it in the last month

sit down if you haven’t

used it in the last week

sit down if you haven’t

used it in the last day

sit down if you haven’t

used it in the last hour

sit down if you haven’t

used it in the last 5 minutes

What is this about?

you’ll build your own language model—from scratch—with just a kids book, pen & paper, and some dice rolling

you’ll learn how language models work by spotting patterns in text to generate new text

AI ML LLMs Claude ChatGPT Gemini DeepSeek

Training

The recipe

walk through your text and tally up which tokens follow which in a grid

Example

“Run, Spot, run. See Spot run.”

after tidying up:

run , spot , run . see spot run .

The empty grid

run,spot,run.seespotrun.
Tokenrun,spot.see
run
,
spot
.
see

Training: run,

run,spot,run.seespotrun.
Tokenrun,spot.see
run|
,
spot
.
see

Training: ,spot

run,spot,run.seespotrun.
Tokenrun,spot.see
run|
,|
spot
.
see

Training: spot,

run,spot,run.seespotrun.
Tokenrun,spot.see
run|
,|
spot|
.
see

Training: ,run

run,spot,run.seespotrun.
Tokenrun,spot.see
run|
,||
spot|
.
see

Training: run.

run,spot,run.seespotrun.
Tokenrun,spot.see
run||
,||
spot|
.
see

Complete model

run,spot,run.seespotrun.
Tokenrun,spot.see
run|||
,||
spot||
.|
see|

A few more tips

the last token in one sentence is followed by the first token in the next (even over the page)

you can ignore quotation/speech marks (")

work in pairs and divide the labour however you like—but don’t neglect the tally marks

Training

10:00

The language of language models

  • model
  • token
  • weights

Generation

The recipe

use your grid to generate new text, rolling dice to choose each next word

Generation: start with see

see spot

one option — no roll needed

Tokenrun,spot.see
run|||
,||
spot||
.|
see|

Generation: from spot

seespot

2 options — roll the die!

Tokenrun,spot.see
run|||
,||
spot||
.|
see|

How the die chooses: spot

spot → ?  roll a d10

012345
run 1 tally
6789
, 1 tally

equal tallies → equal chances

How the die chooses: spot

spot → ?  roll a d10

012345
run 1 tally
6789
, 1 tally

equal tallies → equal chances

rolled 7,

Generation: spot,

seespot ,

rolled 7,

Tokenrun,spot.see
run|||
,||
spot||
.|
see|

Generation: from ,

seespot,

2 options — roll the die!

Tokenrun,spot.see
run|||
,||
spot||
.|
see|

Generation: ,run

seespot, run

rolled 2run

Tokenrun,spot.see
run|||
,||
spot||
.|
see|

Generation: from run

seespot,run

2 options — roll the die!

Tokenrun,spot.see
run|||
,||
spot||
.|
see|

How the die chooses: run

run → ?  roll a d10

0123
, 1 tally
456789
. 2 tallies

more tallies → more faces → more likely

How the die chooses: run

run → ?  roll a d10

0123
, 1 tally
456789
. 2 tallies

more tallies → more faces → more likely

rolled 6.

Generation: run.

seespot,run .

rolled 6.

Tokenrun,spot.see
run|||
,||
spot||
.|
see|

Generation: from .

seespot,run. see

one option — no roll needed

Tokenrun,spot.see
run|||
,||
spot||
.|
see|

Generation: back to see

seespot,run.see

one option — no roll needed

Tokenrun,spot.see
run|||
,||
spot||
.|
see|

Generated text

“see spot, run. see”

a new sentence — not in the training data!

Generation

10:00

Shareback

The language of language models

  • prompt
  • completion
  • hallucination

Pre-trained generation

Why do it again?

same recipe, more training text: yours learnt from a few sentences, this one learnt from a whole novel

first step on the road to the very big ones

The recipe

use the booklet to generate new text, rolling one or more d10s to choose each next word

Pre-trained: start with The

The cat
The ♦♦ 50|cat79|dog99|hat rolled 27
cat 5|sat9|ran
sat .
. The
ran .
dog 6|sat9|ran
hat 5|sat9|ran

Pre-trained: from cat

Thecat sat
The ♦♦ 50|cat79|dog99|hat
cat 5|sat9|ran rolled 4
sat .
. The
ran .
dog 6|sat9|ran
hat 5|sat9|ran

Pre-trained: from sat

Thecatsat .
The ♦♦ 50|cat79|dog99|hat
cat 5|sat9|ran
sat .
. The
ran .
dog 6|sat9|ran
hat 5|sat9|ran

Pre-trained: from .

Thecatsat. The
The ♦♦ 50|cat79|dog99|hat
cat 5|sat9|ran
sat .
. The
ran .
dog 6|sat9|ran
hat 5|sat9|ran

Pre-trained: from The again

Thecatsat.The dog
The ♦♦ 50|cat79|dog99|hat rolled 63
cat 5|sat9|ran
sat .
. The
ran .
dog 6|sat9|ran
hat 5|sat9|ran

Pre-trained: from dog

Thecatsat.Thedog
The ♦♦ 50|cat79|dog99|hat
cat 5|sat9|ran
sat .
. The
ran .
dog 6|sat9|ran
hat 5|sat9|ran

Pre-trained generation

8:00

Shareback

The language of language models

  • open weight model

Agentic AI

The recipe

an agent is a model that can call tools: pause generation, get information from “outside” the model, and continue

generate from your model as before, but every punctuation token triggers a tool call: a real text message to 3 friends/group chats

Worked example

your text so far is “the cat sat”, and the next dice roll gives you . as the next token—pause, that’s a tool call

text What comes next? "the cat sat..." to 3 friends/group chats; the first reply back might be “down by the river”

write down by the river, then the . you rolled anyway, and continue generating from .

Agentic AI

10:00

Shareback

The language of language models

  • agent/agentic loop
  • tool call

Scaling up

every word in the novelone novel, as a grid7,023 distinct words · 49 million cells · 0.08% ever get a tallyyou can't scale a table---it grows faster than you can fill it

Shared numbers

“it wrote something genuinely new—how?”

Tokenthecatsat.randog
the 0.03|| 0.46 0.04 0.02 0.04| 0.41
cat 0.04 0.01| 0.48 0.03| 0.42 0.02
sat 0.12 0.02 0.01|| 0.78 0.04 0.03
.|| 0.82 0.05 0.03 0.02 0.04 0.04
ran 0.11 0.03 0.03| 0.77 0.02 0.04
dog 0.05 0.02| 0.46 0.03 0.40 0.04

the dog ran isn’t in this text, so your grid rates it impossible

a real model rates it 0.40—not more numbers, shared numbers

Now turn everything up

with shared numbers, turning the dials up just keeps working

you ask your grid: “what is the capital of France?” “see spot, run. see”

it isn’t refusing to answer—answering isn’t a thing it does

the fix isn’t a bigger dial—it’s more training, on text made of conversations

still just tokens intokens out

Questions

is there a term you’ve heard—that we haven’t covered?

what new questions do you have about large language models?

how will this change the way you think about and use LLMs in the future?

Next sessions

  • Wednesday 16 September 12:00–14:00
  • Wednesday 25 November 16:00–18:00

Innovation Space, Birch Building, ANU

No public sessions are scheduled right now — get in touch to arrange one.

We’re here to help

this technology isn’t going away—and working out what to do about it is the interesting part

we work with organisations across all of it: how it actually works, how to get something useful out of it, and how to keep people at the centre of the decisions it opens up

lxconvenor.cybernetics@anu.edu.au