Registers of Code
Registers
We know library and application code have different traits. They might even benefit from different styles and language features. An enjoyable REPL experience may require even different things. What else is there?
Sociolinguistics gives us register shaped by 3 dimensions:
field - the social activity the text is embedded in
tenor/relationship between participants
mode (spoken vs. written, monologue vs. dialogue, planned or not…)
creative coding - high level, exploratory gluing lib abstractions together. You want the best error message possible here
- REPL - won’t be read again ever, just try stuff and see the result. Tool of Thought; code’s for thinking not delivery. Can be e.g. exploring data or how to use a library. Good naming, correctness and other virtues are wasted here
- scripts - just gluing components together, correctness=how well it does the job,, who cares about robustness? bugs are often lowcost
- prototypes - built to learn then discarded (…or get shipped), the goal is to answer a specific question
library programmer - should make library abstractions easy to use for creative coding, want to generate good quality error messages when a bad argument’s passed in (at a minimum just predicates and some help, but defensively validate at boundary). Strangers will read this. Correctness is extremely important. May seek to help shrink application code. Backwards compatability/contract. Naur’s Theory Building… code+docs must impart the theory to users. Bugs impact every user
application/system code - long lived, read by strangers (often coworkers), glued to specific (evolving?) requirements. Boundaries and integrations are hard. Richard Gabriel’s Habitability. Many local edits in the future, so consistency/comprehensibility locally is more useful than compacting the whole thing but requiring people to learn the application itself vs. the library’s approach
notebook/literate code - meant for reading somehow documenting a field/domain or showing a narrative, can have weird execution order… fragmentary, scripttish, but…
What are the dimensions here?
- who reads? noone (REPL experiments thrown out, devs read docs and hope they impart enough of the theory), coworkers (to edit it), strangers
- how long it lasts?
- stability (backwards compatability)
- defensveness (is the boundary a friendly thing to help the developer or to protect against missuse?)
- is the code a deliverable or its output?
| register | deliverable | reader | judge (1) | mind. knowledge (2) | activity (3) | bug hits |
|---|---|---|---|---|---|---|
| REPL | mental model, author's | none | author's eye, now | n/a | exploratory design | author, now |
| script | effect | runner | effect happened (P) | lang, libs | transcription | one run |
| prototype | answer | self, collaborators | hypothesis settled (P) | lang, libs | exploratory design | a belief |
| library surface (5) | contract | users, via docs | contract kept (S, drifts E) (4) | theory, contract | (rare) modification | every caller |
| application | capability in the world | maintainers | fits its world today (E) | theory | incrementation, modification | users |
| test (6) | signal | maintainers, learners | fails on behaviour change | theory, impl | incrementation | false confidence |
| notebook | mental model, learner's | learners | learner's model matches author's | lang, domain | exploration then explanation (7) | misunderstanding |
| tutorial | mental model, learner's, one point | learners | the point lands | n/a | transcription | misunderstanding |
| spec, model | definition | implementers | faithful to the intended system | domain | exploratory design | every implementation |
- Who or what decides. Lehman’s S/P/E
- To edit; docs should suffice for use
- Green’s activities
- Hyrum’s law: every observable behaviour becomes contract.
- Lib. impl. is app with frozen spec. See Parnas’s iplementation vs. interface
- Tests, property tests, proofs, formal methods. Proofs additionally require the statement to mean what you think; a wrong theorem misleads everywhere, a hollow test only at its points. Proofs have far higher viscosity than tests.
- Rule, Tabard and Hollan: the notebook switches register mid-file, and most stop at exploration.
- Stage (compile-time vs run-time, Lochbaum) is a property of the language, not of purpose, and cuts across every row.
Purpose could be deliverable, why you’re writing it or why you’re reading it, hm
- Script and prototype seem similar here, but differ on scale and field (effect vs. answer)
- prototype can be big or small (does this API return what I think? At what point do I care to make a hypothesis and iterate vs. subconciously just know what I’m searching for?)
- maybe the issue is what orrectness means: did the script’s side effect happen? do you know something and you can delete the experiment? (prototype with a wrong output can stil succeed)
dialogue is with the changing code, not with a person editing it directly.
User-facing vs contributor-facing library: a library is two registers glued. The public surface is an S-type spec (correct means matches the docs, contract enforced by validation at the boundary). The internals are application register (E-type, edited by people who have learned its theory, contract enforced by review and CI). So drop the distinction as a row and keep it as a footnote: library = spec surface + application interior.
Terseness (e.g. APLers praising notation as a tool of thought while outsiders call it unreadable) is diffuseness, leading to high visibility and juxtaposability K fits in a single screen, for those who hate to scroll. The dimensions are diferent depending on the user’s knowledge: low viscocity for knowers, high error-proneness and hard mental operations for those who don’t know it.
Terseness helps where the reader knows the language and looks at the whole program, but hurts when outsiders need to make small edits without knowing the Theory (Naur) (so terseness is dependent on activity and reader columns) so choose terseness depending on the register
Programmer Tasks
When a programmer is working, there are also various tasks:
- adding a feature - here a good view of the happy path is useful
- fixing a bug - useful to see error handing, validation, perhaps the abstraction layer right above
- getting to know the codebase/domain - a line from start to end, mostly across the happy path but also showing some important definitions or exemplary things is helpful?
This table is in the context of “code views”/what code is relevant for a programmer working on this task. Think of which part of the whole AST/code graph to show and which branch thereof to show.
The shape list is the set of values the shape column takes. I should have said that in one line above the table. Read a row as: for this errand, show this shape of slice, keep these branches inside it, plus the “also shown” items, done when the criterion holds. So orient = one path, happy branches only, plus the definitions along it; fix = a path first (repro to fault), then the neighbourhood of the fault, error and validation branches kept; learn a contract = a surface, no branches at all because no interior is shown; port = the whole. Only five shapes recur across all rows, which is the reason for listing them separately.
Shape column values:
- path: entry to effect, one line through the call graph
- neighbourhood: one unit and everything that touches it (callers, callees, tests, history)
- surface: one unit from outside only
- delta: two versions and what lies between
- whole: everything, in some linear order
Branches kept: happy, error and validation, the ones producing a specific effect, all. Outcome: mental model, change, verdict
| intent | outcome | shape | branches | also shown | hidden | done when |
|---|---|---|---|---|---|---|
| orient | model of the whole | one path | happy | the definitions the path passes through; one exemplary unit of each kind | error handling, config, alternatives | you can guess where a thing lives |
| explain | model of one behaviour | path, backwards from the effect | those producing the effect | conditions and values along the path | sibling branches | you can predict the behaviour on new input |
| learn API by experimenting (without reading implementation) | model of one unit from outside | surface | none | signature, docs, examples, tests read as examples | implementation | you can call it without reading it |
| add feature | change: new capability | path plus the exemplary sibling | happy | extension point: where the sibling is registered or wired | error paths, until the happy path works | sibling and new unit are parallel |
| fix bug | change: remove a behaviour | path from repro to fault, then neighbourhood of the fault | error, validation | the layer above (who set up the invariant that broke); history of the region | rest of the happy path | repro passes and a test pins it |
| refactor | change, behaviour held fixed | neighbourhood | all | every dependent; verification covering the unit; profile if optimising | unrelated paths | dependents behave as before |
| code review | verdict on a change | delta plus neighbourhood of changed code | all in the delta | rationale; tests touching it | rest | accept or reject |
| port | change: same behaviour, new substrate | whole, dependency order | all | nothing | nothing | parity |
diagnose incident logs, traces, error paths runtime traces, roles corrective exploratory understanding promote
optimize profiles, hot path, data layout runtime traces perfective modification
Role of a branch. Every if, cond, try splits code into branches. In
(cond (nil? x) (error "no x")
(invalid? x) (error "bad x")
:else (do-the-thing x))
the first two branches are validation, the third is the happy path. Your original sketch says: for adding a feature show the happy path, for fixing a bug show validation and error handling. To build either view you must know, for every branch, which of those it is. That per-branch label is what I mean by its role. It is not written down anywhere: not in the code, not in the symbol table, not in the call graph. It is only in the reader’s head, the same way register is only in the author’s head. That is why I put it next to the register tag: both are small classifications that decide which half of a text applies, and neither is recorded.
Register x Intent
The views compose from a small set of relations: call graph both ways, data flow, test-to-referent links, history with rationale, and a role per unit. Roles are two: position (entry, boundary) and kind (handler, migration, command, whatever the codebase’s repeated shapes are). Kind gives you the exemplary sibling and the wiring site in one relation, which is most of what add needs and what orient uses to pick its one-of-each.
Transitions:
- REPL to script
- script to cron job
- prototype to application
- application to library
- internals to surface (publishing)
Singeli makes another distinction is compile-time only:
embraces compile-time programming being a different kind of language than runtime, where other frameworks treat the differences as dirty secrets. So, at compile-time, you have immutable “tuples” and a functional each{} that maps over them. At run-time, you allocate memory and loop over it with load and store operations and go-to (though most of the time you’ll use the compile-time features to wrap these in something more structured). A major split is that compile-time code is dynamically typed. There’s no “what type is Type” theoretical jungle. The type i32 is a first-class value at compile-time which can be compared and inspected and constructed and so on, and you can check kind{i32} which gives ’type’ and even kind{kind{i32}} which is ‘symbol’—unlike types, kinds aren’t first-class values and we represent them as text. - Marshal Lochbaum
One observation that comes out of the “knowledge to edit” column: the arguments over terseness (Iverson’s notation vs “readable” code) are really arguments about register. Terse notation is best where the whole thing is read at once and by someone who already has the language (REPL, prototype, spec), and worst where strangers make local edits without the theory (application). Green’s cognitive dimensions formalize this: each notation property (viscosity, hidden dependencies, diffuseness) helps some activities and hurts others
Prior Work
- Brooks, Mythical Man-Month ch.1: program, programming product, programming system, programming systems product. Two axes (generality+docs+tests; integration with other components), each step costs roughly 3x. This is the oldest register taxonomy and it already contains your script vs library vs application split.
- Lehman 1980, S/P/E program types: S = correct against a stated spec, P = correct if the answer is acceptable in the real world, E = embedded in the world, so requirements move and the program must evolve. This is the dimension “what does correct mean”, which you were circling around with script vs prototype.
- Halliday’s field/tenor/mode you already use, but your table packs several things into each column. Tenor should be relationship, not just “who”: author-to-nobody, author-to-later-self, teacher-to-learner, provider-to-consumer, peer-to-peer.
- Sheil 1983 (“Power Tools for Programmers”) on exploratory programming, and Rule/Tabard/Hollan 2018 on notebooks: the notebook register is defined by a tension between exploration (REPL-like) and explanation (literate), and most notebooks are stuck halfway. So it’s not one register but a document that switches register mid-file.
- Naur (theory) and Gabriel (habitability) both say the same thing about application code: it carries a local theory that has to be learned before editing. That is the dimension you proposed at the end (local vs general knowledge), and I think it’s a better column than “edit pattern”.
- Swanson 1976 (corrective/adaptive/perfective, later preventive), Feathers (add feature, fix bug, improve design, optimize), Green & Petre’s cognitive-dimensions activities (incrementation, transcription, modification, exploratory design, searching, exploratory understanding), and Sillito et al. 2006 (44 questions programmers ask, in four groups: find a focus point, expand it, understand a subgraph, compare subgraphs). These give you the task list.