# Reasoning Modules — Tutorial
**Version:** alpha · API 1. This tutorial teaches what a reasoning module is, how to put
one to work on your own facts, and where its answers stop being trustworthy. The
[Reasoning Modules Reference](/ref/dke/reasoning-modules/) is the per-library
lookup — kinds, fields, derived fields and queries, module by module.
**This document as Markdown:** [dke-reasoning-modules-tutorial.md](/tutorial/dke/reasoning-modules/dke-reasoning-modules-tutorial.md) — the same text
this page is rendered from, for readers and tools that would rather have
the source than the page.
It assumes you have read enough of the
[DKE Python Tutorial](/tutorial/dke/python/) to know what a claim is.
---
## Chapter 1 — Two steps, and both of them are yours
Reasoning about anything takes two steps.
1. **Work out what is true, and record it.** You write claims to your store one
at a time, with a write verb — `update` if you mean to run the producer
again, which you usually do (Chapter 3).
2. **Write the DKE Python that reasons over those claims** — the classes, the
standing rules, the queries.
Analysing a codebase is exactly that: read your source, record what you found,
then compile the code that answers questions about it.
**A reasoning module is the second step, already written.** For some domains we
have written that code for you. You download the file — `code` is at
[dke.langsyn.com/source/code/latest/code.dpy](https://dke.langsyn.com/source/code/latest/code.dpy)
— compile it into your store once, and import it from then on:
```python
import code
```
That is the whole of what a reasoning module is. Nothing about it is privileged:
it is an ordinary `.dpy` file, the same kind you could have written, and one you
write yourself behaves identically. You can read every line of it before you
compile it, and you can edit it — it is yours once you have it.
The first step stays yours. **Nothing here knows anything about your subject
matter until you record it** — importing `code` does not scan your repository,
because it has no access to your repository. It installs the vocabulary and the
reasoning; you supply the facts.
---
## Chapter 2 — What you get when you import one
`import code` resolves against your store, and the only modules a store has are
the ones you compiled into it. So the module you get is the one you put there —
whether you downloaded it from us or wrote it yourself.
What arrives is four kinds of declaration:
- **Classes** — the kinds of thing the domain has, and the fields they carry.
`code` brings `Function`, `Class`, `Call`, `UnresolvedCall`, `Rationale`,
`Import`, `Dependency` and `Counterpart`.
- **Derived fields** — values the engine works out from the ones you wrote, and
keeps current as those change. You never write a derived field.
- **Standing rules** — conclusions that hold exactly while their condition holds
over your data, and withdraw themselves when it stops holding.
- **Queries** — `def`s you call by name, which answer over whatever facts you
have loaded.
Names are scoped to the module, so `code.Function` and `finance.Function` are
different kinds that never read each other's data. Two modules can each have a
`Function` without either knowing about the other.
---
## Chapter 3 — Loading facts
`code` reasons over data in your store. Data gets there the way all data does —
one claim at a time — and `code` neither knows nor cares what produced it. There
is no import step and no format to convert to. Analyse your source however you
like, then write what you found:
```python
import code
update(code.Function.parse.name, "parse", "my-analysis")
update(code.Function.parse.module, "compiler", "my-analysis")
update(code.Function.parse.lines, 42, "my-analysis")
update(code.Function.parse.branches, 8, "my-analysis")
update(code.Function.parse.loops, 4, "my-analysis")
update(code.Call.c1.caller, "main", "my-analysis")
update(code.Call.c1.callee, "parse", "my-analysis")
update(code.Function.parse.called, True, "my-analysis")
```
A third party can do the analysis, or a script you write, or your CI. The store
sees claims either way, and the provenance you pass — the third argument —
records where they came from.
**Why `update` and not `remember`.** Both write a claim, and neither is a
fallback for the other. `remember` asserts something new; `update` says *revise
my earlier reading*, and establishes a value on an empty cell so it also works
the first time. A producer you mean to run again — a script, a CI step, anything
pointed at a codebase that changes — means the second, so it should ask for it
by name. `remember` is permissive and a re-run does succeed either way; what you
lose by leaning on that is a reader of your program being able to see which you
meant. Chapter 8 says submitting a module again reconciles it; this is the same
idea one layer down, for the facts.
Supersession is per SOURCE. `update` replaces what *this* source said and leaves
every other source's claim standing, so re-running your producer never quietly
settles a disagreement someone else recorded.
### When a machine writes the facts
Writing the calls by hand is right for a handful of facts, and for learning what
the cells are. When a program produces them — a reader over your source tree, a
CI step, anything you point at data that changes — there is a package that
builds the program for you: **`dkepy`** (distribution `dkepy`, import name
`dke`), published at .
```python
from dke import FactProgram
prog = FactProgram("code", source="my-analysis")
prog.open()
prog.lines.append(prog.put("Function.parse.lines", "42"))
print(prog.render())
```
It emits the same program you would have written, and the output is text you can
read before you run it. What it saves you is the three things that go wrong
quietly when a machine writes facts: a subject that moves between runs, so a
re-run records a second set beside the first instead of revising it; the write
verb, which it always emits as `update`; and a value interpolated into the
program without escaping, where one quote in your data ends the string early and
changes what the REST of the program says.
It also carries readers for this module's own domain, `dke.ingest`, which read
source code and emit `code` facts — both a working producer and the worked
example. Neither the package nor its readers open a socket: producing facts
needs no key, so you can run one over your source without granting it anything,
and submitting the program is a separate, deliberate step.
That last line looks redundant next to the `Call` above it, and it is not. A
rule ranges over one kind, so "does any `Call` name me?" is a question nothing
can conclude from — the answer has to be recorded on the callee to be a fact
about it. Write both for every edge you record: the `Call` is what
`callers_of` and `impact_of` walk, and `called` is what stops `uncalled()`
from returning a function the store plainly shows being called.
### The one mistake that silently breaks everything
`code.Call.caller` and `code.Call.callee` hold the **subject** you recorded a
`Function` under — not its display name.
Above, the `Function` is recorded under the subject `parse`, so the `Call`
records `"parse"`. If you wrote a display name there instead, the queries would
still run and return answers: `uncalled()` compares those fields against
`Function` subjects directly, so every function would look uncalled, with
nothing in the output to tell you why. Getting this right is worth more than any
other single detail on this page.
The same care applies to `code.Function.test`, which names a test that reaches
the function. The `untested` rule reads its **absence**, so write it only where a
test really does reach — otherwise the rule falls silent everywhere instead of
where it should be silent.
---
## Chapter 4 — Asking
Once facts are in, the derived fields are already current — you do not run
anything to refresh them:
```python
import code
def review(fn: string):
who = code.callers_of(fn)
print(str(who.length) + " caller(s)")
for r in code.impact_of(fn):
print(" reaches " + r)
```
`impact_of` answers "what breaks if this changes" by following callers, and
theirs, to three hops. `uncalled()` reports functions with no recorded caller.
`critical()` reports the complicated and unreached — where a defect is likeliest
to survive. The full list is in the
[Reference](/ref/dke/reasoning-modules/#queries).
Those all start from a function you already suspect. `hubs()` is the one that
tells you **where to look first** — the functions most depended on, most first:
```python
import code
def where_to_look():
for h in code.hubs():
print(h)
```
It needs one thing from your producer: `callers` on each `Function`, counting
the **distinct** functions that call it. Record `callees` too if you want the
other direction.
```python
import code
update(code.Function.parse.callers, 12, "my-analysis")
update(code.Function.parse.callees, 3, "my-analysis")
```
Write it for every function including zero — unlike `called`, this one is read
by comparison rather than by absence, so a fan-in of zero is a fact worth
having. Record none of it and nothing breaks; `hubs()` simply stays empty.
Derived conclusions carry a **`derived` provenance** rather than a source, so
you can always tell a conclusion from something you were told.
---
## Chapter 5 — What the answers can and cannot mean
This chapter matters more than the query list, because the failure it describes
is quiet.
**A call graph derived from source is a lower bound.** Python decides calls as it
runs, so anything reading source sees the calls the source makes plain and no
others — a call through a variable, a handler table or `getattr` is invisible to
it. Every edge you record that way is real; edges are **missing**.
That matters most for `uncalled()`. A function reached only dynamically will
appear on it. Read it as *candidates to look at*, never a list to delete.
**You can measure how loose the bound is.** Wherever your producer sees a call
and cannot work out what it reaches, record it:
```python
import code
update(code.UnresolvedCall.u1.caller, "app.dispatch", "ingest")
update(code.UnresolvedCall.u1.site, "app.py:11", "ingest")
update(code.UnresolvedCall.u1.text, "handlers[name]", "ingest")
```
Then `dispatch_sites()` lists them, and you read it beside `uncalled()`:
```python
import code
def audit():
dead = code.uncalled()
sites = code.dispatch_sites()
print(str(dead.length) + " uncalled, " + str(sites.length) + " dispatch site(s)")
for s in sites:
print(" " + s)
```
Two dispatch sites and two hundred are very different reasons to believe the
same `uncalled()` list. Nothing on it is a defect on its own — a plugin table, a
callback and a decorator all work this way — it is where reading stopped telling
you what runs.
If your producer records none of these, nothing changes: `dispatch_sites()`
returns nothing and every other answer is what it was.
**`untested` inherits the same bound**, because which functions your tests reach
is worked out over the same graph — and reaching a function is not executing it,
since a branch may never be taken.
**And structure cannot tell deliberate from overlooked.** Two functions can look
identical to every query on this page — same complexity, neither reached by a
test — while one of them has a comment saying exactly why it is shaped that way
and why testing it is not worth the trouble, and the other says nothing.
`critical()` returns both. That is not a defect in `critical()`; it is that
everything it reads is structure, and the difference is not structural.
So record the reasons too. Most codebases already write them down:
```python
import code
update(code.Rationale.w1.subject, "app.gnarly", "ingest")
update(code.Rationale.w1.marker, "WHY", "ingest")
update(code.Rationale.w1.text, "the arms mirror the wire states", "ingest")
update(code.Rationale.w1.site, "app.py:40", "ingest")
update(code.Function.gnarly.explained, True, "ingest")
```
Then `rationale_for` reads them back, and `unexplained()` is `critical()` with
everything already accounted for taken out:
```python
import code
def triage():
for f in code.unexplained():
print(f)
for r in code.rationale_for(f):
print(" " + r)
```
A short `unexplained()` under a long `critical()` is a good result — it means
the risk is known and written down rather than absent.
**Which comments count is yours to decide**, the same way the counting rule in
Chapter 6 is, and the danger runs the same direction. A convention wide enough
to sweep ordinary comments makes every line in your tree a fact and every
function explained, which silences the rule instead of firing it. Pick something
an author had to type on purpose. Ours records `WHY:` and `NOTE:` at the start
of a comment, uppercase, and nothing else.
`explained` is read by its **absence**, like `test`. Write it only where a
reason really is recorded.
**There is no dependency age**, and that is the schema being honest rather than
an omission. Python's package metadata carries no release date, so an age could
only come from asking an index — which means sending your dependency list to a
third party. A field nothing can honestly fill is worse than no field, because it
reads as a capability.
---
## Chapter 6 — Comparable numbers
`code.Function.hotspot` fires at a complexity above 10. A fixed threshold over a
number *you* compute only means something if we agree on how it is computed — so
the Reference states the counting rule
([§2.3](/ref/dke/reasoning-modules/#how-to-count-branches-and-loops)), and it
is worth following even if you dislike the specific choices.
The rule is: one, plus one for each decision the function makes. `else`,
`default` and a bare catch-all add nothing, because they are where control lands
when none of the decisions above them fired.
**The list of what counts is closed**, and deliberately excludes some constructs
with a fair claim to being a decision — `assert` is the clearest. In some
languages it is a macro or a compiled-out call that a reader working from source
cannot reliably see, and a rule some producers can follow and others cannot is
the problem the rule exists to solve. A number everyone computes the same way is
worth more than a number that is arguably more faithful.
If your store holds more than one language, record `language` on each function.
`branches` is whatever *your* analysis counted, two producers can both be right
on different scales, and `complexity` compares them against one threshold either
way.
---
## Chapter 7 — The module in your store is yours
Compiling `code.dpy` into your store gives you the module as we publish it —
every class, every rule, every query. From that moment it is an ordinary module
in your store, and everything you can do to a module you can do to this one.
**Its rules are removable one at a time.** `forget_rule` on one of them removes
it, and nothing puts it back: your store holds what you compiled into it, and
compiling is something you do. If a rule does not suit your codebase, remove it
— or better, edit the line in your copy of the file and compile again, so the
file and the store agree.
**What you run is what you can read.** The version we publish is the version you
downloaded; the file is the whole of it; and if you have changed it, you know,
because you changed it. That is a stronger guarantee than the one this chapter
used to describe, which was that we would put a rule back if it went missing.
**Check the version before you assume you are current.** The
[Reference](/ref/dke/reasoning-modules/) states the version each module ships
at, and the download URL carries it — `/source/code/1.12.0/code.dpy` is that
exact release for as long as it is the one we publish, while
`/source/code/latest/code.dpy` is whatever is current. Publishing a new version replaces the old one at its
URL rather than keeping both, so link `latest` when a link has to keep working;
the file you downloaded is yours regardless, and nothing reaches into your store.
A version never changes without its content changing, so if the version you
downloaded is the version we publish, your copy is the current one.
---
## Chapter 8 — Writing your own
Nothing distinguishes yours from ours. Compile a file under a name and it
behaves exactly like one of ours:
```python
class Ledger:
balance: int
limit: int
over: bool = balance > limit
def exposed() -> list:
out = []
accounts = list_subjects(Ledger.balance)
for s in accounts:
c = current(Ledger.s.over)
match c:
case active_claim:
out = out + [s]
case empty:
pass
return out
```
Your modules are private to your store.
**Submitting a module again reconciles it.** A field whose value changed is
updated, one that did not is left alone, and an edited rule replaces the earlier
version of itself rather than joining it. So editing and re-submitting is the
ordinary way to change a module, and doing it twice costs nothing the second
time.
One thing to know when a module you import changes: a call is resolved when the
*calling* module is submitted, and the callee's body at that moment is the body
your call runs. If a module you depend on is updated afterwards, submit yours
again to move onto the new version.
---
## Chapter 9 — What it costs
Nothing beyond what you already pay. Reasoning modules are not a separate
subscription: importing one, and everything it then does, is billed by the same
per-operation meter as any other work you send. A module that reasons more costs
more, in proportion, and one you never call costs nothing.
---
## Where to go next
- The [Reasoning Modules Reference](/ref/dke/reasoning-modules/) lists every
kind, field, derived field and query, module by module.
- The [DKE Python Reference §9](/ref/dke/python/#modules) specifies module
semantics in general — resolution, namespacing, reconciliation.
- The [DKE Python Tutorial](/tutorial/dke/python/) teaches the language modules
are written in.