Implementation Guide

dirgo A fast, interactive terminal directory analyzer — how it is built, how it stays responsive, and how it uses Go's concurrency model to keep a single-threaded UI honest.

Go 1.25 Bubble Tea (MVU) Lipgloss Bubbles goroutines + semaphores atomics sync.Pool LRU cache macOS · Linux · Windows
New here? Start with the tour

This page is the reference — precise, dense, and organised for lookup. If you would rather learn by doing, the interactive tour lets you drive a working replica of the TUI, step through the message loop one frame at a time, run the concurrent scanner with adjustable core counts, and click through a map of the codebase. Come back here for the details.

01Overview

dirgo is a terminal UI that answers one question fast: what is eating my disk? It lists every child of a directory sorted by size, with proportional bars, recursive roll-ups, line counts, fuzzy search, and one-key navigation. (Try the replica if you have not seen it running.)

The engineering problem is not "walk a filesystem". It is: walking a filesystem is slow and unbounded, but a UI must repaint in milliseconds. Every design decision in the codebase falls out of that tension:

  • The UI never blocks. All filesystem work happens in goroutines detached from the render loop.
  • Work is parallel but bounded. Directory recursion fans out across cores, capped by a semaphore so the OS is not drowned in concurrent syscalls.
  • Progress is observable while it happens. Scanner goroutines publish counters through atomics that the UI samples on each spinner tick.
  • Nothing is scanned twice if it can be helped. An LRU cache plus a modtime check make re-navigation and refresh effectively free.
  • Rendering is allocation-averse. Reused builders, pre-computed styles, and slice reuse keep the per-frame cost flat.
9
Go source files, single main package
~1.6k
lines of application code
3
direct dependencies (all Charm)
1
goroutine that may touch Model
Mental model

Think of dirgo as a single-threaded state machine surrounded by a pool of disposable workers. Workers do I/O and return immutable snapshots; the state machine folds those snapshots into the model and repaints. There is no shared mutable state between the two halves except a handful of atomic counters.

02Stack & dependencies

DependencyRole in dirgo
charmbracelet/bubbleteaThe MVU runtime: owns the terminal, the event loop, the message queue, and the goroutine that executes tea.Cmd functions. Also provides tea.ExecProcess to hand the terminal to a child process (hex view) and give it back.
charmbracelet/bubblesOff-the-shelf sub-components: spinner (loading animation + tick messages), textinput (search and cd prompts), key (declarative binding matcher used by KeyMap).
charmbracelet/lipglossStyling and layout: ANSI colours, padding, borders, and — critically — lipgloss.Width(), which measures visual width so emoji and CJK filenames do not break column alignment.
stdlib only, elsewhereos, path/filepath, sync, sync/atomic, container/list, bytes, runtime/pprof, os/exec. No filesystem-walk library, no logging framework, no config parser.

Distribution is dual-track: a Go module installable with go install / Homebrew, and a thin Python wrapper package (dirgo_python) that downloads the correct prebuilt binary for the host platform on first run and execs it — so pip install dirgo works without a Go toolchain.

03Architecture: the MVU loop

dirgo follows The Elm Architecture as implemented by Bubble Tea. Three methods define the whole application surface:

MethodSignatureContract
Init() tea.CmdReturns the initial side effects. dirgo batches the first directory scan with the spinner tick.
Update(tea.Msg) (tea.Model, tea.Cmd)Pure-ish state transition. Takes a value receiver: it is handed a copy of the model and returns a new one. Never performs blocking I/O.
View() stringRenders the entire screen as one string. Also a value receiver — it cannot mutate anything the runtime will observe.

The value receiver is the load-bearing detail. Because Model is copied on every Update, and because the runtime serialises all message delivery onto one goroutine, no lock is ever needed to protect application state. Concurrency is pushed entirely into tea.Cmd closures, which communicate only by returning a message.

UI GOROUTINE — strictly serial, no locks Msg queue keys · mouse · resize Update(msg) value receiver View() string → diff → tty Model′ immutable copy repaint COMMAND GOROUTINES — parallel I/O scanDirectory fan-out, bounded by NumCPU smartRefreshCmd modtime probe → maybe rescan countLinesCmd selected file only countAllLinesCmd worker pool + mutex map trashCmd osascript / gio / powershell spinner.Tick tea.Cmd → new goroutine returns tea.Msg → pushed onto queue
One serial state machine · N disposable workers · communication only via messages
model.gothe entire concurrency contract, in three lines
func (m Model) Init() tea.Cmd {
	return tea.Batch(scanDirectory(m.path, m.scanProg), m.spinner.Tick)
}

tea.Batch runs both commands concurrently: the scan may take seconds, while the spinner ticks every ~100 ms and gives Update a chance to sample progress and repaint. Step through the full sequence — launch, scan, keypress, line count — one frame at a time in the interactive stepper.

04Component map

Nine files, one main package. For an interactive version with per-file symbols, dependency highlighting and "change this when you want to…" pointers, see the component explorer.

FileResponsibilityNotable internals
main.goEntry pointFlag parsing (--version, --profile), path validation via os.Stat, CPU profile lifecycle, program construction with tea.WithAltScreen() and tea.WithMouseCellMotion().
model.goState + Update + View~40 fields of UI state, the message switch, navigation (navigateIn/Up/To), external process launchers (open, Quick Look, hex view), and the trash implementations per OS.
scanner.goAll filesystem traversalscanDirectory, dirSizeRecursive, smartRefreshCmd, countLinesCmd, countAllLinesCmd, and the ScanProgress atomic counter block.
cache.goBounded LRUcontainer/list + map[string]*list.Element behind a sync.Mutex; O(1) get/put/evict.
entry.goDomain modelFileEntry, ViewFilter, size sorting, allocation-reusing filterEntriesInto, ASCII-fast fuzzyMatch.
render.goPresentationHeader (adaptive 1–2 lines), row layout arithmetic, footer key bar, centred help overlay.
styles.goThemePre-allocated Lipgloss styles — deliberately package-level so no style object is constructed per row.
keys.goInput bindingsDeclarative KeyMap struct; every binding carries its own help text, so help and behaviour cannot drift.
utils.goPrimitivesSize/count formatting, visual-width-aware truncation and padding, pooled-buffer line counting, binary sniffing, bar rendering.

05Data model

Everything the UI draws comes from one flat struct. There is no tree in memory — dirgo holds exactly one directory level at a time, with recursive aggregates rolled into each directory row. That is what keeps memory flat regardless of how deep the tree is.

entry.goFileEntry
type FileEntry struct {
	Name       string
	Size       int64     // for dirs: recursive total of all descendants
	IsDir      bool
	IsHidden   bool      // name starts with '.'
	IsBinary   bool      // decided by extension at scan time, cheap
	IsSymlink  bool
	LineCount  int       // 0 = unknown / binary / directory; filled in lazily
	Percentage float64   // share of the parent's total size, drives the bar
	ChildFiles int       // recursive descendant file count  (dirs only)
	ChildDirs  int       // recursive descendant dir count   (dirs only)
	ModTime    time.Time
}

Two derived views are maintained on the model:

  • entries — the full scan result, sorted by size descending (stable sort, so equal-size entries keep directory-listing order).
  • filtered — the subset currently visible after hidden toggle → type filter → fuzzy search → top-10 clamp. It is rebuilt on every keystroke while searching.
entry.go / model.gorebuilding `filtered` without allocating
func (m *Model) applyFilter() {
	search := m.searchInput.Value()
	// Reuse the underlying array to reduce GC pressure: [:0] keeps capacity.
	m.filtered = filterEntriesInto(m.filtered[:0], m.entries, m.showHidden, m.viewFilter, search)
	if m.topMode && len(m.filtered) > 10 {
		m.filtered = m.filtered[:10]
	}
}

Search uses a subsequence match rather than substring, so mgo finds model.go and rdm finds README.md. The comparison lowercases with a hand-rolled ASCII branch instead of strings.ToLower — no allocation, and it runs once per character per entry on every keystroke.

entry.goallocation-free fuzzy match
func fuzzyMatch(name, pattern string) bool {
	pi := 0
	for ni := 0; ni < len(name) && pi < len(pattern); ni++ {
		if toLower(name[ni]) == toLower(pattern[pi]) {
			pi++
		}
	}
	return pi == len(pattern)
}

06Message catalog

Messages are the only channel through which the outside world reaches application state. They are plain structs — no interfaces, no channels exposed to application code.

MessageProduced byEffect on the model
scanResultMsgscanDirectory goroutinePopulates entries/totals, caches by path, computes deep totals, clears loading, releases the ScanProgress block, restores a remembered cursor, and fires a line-count command for the newly selected file.
scanUpToDateMsgsmartRefreshCmdJust clears loading. The screen never flickers because nothing was replaced.
scanErrorMsgscan / hex viewStores the error for the centred error panel.
lineCountMsgcountLinesCmdPatches one entry in entries, filtered, and the cached copy — a write-back so re-navigation keeps the count.
batchLineCountMsgcountAllLinesCmdBulk-applies a map[name]int across all three of the same locations.
trashResultMsgtrashCmdRemoves the row locally, adjusts totals, recomputes every percentage, and invalidates the cache entry for the current path — deliberately avoiding a full rescan.
spinner.TickMsgbubbles/spinnerAdvances the animation and samples the atomic progress counters into display fields.
tea.KeyMsgruntimeRouted through a mode cascade: goto → search → help → normal.
tea.MouseMsgruntimeWheel up/down moves the cursor 3 rows and triggers a line count.
tea.WindowSizeMsgruntimeStores dimensions and rebuilds the cached separator string — see §11 for why this lives in Update and not View.

07The scanning pipeline

A scan of directory D produces one scanResultMsg. Internally it runs in five phases.

1 · Read os.ReadDir(D) one getdirentries 2 · Partition dirs vs files symlink resolution 3 · Dir sizes parallel recursion sem = min(NumCPU,16) 4 · File stat parallel if > 20 files sem = NumCPU 5 · Finalise percentages SortBySize PHASE 3 DETAIL — one goroutine per immediate subdirectory g₀ dir A g₁ dir B gₙ … blocked semaphore chan struct{} results[i] no mutex
Phases 1–2 are serial and cheap · phase 3 is where the wall-clock time lives

Phase 1–2 — one syscall, then partition

os.ReadDir returns []os.DirEntry from a single getdirentries syscall. Crucially, DirEntry.IsDir() and .Type() are answered from data the kernel already returned — no per-entry lstat. Symlinks get one extra os.Stat to decide whether they point at a directory (and should therefore be recursed into for sizing).

Phase 3 — recursive sizing, in parallel

Each immediate subdirectory is sized by dirSizeRecursive, which is a deliberate reimplementation of filepath.WalkDir:

scanner.gowhy not filepath.WalkDir
// dirSizeRecursive computes total size, file count, and subdirectory count using
// os.ReadDir + manual recursion. More efficient than filepath.WalkDir because
// os.ReadDir uses a single getdirentries syscall per directory, and we only call
// Info() on files (not dirs) since we only need file sizes.
func dirSizeRecursive(path string, prog *ScanProgress) (size int64, files int, dirs int) {
	entries, err := os.ReadDir(path)
	if err != nil {
		return 0, 0, 0 // unreadable dir contributes zero; never aborts the scan
	}
	for _, e := range entries {
		if e.IsDir() {
			dirs++
			if prog != nil {
				prog.Dirs.Add(1)
			}
			s, f, d := dirSizeRecursive(filepath.Join(path, e.Name()), prog)
			size += s
			files += f
			dirs += d
		} else {
			files++
			if info, err := e.Info(); err == nil {
				size += info.Size()
				if prog != nil {
					prog.Files.Add(1)
					prog.Size.Add(info.Size())
				}
			}
		}
	}
	return size, files, dirs
}

Three things matter here:

  1. Errors are swallowed, not propagated. A permission-denied subtree contributes 0 and the scan continues. A disk analyzer that dies on /proc or a locked folder is useless.
  2. Recursion is sequential inside each goroutine. It does not spawn a goroutine per nested directory. See §8.2 for why that restraint is essential.
  3. Progress is published inline. The atomic adds happen at the leaves, so the counter climbs smoothly even on a single enormous subtree.

Phase 4–5 — stat, percentage, sort

Files are stat'ed in parallel only when there are more than 20 of them (§8.9). Then each entry's Percentage is computed against the directory total, and sort.SliceStable orders by size descending. The result is packed into one immutable scanResultMsg and returned — which is how it crosses the goroutine boundary back into the UI.

08Concurrency deep dive

This is the part of dirgo worth studying. The program is small, but it exercises six distinct concurrency patterns, each chosen because a cheaper one would not have worked — and, just as importantly, it declines to use concurrency in several places where it would have been a net loss.

The governing rule

Exactly one goroutine may touch Model. Everything else operates on values it owns exclusively, and publishes results by returning a message. If you find yourself wanting a mutex around model state, the design has been violated.

8.1 · Primitives in use

PrimitiveWhereWhy this one
goroutine + tea.Cmdevery I/O operationThe runtime already spawns a goroutine per command and funnels the return value into the message queue. Free structured concurrency with a serialisation point.
chan struct{} (semaphore)dir sizing, file stat, batch line countCaps in-flight syscalls. A buffered channel is the cheapest counting semaphore in Go — no allocation per acquire, and it blocks rather than spins.
sync.WaitGroupall three fan-outsJoin barrier. The parent must not read results until every writer has finished — Wait() provides the happens-before edge that makes those reads legal.
atomic.Int64ScanProgressMany writers, one reader, high frequency, and a torn read is harmless. A mutex here would serialise the hot leaf loop for no benefit.
sync.Mutexbatch line-count map; LRU cacheGo maps and container/list are not concurrency-safe, and both are touched under low contention — a mutex is simpler and faster than sharding or sync.Map at this scale.
sync.Pool64 KB line-count buffersThe batch counter can open hundreds of files; without pooling that is hundreds of 64 KB allocations churning the heap in a burst.
index-disjoint slice writesdir results, file stat resultsDistinct elements of a slice are distinct memory — concurrent writes to different indices need no synchronisation at all. Zero-cost "sharding".

8.2 · Bounded fan-out: the semaphore pattern

The core of phase 3. Read it closely — the ordering of wg.Add, go, and the semaphore acquire is intentional. (The scan simulator lets you watch this exact loop execute with adjustable core counts and work distributions.)

scanner.goparallel directory sizing
results := make([]dirResult, len(dirEntryIndices))
var wg sync.WaitGroup
sem := make(chan struct{}, minInt(runtime.NumCPU(), 16))

for ri, di := range dirEntryIndices {
	wg.Add(1)                                  // registered BEFORE the goroutine starts
	go func(resultIdx int, info dirInfo2) {
		defer wg.Done()
		sem <- struct{}{}                      // acquire — blocks when N are in flight
		defer func() { <-sem }()               // release

		dirPath := filepath.Join(absPath, info.name)
		size, files, dirs := dirSizeRecursive(dirPath, prog)
		results[resultIdx] = dirResult{index: info.index, size: size,
			childFiles: files, childDirs: dirs}
	}(ri, di)                                  // params copied — no loop-variable capture
}
wg.Wait()                                      // join: all writes to results happen-before here

for _, r := range results {
	entries[r.index].Size = r.size
	entries[r.index].ChildFiles = r.childFiles
	entries[r.index].ChildDirs = r.childDirs
	totalSize += r.size
}

Why the cap is min(NumCPU, 16)

Directory sizing is I/O bound with a CPU-bound decoding component. Two ceilings apply:

  • NumCPU — beyond that, additional goroutines only add scheduler churn for the parsing half of the work.
  • 16 — a hard cap for large machines. On a 64-core box, 64 concurrent directory walks would thrash the page cache and, on spinning or network-backed storage, degrade badly through seek contention. Filesystems do not scale linearly with concurrent readers.

Why goroutines are spawned eagerly but gated inside

All N goroutines are created immediately; the semaphore is acquired inside the goroutine body. A goroutine parked on a channel send costs ~2 KB of stack and no CPU, so for the realistic case (tens to low hundreds of subdirectories) this is cheaper and far simpler than a worker-pool-with-job-channel. The pattern is: unbounded goroutines, bounded work.

Why the recursion does not fan out further

dirSizeRecursive is plain recursion. If it spawned a goroutine per nested directory, a node_modules tree would create tens of thousands of goroutines, all queued on the same semaphore, with stack memory proportional to the tree size. Instead concurrency is applied at exactly one level — the widest, cheapest place to parallelise — and depth stays sequential.

Load-balancing consequence

Work is partitioned by top-level subdirectory, so a directory containing one giant node_modules and nine tiny folders gets no speedup: nine goroutines finish instantly and one carries the whole scan. This is the accepted trade-off for avoiding recursive fan-out. A work-stealing deque would fix it at a large complexity cost, and the progress counters keep the UI informative in the meantime.

8.3 · Lock-free result collection

Notice what is absent from the snippet above: there is no mutex around results, and no channel collecting outputs. Each goroutine receives a unique resultIdx and writes only to results[resultIdx].

This is safe because the Go memory model treats distinct slice elements as distinct memory locations. Concurrent writes to different indices of the same slice are not a data race. The slice header itself is never mutated (no append), and wg.Wait() establishes the happens-before edge that makes the subsequent sequential read of every element well-defined.

chosenindex write · 0 sync
results[i] = dirResult{...}
// … wg.Wait() …
for _, r := range results { … }
rejectedchannel collect · N allocs + select
ch := make(chan dirResult, n)
ch <- dirResult{...}
for r := range ch { … }   // + closer goroutine

The same technique appears in phase 4 for parallel file stat, where results are collected into []FileEntry by index and then appended to entries in order — which also makes the parallel and sequential paths produce byte-identical output.

8.4 · Atomic progress counters

The scan is opaque by nature: you cannot know the total ahead of time without doing the scan. dirgo solves this by streaming counters out of the workers while they run.

scanner.gothe shared progress block
// ScanProgress holds live progress counters updated by the scanner goroutines.
// Read via atomic loads from the UI goroutine (spinner tick).
type ScanProgress struct {
	Files atomic.Int64
	Dirs  atomic.Int64
	Size  atomic.Int64
}
model.gosampling from the UI goroutine
case spinner.TickMsg:
	if m.loading {
		var cmd tea.Cmd
		m.spinner, cmd = m.spinner.Update(msg)
		if m.scanProg != nil {
			m.scanProgFiles = m.scanProg.Files.Load()   // lock-free read
			m.scanProgDirs  = m.scanProg.Dirs.Load()
			m.scanProgSize  = m.scanProg.Size.Load()
		}
		cmds = append(cmds, cmd)
	}
	return m, tea.Batch(cmds...)

This is a textbook multi-writer / single-reader telemetry arrangement:

  • Writers — every goroutine in the phase-3 fan-out, calling Add(1) at leaf granularity, potentially millions of times.
  • Reader — the UI goroutine, ~10 times per second.
  • Consistency — the three counters are read independently, so a sample can show 1,000 files but a size that lags by a few entries. For a progress indicator this is invisible and irrelevant. Buying atomicity across all three (a mutex or a seqlock) would cost far more than the inconsistency does.

atomic.Int64 (the Go 1.19+ struct type) is used rather than the free functions atomic.AddInt64(&x, 1). The struct form makes misuse impossible — you cannot accidentally read the field non-atomically — and it embeds align64, guaranteeing correct alignment on 32-bit platforms where a misaligned 64-bit atomic would panic.

Lifecycle: how the pointer is retired safely

The model holds scanProg *ScanProgress. On scanResultMsg it does m.scanProg = nil. This looks dangerous — are workers still writing through that pointer? — but it is sound:

  1. The command closure captured its own copy of the pointer when the scan started.
  2. scanResultMsg is only produced after wg.Wait(), so every writer has already returned.
  3. Setting the model field to nil only drops the UI's reference; the object becomes garbage once the closure is done with it.
  4. Each new scan allocates a fresh &ScanProgress{}, so counters never need resetting and a late-finishing previous scan can never contaminate the new one's numbers.
Pattern worth stealing

Allocate a fresh telemetry block per operation instead of resetting a shared one. It removes an entire class of "whose counter is this?" bugs at the cost of one tiny allocation per scan.

8.5 · Mutex-guarded map: batch line counting

Pressing s counts lines in every visible text file. Here the output is keyed by filename rather than index, so a map is the natural container — and Go maps are not safe for concurrent writes.

scanner.gocountAllLinesCmd
counts := make(map[string]int)
var mu sync.Mutex
var wg sync.WaitGroup
sem := make(chan struct{}, runtime.NumCPU())

for _, e := range entries {
	if e.IsDir || e.IsBinary {
		continue                        // filtered before spawning — no wasted goroutine
	}
	wg.Add(1)
	go func(name string) {
		defer wg.Done()
		sem <- struct{}{}
		defer func() { <-sem }()

		lines, isBin, _ := countLines(filepath.Join(dir, name), 10*1024*1024)
		if !isBin && lines > 0 {
			mu.Lock()                   // critical section = one map assignment
			counts[name] = lines
			mu.Unlock()
		}
	}(e.Name)
}
wg.Wait()
return batchLineCountMsg{Counts: counts}

The critical section is a single map store, held for nanoseconds, while the work outside it — opening a file and scanning up to 10 MB — takes microseconds to milliseconds. Contention is negligible, so a plain sync.Mutex beats sync.Map (which is optimised for read-mostly workloads and would be slower for this write-heavy burst) and beats per-goroutine maps with a merge step.

Note the two filters that keep the pool efficient: directories and known-binary extensions are skipped before spawning, and countLines aborts on any file over 10 MB. A pathological directory cannot turn s into a multi-second freeze — and even if it did, it would freeze a worker, not the UI.

8.6 · Buffer pooling under burst load

utils.gosync.Pool + chunked counting
var bufPool = sync.Pool{
	New: func() interface{} { return make([]byte, 64*1024) },
}

func countLines(path string, maxSize int64) (int, bool, error) {
	// … stat guards: skip dirs, empty files, and anything over maxSize …
	f, err := os.Open(path)
	if err != nil {
		return 0, false, err
	}
	defer f.Close()

	buf := bufPool.Get().([]byte)
	defer bufPool.Put(buf)

	// First chunk doubles as the binary check — no seek, no second read.
	n, err := f.Read(buf)
	if n > 0 && isBinaryContent(buf[:n]) {
		return 0, true, nil
	}
	count := 0
	if n > 0 {
		count = bytes.Count(buf[:n], []byte{'\n'})
	}
	totalRead := int64(n)
	for totalRead < maxSize && err == nil {
		n, err = f.Read(buf)
		if n > 0 {
			count += bytes.Count(buf[:n], []byte{'\n'})
			totalRead += int64(n)
		}
	}
	return count, false, nil
}

Three compounding optimisations:

  • sync.Pool — with NumCPU workers, at most NumCPU buffers exist at once, and they are recycled across hundreds of files. sync.Pool is per-P internally, so Get/Put are usually uncontended pointer swaps.
  • bytes.Count over bufio.Scanner — counting newlines needs no line materialisation. bytes.Count is SIMD-accelerated on amd64/arm64; a Scanner would allocate and copy every line.
  • Binary detection fused into the first read — a NUL byte in the first 512 bytes (the same heuristic HTTP content sniffing uses) marks the file binary, and the function returns immediately without reading the rest.

8.7 · Subprocesses: reaping and terminal hand-off

dirgo shells out for four things: opening files, Quick Look previews, moving to trash, and hex dumps. Two distinct concurrency concerns arise.

Fire-and-forget, but never zombies

model.goopenPath
cmd.Stdout = nil
cmd.Stderr = nil
if err := cmd.Start(); err == nil {
	go cmd.Wait() // reap the child process to avoid zombies
}

Start() without a matching Wait() leaves a zombie entry in the process table for the lifetime of dirgo, plus leaked file descriptors and an os/exec bookkeeping goroutine. Since dirgo must not block waiting for a GUI app that might stay open for hours, the reaping Wait() is moved into a throwaway goroutine. It is one line, and it is the difference between a clean process table and a slow leak in a long-lived session.

Handing the terminal over, and taking it back

model.gohex view via tea.ExecProcess
shellCmd := fmt.Sprintf("%s %q | %s", hexCmd, targetPath, pager) // xxd file | less
c = exec.Command("sh", "-c", shellCmd)

return m, tea.ExecProcess(c, func(err error) tea.Msg {
	if err != nil {
		return scanErrorMsg{err: fmt.Errorf("hex view failed: %w", err)}
	}
	return nil
})

tea.ExecProcess is the one place where the runtime deliberately pauses the event loop: it releases raw mode, restores the terminal, runs the child in the foreground with inherited stdio, and on exit re-enters the alt screen, re-enables raw mode, forces a full repaint, and delivers the callback's message. Running less with a naive cmd.Run() instead would fight the TUI for the terminal and corrupt the display.

8.8 · Race-safety analysis

A useful way to audit a Go program: enumerate every piece of memory reachable by more than one goroutine, and name its protection.

Shared stateWritersReadersProtection
Model (all fields)UI goroutineUI goroutinenone needed value receiver + serial message delivery
results[i] (scan)N workers, disjoint indicesparent, after joinnone needed disjoint memory + WaitGroup edge
ScanProgressN workersUI goroutineatomic.Int64
counts mapN workersparent, after joinsync.Mutex
lruCacheUI goroutine only todayUI goroutinesync.Mutex — defensive, keeps future off-thread use safe
bufPoolN workersN workerssync.Pool (internally per-P)
Lipgloss style varsinit onlyUI goroutineimmutable after init

The test suite runs under the race detector by default, so violations of this table fail CI rather than surfacing as heisenbugs:

Makefilerace detection is the default, not an option
test:
	go test -race -v ./...

A closure detail that is easy to get wrong

Every fan-out in dirgo passes loop variables as goroutine parameters (}(ri, di), }(idx, de), }(e.Name)) rather than capturing them. Go 1.22 changed loop-variable scoping to make capture safe, and this module targets Go 1.25 — but passing explicitly documents the ownership transfer and keeps the code correct if it is ever copied into an older codebase.

8.9 · Tuning: where the thresholds come from

ConstantValueReasoning
Directory-sizing concurrencymin(NumCPU, 16)Balances core count against filesystem seek contention; the hard cap protects large-core machines from thrashing storage.
File-stat concurrencyNumCPUstat is short and cache-friendly; no benefit from oversubscription.
Parallel-stat threshold> 20 filesBelow this, goroutine creation + WaitGroup + semaphore overhead exceeds the syscall cost. dirgo keeps a literal sequential fallback branch for the small case.
Line-count concurrencyNumCPUMixed I/O and SIMD scanning; core count is the right ceiling.
Line-count size limit10 MBBounds worst-case time per file. Larger files simply report no count.
Read buffer64 KBComfortably above typical filesystem block/readahead sizes; large enough that syscall overhead is amortised, small enough to pool cheaply.
Binary sniff window512 bytesMatches the HTTP content-sniffing convention; a NUL in the first block is a reliable binary signal.
LRU capacity100 dirsDeep navigation sessions stay fully cached while memory remains bounded.
Cursor history500 pathsPrevents unbounded map growth; eviction exploits Go's randomised map iteration to drop an arbitrary key in O(1).

The sequential fallback, verbatim

scanner.goconcurrency is opt-in, by size
if len(fileEntries) > 20 {
	// … parallel stat with WaitGroup + semaphore …
} else {
	for _, de := range fileEntries {
		// … plain sequential stat …
	}
}
Principle

Duplicating a small loop to avoid parallelism overhead on small inputs is not premature optimisation — it is recognising that parallelism has a fixed cost, and that most directories on a real machine are small.

09Cache & smart refresh

The LRU

A classic intrusive LRU: a doubly linked list for recency order, a map for O(1) lookup, one mutex for both.

cache.go
type lruCache struct {
	maxEntries int
	ll         *list.List               // front = most recently used
	items      map[string]*list.Element // path → node
	mu         sync.Mutex
}

func (c *lruCache) Get(key string) (scanResultMsg, bool) {
	c.mu.Lock()
	defer c.mu.Unlock()
	if el, ok := c.items[key]; ok {
		c.ll.MoveToFront(el)             // O(1) pointer surgery, no reallocation
		return el.Value.(*cacheItem).value, true
	}
	return scanResultMsg{}, false
}

Cached values are whole scanResultMsg structs — which means a cache hit needs no recomputation at all. Get mutates recency, hence a full Lock rather than an RWMutex: an RWMutex would be strictly worse here because every read is also a write.

The cache is currently only touched from the UI goroutine, so the mutex is defensive. It costs an uncontended atomic compare-and-swap per access — nothing measurable — and it means moving cache warming into a background command later cannot introduce a race.

Write-back on lazy fields

Line counts arrive after the scan, so they must be written into all three places the data lives — entries, filtered, and the cached snapshot — or navigating away and back would lose them:

model.golineCountMsg write-back
if cached, ok := m.cache.Get(m.path); ok {
	for i := range cached.entries {
		if cached.entries[i].Name == msg.name {
			cached.entries[i].LineCount = msg.lines
			m.cache.Put(m.path, cached)
			break
		}
	}
}

Smart refresh

Pressing r does not blindly rescan. It first compares the directory's modtime against the cached value:

scanner.goone stat instead of a full walk
func smartRefreshCmd(path string, cached scanResultMsg, prog *ScanProgress) tea.Cmd {
	return func() tea.Msg {
		info, err := os.Stat(path)
		if err != nil {
			return scanDirectory(path, prog)()      // stat failed → full scan
		}
		if info.ModTime().Equal(cached.dirModTime) {
			return scanUpToDateMsg{path: path}      // ~microseconds, zero repaint churn
		}
		return scanDirectory(path, prog)()          // changed → full rescan
	}
}
Known limitation

A directory's mtime changes when entries are added or removed, not when a file inside it merely grows. So refreshing after appending to a log file will report "up to date" even though the size changed. This is a conscious trade: the check is essentially free and correct for the common case (files created, deleted, renamed). Deletion via d sidesteps it entirely by invalidating the cache entry directly.

Deletion without rescanning

When trashResultMsg arrives, dirgo patches state in place rather than re-walking the directory: it removes the row, decrements the totals, recomputes every percentage against the new total, recomputes deep totals, re-applies the filter, clamps the cursor into range, and deletes the now-stale cache entry. A rescan would be correct but would cost seconds on a large tree and would scroll the user's position away.

10Line counting & binary detection

Two paths, both flowing through the same countLines primitive from §8.6:

TriggerScopeConcurrencyGuard
Cursor moves onto a text fileone filesingle command goroutineSkipped entirely if the entry is a directory, has a binary extension, or already has a count — so scrolling fast does not queue redundant work.
s — count allevery visible text fileNumCPU workers + mutex mapDirectories and binary extensions filtered before spawn; 10 MB cap per file.
model.gothe "don't queue redundant work" guard
func (m Model) lineCountForSelected() tea.Cmd {
	if len(m.filtered) == 0 {
		return nil
	}
	e := m.filtered[m.cursor]
	if e.IsDir || e.IsBinary || e.LineCount > 0 {
		return nil            // returning a nil Cmd is a no-op for the runtime
	}
	return countLinesCmd(m.path, e.Name)
}

Binary detection is deliberately two-tier: a free extension check (isBinaryExt covers ~45 extensions across images, media, archives, executables, documents, fonts, bytecode and databases) decides whether to bother opening the file at all, and a content check (NUL byte in the first 512 bytes) catches extensionless binaries once the file is open.

11Render pipeline

View() rebuilds the entire screen as a single string on every message. Bubble Tea diffs that string against the previous frame and writes only what changed. This means View is extremely hot — potentially every keystroke, every spinner tick, every mouse wheel notch — so it is written to allocate as little as possible.

Only the visible window is rendered

model.gowindowed rendering
visibleEnd := minInt(m.offset+listHeight, len(m.filtered))
for i := m.offset; i < visibleEnd; i++ {
	m.viewBuf.WriteString(renderRow(m, i, m.filtered[i], i == m.cursor))
	m.viewBuf.WriteString("\n")
}

A directory with 100,000 entries costs the same per frame as one with 40 — only listHeight rows are ever formatted.

The separator-cache trick

model.gocaching in Update because View cannot
case tea.WindowSizeMsg:
	m.width = msg.Width
	m.height = msg.Height
	// Rebuild separator cache. View() has a value receiver, so anything it
	// memoises is discarded — the cache must be built where mutation persists.
	m.cachedSep = lipgloss.NewStyle().Foreground(colorDimmer).
		Width(m.width).Render(strings.Repeat("─", m.width))
	m.cachedSepWidth = m.width
	return m, nil

This is a neat illustration of an MVU consequence: because View receives a copy, memoisation inside it is pointless. Derived values must be computed in Update, where the returned model carries them forward.

Row layout arithmetic

Each row is a fixed 36 columns of chrome plus two elastic columns:

render.goelastic bar and name widths
// pointer(2) + num(4) + sp + bar + sp + pct(6) + sp + sep(1) + sp + icon(2)
//   + name + sp + size(9) + sp + meta(6)  ⇒ 36 fixed columns
const fixedNonBar = 36
barMaxWidth  := maxInt(6, minInt(28, (w-fixedNonBar-16)/3))
nameMaxWidth := maxInt(8, w-fixedNonBar-barMaxWidth)

The bar takes up to a third of the leftover space (clamped to 6–28 columns) and the filename gets the rest, so the layout degrades gracefully from an 80-column terminal to an ultrawide one.

Allocation discipline

  • Reused strings.Builder stored on the model, Reset() per frame — the backing array is allocated once and reused for the lifetime of the process.
  • Package-level styles in styles.go, including a pre-built map of bar-colour styles, so no Lipgloss style is constructed per row.
  • strconv over fmt.Sprintf in the row hot path — Sprintf allocates and reflects; strconv.FormatFloat does neither.
  • Concatenation over Sprintf when assembling row segments.
  • b.Grow() in barString, sized for UTF-8 block characters (█ is 3 bytes) so the builder never re-grows.

Unicode correctness

Column alignment uses lipgloss.Width(), not len() and not len([]rune). A filename containing an emoji or CJK characters occupies two terminal cells per glyph; measuring in bytes or runes would shear the columns. truncateStrVisual and padRightVisual both binary-walk down to a candidate that fits the true visual width.

Key handling is a cascade of modal guards, evaluated in priority order. Each mode returns early, so lower-priority handlers cannot see the event.

tea.KeyMsg arrives gotoMode? cd prompt searchMode? filter as you type helpMode? any key closes normal KeyMap match nav · filters · actions consumed consumed consumed each guard returns early — no fallthrough, no ambiguity

Cursor memory

Navigating into a directory and back out restores your position. Two mechanisms cooperate:

  • cursorHistory map[string]string — path → last selected entry name, capped at 500 with O(1) arbitrary eviction.
  • pendingCursorEntry — because a scan is asynchronous, the restore target is stashed on the model and applied when scanResultMsg lands. On a cache hit the restore happens immediately instead.

Names are stored rather than indices, so the cursor lands on the right entry even if sizes changed and the sort order shifted.

Navigation entry points

FunctionTriggerBehaviour
navigateIn→ l EnterDirectory → descend (cache hit renders instantly, marked ⚡cached). File → open with the OS default handler, unless it is a binary over 100 MB, which is refused with a pointer to hex view.
navigateUp← BackspaceAscends to the parent and highlights the directory you came from. No-op at filesystem root.
navigateTocFree-form path entry: expands ~, resolves relative paths against the current directory, filepath.Cleans, then validates with os.Stat before committing.

13Keybindings

All bindings live in one KeyMap struct built by DefaultKeyMap(). Each carries its keys and its help text together, so the footer, the help overlay, and the actual behaviour are driven from a single source.

KeyActionNotes
↑ / k · ↓ / jMove cursorTriggers a lazy line count for the newly selected file
wheelScroll3 rows per notch
PgUp / ^U · PgDn / ^DPagePage size derives from terminal height
g · GTop / bottom
→ / l / EnterOpenDescend into dir, or open file with OS handler
← / BackspaceParentRestores previous cursor position
SpaceQuick Lookqlmanage · xdg-open · start
rSmart refreshmodtime probe first
tTop-10 viewClamps filtered to 10 rows
oReveal in file manager
/Fuzzy searchSubsequence match, live
ccd to path~ expansion + validation
hToggle hiddenOn by default
fCycle filterall → dirs → files
sCount all linesParallel worker pool
xHex viewxxd/hexdump piped to $PAGER; 50 MB cap
dMove to trashRecoverable — never unlink
?Help overlay
EscCancelExits search / help / top-10
q / ^CQuit

14Performance techniques, consolidated

TechniqueLocationPayoff
os.ReadDir + manual recursionscannerOne getdirentries per directory instead of WalkDir's per-entry lstat; Info() is called on files only.
Bounded parallel recursionscannerNear-linear speedup on multi-core machines without saturating the I/O queue.
Adaptive parallelism (>20 files)scannerAvoids paying goroutine overhead on small directories.
Index-disjoint result writesscannerZero synchronisation on the collection path.
Atomic progress countersscanner ↔ modelLive feedback at effectively zero cost to the workers.
LRU of full scan resultscacheRe-navigation is instant and allocation-free.
Modtime-gated refreshscannerTurns a full walk into a single stat when nothing changed.
In-place delete patchingmodelAvoids a full rescan after d.
Slice reuse (filtered[:0])model / entryNo allocation per keystroke while searching.
Persistent strings.BuildermodelOne backing buffer for the whole session.
Pre-built Lipgloss stylesstyles / renderNo style construction per row.
strconv + concatenationrenderRemoves fmt reflection and allocation from the hot row path.
sync.Pool read buffersutilsBounded memory during batch line counting.
bytes.Count newline scanutilsSIMD-accelerated; no line materialisation.
Windowed row renderingmodelFrame cost is independent of directory size.
ASCII-only toLowerentryAllocation-free fuzzy matching per keystroke.
Built-in --profilemainCPU profile of a real session, not just a benchmark.

15Safety & hardening

dirgo executes external commands with user-controlled filenames, which is exactly where TUI tools tend to get command injection wrong.

RiskMitigation
AppleScript injection via filename (macOS trash)Backslashes and double quotes are escaped before interpolation into the osascript string.
PowerShell injection (Windows trash)Single quotes are doubled per PowerShell literal-string escaping rules.
Shell injection in hex viewThe path is interpolated with %q, producing a quoted, escaped Go string literal that the shell parses safely.
Direct exec pathsopen, xdg-open, gio, qlmanage are invoked via exec.Command with the path as a discrete argv element — no shell involved, so no quoting concerns at all.
Irreversible deletiond moves to the system trash. macOS delegates to Finder; Linux tries gio trash, then trash-put, then a hand-rolled XDG Trash implementation with .trashinfo metadata, collision suffixes, and a cross-device copy+remove fallback; Windows uses the VisualBasic FileSystem recycle-bin API.
Freezing on huge filesLayered caps: 100 MB for opening binaries, 50 MB for hex view, 10 MB for line counting.
Unreadable subtreesErrors from ReadDir/Info are absorbed; the scan always completes.
Unbounded memory growthBoth long-lived maps are capped: LRU at 100, cursor history at 500.
model.goescaping before AppleScript interpolation
// Escape backslashes and double quotes to prevent injection.
escaped := strings.ReplaceAll(path, `\`, `\\`)
escaped  = strings.ReplaceAll(escaped, `"`, `\"`)
script  := fmt.Sprintf(`tell application "Finder" to delete POSIX file "%s"`, escaped)

16Build, test, ship

Makefilethe whole workflow
make build          # go build with -s -w and version injected via -ldflags
make test           # go test -race -v ./...   (race detector always on)
make bench          # go test -bench=. -benchmem ./...
make profile-cpu    # benchmark → cpu.prof → pprof web UI on :8080
make profile-mem    # benchmark → mem.prof → pprof web UI on :8080
make release        # cross-compile darwin/linux (amd64+arm64) and windows/amd64

The version string is compiled in rather than hardcoded:

Makefile / main.go
LDFLAGS = -ldflags="-s -w -X main.version=$(VERSION)"

-s -w strip the symbol table and DWARF data for a smaller binary; -X main.version overwrites the var version = "dev" package variable at link time.

Runtime profiling

dirgo --profile /some/path wraps the entire interactive session in pprof.StartCPUProfile and writes cpu.prof on exit. This profiles the real workload — scanning, rendering, and input handling together — rather than a synthetic benchmark, and is the right tool for investigating goroutine scheduling and syscall behaviour under actual use.

Test surface

Four test files cover the pieces that are pure enough to assert on: cache_test.go (LRU semantics and eviction), scanner_test.go (scan correctness and benchmarks), render_test.go (layout arithmetic), utils_test.go (formatting, truncation, line counting). Running under -race means the concurrency patterns in §8 are continuously validated by the scanner tests.

Distribution matrix

ChannelMechanism
go installStandard module install from source.
brewCustom tap pointing at released binaries.
pip / uvdirgo_python resolves platform.system() + platform.machine() to a release asset, downloads and extracts it on first run into a package-local _bin/, sets the executable bit, and forwards argv.
GitHub ReleasesFive prebuilt targets produced by make release.

17Trade-offs & caveats

Honest accounting of where the current design gives something up.

AreaCurrent behaviourConsequence / possible direction
Scan cancellationCommands take no context.Context. Navigating away mid-scan leaves the old scan running to completion.Wasted work, and a late scanResultMsg sets m.path = msg.path — so a slow scan can pull the view back to a directory you already left. A generation counter on the model, or a context.WithCancel stored per scan, would fix both.
Uneven work distributionParallelism is one goroutine per top-level subdirectory.A single dominant subtree gets no parallelism at all. Recursive fan-out with a shared semaphore, or a work-stealing queue, would balance it — at meaningful complexity cost.
mtime-only change detectionSmart refresh compares directory mtime.Growth of an existing file inside the directory is missed. Hashing the entry list or comparing aggregate size would catch more, for more cost.
Hard links & sparse filesSizes come from Info().Size().Hard-linked files are counted once per link and sparse files report apparent rather than allocated size. Reading st_blocks via syscall.Stat_t would give true on-disk usage, at the price of platform-specific code.
Symlink loopsOnly immediate symlinks are resolved to decide dir-ness; recursion follows real directories.Adequate in practice, but a directory symlink cycle within a scanned subtree could in principle be traversed repeatedly. A visited-inode set would make it airtight.
No disk-persisted cacheCache lives only for the session.A deliberate simplicity choice, documented in the README: gob persistence was considered and dropped to keep behaviour deterministic and avoid stale-state bugs across sessions.
Single sort orderAlways size-descending.Sorting by name, mtime, or line count would need a sort-mode field and a keybinding, but no architectural change.
Takeaway

The reason dirgo stays fast is not any single trick — it is the discipline of keeping one goroutine authoritative over state, pushing every unbounded operation behind a message boundary, and bounding every resource: goroutine concurrency, cache size, map growth, file sizes, and per-frame work. Concurrency is used where it pays, and deliberately declined where it does not.