Specula · User guide
Every mode, panel, setting, and CLI flag, on one page.
⌘O opens the standard file picker. Multiple-selection is allowed: the first file becomes the primary and the rest open as compare slots. Specula loads each file into RAM as a Float32 non-interleaved buffer. Pre-computed waveform peaks and clip regions are calculated during load.
Other ways to open files:
open path/to/file.wav from Terminal works for the same reason: it routes through Launch Services.Opening files always starts a clean comparison. If two files were already in compare, opening a single new file replaces the whole set; opening three new files replaces the whole set with those three. To append to the current comparison without starting fresh, use File → Add Files to Compare… (⇧⌘O) or the Compare button in the file dock.
A decoder under test emits raw samples. A logic analyzer dumps an I2S capture. A test harness writes a bare array of floats. None of these files carry a header, so nothing in them says how to read them. When a file can't be read as a known audio format (any extension, through every open path above), Specula asks for the format instead of refusing the file. Declare it once and Specula treats it like any other audio file from there: measure it, compare it, edit it, report on it.
The import sheet declares:
Rank likely formats decodes a few seconds under every byte format that fits and scores each by how much the result behaves like audio; the top candidates apply with a click. Ranking covers the byte axes only; the rate stays yours to declare. The preview draws the decoded waveform, shows the same plausibility score for the current declaration, and plays a few seconds. A wrong declaration is unmistakable: wrong byte order is full-scale noise, a wrong channel count combs the waveform.
After a raw file opens:
--raw-preset).Saving from Edit mode writes a normal WAV or CAF, as always: the declaration applies to reading only.
| Space | Play / pause |
| ⌘. | Stop and return to start |
| ← / → | Skip ±5 seconds |
| ⇧← / ⇧→ | Skip ±1 second |
| ⌘← | Jump to start |
| L | Toggle loop. Loops the current selection if one exists, otherwise the whole file. |
Next to Loop in the transport bar is a small ↺ Hold play-start toggle. Default: off. Pause stops playback at the current position (classic transport behaviour). Turn it on (the icon fills, accent-tinted) and Pause (Space or the play / pause toggle) returns the playhead to where Play was last started, or where you last manually placed the cursor with a waveform click. Tap Space again and you re-play the same passage from the same starting point. Useful when you're A/B'ing a passage by ear without wanting to re-position the cursor each time.
In compare mode the toggle also makes nudge buttons (±1 sample / ±1 ms) re-seek to that anchor, so the comparison point doesn't drift forward while you click alignment buttons.
Stop is unchanged either way: it always rewinds to the beginning of the file.
The floating info panel on the left holds the File section, the loudness and channel readouts, the mode picker, and the output selector. Hide it with the sidebar button at the far left of the transport bar (accent-tinted while the panel is showing), with ⌘I, or from View → Show Info Panel, to give the waveform and analysis views the full window width. The choice is remembered between launches.
Variable rate from 0.25× to 2×, with optional pitch preservation (timestretch).
| [ | Step rate down |
| ] | Step rate up |
| \ | Reset rate to 1× |
The Transport bar exposes the rate menu (0.25× · 0.5× · 0.75× · 1× · 1.25× · 1.5× · 2×) and the "preserve pitch" toggle.
The current output device shows as a compact button at the bottom of the info panel (the left sidebar). Click it (or press ⌥⌘O, or pick Window → Output Panel) to open the Output window. The window holds a device dropdown (with channel count next to each name; the Refresh Devices item lives at the bottom of the menu) above the per-channel routing matrix. Switching device rebuilds the audio engine. Specula is happy with any device, from built-in speakers to multichannel pro audio interfaces.
The Output window is hidden by default and remembers its size and position across launches. On macOS 26 the window itself uses Liquid Glass, so it reads as a floating utility panel over the main app (the same Liquid Glass opt-out toggle covers it). Routing details are in Channel routing & monitor mode.
Renders all channels of the loaded file simultaneously, scaled to fill the section height. Each channel uses its layout colour - mono, stereo, quad, 5.1, 7.1 each have a distinct palette.
A time ruler runs above the waveform. Click the ruler-mode toggle in the waveform control strip to swap between Time (HH:MM:SS.mmm) and Samples (frame count). Useful when you need to align two takes by sample.
The control strip below the waveform has an amplitude zoom control (−/+ buttons) - useful for inspecting low-level detail without it being lost in the noise floor. Default is 1× (full ±0 dBFS range); double-click the readout between the buttons to return to it.
During load, Specula scans every channel for sample values ≥ ±1.0. Each detected clip region is highlighted in red on the waveform and counted in the Channel Info sidebar.
The waveform draws horizontal grid lines at 0, −3, −6, −9, −12, −18, −24, −30, −36, −48, −60, −72, −84, −96 dBFS. Lines auto-cull when the channel is too short to label them legibly.
After file load, Silero VAD runs in the background and produces a per-100 ms speech / non-speech timeline. The waveform's control strip has an SG toggle - when enabled, detected speech blocks are shaded in teal. Useful for visually verifying the VAD result before relying on speech-gated loudness numbers.
Click a channel label in the Channels sidebar to mute that channel for playback. Useful for soloing one side of a stereo file or isolating a single channel of a multi-channel master.
A selection is a [start..end] range on the waveform. Selections drive offline analysis, edit-mode operations, and loop playback.
Typed values commit on Return or when you click anywhere outside the field, which holds for every text field in the app: the IN / OUT / LEN selection fields, violation thresholds, chapter names, routing gain cells, and popover parameters. Leaving a field also hands the keyboard back to the app, so single-key shortcuts work again straight away.
| Esc | Clear selection (waveform + FFT) |
| ⌘A | Select all |
| ⌘Return | Run offline selection analysis |
Most of the offline numbers populate the moment a file finishes loading, with no key press. A detached background pass over the whole file produces:
These numbers feed the info panel's Offline tab and both the PDF and JSON reports immediately after open. A typical music track finishes the pass in a few hundred milliseconds. Long-form audio takes proportionally longer.
⌘Return reruns the full analysis pipeline on the current selection rather than the whole file. Use it when you want metrics scoped to a specific region (a chorus, a problem section, a candidate edit point), or when you want the averaged FFT spectrum view.
Selection-scoped output:
A spinner in the Loudness sidebar indicates the analysis is in flight; results land when it completes (typically ~1-3 s for a multi-minute selection). Live measurement keeps running in the background regardless.
⌘Return is not required to populate true peak or per-channel stats on a freshly loaded file: those land in the Offline tab and the report automatically. Reach for ⌘Return when the question is about a region of the file, or when you want the averaged spectrum on the FFT view.
The FFT spectrum view has its own horizontal selection on the frequency axis - and it's an active playback filter, not a measurement readout. See FFT spectrum for details. Esc clears the FFT selection too.
Specula implements full ITU-R BS.1770-4 / EBU R128:
| Metric | Description |
|---|---|
| Momentary (M) | 400 ms window, updates 10×/sec |
| Short-term (S) | 3 s window |
| Integrated (I) | Whole-file integrated LUFS, dual-gated. Resets on file load and on play-from-zero. |
| LRA | The 10-95 percentile range of the gated short-term distribution |
| True Peak (dBTP) | Highest 4×-oversampled inter-sample peak. Turns red above −1 dBTP. |
| Sample Peak (dBFS) | Highest absolute sample value across the whole file. |
The sidebar's Analysis section, between Mode and Channels, holds a full-width [Live] [Offline] switch that selects the analysis source for the whole main view: the Loudness section below it and the spectrogram both follow it. Live shows the real-time measurement updating during playback. Offline shows the whole-file results from the load-time analysis pass, populated immediately after the file finishes loading; running ⌘Return on a selection narrows the same metrics to that selection.
The same rows appear in both modes from the chosen source: Integrated, LRA, Speech-Gated, Speech %, Momentary / Short-term (labelled Max Momentary / Max Short-term in Offline since they're whole-file or whole-selection maxima rather than instantaneous), True Peak, Sample Peak, Loudness Targets. Offline adds Time range, Dynamic Range, Headroom, Noise Floor (Podcast mode), and per-channel Peak / RMS / Crest / DC below. The choice persists across launches. Live measurement keeps running in the background regardless of mode.
The info panel has a top-level Mode picker right under the File section. Mode is the per-file analysis context - switch it to reframe the same file as a different deliverable. It drives both the Loudness Targets panel and which metric rows appear in the sidebar:
| Metric | Music | Podcast | VOD | Broadcast |
|---|---|---|---|---|
| Integrated, LRA, M/S, True/Sample Peak, Targets, Per-channel | ✓ | ✓ | ✓ | ✓ |
| Speech-Gated I, Speech % | - | ✓ | ✓ | ✓ |
| Noise Floor (ACX) | - | ✓ | - | - |
| Dynamic Range, Headroom | ✓ | - | - | ✓ |
Settings → Targets is a per-mode catalog editor - four sections, one per mode, each with its target toggles. Use it to decide which targets show up next to a loaded file when that mode is active.
When integrated LUFS is finite, Specula shows the Loudness Targets panel: one row per selected target with the verdict for the active mode. Pick the mode and up to four targets per mode in Settings → Targets.
| Mode | Treatment | Source |
|---|---|---|
| Music / Streaming | Penalty: shows the gain the platform will apply (e.g. Spotify −4 dB) | Integrated LUFS |
| Podcast / Spoken Word | Penalty for streaming podcast targets; compliance for ACX | Dialog-gated (streaming) / RMS (ACX) |
| VOD | Compliance: green dot or FAIL L / FAIL TP against the dialog-gated band | Dialog-gated LUFS |
| Broadcast | Compliance against the integrated band and the true-peak ceiling | Integrated LUFS |
Custom targets. Settings → Targets → Custom targets holds your own delivery specs: give a spec a name, a mode, a measurement source (integrated, dialog-gated, or un-weighted RMS, with the mode's natural source pre-selected), a reference level, a tolerance band, and a true-peak ceiling. A custom target joins its mode's tick list and works everywhere a built-in target does: the verdict badge, the hover breakdown, Match, Normalize to target, and the report exports. Custom targets are compliance targets; the penalty model describes how a specific platform's normalizer behaves, which is not something a house spec defines. A typical use: a client delivers to -18 LUFS ±1 under -1 dBTP, so that becomes a custom Podcast target sitting next to Apple Podcasts in the list.
Music targets report the gain the platform will apply rather than a hard verdict. A mix at −10 LUFS reads "Spotify −4 dB" (the platform will turn it down 4 dB to hit its −14 LUFS reference). A quiet mix at −18 LUFS reads "+4 dB" on symmetric platforms, "as-is" on asymmetric ones (see below). The dot is green when the gain is within ±1.5 dB of zero, yellow within ±4 dB, orange beyond. A separate triangle warns when the true peak breaches the platform's ceiling.
VOD and broadcast targets are hard pass/fail. The dot is green when the loudness sits within the tolerance band AND the true peak is at or below the ceiling. FAIL L means the loudness is outside the band; FAIL TP means the loudness is in band but the true peak is over the ceiling; FAIL S↑ means the maximum short-term loudness exceeded the spec's short-term ceiling (used by EBU R128 S1); FAIL NF means the noise floor exceeded the spec's RMS ceiling (used by ACX). Loudness fails take precedence; the loudness reason is reported first even when several criteria are off.
Dialog-gated targets (Netflix and peers, streaming podcasts, ATSC A/85) evaluate against speech-gated LUFS. Files with no detected speech show "no speech" rather than a misleading verdict. Target reference lines lead with dialog · whenever a target reads from the speech-gated path, so you can scan the measurement source without parsing LUFS-vs-LKFS unit suffixes.
ACX noise floor. Specula measures the dB RMS across all VAD-classified non-speech samples and shows it in the Speech-Gated section of the Info panel. ACX rejects audiobooks above −60 dB RMS, so the reading turns red when it exceeds that. The ACX preset in the loudness-targets panel automatically picks up the measurement and shows FAIL NF when the noise floor breaches the ceiling. Files with no non-speech samples show −∞. The noise floor is undefined in that case and the criterion passes trivially.
Each row in the Loudness Targets panel maps to a verdict object on the corresponding target in the JSON export, so downstream scripts can branch on pass / fail without recomputing against the target spec. The PDF and HTML reports render the same verdict as the status badge above each target card.
| Field | Type | When populated |
|---|---|---|
status | compliant · penalty · notCompliant · unavailable | Always. |
actualLUFS | number | The integrated or dialog-gated reading the verdict was computed against. |
penaltyGainDB | number | Streaming-platform targets (Music + streaming Podcast). |
penaltyApplies | boolean | Streaming-platform targets. False for "as-is" on asymmetric platforms. |
truePeakClean | boolean | Streaming-platform targets. False when the master's true peak breaches the platform's ceiling. |
nonComplianceReason | loudness · truePeak · shortTerm · noiseFloor | Hard-spec targets (VOD / Broadcast / ACX) when status is notCompliant. |
nonComplianceActual | number | The actual reading on the failing criterion. |
nonComplianceLimit | number | The ceiling or band edge the failing criterion exceeded. |
JSON consumers that ignore the field keep working. Useful for batch QC scripts, CI gates on render farms, and Shortcuts workflows that branch on whether a master will pass a given platform without re-running the math.
Different specs measure loudness differently and Specula computes both in parallel so the right one feeds each target.
The Loudness section shows both values side by side. Each target row's reference line is prefixed with dialog · when the target evaluates against the dialog-gated reading. Files that contain no detected speech show the dialog-gated reading as "-" and dialog-gated target verdicts as "no speech".
Not every streaming platform applies the gain it computes. Apple Music's Sound Check, YouTube, and Tidal only turn loud tracks down; they leave quieter-than-reference tracks at their original level. So if your master sits at −18 LUFS against Apple Music's −16 reference, the platform plays it at −18, not boosted to −16. Specula reflects this:
Settings → Targets surfaces the behaviour in each target's subtitle ("penalty, turns down only" vs "penalty"). The match-button tooltip on each row also states which measurement source the match will use.
Each target row has a small waveform-circle toggle on the right. Press it and Specula matches playback to that target:
target reference − source LUFS, where source is integrated LUFS for integrated targets (most music + broadcast) or dialog-gated LUFS for dialog-gated targets (Netflix and peers, ATSC A/85, streaming podcasts).dBTP ceiling. Reads the 4× oversampled inter-sample peak (ITU-R BS.1770 polyphase FIR), so the ceiling is enforced against the analog-reconstructed signal rather than just the raw sample magnitude. The limiter only acts when the gain would push peaks above the ceiling; below it the limiter is transparent.Hit play and you'll hear the file the way the platform plays it. Press the toggle again (or pick another target) to release. Nothing in the loaded buffer is modified - match is purely a playback gain stage with a limiter after it.
The match follows the loudness display mode: switching Live → Offline (or back) while a match is active recomputes the gain against the new source. If a matched target is no longer in the active mode's selected targets (you switched mode, or unticked it in Settings), the match auto-disables so you can never apply gain you can no longer see the verdict for.
5 ms playback latency while match is active. Dialog-gated targets can be matched too - the gain is computed against the speech-gated reading. Targets disable with a tooltip explaining which reading is missing (typically "no detected speech" for dialog-gated targets on speech-free files).
Defaults are marked ✓.
| Mode | Target | Reference | Tolerance | True-peak ceiling |
|---|---|---|---|---|
| Music | Spotify ✓ | −14 LUFS | penalty (symmetric) | −1 dBTP |
| Music | Apple Music ✓ | −16 LUFS | penalty (turns down only) | −1 dBTP |
| Music | YouTube ✓ | −14 LUFS | penalty (turns down only) | −1 dBTP |
| Music | Tidal ✓ | −14 LUFS | penalty (turns down only) | −1 dBTP |
| Music | Amazon Music | −14 LUFS | penalty (symmetric) | −2 dBTP |
| Music | Deezer | −15 LUFS | penalty (symmetric) | −1 dBTP |
| Music | SoundCloud | −14 LUFS | penalty (symmetric) | −1 dBTP |
| Music | AES TD1008 (streaming) | −18 LUFS | penalty (symmetric, vendor-neutral) | −1 dBTP |
| Podcast | Apple Podcasts ✓ | −16 LKFS dialog-gated | penalty | −1 dBTP |
| Podcast | Spotify (Podcast) ✓ | −14 LKFS dialog-gated | penalty | −1 dBTP |
| Podcast | ACX (Audiobook) | −20.5 dB RMS | ±2.5 dB | −3 dBTP (plus noise floor) |
| VOD | Netflix ✓ | −27 LKFS dialog-gated | ±2 LU | −2 dBTP |
| VOD | Prime Video ✓ | −27 LKFS dialog-gated | ±2 LU | −2 dBTP |
| VOD | Apple TV+ ✓ | −27 LKFS dialog-gated | ±2 LU | −2 dBTP |
| VOD | Disney+ | −27 LKFS dialog-gated | ±2 LU | −2 dBTP |
| VOD | Max ✓ | −27 LKFS dialog-gated | ±2 LU | −2 dBTP |
| Broadcast | EBU R128 (EU) ✓ | −23 LUFS | ±0.5 LU | −1 dBTP |
| Broadcast | EBU R128 (EU, live) | −23 LUFS | ±1 LU | −1 dBTP |
| Broadcast | EBU R128 S1 (EU, short-form) | −23 LUFS | ±0.2 LU + Max Short-term ≤ −18 LUFS | −1 dBTP |
| Broadcast | ATSC A/85 (CALM Act) ✓ | −24 LKFS dialog-gated | ±2 LU | −2 dBTP |
| Broadcast | ARIB TR-B32 (Japan) | −24 LUFS | ±1 LU | −1 dBTP |
| Broadcast | OP-59 (Australia) | −24 LUFS | ±1 LU | −1 dBTP |
Where Match previews a target by ear, Normalize commits to it. Integrated targets (music streaming, broadcast) and dialog-gated targets (VOD, streaming podcasts) carry a second button on the row, the up-to-line icon next to the Match toggle. Press it and Specula opens Edit mode with the Normalize section pre-filled to that target, then click Apply to normalize the loaded buffer.
The Normalize panel states its Basis so you always know which loudness is being targeted, and shows the move before you commit:
dBTP true-peak limiter ceiling, pre-filled to the target's.This is a normal edit: it lands on the 16-level undo stack, and exporting writes a new file rather than overwriting the source. Integrated targets re-measure and iterate to the reference; dialog-gated targets apply the uniform gain that lands the speech-gated reading on the reference. Both finish with the same two-pass true-peak limiter at the ceiling.
The ACX RMS target is the one exception: it stays compliance-only with no Normalize button. ACX is a hard delivery spec whose noise-floor requirement a gain can't satisfy (raising RMS into the window raises the noise floor with it), so normalizing its level alone could read as "now compliant" when it isn't. The button is also disabled (with a tooltip) until a loudness reading is available, and unavailable in multi-file Compare, since Edit and Compare are mutually exclusive.
Normalize is the destructive counterpart to Match: Match is a live playback gain you can release at any time, Normalize bakes the gain into a new file. Both target the same basis the row is scored against and share the same true-peak limiter math; the difference is whether the loaded buffer is rewritten.
Where Normalize fixes a loudness miss (FAIL L), Limit TP fixes a true-peak overage (FAIL TP). When the file's true peak sits above a target's ceiling, that row shows a Limit TP button (the waveform-pulse icon) alongside Normalize. Press it and Specula opens Edit mode with the Limit TP section pre-filled to that target's ceiling; click Apply to cap the inter-sample peaks. Loudness is left untouched, so this is the right fix when the level is already where you want it and only the peaks breach the ceiling, the one-click counterpart to the FAIL L → Normalize button.
The button is contextual: it appears only when there's a true peak over the ceiling to fix. That covers a hard FAIL TP (loudness in band, peak over), a loudness-compliant row carrying the orange over-ceiling triangle, and a music target whose peaks still clip its ceiling. It's offered on any target type, ACX included, because every target carries a true-peak ceiling and capping peaks never raises the noise floor (unlike Normalize). FAIL S↑ and FAIL NF get no fix button, since no peak limit corrects a short-term or noise-floor miss. Like Normalize, it lands on the undo stack, saves as a new file, and is unavailable in multi-file Compare.
Specula uses a two-tier voice activity detection pipeline to compute integrated loudness over speech-only blocks. This matters for podcast / dialog / film mixing where you want to ignore room tone, music beds, and silence.
Runs the Silero VAD model via FluidAudio. On file load, the entire file is downsampled to 16 kHz mono and classified. The resulting per-100 ms speech / non-speech history feeds the loudness measurement. Selection analysis re-runs Silero on just the selection.
Used when Silero is unavailable (it fails to load or returns an error). Four spectral features:
- is shown for content with < 0.5 % speech.−∞.
Three sliders bias the detector for the content you work with. Edits trigger a re-analysis in roughly half a second, and the new history feeds straight back into the loudness measurement.
For programme material where the VAD's false positives or misses matter, ⌘4 (or the Dialogue button on the transport bar) enters Dialogue authoring mode - one of the six modal authoring modes (Edit, Compare, Chapter, Dialogue, Spectral, Record) the app exposes, mutually exclusive with the other five. Every detected region is visible on the waveform while the mode is active (the SG overlay is forced on so you don't need to discover the toggle).
Region painting on the waveform:
Toolbar:
Settings → Speech slider edits always re-run detection. If you have manual edits when you move a slider, they're pushed onto the undo stack first so you can ⌘Z back to them.
While Dialogue mode is active, each region's waveform tint tells you where it came from:
Edited regions are auto-saved to <audio>.dlg.json next to the audio file (debounced 500 ms after the last edit). Sidecar files round-trip cleanly: re-opening the same audio re-applies your edits, and the schema is versioned (currently v1) so future format changes won't silently corrupt old work.
The regions aren't only a measurement input. Once they're right, switch to Edit mode (⌘1) and use Level Dialogue to act on them. It ducks the room tone in the non-speech regions (a downward expander with a fast attack and slow release, so word onsets and tails stay clean) and can optionally lift the dialogue to a target in the same pass. Gating the silence before applying the makeup gain is what lets dialogue normalize without the noise floor riding up with it, the fix for a common ACX rejection. When you lift to a target you also set a true-peak ceiling, and the makeup boost is capped there by the same limiter Normalize uses. It's a gate, not a denoiser, so it helps a marginal room tone rather than a hissy recording.
A dedicated section that plots per-100 ms momentary and short-term LUFS over time, with the time axis aligned to the waveform's. Toggle with ⌃2.
After offline selection analysis, the Loudness Curve view fills in for the analysed range; live playback adds new samples as they're measured.
The exported JSON and PDF reports plot three traces: Momentary (yellow), Short-term (blue), and a running Integrated trace (green) that converges to the file's final integrated reading at the tail. The third trace is the easiest way to see how a master "settled" against its target during the run.
Set programme-loudness thresholds and watch the waveform highlight every block that exceeds them. In the Loudness Curve / control strip:
Any 100 ms block above its threshold is highlighted on the waveform in real time during playback and after offline selection analysis. Useful for spotting the exact moment a master clips a streaming target.
The Violations toggle in Settings → Report controls whether they're included in JSON / PDF exports.
Real-time logarithmic-frequency spectrum from 20 Hz - 20 kHz. Toggle with ⌃4. Scroll to zoom the frequency axis around the cursor; double-click the spectrum (or click the reset button that appears in the control strip) to restore the full range.
Five windows, switchable from Settings → FFT / Analysis:
| Window | Side-lobe rejection | Best for |
|---|---|---|
| Hann | −31 dB | General purpose. Good balance. |
| Hamming | (slightly higher) | Slightly narrower main lobe than Hann; higher far side-lobes. |
| Blackman | −58 dB | Excellent side-lobe rejection. Slightly wider main lobe. |
| Blackman-Harris | −92 dB | Best side-lobe rejection. Widest main lobe. Closely-spaced harmonics. |
| Flat Top | (very wide main lobe) | Most accurate amplitude reading. Calibration & level measurement. |
Hover anywhere on the spectrum and the cursor reads out frequency in Hz, magnitude in dB, and the closest musical note + cents deviation. Toggle on/off via the NOTES button in the FFT control strip.
⌘-drag horizontally on the spectrum to capture a frequency range. The selection is an active playback filter, not a measurement readout - playback is routed through it so you can hear what's inside or outside the band.
A BAND / NOTCH toggle in the FFT control strip switches the filter mode:
The filter is live - adjust the selection edges and the audio responds in real time. Drag the body of the selection to slide the band across the spectrum while keeping its width.
Stereo files use a highpass + lowpass pair for BAND and a parametric −40 dB cut for NOTCH; multichannel files apply the same filtering per channel.
Esc clears the selection and bypasses the filter.
The FFT control strip exposes a smoothing slider - exponential averaging across frames. Higher values produce a calmer trace; lower values respond faster.
Pick which channel feeds the FFT (or use a downmix). Useful for inspecting one channel of a multichannel file in isolation.
After ⌘Return offline analysis, the FFT view switches to the averaged spectrum of the selection - a much smoother, more accurate read than the rolling live spectrum. Clearing the selection (Esc) returns to live mode.
On long selections the average samples up to 65,536 FFT windows spread evenly across the selection, which is statistically equivalent to averaging every overlapping window and much faster. Settings → FFT / Analysis → Average every window (exact) switches to accumulating every window instead; on hours-long audio at high overlap that can take minutes. Loudness measurements and the spectrogram are unaffected by this choice.
While playback is stopped, clicking anywhere on the waveform or spectrogram recomputes the FFT for that position and updates the spectrum view to match. The FFT window ends at the cursor, so the spectrum you see lines up with the spectrogram column under the playhead. Useful for inspecting a specific moment without scrubbing: click, read, click again, read. Works for keyboard skips too (←, →, ⇧←, ⇧→).
Rolling time-frequency display. Toggle with ⌃3. All settings live in Settings → Spectrogram.
The spectrogram follows the [Live] [Offline] switch in the sidebar's Analysis section, the same source switch the Loudness section uses. Live paints columns as the file plays. Offline shows the whole-file (or whole-selection) spectrogram produced by Analyse (⌘Return), or built automatically at load, so it is there before you press play.
Two sliders set the colour-mapping window:
Anything quieter than the floor is mapped to the bottom colour; anything louder than the ceiling clips to the top colour. Tighten the window to highlight subtle detail; widen it to see the full dynamic range.
Each spectrogram column is one FFT. Overlap controls how often that FFT is recomputed per second, which sets the time-axis density. Higher overlap means columns are closer together: less blocky at high zoom, smoother gradients in time.
| Overlap | Hop (fftSize=4096, 48 kHz) | Columns/sec |
|---|---|---|
| 75% (default) | 21.3 ms | ~47 |
| 87.5% | 10.7 ms | ~94 |
| 93.75% | 5.3 ms | ~188 |
| 96.875% | 2.7 ms | ~375 |
| 98.4375% | 1.3 ms | ~750 |
96.875% and 98.4375% can substantially lengthen the offline Analyse pass at large FFT sizes (16k or 32k), since the column count scales with the overlap factor. Use them when you need maximum smoothness; 75% is the standard pro-audio default.
The spectrogram follows the same [Live] [Offline] toggle the loudness section uses.
The two stores are independent: live playback never overwrites the offline analysis, and Analyse never disturbs the live history.
Settings → Spectrogram → Performance → Auto-compute spectrogram on load. Off by default. When on, the whole-file spectrogram is built during the load-time background pass instead of waiting for ⌘Return on a selection. Sub-second for typical music tracks; the cost scales with file length, so long-form audio takes proportionally longer.
The setting is independent of the FFT spectrum view, which still requires a selection because its averaging is selection-scoped. Turn it on for sessions that mainly work with shorter material (stereo mixes, single tracks) and want the spectrogram immediately on load; leave it off for long-form audio when you'd rather pick a region with ⌘Return.
Pick a single channel or a downmix. The spectrogram persists per-slot in compare mode, so switching slots restores that slot's spectrogram history (both live and offline).
Colormap, frequency scale, dB Floor, dB Ceiling, and FFT overlap all persist across launches. Tune the view once and Specula remembers.
Computes the correlation between the left and right channels over rolling windows, using the standard correlation-meter convention (zero-lag normalized product). Toggle with ⌃5.
The control strip has a Correlation / Width % toggle:
(1 − corr) / 2 × 100. 0 % = mono, 50 % = decorrelated, 100 % = out of phase.Optional shading of the correlation curve:
Mouse over any point on the curve to see a plain-English label ("narrow stereo image", "natural stereo", "potential mono sum issue", etc.) plus the raw correlation value.
Lissajous correlation meter - plots L on the X axis and R on the Y axis as a 2D scatter. Useful for catching out-of-phase content at a glance.
Hidden by default in the FFT panel area; can be enabled via Layout settings or a panel toggle.
Per-channel vertical bars in the Channels sidebar. Each bar shows:
Click a channel label to toggle mute on that channel for playback. The meter still shows the muted-source level - useful for inspecting muted channels.
The Routing panel lives in the dedicated Output window (Window → Output Panel, ⌥⌘O, or click the Output pill in the Transport bar). It is an N → M matrix that maps each input/source channel to one or more output channels with per-route gain trim. The Output window also holds the device selector, so the channel count the matrix targets and the device driving it are always co-located.
The window is hidden by default. Open it when you need to reroute, then close it again. Routing settings persist for the loaded file.
The routing section has a List / Matrix toggle. Same data, two presentations.
Your preference persists across launches.
A horizontally scrolling preset bar (visible in both views) applies a complete mapping in one click. Presets are filtered to the ones that apply to the current file and device channel counts.
Applying a preset is destructive: it overwrites the routing for every output channel, including any custom gain you set. Outputs the preset doesn't address are muted, so the preset fully replaces the routing.
Coefficients follow ITU-R BS.775-3. The −3 dB factor on summed channels is 1/sqrt(2), the equal-power gain for correlated channel summation. The Pro Logic II coefficients (−1.2 dB and −6.2 dB on surrounds, with phase inversion on the left total) match the Dolby specification.
Above the built-in row is a Saved row for your own routings.
Polarity inversion is a per-source-slot flag, separate from gain. List view has a small ϕ button next to each source's gain knob; Matrix view has the dedicated ϕ click zone in every routed cell (orange glyph when active). Needed for Pro Logic II Lt/Rt encoding and useful for quick A/B phase checks across any single route.
A 5.1 master with the wrong channel order is a routing problem, not an EQ problem. Specula lets you reroute channels for playback without modifying the file, so you can verify L / R / C / LFE / Ls / Rs ordering against your monitor system. The preset library covers the common stereo and surround downmix workflows in one click; the matrix view makes any custom mapping obvious at a glance.
Once a downmix, reorder, or custom matrix is right, Render to File… (in the Routing section header) renders it to a new audio file, so the routing is a deliverable, not just an audition. Every per-route gain, polarity, and Pro Logic II ±90° phase shift is applied exactly as it sounds on playback, computed offline.
Only the outputs you've actually routed are written. A 5.1 → stereo downmix on an 8-channel device renders a 2-channel file, not eight channels padded with silence (outputs with no source are dropped). The result writes as a new file (WAV, or CAF when the output channel layout can't be stored in WAV order) and never overwrites the source.
It bakes the routing only: the playback-side loudness match and the channel solo / mute are not baked, so the output is exactly what the matrix defines. To drop a channel from the bake, mute its route in the matrix (set its source to none) rather than relying on the monitor mute.
A PRE / POST toggle in the Channels sidebar header decides which signal feeds the loudness measurement:
For most workflows, PRE is what you want - the file's actual loudness. POST is useful when you've reduced channel count via the matrix (e.g. a 5.1 → stereo downmix) and need the loudness of the resulting stereo pair.
Load up to 6 files as slots A through F, switch between them instantly, level-match them by integrated LUFS, and compute sample-accurate residuals between any two.
Compare is a toggled mode (⌘2, or the Compare button in the transport bar), one of the six mutually exclusive modes (Edit, Compare, Chapter, Dialogue, Spectral, Record) Specula exposes. The mode carries its own sky-blue accent colour and its own dedicated toolbar: the compare toolbar sits below the transport with the per-slot offset / gain / polarity controls, the diff toggle, and the metrics popover - the same shape the other five mode toolbars use, so muscle memory carries between modes.
Three ways in:
Compare is mutually exclusive with Edit, Chapter, Dialogue, Spectral, and Record. Entering one mode exits the others; the Compare-mode-only controls (per-slot toolbar, diff toggle, slot-switch keys) disappear when you leave the mode and reappear when you return.
After the first file is open, drag additional files onto the window, drop them on the Compare toggle, or use the + button on the file dock. New files load in detached background tasks; their per-file metrics (integrated LUFS, true peak, sample peak, LRA, DR, speech stats, loudness curve, stereo width curve, speech history) are computed on load.
| 1 - 6 | Switch to slot A - F (only slots that exist). Bare digits, no modifier. |
Slot switching uses two paths. Fast path: when the new slot's audio format matches the current one, the engine performs an atomic buffer swap. Sub-millisecond. Slow path: when sample rate or channel count differs, the engine is rebuilt. ~50-100 ms.
Drag any slot chip in the file dock to a different position to reorder. The slot you drop on becomes the new home of the dragged slot, and the chip you targeted (plus everything between) shifts to make room. Dragging a slot to position A (the leftmost) promotes it to the reference file; existing residuals are invalidated at that point because they were computed against the previous A. Use this when your file came in from a Dock drop or Finder drag and Specula didn't pick the file you wanted as the primary. Finder serialises drag selections in display order, not click order, so the file you wanted as the reference may not be the one Specula picked.
Each slot's chip has an L toggle. When on, Specula applies a gain offset that brings the slot's integrated LUFS to match slot A's. Hear two masters at the same loudness instead of "the louder one always wins".
Each slot has a ϕ toggle that flips polarity. Useful for catching wiring errors and for diff workflows where flipping one signal makes the residual sit closer to zero.
The compare toolbar's offset (samples + ms) and gain (dB) readouts are editable TextFields, not just display labels. Click into one, type the value you want, press Return to commit (or Esc to cancel). Each field has a small ↺ reset button that appears when the value is non-zero - one click returns it to zero. Right-click either field for the same Reset entry.
The ±1 sample / ±1 ms / ±0.1 dB / ±1 dB nudge buttons sit next to the fields for fine adjustments by ear, folded together with level-match into one gain stage. The transport bar's ↺ Hold play-start toggle (full description in Loading & playback → Transport) extends to compare-mode nudges: with it on, every nudge re-seeks playback to where Play was last started, so the comparison point stays fixed while you click ±1 sample / ±1 ms.
Click the auto-align button on any non-A slot. Specula cross-correlates slot A and the active slot in two passes, a decimated pass to locate the offset coarsely and a full-rate pass to land on the exact sample, then applies the offset. Works on takes that are within ~10 seconds of each other.
Diff is a per-slot listening mode, not a separate slot. The active non-A slot's toolbar carries a prominent Listen to Diff toggle: press it and playback switches to the cached A − slot residual through the same waveform, spectrogram, and export paths the source uses. Press the toggle again (or pick another slot) to return to the source instantly.
Bare D (no modifier) toggles diff view on the active compare slot - a one-key A/B between source and residual that fits the muscle memory of bare-digit slot switching.
The residual:
The Waveform A Overlay toggle (Edit menu → "Waveform A Overlay" or in the Compare panel) draws slot A's waveform as a ghost behind whichever slot is active - direct visual A/B at the sample level.
The Compare metrics button (toolbar) opens a popover with a side-by-side metrics table for every loaded slot - integrated LUFS, true peak, sample peak, LRA, DR, speech stats, etc.
Non-destructive editing. Always saves as a new file - never overwrites the original.
Edit, Chapter, Dialogue, Spectral, and Record share one editing session: the edited audio, the undo history, and the unsaved-changes state follow you across all five modes, so a workflow that moves between them is continuous and a single Save at the end writes everything. Switching modes never discards edits; loading a different file, or Discard Edits, is what clears the session. Multi-file Compare is separate: entering it from unsaved edits prompts to Save or Discard first.
⌘1 toggles edit mode. Edit is one of the six mutually exclusive modes (Edit, Compare, Chapter, Dialogue, Spectral, Record); entering it exits any other active mode. The mode carries the amber accent colour - matching the EDIT badge that lights on the waveform and the tint of the edit toolbar.
Each operation pauses playback, captures an undo snapshot (a 16-level disk-backed stack), applies the change, and resumes.
| Operation | Notes |
|---|---|
| Trim | Crop to the current selection (keep the selection, discard the rest). |
| Cut | Remove the selection and join the audio on either side - the opposite of Trim. Use it to take a too-long pause or a flubbed line out of the middle. An equal-power crossfade at the join (default 10 ms, set it or 0 for a hard cut) hides the splice click; it's clamped to the audio kept on each side. Undoable, saves as a new file, and large files stream the result. |
| Add Silence | Insert at the playhead (no selection needed - park the cursor where the gap belongs, say 500 ms for a video edit, and Apply), insert before / after or replace a selection, or insert at the file start / end. Units: seconds or samples. |
| Change Level | Linear gain in dB, applied to the selection or the whole file. |
| Invert Phase | Multiply by −1. |
| Fade | Linear / Logarithmic / Equal-power × In / Out - 6 fade curves total. |
| Normalize Peak | Target dBFS. Applies to the selection when one is set, otherwise the whole file. |
| Normalize LUFS | Target LUFS + ceiling (interpreted as dBTP). Auto-applies a 5 ms lookahead true-peak limiter when normalization increases loudness so inter-sample peaks stay under the ceiling, not just sample peaks. Two-pass offline: pass 1 builds the gain envelope from the input's oversampled peaks; pass 2 catches any residual peaks the release stage might leave fractionally over the ceiling. Scopes to the selection when one is set (measure and gain over that range only), otherwise the whole file. |
| Limit TP | True-peak limiter only, ceiling in dBTP. Loudness otherwise unchanged. Use when integrated LUFS is already where you want it but inter-sample peaks need to be capped under a platform ceiling. Same DSP as the limiter Normalize LUFS uses; just skips the loudness measurement and gain stage. Scopes to the selection when one is set, otherwise the whole file. |
| Level Dialogue | A region-aware downward expander keyed to the detected speech regions (the ones you can tune in Dialogue mode). It ducks the room tone between phrases with a fast attack and slow release, so word onsets and tails stay clean, no clipped starts, no chopped ends. Tick Lift dialogue to target and it also gains the speech to a loudness target in the same pass, capping inter-sample peaks at a true-peak ceiling. Because the silence is gated before the makeup gain goes on, the floor ends up lower, not higher: that's the move a plain dialog-gated Normalize can't make, since a uniform gain lifts the floor along with the voice. It's a gate, not a denoiser, so it helps a marginal floor (a quiet room tone just under the gate), not a hissy recording. Undoable, saves as a new file. |
| Remove DC Offset | Subtracts each channel's mean. |
| Swap Channels | Stereo L↔R swap. |
| Split to Mono | Writes N mono WAV files (one per channel) to a directory you pick. Channel labels in the filename. |
Edit mode adds a set of single-key edits, active only in Edit mode so they don't shadow the transport keys, and listed in the Edit menu so the keys are always visible. The first group acts on the selection; the playhead group acts at the cursor. Cut uses a 10 ms equal-power crossfade by default; the fade keys use the equal-power curve.
| Key | Action | Acts on |
|---|---|---|
| ⌫ | Cut (remove and crossfade-join) | selection |
| ⌥⌫ | Silence selection (replace, keep length) | selection |
| ⌘T | Trim to selection (keep it, discard the rest) | selection |
| ⌘F / ⌥⌘F | Fade In / Fade Out | selection |
| ⌃↑ / ⌃↓ | Gain +1 / −1 dB | selection |
| ⌘⌫ | Cut to start (remove start → playhead) | playhead |
| ⇧⌘⌫ | Cut to end (remove playhead → end) | playhead |
| ⌥[ / ⌥] | Fade in to / out from the playhead | playhead |
⌘Z steps back through the session's edits, including edits applied from Chapter mode (Level chapters), the Level Dialogue pass, Spectral mode operations, and takes landed from Record mode (each Append, Insert, or Punch landing is one undo step), since the modes share one undo history. ⇧⌘Z redoes. Undo is disk-backed (each level is a temp file, not a second copy in RAM), so deep history doesn't multiply memory. In Dialogue mode, ⌘Z targets region edits (a separate 50-step history); audio-buffer undo is reached from Edit, Chapter, or Spectral.
⇧⌘S opens a save panel. Specula writes the current audio to a new file (it never overwrites the source). Format defaults to WAV at the file's native bit depth and sample rate. Saving makes the written file the working document in place: your chapters, dialogue regions, and playhead stay put, and you stay in the mode you were in.
Discard Edits (Edit menu) reverts the buffer to the last saved version in one step, the explicit counterpart to Save. Save As, Discard, and the unsaved-changes dot are reachable from Edit, Chapter, Dialogue, Spectral, and Record alike, because those five modes share one editing session.
Quitting or opening over unsaved work asks first. ⌘Q with unsaved edits prompts to save them, quit without saving, or cancel. If a take is still recording, Specula asks before stopping it, and if the loaded audio is a recorded take that hasn't been saved out of the takes folder, it offers Save Take As… before quitting. Opening a new file runs the same checks, whatever the path: the picker, a drop, the Dock, Open With, or Open Recent.
While in edit mode, a horizontal toolbar appears below the transport. Operations with parameters (Cut, Silence, Gain, Fade, Normalize, Limit TP, Level Dialogue) open small popovers - set the parameters, click Apply. The Limit TP popover takes a ceiling in dBTP with presets at -2 / -1 / -0.5 / -0.1. The Cut popover takes the crossfade length (default 10 ms, 0 for a hard cut). The Level Dialogue popover sets the duck amount, an optional Lift dialogue to target with a target and true-peak ceiling.
A second row under the edit toolbar hosts a chain of Audio Unit effects. It appears in Edit mode for mono and stereo files, whatever output device you monitor through (it is hidden only for multichannel files). On a multichannel interface, the output routing chooses which outputs carry the processed signal, so monitoring on outputs 3/4, or on several pairs at once, works; per-output gain and phase adjustments apply outside Edit mode. Click Add to browse the effects installed on your Mac, grouped by manufacturer and searchable by name, then pick one to append it to the chain. Both AUv2 and AUv3 effects are listed; instruments and other non-effect units aren't, since the chain processes an existing file.
Each plugin sits in its own chip showing the effect name and maker. While the file plays you tune it live: click the sliders icon to open the plugin's own window (or a generic parameter panel for plugins that ship no interface of their own), and because the analysis tap sits after the chain, every live meter, LUFS, true peak, FFT, spectrogram, and phase, responds to the processed sound as you turn knobs. The power icon bypasses a plugin, the chevrons reorder it, and ✕ removes it. A plugin that fails to load shows a red warning with the reason on hover. Parameter tweaks and bypass are gapless; adding, removing, or reordering rebuilds the playback graph with a brief gap, the same as switching output device.
Render bakes the chain into the audio. It renders offline through fresh plugin instances, compensates for each plugin's latency so the result stays sample-aligned with the original, and commits the result as a normal edit you can Undo. Rendering consumes the chain: it clears, and the snapshot is remembered, so Restore Last brings it back. The tweak-again loop is Undo, Restore Last, adjust, Render again, and Restore Last survives quitting the app. An Include effect tail checkbox (off by default) appends reverb and delay ring-out; with it off, the rendered audio is exactly as long as the input, which is what A/B comparison and auto-align rely on.
Render to Compare Slot renders the same way but loads the result as a new slot next to the original and switches to Compare mode, so you can measure and A/B the processed version against the source with real numbers.
Leaving Edit mode with plugins still in the chain asks whether to render them or discard them first. Loading a new file starts a fresh chain; the previous one stays available under Restore Last, even across app restarts. Under Settings, Plugins, Prefer out-of-process Audio Units (on by default) keeps plugins isolated from the app so a crash in a third-party plugin doesn't take Specula down; turning it off loads them in-process for lower latency.
See it, select it, take it out. Spectral mode (⌘5, or the Spectral button in the transport bar) turns the main canvas into a full resolution, zoomable spectrogram of the loaded file, where you select time and frequency regions and process just those. A cough inside a piano chord, a chair squeak under speech, a bite of mains hum: paint over the blemish and remove it, leaving everything around it untouched. Works on any channel count, from mono to surround beds; edits share Edit mode's undo history, unsaved-changes dot, and Save As.
The spectrogram is computed from the file at full resolution as you navigate; it uses the same color map and frequency scale (log or mel) as the analysis spectrogram. A time ruler runs along the top (click it to seek), a frequency scale with gridlines runs along the right edge of each channel row, and a small waveform strip below the ruler can be toggled from the toolbar. The strip is the transport surface: click it to place the playhead, or hold and drag to scrub, since clicks on the canvas itself are selection tools. Scroll to zoom time around the cursor (pinch works too), ⌥-scroll to zoom frequency around the cursor, ⇧-scroll to pan. The readout under the cursor shows time, frequency, and level. For stereo files a channel switch in the toolbar shows the left channel, the right, or Both stacked, so you can see each channel's spectrogram on its own and spot content that lives in only one side. Files with more channels show one channel at a time: pick it from the toolbar's channel menu, labelled with the file's layout (L, R, C, LFE, Ls, Rs, and so on).
To get back out, ⌘0 returns both axes to the whole file across the full frequency range. ⇧⌘0 fits time alone and ⌥⌘0 restores the frequency range alone, so you can fit the whole file while the view stays on a 3 kHz band. Each axis also gets a small button in the top right corner of the canvas, shown while that axis is zoomed.
Four tools, one selection: Rectangle (R) for straight time and frequency ranges, Lasso (L) for free shapes, Brush (B) for painting over curved traces like a bird chirp or a whistle, and Wand (W) for shape matching: click a spot and the selection grows across the connected region at a similar level, or drag to sweep the wand along a blemish, seeding as it goes. Two sliders shape it, and adjusting either re-tunes the last stroke live: the dB tolerance sets how far below the clicked level the selection may grow, and the time reach caps how far from the click it may spread, so a click on a hum line grabs the local stretch instead of the whole file. With the rectangle and lasso, plain dragging starts a new selection; hold ⇧ to add to it, ⌥ to subtract from it. The brush and wand paint: every stroke adds to the selection, and an ⌥-stroke removes. The brush size slider (or ⇧[ and ⇧]; bare [ and ] keep stepping the playback rate) sets the brush; on the log frequency axis the brush covers what it visually covers, wider in tone terms at the top than the bottom. Esc clears the selection.
| Operation | Notes |
|---|---|
| Attenuate | Lowers (or raises) the selection by a dB amount you choose, default −18 dB. |
| Erase | (Delete) Pushes the selection down to the noise floor learned from the audio just before and after it, minus an adjustable erase depth (in the Feather popover, default 12 dB). At depth 0 the hole matches its surroundings; more depth removes more decisively, which matters inside busy material where the surroundings are the music itself. |
| Repair | Bridges the selection in time: each frequency's level is interpolated from the clean audio on both sides. Best for short blemishes crossing otherwise steady material. |
Every operation feathers its edges in both time and frequency (the Feather popover sets how much), applies to the selected channels only (All channels by default on mono and stereo; editing one side of a stereo file can skew the stereo image), and lands as a normal edit: Undo brings the audio back, the selection stays so you can iterate, and Save As writes a new file, never overwriting the original. Play span plays the selection's time range for a quick audition; with a selection active, the spacebar does the same. Solo (⇧Space) plays only the selection itself: everything around it is silenced, with the same feathered edges the operations use, so you hear exactly what an Erase or Attenuate would touch. Result (⌥Space) plays the opposite: the span with the selection removed, an audition of the outcome before you commit anything.
On a file with more than two channels, every operation applies to the channel on screen and only that channel: a squeak in the left surround of a 5.1 bed comes out of Ls alone, and the other five channels stay untouched, sample for sample. The selection persists when you switch channels, so a blemish that spans channels can be fixed with the same region, channel by channel, and the confirmation names the channel it edited. Solo and Result audition the shown channel by itself, so what you hear is exactly the channel the operation will touch.
The analysis window size (1024 to 8192) trades time detail against frequency detail; changing it clears the selection since the selection grid changes with it. Selections longer than 10 minutes are refused; select a shorter range and apply in passes.
A dedicated mode for audiobooks and long-form podcasts. Specula segments the file at long silences, scores every chapter against the full ACX measurement set, and surfaces both the per-chapter ACX gates and each chapter's deviation from the book's own median. Mutually exclusive with the other five modes (Edit, Compare, Dialogue, Spectral, Record) so the workspace stays focused.
A Chapter button sits in the transport bar next to Edit, Compare, Dialogue, Spectral, and Record. Click it to enter or exit chapter mode; ⌘3 does the same from the keyboard. Chapter mode carries the green accent colour - the chapter toolbar that slots in below the transport, the playing-chapter fill on the ribbon, and the Chapter button itself all match. Mutually exclusive with Edit, Compare, Dialogue, Spectral, and Record.
Detect scans the loaded mono mix for silences longer than the minimum duration (default 2 s) below the silence threshold (default −55 dB). Each silence becomes a boundary; the audio between two silences becomes a chapter. Chapters whose total non-silent content runs under 1 s are dropped so a stray cough doesn't fragment the result.
The detector absorbs leading and trailing silence into the first and last chapters. Chapter 1 always starts at 0 s, the last chapter always ends at file duration, and the ribbon covers the full file timeline with no gaps at the edges.
Boundary markers are drawn inside every time-axis section that has one open: waveform, loudness curve, spectrogram, and stereo width. Each section uses its own time-to-x mapping, so the boundaries stay aligned across stacked views at any zoom level. The chapter ribbon mirrors the waveform's zoom and horizontal scroll, and narrow chapters keep their width when neighbours dominate the visible range.
Each chapter boundary draws as a vertical teal line with a "#N" badge labelling the chapter that starts at that boundary. Drag a boundary on the waveform to move it. Hit zone is 6 px; the cursor switches to the macOS resize cursor when it lands on a boundary. Drag clamps so neither neighbouring chapter shrinks below 1 s of content. A time + dBFS hover readout tracks the boundary's new position while the drag is in progress.
Add drops a new boundary at the playhead. The chapter ribbon and the per-chapter measurements update as you drag.
Click a slot in the ribbon to select it. The selected slot takes a 2 pt accent border; unselected slots keep a 1 pt white border. Selection and playback are independent: the currently playing chapter is shown by an accent-tinted fill (28 %), regardless of which chapter is selected. Double-click a slot to seek the playhead to its start.
Click a boundary line in the waveform (not a slot in the ribbon) to flag it. The flagged boundary draws thicker in amber with a "✕" marker next to its number, and the chapter that starts at that boundary becomes the selected chapter in the ribbon. The next press of Remove in the chapter toolbar deletes that boundary, merging its two chapters into one. This is the precise way to pick which split to undo; clicking a slot in the ribbon still selects the same chapter without changing which boundary is highlighted.
Double-click a chapter name in the ribbon to rename. Default names ("Chapter 01", "Chapter 02"…) follow the index after edits; custom names ("Foreword", "Walk through the woods") stay put.
Analyse Loudness in the toolbar runs an offline BS.1770 pass per chapter and populates the full measurement set. For each chapter:
Chapters shorter than 10 s skip the BS.1770 measurements entirely (the absolute-gate and relative-gate machinery isn't reliable below that). RMS and sample peak still populate, since plain RMS and peak are valid at any length, so the ACX loudness gate and clipping don't slip through on short chapters.
Each slot is four lines tall:
#NN · Chapter Name m:ss
-20.4 dB RMS -20.6 LUFS +0.0 LU
TP -2.1 dBTP NF -62 dB
LRA 4.2 LU Dlg -19.8 LUFS Sp 87%
Missing metrics render as muted - so the row positions stay stable across chapters that haven't been analysed yet, are too short for an integrated reading, or fall outside the VAD's coverage.
Each chapter is independently checked against three ACX delivery limits. Failures render the metric in red on the ribbon slot and on the report's chapter table:
These three gates are independent of the deviation-from-median flag, so each chapter answers two separate questions: will ACX accept this chapter on its own? (the three red/no-red signals) and is this chapter consistent with the rest of the book? (the deviation badge, controlled by the ±2 LU threshold in Settings → Targets → Chapter detection).
The consistency metric (RMS or integrated LUFS) and the deviation threshold (default ±2.0, in that metric's unit) drive the red-tint flag on outlier slots and the headline banner at the top of the ribbon ("Chapter 07 is 2.3 dB above median"). RMS is the default, since it's what ACX delivery is judged on. Both live in Settings → Targets → Chapter detection (the metric is also set in the Level… popover).
The fix for the outlier flag, and for the consistency deviation across a book. Level… in the toolbar opens a popover with two choices:
It applies a per-chapter uniform gain, then re-measures so the ribbon updates. RMS covers every chapter; LUFS skips chapters under 10 s (no integrated reading), as it does chapters already on target. Run Analyse Loudness first; the button stays disabled until there's a reading.
It's a single undoable edit (⌘Z). Chapter mode shows a dirty dot and a Save As button when the buffer has unsaved edits, and ⇧⌘S writes the leveled file without leaving Chapter mode.
A true-peak ceiling is applied in the same edit, on by default, so a boosted chapter can't run past the delivery ceiling. It defaults to −3 dBTP for RMS/ACX and −1 dBTP for LUFS (tracking the metric), and the Level popover lets you change it or switch it off. One caveat remains: a uniform per-chapter gain moves that chapter's room-tone noise floor with its level, so for a floor-aware pass, gate the room tone first with Level Dialogue, then level.
Restores Specula's silence-detection result, dropping all manual boundary edits, adds, and removes. Renames are kept.
Export… writes the current chapters (name + start + end) to a JSON sidecar. Import… loads one back, replacing the current chapter list. The exported JSON is plain Specula JSON; the importer also accepts a slimmer format with just start and end per entry, so you can hand-craft one.
Export CSV… writes the per-chapter metrics table (RMS, integrated LUFS, deviation from median, true peak, sample peak, LRA, max momentary / short-term, dialog LUFS, noise floor, speech %) as a spreadsheet-ready CSV, for the producers who live in spreadsheets. Run Analyse Loudness first so the metric columns are populated.
When you load a file whose fingerprint (filename + duration + sample rate + channels) matches a prior chapter setup saved in ~/Library/Application Support/Specula/chapters/, a Recall button appears in the toolbar. Click to restore the saved chapters. Toggle the prompt in Settings → Targets → Chapter detection.
The sixth section at the bottom of the window. Time-aligned to the waveform's zoom and scroll. Click a slot to select; double-click to seek the playhead to its start. Toggle the section's visibility with ⌃6 (independent of whether chapter mode is active).
When chapters exist on the file, the PDF / JSON report includes a Chapters section. The table carries every per-chapter metric the ribbon shows: RMS, integrated LUFS, deviation from median, true peak, sample peak, LRA, max momentary, max short-term, dialog LUFS, noise floor, and speech percentage. Cells fail-red on the same three ACX gates the ribbon uses (RMS outside [−23, −18] dB, TP > −3 dBTP, NF > −60 dB). The caption above the table documents all three gates and the configured deviation threshold so a producer reading the PDF can interpret it without opening the app.
Record straight into the app. Record mode (⌘6, the Record button in the transport bar, or the Record audio tile in the empty window) captures from any input device on your Mac: a new file when nothing is loaded, a continuation appended to the loaded file, an insert at the playhead that pushes the rest later, or a punch that records over existing audio. It is the one mode that works with no file open, because it is how a file gets created. The first time you enter it, macOS asks for microphone permission.
Three pickers in the toolbar set up the capture. The input device list shows every device with input channels; the layout picker sets how many channels the take records: Mono, Stereo, or 5.1, always stored as one file. The source picker chooses which hardware inputs feed the take: any single input for mono (a microphone on input 3 records from input 3), any adjacent pair for stereo, the first six for 5.1. Layouts a device can't feed don't appear, and the layout adapts automatically when you pick a device with fewer inputs: selecting a one-channel microphone switches the take to mono on its own. Recording 5.1 needs an interface with at least six inputs.
Specula records the signal as the device delivers it: no noise suppression, no automatic gain, no EQ. Call apps process the same microphone heavily (FaceTime and its relatives add noise reduction, a low cut, and automatic leveling), so a built-in or display microphone can sound noticeably bassier and quieter here than it does on a call. The take holds what the microphone actually picked up, and the meters measure the same thing. Set the recording level with the input volume in System Settings › Sound › Input, and shape the sound afterwards in Edit mode if you want to.
Entering the mode arms the input: live meters run before anything is recorded. Per-channel level bars, momentary loudness, and true peak track the incoming signal, and a floor readout shows the quietest level heard over the last five seconds, so you can verify the room and chain before committing a take. Spoken-word specs care about exactly this: ACX wants the noise floor under −60 dB, and the readout turns orange when the room is louder than that. When recording starts, the loudness and true-peak meters reset so they describe the take, not the level check before it. Live values are for steering the session; the definitive numbers come from the analysis pass that runs when the take lands.
With New selected, Record (or R) starts the take and Stop ends it. While recording, the take draws itself as a live waveform across the main canvas, growing as it is captured and scrolling once the window fills, so a dead input is obvious within a second. The take loads immediately when you stop, and everything that happens on any file load happens to it: waveform, loudness, speech detection, per-channel stats, report inputs. Takes are written to Specula's takes folder as they record (~/Library/Application Support/Specula/Recordings); Save Take As… writes the loaded audio to a location of your choosing, and until then unsaved takes are kept in that folder for seven days, so yesterday's unsaved take is recoverable from there in Finder. Recording streams to disk as it happens, so take length is bounded by free disk space, not memory.
With a file loaded, three targets place the take in the existing audio. Append joins it to the end of the file. Insert puts it in at the playhead and pushes everything after that point later: nothing is recorded over, the file just grows by the take. Punch records over existing audio: with a selection, the take replaces the selected range, free to be longer or shorter than what it replaces; without one, the take records over the audio from the playhead for as long as you keep going, tape style, extending the file if it runs past the end. That's the flubbed-word workflow: select the word (or just park the playhead on it), hit Record, say it again, hit Stop; use Insert instead when the file is missing a sentence rather than containing a wrong one. Where a take lands is fixed the moment recording starts, so moving the selection or the playhead mid-take doesn't move the landing. All joins are smoothed with an equal-power crossfade (the Splice popover sets its length, default 10 ms; 0 ms is a hard splice), and each landing is one edit: Undo (⌘Z, or the toolbar's Undo button) restores the audio exactly as it was, right from Record mode, and Redo brings it back. An insert that didn't work out is one click gone. The undo history is the same one Edit mode uses, so record, edit, record, edit stays one timeline of changes. After an insert or punch the selection moves to the landed take, so listening back is one Space away. New is the resting state: clicking the active Append, Insert, or Punch button switches back to it.
A take records at the input device's sample rate and is converted to the file's rate when it lands. Mono takes adapt to any file (the same audio on every channel) and stereo takes fold to mono files; other channel mismatches are refused with a message, so a 5.1 file takes a 5.1 take.
Recording shares the same session as Edit mode: record a chapter, trim its head in Edit, level it, then come back and re-record the one sentence you don't like, all in one undo history, and Save As always writes a new file. If the input device changes or disappears mid-take, the recording stops and everything captured up to that moment is kept.
To Compare sends the last finished take to a Compare slot. Record the same part twice, load both takes side by side, and A/B them level-matched, or diff them, the way you'd compare two masters. Useful for microphone shootouts, performance picks, and before/after chain checks.
Build test signals from scratch and write them to a file. The Signal Generator window (⌘G, the Generate button next to the mode selector, or View → Signal Generator) is a standalone builder: it synthesizes audio rather than measuring it, so it opens in its own window and needs nothing loaded. With no file open, the main window also offers a Generate a test signal tile beside the drop zone. The builder keeps its state, so a signal you were working on is still there after closing the window or relaunching the app.
A signal is made of one or more modules. A module is one sound with a level, a start time, and a length:
A module's level is peak dBFS for the tones, sweeps, and impulse, and RMS dBFS for noise, so "−6 dBFS sine" and "−20 dBFS pink noise" both mean what they say.
Give a channel several modules and you get both layering and timelines from the same list:
Under Modifiers, each module can add click-free fade in and fade out edges, a level sweep (ramp the level from one dB value to another, linear or in dB), and bursts (gate the module on and off, so a tone becomes tone bursts and an impulse becomes an impulse train; set the on and off lengths and how many pulses; for an impulse, on plus off set the spacing between impulses).
Same signal on all channels builds one signal and copies it to every channel; noise is decorrelated per channel, so stereo noise is real stereo. Turn it off to give each channel its own modules. Channel count runs from mono up.
Set the sample rate, bit depth (16 or 24-bit integer, or 32-bit float), and container (WAV, AIFF, or CAF). AIFF stores integer PCM only, so 32-bit float needs WAV or CAF. Set total length fixes the file's length, padding or trimming the modules to fit; leave it off and the length follows the last module's end. An optional peak ceiling scales the whole render down so its loudest sample sits at the level you choose, useful when stacked tones would sum past 0 dBFS.
A live waveform and spectrum of the current signal sits at the top of the window and updates as you edit, with running peak, RMS, and integrated-loudness readouts. Preview in Specula renders the current signal to a temporary file and loads it in the main window for the full analysis set (waveform, FFT, spectrogram, loudness). For long signals the in-window preview shows the first 30 seconds.
Generate… (⌘Return) asks where to save, writes the file, and reports its peak, RMS, and integrated loudness. Reveal shows it in the Finder; Open in Specula loads it into the main window for analysis. The same generator is on the command line as specula generate (see the CLI section) for scripted batches.
| Key | Action | Output |
|---|---|---|
| ⇧⌘E | Export JSON Report | *.json - all metrics, both live and selection analysis, channel info, file metadata |
| ⌥⌘E | Export PDF Report | *.pdf - formatted report with charts |
| ⌃⌘E | Export Diff Audio | *.wav - the active slot's cached A − slot residual buffer (compare mode only). Same output as the Export Diff button next to the Listen to Diff toggle. |
In edit mode, ⇧⌘S is "Save Edited Audio As", not the JSON export - the menu item swaps based on context.
The Settings → Report tab decides what's included:
JSON is a stable, machine-readable schema. PDF is the same content rendered for human reading. For the audiobook and podcast producers who live in spreadsheets, Chapter mode → Export CSV… writes the per-chapter metrics table as a CSV (one row per chapter), and the CLI's specula report <file> --format csv does the same from the command line. See Chapter mode.
The full-file report (no selection) carries integrated LUFS, true peak (4× oversampled dBTP), sample peak, LRA, Max Momentary, Max Short-Term, DR, speech-gated LUFS, speech %, stereo correlation, per-channel stats, and loudness-target verdicts - every metric the info panel shows.
Each entry in loudnessTargets.targets[] in the JSON envelope carries an inline verdict object so a Shortcut, CLI pipeline, or batch QC script can branch on pass / fail without recomputing it against the target spec. Fields: status (compliant / penalty / notCompliant / unavailable), actualLUFS, plus penaltyGainDB / penaltyApplies / truePeakClean for streaming targets and nonComplianceReason / nonComplianceActual / nonComplianceLimit for hard-spec fails. The PDF and HTML reports render the same verdict as a status badge. Full field reference in Loudness targets.
specula) #For batch QC, shell scripts, and Shortcuts that walk a folder of files, Specula ships a CLI built on the same measurement engine the app uses. Same numbers, no window.
The CLI ships inside the app bundle. Pick Specula → Install Command-Line Tool… and Specula links specula into /usr/local/bin (one admin prompt if that folder needs it). The link points at the copy inside Specula.app, so app updates keep the installed command current.
Manual alternative (scripted setups):
sudo ln -sfh "/Applications/Specula.app/Contents/Helpers/specula" /usr/local/bin/specula
Verify with specula --version.
The CLI shares the app's license and trial. It runs through the 7-day trial (a days-left note prints on stderr) and requires the license activated in the app after that (Specula → Manage License…); an expired, unlicensed install exits with code 77 and an explanation on stderr. --help and --version always work. The first licensed run asks once for access to the license in the Keychain - click Always Allow. Do that once in a local Terminal window before using the tool over SSH, where the prompt can't appear.
specula analyze <FILE>Headline numbers (integrated LUFS, true peak, sample peak, loudness range, speech-gated LUFS, stereo correlation) as JSON.
specula analyze mix-v2.wav
specula analyze mix-v2.wav --no-vad # skip Silero (faster)
specula analyze mix-v2.wav --no-stereo # skip the per-block correlation pass
specula compare <A> <B>Two files side-by-side with B − A deltas for every metric.
specula compare master-v1.wav master-v2.wav
specula compare master-v1.wav master-v2.wav --match-loudness
specula compare master-v1.wav master-v2.wav --no-vad # skip Silero on both files
--match-loudness subtracts the integrated-LUFS delta from B's peak / speech-gated readings so the deltas surface spectral or shape differences instead of being dominated by level mismatch. Useful when you've rebalanced a mix but want to know whether the tonality really changed.
--out-diff <PATH> additionally writes the A − gainB·B residual to that path as a Float32 WAV (paired with --match-loudness to net out pure level differences before subtracting). The compare JSON gains a residual block carrying the residual's peak / RMS / applied gain so a script can branch on the residual energy without re-loading the file. Sample rate + channel count must match between A and B - the diff refuses structural mismatches rather than silently resampling.
specula edit <FILE> --out <OUT> …Applies edit operations and writes the result to --out (-o for short). --out is required and must not resolve to the same path as the input - Specula never writes in place, the same contract the app's Save As enforces.
specula edit mix.wav --out mix-norm.wav --normalize-lufs=-14 --lufs-ceiling=-1
specula edit mix.wav --out mix-trimmed.wav --trim 5.0:30.5
specula edit mix.wav --out mix-louder.wav --gain 3
Negative values need the --option=value form (e.g. --gain=-3, --normalize-lufs=-14). The CLI otherwise reads a bare leading - as an option name and rejects the value.
Supported operations (applied in this order if multiple are passed):
--trim START:END - keep only the range, in seconds (e.g. 5.0:30.5).--dc-remove - subtract per-channel mean.--invert-phase - flip polarity on every channel.--gain DB - uniform gain.--normalize-peak DBFS or --normalize-lufs LUFS [--lufs-ceiling DBTP] - mutually exclusive.--limit-tp DBTP - apply the true-peak limiter at this ceiling without changing loudness; runs last so it can catch any residual peaks left by an upstream normalize.LUFS normalization and --limit-tp use the same 5 ms two-pass true-peak limiter (BS.1770 4× polyphase oversampled detection).
specula generate --out <OUT> …Synthesize a test signal and write it to --out. One source per run, across N identical channels:
specula generate --sine 1000 --duration 5 --channels 2 --out tone.wav
specula generate --noise pink --level=-20 --duration 10 --out pink.wav
specula generate --sweep 20:20000 --sweep-shape log --duration 3 --bit-depth 24 --out sweep.wav
Pick exactly one source: --sine HZ, --square HZ, --saw HZ, --triangle HZ, --noise white|pink|brown, --sweep START:END, --impulse, or --silence. --level is peak dBFS for tones and RMS dBFS for noise (default −6 / −20); pass negative values with the --level=-20 form. Format flags: --sample-rate (default 48000), --bit-depth 16|24|32f (default 24), --container wav|aiff|caf (inferred from the --out extension otherwise), and --peak-ceiling to scale the render under a ceiling. The square, sawtooth, and triangle are band-limited, matching the app.
For layered or time-sequenced multichannel signals, describe the whole signal in JSON and pass it with --spec-file. The JSON mirrors the window's builder (channels holding timed modules), and omitted fields take the builder's defaults, so a minimal spec is:
{"channels": [{"name": "L", "modules": [{"source": {"sine": {"frequency": 1000}}, "duration": 5}]}]}
Module fields: source (required; sine/square/sawtooth/triangle with frequency, noise with a colour, sweep with start/end/shape, impulse, silence), start, duration, gainDB, fadeInSeconds, fadeOutSeconds, levelSweep, burst. Spec fields: sampleRate, bitDepth (int16/int24/float32), container, channels, duration, peakCeilingDB. The stdout receipt reports the written format and the measured peak, RMS, and integrated LUFS.
specula report <FILE>Full Specula report through the headless pipeline, in your choice of format.
specula report mix.wav --mode music # JSON to stdout
specula report mix.wav --mode podcast --out report.json # JSON to file
specula report mix.wav --mode music --format html --out report.html # the same dark-themed HTML the app's preview window renders
specula report mix.wav --mode music --format pdf --out report.pdf # rasterised PDF
specula report book.wav --mode podcast --format csv --out chapters.csv # per-chapter metrics table as CSV
specula report mix.wav --mode broadcast --no-curve # drop the loudness curve for smaller batch output
--format selects the output shape: json (default), html, pdf, or csv. PDF requires --out (binary data on stdout isn't useful), and the receipt printed to stdout on success carries the written byte count + mode so a pipeline can confirm the file landed. csv emits the per-chapter metrics table when the file has chapters, otherwise a flat metric,value table of the file-level numbers. --mode selects which loudness-target catalog drives the evaluation: music, podcast, vod, broadcast. --no-curve omits the per-100 ms loudness curve when the receipt only needs the summary numbers. --no-vad skips Silero speech detection (the speech-gated fields read -inf) and --no-stereo skips the correlation pass, both for faster unattended batches.
--raw-*)analyze, report, and edit accept headerless raw PCM when you declare the format (see Open raw PCM):
specula analyze dump.pcm --raw-rate 48000 --raw-channels 2 --raw-format s24le
specula report capture.bin --raw-rate 96000 --raw-channels 1 --raw-format f32le --format pdf --out capture.pdf
specula edit take.raw --raw-rate 48000 --raw-channels 2 --raw-format s16le --gain=-3 --out take-quiet.wav
specula analyze dump.pcm --raw-preset "Apollo capture"
--raw-rate, --raw-channels, and --raw-format are all required for a raw input. There is no default rate: a silently assumed one would corrupt every measurement without an error.--raw-format tokens: u8, s16le, s16be, s24le, s24be (packed 3-byte), s24_32le, s24_32be (24-bit low-aligned in a 32-bit container), s32le, s32be, f32le, f32be, f64le, f64be.--raw-skip N ignores N header bytes. --raw-tolerance N allows N trailing bytes after the last whole frame. --raw-planar reads each channel as one contiguous block instead of interleaved.--raw-preset NAME uses a preset saved in the app's import sheet; any other --raw-* flag refines it.file block records the full declaration under rawFormat.Every subcommand emits pretty-printed JSON with sorted keys, so diffs across runs stay deterministic. Non-finite floats (silence → -∞) render as "-inf" strings rather than blowing up the encoder.
Specula surfaces eight App Intents that the system Shortcuts app and Siri can call directly. Same engine as the CLI; same numbers. No file-handling shell glue required, output is chainable. The four core actions ship in two variants each (file-input and text-path) so they fit whichever Shortcuts workflow shape you have.
| Action | File-input | Path-input | Returns |
|---|---|---|---|
| Measurements | Get Measurements | Get Measurements (from Path) | Specula Measurements (I, M, S, LRA, TP, SP, DR, Speech-Gated, Speech %) |
| Numeric compare | Compare Files | Compare Files (from Paths) | Specula Comparison (both sides + B − A deltas) |
| Audio compare | Get Compare Diff | Get Compare Diff (from Paths) | residual WAV file (A − gainB·B) |
| Report | Get Report | Get Report (from Path) | report file (JSON / HTML / PDF) |
Each action's behaviour:
--mode. Pipe the result straight into AirDrop, Save File, Mail, or any Shortcut file step.Each intent's file parameter shows a Choose button followed by a small … menu. The Choose button opens a static file picker (one-off use). The … menu is where the magic-variable picker lives - pick Shortcut Input to bind the file that came in from a Finder Quick Action / Share Sheet / prior action's output. Don't pick "Clear" - that wipes the binding.
For text-path workflows (clipboard text, Ask for Text, paths pasted from a shell), each intent has a sibling (from Path) variant that takes the audio file as a text string instead of a File. The text variants accept:
/Users/me/foo.wav/Users/me/My Mix.wav (no quoting needed)~/Music/foo.wav'/Users/me/My Mix.wav' (terminal realpath output)"/Users/me/My Mix.wav"'/Users/me/My Mix.wav'file:///Users/me/My%20Mix.wav (percent-encoded)Quotes are stripped automatically; tilde gets expanded against your home directory. So a typical clipboard-driven shortcut is just Get Clipboard → Get Report (from Path) with the clipboard variable bound to Audio File Path.
Each file-input action also answers to a second phrasing: "Get measurements from Specula", "Compare audio in Specula", "Get Specula report".
Type one into Spotlight (⌘Space). Intents with required file parameters will prompt for the file via a picker; for variable-driven workflows the Shortcuts app is the better surface.
The six measurement, compare, and report intents default to speech detection off. Flip the "Enable Speech Detection" toggle in the Shortcut step when you need speech-gated metrics; the first speech-gated run may take a moment as the model loads. The two Get Compare Diff intents return a pure residual, so they carry no speech toggle.
Three services install automatically and appear under Services when you right-click one or more audio files in Finder. They work out of the box, with nothing to build in Shortcuts or configure in Automator.
Specula Report - <filename>.pdf next to the source (auto-numbered if that name exists), falling back to ~/Downloads if the source folder isn't writable, then reveals it in Finder. If Specula wasn't running it stays hidden the whole time, so no empty window pops up.If the entries don't appear right after first install, log out and back in (or reboot) so Launch Services indexes them.
Open with ⌘,. Nine tabs: Layout · Spectrogram · FFT / Analysis · Speech · Targets · Report · Updates · License · Acknowledgements.
Defaults for which sections show on launch:
Each section can still be toggled live with ⌃1-⌃5.
Text size scales the text across the main window: the info panel (file details, channel meter labels, every loudness readout), the module control strips (the waveform's channel chips, zoom and amplitude clusters, the IN / OUT / LEN selection fields, the loudness-curve, FFT, and stereo-width strips), the transport bar (timecode display, Analyse, Generate, the mode selector), the segmented switches, and the axis labels. Default, Large, or Extra Large. The info panel also honors the system Increase Contrast setting (System Settings → Accessibility → Display) by rendering its dimmed labels brighter, and its faintest labels were brightened across the board.
Loudness readout sets how often the live Momentary, Short-term, and Integrated numbers repaint: 10 Hz (EBU Tech 3341), the default, or 5 Hz and 2 Hz for easier reading. All three sit above the 1 Hz the standard asks of integrated, and 10 Hz is the minimum the standard requires of a live short-term meter. A separate Rolling digits toggle controls the digit animation on its own. Both settings are display-only: no measurement changes.
Liquid Glass chrome (macOS 26+ only). On macOS 26 the transport bar, file dock, info panel, seek bar, and section-toggle strip use the system Liquid Glass material so they feel native alongside Finder, Safari, and other macOS 26 apps. Default on. Turn off in Settings → Layout → Appearance for solid dark chrome, useful when recording screenshots or screencasts where the depth-aware refraction would change with whatever's underneath the window, or simply if you prefer the flat look. The toggle is hidden on macOS 14-25 (Liquid Glass isn't available there; chrome stays solid regardless).
With the toggle on, each chrome strip renders as a rounded floating pill (radius 6, with a small dark gutter between strips and the window edges), the macOS 26 native pattern Music, Photos, and Calendar use. The Liquid Glass material's edge highlights become the rim of each pill rather than horizontal lines crossing the window. The seek bar's accent fill also picks up the glass material when the toggle is on (otherwise solid). With the toggle off, chrome is edge-to-edge square strips; the rounding belongs to the Liquid Glass look.
The seek bar's track and fill are 8 px tall regardless of the toggle, a substantial slider lane that reads well whether it's filled with glass or solid colour.
All six values (colormap, frequency scale, Floor, Ceiling, FFT overlap, auto-compute on load) persist across launches.
Edits re-run detection in about half a second and the new history feeds back into the live loudness path. If you have manual edits in Dialogue mode when you move a slider, they're pushed onto the Dialogue mode undo stack first so you can ⌘Z back to them. See Speech-gated loudness + Dialogue mode for the full workflow.
| ⌘O | Open audio file |
| ⌘, | Open Preferences |
| Space | Play / pause |
| ⌘. | Stop and return to start |
| ← / → | Skip ±5 seconds |
| ⇧← / ⇧→ | Skip ±1 second |
| ⌘← | Jump to start |
| L | Toggle loop |
| [ | Step rate down |
| ] | Step rate up |
| \ | Reset rate to 1× |
| Gesture | Action |
|---|---|
| Click / plain drag | Move the playhead (drag scrubs) |
| ⌘-drag | Create a selection region |
| Drag a selection's edge / body | Resize / move it |
| Scroll (wheel or two-finger) | Horizontal zoom (cursor-anchored) |
| ⇧-scroll | Pan horizontally |
| ⌥-scroll | Amplitude (vertical) zoom |
| Esc | Clear selection (waveform + FFT) |
| ⌘A | Select all |
| ⌘Return | Analyse selection (offline) |
| ⌘1 | Enter / exit Edit mode (accent: amber) |
| ⌘2 | Enter / exit Compare mode (accent: sky-blue) |
| ⌘3 | Enter / exit Chapter mode (accent: green) |
| ⌘4 | Enter / exit Dialogue mode (accent: rose) |
| ⌘5 | Enter / exit Spectral mode (accent: violet) |
| ⌘6 | Enter / exit Record mode (accent: red; the only mode that works with no file loaded) |
| 1 - 6 | Switch to compare slot A - F (bare digits, Compare mode only) |
| D | Toggle Listen to Diff on the active slot (bare D, Compare mode only) |
These operate only in Edit mode (so they don't shadow the transport keys) and are listed in the Edit menu. The first group needs a selection; the playhead group acts at the cursor. Defaults: Cut uses a 10 ms equal-power crossfade; the fade keys use the equal-power curve.
| ⌘Z / ⇧⌘Z | Undo / redo last edit |
| ⇧⌘S | Save edited audio (new file) |
| ⌫ | Cut (remove and crossfade-join) - selection |
| ⌥⌫ | Silence selection (replace, keep length) - selection |
| ⌘T | Trim to selection (keep it, discard the rest) - selection |
| ⌘F / ⌥⌘F | Fade In / Fade Out - selection |
| ⌃↑ / ⌃↓ | Gain +1 / −1 dB - selection |
| ⌘⌫ | Cut to start (remove start → playhead) - playhead |
| ⇧⌘⌫ | Cut to end (remove playhead → end) - playhead |
| ⌥[ / ⌥] | Fade in to / out from the playhead - playhead |
| ⌃6 | Toggle chapter ribbon section visibility |
| I | Mark In (anchor at playhead, or set start of selected region) |
| O | Mark Out (close pending region, or set end of selected region) |
| S | Split selected region at playhead |
| Delete | Delete selected region |
| ⌘Z / ⇧⌘Z | Undo / redo region edit (50-step history) |
| R / L / B / W | Rectangle / Lasso / Brush / Wand selection tool |
| ⇧[ / ⇧] | Brush size down / up (bare [ / ] step the playback rate, as everywhere) |
| ⌘0 | Reset zoom on both axes |
| ⇧⌘0 | Fit the whole file in view (time axis only) |
| ⌥⌘0 | Show the full frequency range (frequency axis only) |
| Space | Play the selection's span (play / pause without one) |
| ⇧Space | Solo the selection (play only what is selected) |
| ⌥Space | Result preview (play the span with the selection removed) |
| ⇧-drag / ⌥-drag | Add to / subtract from the selection |
| Esc | Clear the selection |
| Delete | Erase the selection down to the learned noise floor |
| ⌘Z / ⇧⌘Z | Undo / redo (shared with Edit mode's history) |
| ⇧⌘S | Save As |
While Spectral mode is active, L selects the Lasso tool; the loop toggle stays available on the transport bar.
| R | Start / stop recording |
| ⌘Z / ⇧⌘Z | Undo / redo landed takes and edits (shared with Edit mode; not while a take is rolling) |
| ⇧⌘S | Save As (a landed Append / Insert / Punch, like any edit) |
While Record mode is active, R starts and stops the take; in Spectral mode the same key is the Rectangle tool. The two never conflict because only one mode is active at a time.
| ⌘G | Open the Signal Generator window |
| ⌘Return | Generate (in the generator window) |
| ⇧⌘E | Export JSON report |
| ⌥⌘E | Export PDF report |
| ⌃⌘E | Export diff audio (WAV) |
| ⌘I | Show / hide the info panel |
| ⌃1 | Toggle Waveform |
| ⌃2 | Toggle Loudness Curve |
| ⌃3 | Toggle Spectrogram |
| ⌃4 | Toggle FFT |
| ⌃5 | Toggle Stereo Width |
| ⌃6 | Toggle Chapter Ribbon |
| ⌃⌥1 | Focus Waveform |
| ⌃⌥2 | Focus Loudness Curve |
| ⌃⌥3 | Focus Spectrogram |
| ⌃⌥4 | Focus FFT |
| ⌃⌥5 | Focus Stereo Width |
| Standard | Where it appears |
|---|---|
| ITU-R BS.1770-4 | K-weighting filter, 100 ms gating blocks, dual gating, 4× true-peak oversampling |
| EBU R128 | Momentary / Short-term / Integrated LUFS, LRA, True Peak, recommended programme loudness |
| Silero VAD (MIT) | Speech detection model |
| FluidAudio (Apache 2.0) | Swift wrapper for Silero on Apple platforms |
Specula's loudness implementation is validated against the EBU R128 loudness test set (EBU Tech 3341 / 3342), to the EBU Tech 3341 ±0.1 LU tolerance. The underlying measurement algorithm is ITU-R BS.1770-4, which EBU R128 builds on.
EBU R 128