Γ-W1-T5: a real Preserve time-stretcher — write rate is duration, tap rate is pitch

Generalizes the correlation-aligned SOLA delay line so the feed and the shift are
independent rates over one ring. Unity is bit-identical to the shipped read, asserted
against a hash baseline captured pre-change.
This commit is contained in:
2026-08-01 19:06:06 -04:00
parent e589addc54
commit 589a8e078b
10 changed files with 873 additions and 71 deletions
+3 -1
View File
@@ -287,7 +287,9 @@ anything for a trigger shape.
- `voice.h` / `voice.cpp` — one voice. The per-SAMPLE render half (`advanceFrame` and everything it calls) is INLINE IN THE HEADER by RT constraint; the per-NOTE half (note-on setup incl. the Preserve ring prime, legato retune, gate-off, the off-thread shifter presize) is out of line in the TU. The voice owns its own `VoiceFilter` and filter envelope, run between the pitch stage and the amp multiply — see `engine/filter/CLAUDE.md`. **Documented ~600-line-ceiling exception** (root `CLAUDE.md` structural heuristic 1): `voice.h` sits over the ceiling because `advanceFrame`'s RT-inline constraint forbids the seam a split would need — a documented exception, not silent overshoot.
- `voice_engine.h` / `voice_engine.cpp` — `VoiceEngine`: note routing, bounded-stealing allocation, user-parameterized voice count (132, default 16), `VoiceMode` Poly/Mono (last-note held-note stack, `MonoTrigger` Retrigger/Legato), two-tier panic (CC 123 = all-notes-off release, CC 120 = immediate hard-stop including Trigger one-shots), and the block render loops. Preview injects a synthetic note-on at the loaded capture's root note into the main `VoiceEngine` — no dedicated `PreviewCard`; preview obeys polyphony/mono/voice-stealing/envelopes.
- `engine/loop/` — the sustain loop's ONE validity/clamp fold (`resolveLoop`) plus its pre-seam crossfade geometry and the editor's default handle span; see `engine/loop/CLAUDE.md`. The voice folds it once at note-on; the crossfade weight is header-inline because it rides the per-sample read.
- `pitch_shift` — hand-rolled **correlation-aligned SOLA** (splice-overlap-add) pitch shifter for the Preserve playback mode: one active read tap chases the write head at the shift ratio; each splice jump is refined by a cross-correlation search so the new read point is waveform-aligned, then old and new taps are crossfaded (raised-cosine, amplitude-complementary). Replaces the prior dual-tap OLA whose fixed half-window tap offset caused anti-phase cancellation on many source frequencies. **GA2:** ring buffer **primed with the actual upcoming source** at note-on (was zero-filled) → gap-free frame-0 onset, ~25 ms Preserve onset latency eliminated (Preserve now speaks on frame 0, matching Varispeed), and real-content-bounded tail (last-window tail-truncation gone). No third-party dependencies; RT-discipline: no allocation in `process()`.
- `pitch_shift` — hand-rolled **correlation-aligned SOLA** (splice-overlap-add) pitch shifter AND time-stretcher for the Preserve playback mode: one active read tap chases the write head at the shift ratio; each splice jump is refined by a cross-correlation search so the new read point is waveform-aligned, then old and new taps are crossfaded (raised-cosine, amplitude-complementary). Replaces the prior dual-tap OLA whose fixed half-window tap offset caused anti-phase cancellation on many source frequencies. **GA2:** ring buffer **primed with the actual upcoming source** at note-on (was zero-filled) → gap-free frame-0 onset, ~25 ms Preserve onset latency eliminated (Preserve now speaks on frame 0, matching Varispeed), and real-content-bounded tail (last-window tail-truncation gone). No third-party dependencies; RT-discipline: no allocation in `process()`.
- **The WRITE rate (duration) and the TAP rate (pitch) are independent, and that is the whole time-stretcher** — `writeFrame` for a surplus source frame, `processNoInput` for a starved output frame, plain `process` for the 1:1 case, `setShiftRatio` for pitch, and `setFeedRate` so the splice crossfade is sized against the real drain rate. The header owns the argument, including why this is not the resampled-read-with-a-cancelling-shift the `WDL_Resampler` invariant above forbids.
- `time_stretch` — the TIME half beside `pitch_shift`'s PITCH half, header-only: `StretchCursor`, the per-output-frame source-feed schedule (a fractional cursor carrying its rate debt, loop-wrapped), plus the rate bounds and their clamp. Rate 1.0 is exactly one source frame per output frame with no residue, which is what makes the unity Preserve read bit-identical to the pre-stretch engine. The bounds are **measured**, not arbitrary — see the header.
- `velocity_curve` — THE monotone spline, shared by every consumer: the three velocity transfer curves and the three spline EGs. `VelocityCurve` is evaluated as ONE OR MORE FritschCarlson monotone cubic Hermite splines joined at its HARD points — a hard knot is a sub-curve boundary for tangent purposes (exactly what the point array's own ends already are), so the two adjacent segments meet at their natural angle instead of a shared derivative and the no-overshoot guarantee holds PER SEGMENT rather than globally. Points are smooth by default; the ceiling is `kMaxCurvePoints` = 128, a MUSICAL bound (long rhythmic phrases, ~two points per articulation event) and not a performance one — **do not lower it**. `eval(velocity)` is the COLD reader, called once per note-on or once per drawn pixel column; `SplineCursor` is the RT one, an indexed segment search plus one Hermite evaluation with the segment and its tangents cached across samples. Both share the same `segmentTangents`/`hermiteAt` free functions, so there is one spline and not two. It carries its own y `CurveDomain`: UNIPOLAR [0,1] is the amp's GAIN, defaulting to `flat()` (y=1, every velocity→unity — a deliberate non-back-compat replacement of the old fixed `velocity/127` path, Daniel-approved); BIPOLAR [1,1] is the signed modulation shape for pitch and filter, defaulting to `zero()` so velocity modulates neither until a curve is drawn. A bipolar curve does not imply the absence of a depth beside it: the filter keeps its `velAmount` knob and the two compose multiplicatively (`velAmount × curve.eval(v)`, `play_params.h`), while the pitch curve's throw is the fixed `kVelocityPitchRangeSemitones`.
- `master_gain` — pure dB↔linear taper math (FB1): normalized [0,1] ↔ dB ↔ linear for the post-mixer master gain control (−∞…+24 dB, norm 0 = true silence, unity ≈ 0.714). Shared by the editor knob and the processor multiply so the needle, persisted value, and audio multiply cannot drift.
+9 -1
View File
@@ -32,7 +32,8 @@ reasampler_test(live_params LINK live_params)
# boundary costs the hot path nothing.
reasampler_pure_library(sampler_core
SOURCES voice.cpp voice_engine.cpp
LINK PUBLIC peaks pitch_shift velocity_curve filter live_params curve_law loop_span)
LINK PUBLIC peaks pitch_shift velocity_curve filter live_params curve_law loop_span
time_stretch)
# Links only sampler_core: linking more would break the plain-data-boundary proof — a VST3
# or REAPER type reaching the core would fail to compile or link here.
reasampler_test(sampler_core LINK sampler_core)
@@ -48,3 +49,10 @@ reasampler_test(live_delivery LINK sampler_core)
# The staged-envelope system across the same engine: per-segment curves, the sustain-less AHD
# both mode shapes share, and the Trigger tail's terminal behaviour.
reasampler_test(staged_envelopes LINK sampler_core)
# The Preserve read's source-feed schedule — the TIME half beside pitch_shift's PITCH half.
# Header-only (it sits on the per-sample feed), hence INTERFACE.
add_library(time_stretch INTERFACE)
target_include_directories(time_stretch INTERFACE ${REASAMPLER_SRC_DIR})
target_link_libraries(time_stretch INTERFACE loop_span)
reasampler_test(time_stretch LINK time_stretch)
+44 -21
View File
@@ -1,8 +1,9 @@
// pitch_shift — pure implementation. See pitch_shift.h for the contract and regression history.
//
// Algorithm: a delay ring of 2*window frames. The write head advances one frame per input
// sample (source rate, duration preserved). One active read tap advances by the shift
// `ratio_` per frame, so its delay behind the writer drifts at (1 - ratio) per frame. When
// Algorithm: a delay ring of 2*window frames. The write head advances one frame per source
// frame the caller feeds; the active read tap advances by the shift `ratio_` per OUTPUT frame,
// so its delay behind the writer drifts at (feedRate - ratio) per frame — one frame in, one
// frame out (`feedRate == 1`) preserves duration, and any other feed cadence stretches it. When
// that delay leaves the safe band [dLow, dHigh], the tap is relocated by a nominal jump of
// one window — clamped to the filled span so it never lands in unwritten silence — refined
// by a cross-correlation search over +/- maxLag plus a parabolic peak interpolation for a
@@ -38,6 +39,7 @@ void PitchShifter::configure(std::int64_t windowFrames) {
fadeFrames_ = fadeLen_ = maxLag_ = corrFrames_ = dLow_ = dHigh_ = 0;
filled_ = 0;
ratio_ = 1.0;
feedRate_ = 1.0;
tailFrozen_ = false;
lastSplice_ = SpliceEvent{};
return;
@@ -85,6 +87,7 @@ void PitchShifter::reset() {
}
filled_ = 0;
ratio_ = 1.0;
feedRate_ = 1.0;
tailFrozen_ = false;
lastSplice_ = SpliceEvent{};
}
@@ -93,7 +96,7 @@ void PitchShifter::freezeTail() {
if (window_ <= 1 || tailFrozen_) return;
tailFrozen_ = true;
// An in-flight crossfade was sized for a retreating writer (outgoing tap drains at
// ratio-1 per frame); frozen, it closes at the full ratio instead. Cap the live fade so
// ratio-feedRate per frame); frozen, it closes at the full ratio instead. Cap the live fade so
// it completes before tap B reaches the parked writer and reads lapped content mid-fade.
if (fading_) {
// Preserve t = fadePos_/fadeLen_ across the shortening so gNew is continuous at the
@@ -158,6 +161,10 @@ void PitchShifter::setShiftRatio(double ratio) {
if (ratio > 0.0) ratio_ = ratio; // ignore non-positive (never run the tap backward/stall)
}
void PitchShifter::setFeedRate(double rate) {
if (rate > 0.0) feedRate_ = rate;
}
double PitchShifter::readTap(double pos) const {
// Fractional linear interpolation with ring wrap.
double p = pos;
@@ -266,18 +273,19 @@ void PitchShifter::splice(std::int64_t nominalJump, double delay) {
while (p >= len) p -= len;
posA_ = p;
// Ratio-scaled fade length. At an up-splice the outgoing tap keeps draining toward the
// writer at (ratio - 1) per frame; the nominal window/4 fade only keeps it behind the
// writer for ratios up to 2 — beyond that (e.g. +24 st = ratio 4) it would cross mid-fade
// and play stale read-ahead data. Cap the live fade at the drain headroom actually
// available, minus 2 (trigger undershoot + interpolator read-ahead margin). Down-shifts
// drain at (1 - ratio) < 1 per frame and can't reach the ring end within window/4 frames,
// so they always keep the full fade.
// writer at (ratio - feedRate) per frame; the nominal window/4 fade only keeps it behind
// the writer while that rate stays under ~1 — beyond that (e.g. +24 st = ratio 4, or a
// half-speed feed under any up-shift) it would cross mid-fade and play stale read-ahead
// data. Cap the live fade at the drain headroom actually available, minus 2 (trigger
// undershoot + interpolator read-ahead margin). A drain rate at or below zero (down-shifts,
// and up-shifts the feed outruns) can't reach the ring end within window/4 frames, so those
// always keep the full fade.
//
// Tail-frozen: with the writer parked, the outgoing tap closes on it at the full ratio in
// either shift direction, so the drain rate is ratio_ instead of (ratio_ - 1) and the cap
// applies at every ratio (including unity, since delay now drains at unity too).
// either shift direction, so the drain rate is ratio_ regardless of feed and the cap applies
// at every ratio (including unity, since delay now drains at unity too).
fadeLen_ = fadeFrames_;
const double drainRate = tailFrozen_ ? ratio_ : (ratio_ - 1.0);
const double drainRate = tailFrozen_ ? ratio_ : (ratio_ - feedRate_);
if (drainRate > 0.0) {
const double headroom = static_cast<double>(dLow_) - drainRate - 2.0;
// Clamp in double before the int64 cast to avoid UB at pathological near-unity ratios
@@ -310,14 +318,27 @@ void PitchShifter::applySplice(const SpliceEvent& ev) {
lastSplice_ = ev; // observable mirror (tests assert follower == master per frame)
}
AudioSample PitchShifter::process(AudioSample in) { return processImpl(in, nullptr); }
AudioSample PitchShifter::process(AudioSample in) { return processImpl(in, nullptr, true); }
AudioSample PitchShifter::processLinked(AudioSample in, const SpliceEvent& master) {
return processImpl(in, &master);
return processImpl(in, &master, true);
}
AudioSample PitchShifter::processImpl(AudioSample in, const SpliceEvent* linked) {
if (window_ <= 1) return in; // pass-through (unconfigured / degenerate)
AudioSample PitchShifter::processNoInput() { return processImpl(0.0f, nullptr, false); }
AudioSample PitchShifter::processNoInputLinked(const SpliceEvent& master) {
return processImpl(0.0f, &master, false);
}
void PitchShifter::writeFrame(AudioSample in) {
if (window_ <= 1 || tailFrozen_) return;
ring_[static_cast<std::size_t>(writePos_)] = in;
if (filled_ < ringLen_) ++filled_;
if (++writePos_ >= ringLen_) writePos_ = 0;
}
AudioSample PitchShifter::processImpl(AudioSample in, const SpliceEvent* linked, bool write) {
if (window_ <= 1) return write ? in : 0.0f; // pass-through (unconfigured / degenerate)
// Copy the linked decision before clearing lastSplice_ (guards a self-aliased pointer).
const SpliceEvent linkedEv = linked != nullptr ? *linked : SpliceEvent{};
@@ -325,8 +346,9 @@ AudioSample PitchShifter::processImpl(AudioSample in, const SpliceEvent* linked)
// Tail-frozen: the source is exhausted, `in` is padding, not stream — write nothing (the
// ring keeps its all-real final two windows) and hold the write head; read/splice/fade
// below run unchanged over the frozen content.
if (!tailFrozen_) {
// below run unchanged over the frozen content. A starved stretch frame (`write` false) takes
// the identical shape: no input was due this output frame, so there is nothing to write.
if (write && !tailFrozen_) {
ring_[static_cast<std::size_t>(writePos_)] = in;
if (filled_ < ringLen_) ++filled_;
}
@@ -378,8 +400,9 @@ AudioSample PitchShifter::processImpl(AudioSample in, const SpliceEvent* linked)
}
}
// Advance heads: write head one frame (parked while tail-frozen), tap(s) by the shift ratio.
if (!tailFrozen_) {
// Advance heads: write head one frame (parked while tail-frozen or starved), tap(s) by the
// shift ratio.
if (write && !tailFrozen_) {
++writePos_;
if (writePos_ >= ringLen_) writePos_ = 0;
}
+34 -5
View File
@@ -1,11 +1,16 @@
#pragma once
// pitch_shift — per-voice, duration-preserving pitch shifter (the Preserve engine's DSP core).
// pitch_shift — per-voice pitch shifter and time-stretcher (the Preserve engine's DSP core).
// Time-domain delay-line with correlation-aligned splices (SOLA-style): one active read tap
// chases the write head at the shift ratio; when it drifts out of its safe delay band it is
// relocated by a nominal window jump, refined by a cross-correlation search so the new read
// point is waveform-aligned, then old/new taps crossfade (raised-cosine). Source and output are
// both consumed/produced 1:1 — only pitch changes, duration is held (unlike the Varispeed
// `readPos_ += ratio_` resample path).
// point is waveform-aligned, then old/new taps crossfade (raised-cosine).
//
// The WRITE rate (how fast source is consumed = duration) and the TAP rate (setShiftRatio =
// pitch) are INDEPENDENT, and only their difference drives the splice cadence. Feeding 1:1 via
// process() holds duration and moves pitch; feeding faster/slower via writeFrame() /
// processNoInput() moves duration at whatever pitch the tap is set to. Nothing here resamples
// to preserve duration — the splice/overlap-add IS the pitch-preserving mechanism, which is
// what the "WDL_Resampler is not a Preserve engine" invariant asks for.
//
// Regression history — do not revert any of these:
// - Correlated splices, vs. the original two-tap OLA (taps hard-locked w/2 apart, Hann
@@ -89,6 +94,13 @@ public:
// ratio) so a bad input never runs the tap backward or stalls it.
void setShiftRatio(double ratio);
// Source frames written per output frame — 1.0 unless the caller is stretching. Used ONLY
// to size a splice crossfade safely: the outgoing tap closes on the write head at
// (ratio - feedRate) per frame, so a fade sized against an assumed 1.0 overruns when the
// source is fed slower than the output runs and the tail of the fade reads lapped content.
// Values <= 0 are ignored. Exactly 1.0 reproduces the 1:1 geometry bit for bit.
void setFeedRate(double rate);
// Transforms one input frame into one output frame (1 in, 1 out). RT-safe: reads/writes the
// pre-sized ring only, no allocation, no lock. Unconfigured returns `in` unchanged. Otherwise
// writes `in` at the write head, reads the active tap (crossfading against the outgoing tap
@@ -103,6 +115,20 @@ public:
// their ring state advances in lockstep. RT-safe: same guarantees as process().
AudioSample processLinked(AudioSample in, const SpliceEvent& master);
// Writes one source frame WITHOUT producing an output frame — the stretch path's surplus
// input when the source is consumed faster than the output runs. No splice can fire here:
// splices are decided on the read side. No-op while unconfigured or tail-frozen, and it
// deliberately leaves lastSplice_ alone so a linked follower's schedule is unaffected.
// RT-safe.
void writeFrame(AudioSample in);
// Produces one output frame WITHOUT consuming a source frame — the stretch path's starved
// output frame when the source is consumed slower than the output runs. Identical to
// process()/processLinked() in every other respect. Returns 0 while unconfigured (there is
// no input to pass through). RT-safe.
AudioSample processNoInput();
AudioSample processNoInputLinked(const SpliceEvent& master);
const SpliceEvent& lastSplice() const { return lastSplice_; }
// Call once the source stream is exhausted — no real frame remains to feed process().
@@ -137,7 +163,9 @@ private:
void applySplice(const SpliceEvent& ev);
// Shared body of process()/processLinked(); `linked` null = master mode (own trigger +
// search), non-null = follower mode (splice iff linked->fired, with linked's decision).
AudioSample processImpl(AudioSample in, const SpliceEvent* linked);
// `write` false is the starved stretch frame: read/splice/advance the taps, but consume no
// input and hold the write head (the same shape tail-freezing already takes).
AudioSample processImpl(AudioSample in, const SpliceEvent* linked, bool write);
std::vector<AudioSample> ring_; // delay line, length `ringLen_` == 2 * window_
std::int64_t window_ = 0; // nominal splice jump in frames; <= 1 = pass-through
@@ -161,6 +189,7 @@ private:
// clamps its up-jump to this so it never lands in
// unwritten silence
double ratio_ = 1.0; // current shift ratio (>0)
double feedRate_ = 1.0; // source frames written per output frame; splice-fade only
SpliceEvent lastSplice_{}; // decision of the most recent process*() frame; cleared
// at the top of every frame, set on a splice
bool tailFrozen_ = false; // writer frozen (source exhausted); tap recycles the
+73
View File
@@ -0,0 +1,73 @@
#pragma once
// time_stretch — the Preserve engine's TIME half: how fast the source is consumed, given a
// playback rate. It pairs with pitch_shift's PITCH half (how fast the ring's read tap runs);
// the two rates are independent over one delay ring, and only their difference reaches the
// splice machinery. Header-inline: every member sits on the per-voice-per-sample feed.
#include <cstdint>
#include "core/instrument/engine/loop/loop_span.h"
namespace reasampler::instrument::engine {
// The playback rates the Preserve DSP is measured over, and therefore the only ones it
// accepts. Two independent reasons they are here and not wider:
// - the ceiling is what bounds a voice's per-output-frame feed loop (kMaxFeedPerFrame source
// frames), which is the RT-safety argument for feeding a variable count at all;
// - the splice search can only align a period it can see. The tap's delay drifts at
// |rate - shift| per frame, so a wide rate over a deep DOWN-shift splices faster than one
// period of the output tone and the correlation stops holding the pitch: measured at rate
// 4.0 with -24 st, the observed period came out 539 frames against 785 wanted.
inline constexpr double kStretchRateMin = 0.5;
inline constexpr double kStretchRateMax = 2.0;
inline constexpr int kMaxFeedPerFrame = 2; // ceil(kStretchRateMax)
// Non-positive and NaN fold to unity rather than to the minimum: an unusable rate should leave
// playback alone, not silently quarter-speed it (the same stance as setShiftRatio's refusal to
// run the tap backward). 1.0 in gives exactly 1.0 out, which is what keeps the unity read
// bit-identical.
inline double clampStretchRate(double rate) {
if (!(rate > 0.0)) return 1.0;
if (rate < kStretchRateMin) return kStretchRateMin;
return rate > kStretchRateMax ? kStretchRateMax : rate;
}
// One Preserve voice's source-feed schedule: a fractional source cursor answering, per OUTPUT
// frame, which whole source frames fall due. At rate 1.0 that is exactly one frame per output
// frame with no residue carried — bit for bit the pre-stretch feed.
class StretchCursor {
public:
// `frame` is where the ring prime stopped; the per-frame feed continues there.
void start(std::int64_t frame) {
frame_ = frame;
debt_ = 0.0;
}
// Adds one output frame's worth of source at `rate` and returns how many whole source
// frames are now due, in [0, kMaxFeedPerFrame]. Take each of them with next(). The clamp
// lives here rather than at the caller because this return value is the loop bound.
std::int64_t due(double rate) {
debt_ += clampStretchRate(rate);
const std::int64_t whole = static_cast<std::int64_t>(debt_); // debt_ >= 0: trunc = floor
debt_ -= static_cast<double>(whole);
return whole;
}
// The next due source frame, wrapped into the sustain loop, advancing the cursor past it.
// Advances even past the playable span — the caller freezes the shifter's writer there, and
// a cursor that stalled instead would re-feed one frame forever.
std::int64_t next(const loop::ResolvedLoop& lp) {
if (lp.active) {
while (frame_ >= lp.end) frame_ -= lp.length;
}
return frame_++;
}
std::int64_t frame() const { return frame_; }
private:
std::int64_t frame_ = 0;
double debt_ = 0.0; // fractional source frames carried into the next output frame
};
} // namespace reasampler::instrument::engine
+8 -2
View File
@@ -18,7 +18,8 @@ void Voice::presizePreserveShifters(std::int64_t windowFrames) {
primeBuf_.assign(windowFrames > 1 ? static_cast<std::size_t>(windowFrames) : 0, 0.0f);
}
void Voice::start(int note, int velocity, const SampleData& sample, bool declickTakeover) {
void Voice::start(int note, int velocity, const SampleData& sample, bool declickTakeover,
double stretchRate) {
// Before any state reset, record the pre-cut reference (last rendered output) and mark
// the compensation pending iff this start is a takeover/steal of a sounding voice and the
// caller opted in. The ramp is seeded on the first frame rendered after the restart, from
@@ -59,6 +60,9 @@ void Voice::start(int note, int velocity, const SampleData& sample, bool declick
baseRatio_ = keyTrackedRatio(note, sample.rootNote, sample.keyTrack) * velPitchRatio_;
playMode_ = p.playMode;
pitchEngine_ = p.pitchEngine;
// Clamped once here so the read head's increment and the feed cursor's debt accumulate the
// SAME value — they must stay exactly one window apart for the note's whole life.
stretchRate_ = instrument::engine::clampStretchRate(stretchRate);
// Clamp into [0, frames): a start at or past the end degrades to 0 (play from the top)
// rather than starting a voice already off the end.
@@ -234,7 +238,9 @@ void Voice::start(int note, int velocity, const SampleData& sample, bool declick
}
// Per-frame feed continues at `p` (the feed bound when the prime exhausted the
// playable span).
feedPos_ = p;
stretch_.start(p);
shiftL_.setFeedRate(stretchRate_);
shiftR_.setFeedRate(stretchRate_);
if (!loopWrap && primeCount < w) {
// Sub-window playable span: the source is already exhausted at prime time.
shiftL_.freezeTail();
+71 -41
View File
@@ -19,6 +19,7 @@
#include "core/instrument/engine/loop/loop_span.h"
#include "core/instrument/engine/pitch_shift.h"
#include "core/instrument/engine/play_params.h"
#include "core/instrument/engine/time_stretch.h"
#include "core/instrument/engine/velocity_curve.h"
namespace reasampler {
@@ -108,7 +109,14 @@ public:
// and this voice is currently active (a takeover/steal restart, not a fresh start), arms
// the difference-seeded declick compensation on the first frame after the restart (see
// kDeclickDecay above). A fresh start never declicks.
void start(int note, int velocity, const SampleData& sample, bool declickTakeover = false);
//
// `stretchRate` is the PRESERVE playback rate — source frames consumed per output frame,
// clamped to [kStretchRateMin, kStretchRateMax]. It is a note-on latch by construction (an
// argument, not a member set separately) because the loop fold and the contour scale it
// composes with are both note-on folds. Varispeed ignores it: there, rate is a factor of the
// read increment, not a second rate. 1.0 is the shipped Preserve read, bit for bit.
void start(int note, int velocity, const SampleData& sample, bool declickTakeover = false,
double stretchRate = 1.0);
// Mono legato takeover: re-pitch this active voice to `note` without touching the
// amplitude envelope, read position, or shifter state — pitch moves, no re-attack. Both
@@ -441,56 +449,75 @@ private:
// and the amp envelope shapes the filtered result (drive included).
double outL, outRlocal = 0.0;
if (pitchEngine_ == PitchEngine::Preserve && shiftL_.configured()) {
// Feed the shifters the source stream at unity rate (duration held) and transpose
// the output by 2^((note-root + pitchEnvSemis)/12) — pitch envelope adds to the
// shift amount, not the read rate. The feed runs one window ahead of readPos_ (the
// rings were primed with that window at start()), under the same sustain-loop wrap
// rule, reading integer source frames (nothing to interpolate). Past the last real
// frame the shifter's writer is frozen — it recycles the real tail it already holds.
if (loop.active) {
while (feedPos_ >= loop.end) feedPos_ -= loop.length;
}
// feedPos_ runs one window ahead of readPos_; the last real source frame is
// playEnd_-1 for Trigger or frameCount-1 for Gate. Once feedPos_ reaches that bound
// the source is exhausted — feeding the held last sample instead would give the
// splice correlation a DC plateau it can't align on (periodic troughs at the splice
// cadence, growing toward the note end). Freezing the shifter's writer means no
// padding ever enters the ring, so the splice machinery keeps recycling the frozen
// all-real tail — a continuous tone through the voice's own end. The sustain-loop
// path never gets here: the wrap above keeps feedPos_ < loop.end forever.
// The two rates the shifter takes (pitch_shift.h owns why they are independent):
// the source is FED at stretchRate_, and the tap is SHIFTED by
// 2^((note-root + pitchEnvSemis)/12) — the pitch envelope adds to the shift amount,
// never to the read rate. The feed runs one window ahead of readPos_ (the rings were
// primed with that window at start()), under the same sustain-loop wrap rule,
// reading integer source frames (nothing to interpolate).
const bool stereoOut = stereo && haveR && shiftR_.configured();
// The last real source frame is playEnd_-1 for Trigger or frameCount-1 for Gate.
// Once the feed reaches that bound the source is exhausted — feeding the held last
// sample instead would give the splice correlation a DC plateau it can't align on
// (periodic troughs at the splice cadence, growing toward the note end). Freezing the
// shifter's writer means no padding ever enters the ring, so the splice machinery
// keeps recycling the frozen all-real tail — a continuous tone through the voice's
// own end. The sustain-loop path never gets here: the wrap keeps the cursor inside
// the loop forever.
const std::int64_t feedBound =
(playMode_ == PlayMode::Trigger && playEnd_ > 0 && playEnd_ < frameCount)
? playEnd_ : frameCount;
const bool exhausted = feedPos_ >= feedBound;
if (exhausted) shiftL_.freezeTail(); // idempotent; input ignored while frozen
const bool feedOk = (!exhausted && feedPos_ >= 0 && feedPos_ < frameCount);
// Crossfaded on the way IN to the shifter, not on the way out: loop the source,
// shift the output.
const double feedXw = crossfadeWeight(loop, static_cast<double>(feedPos_));
const AudioSample feedL =
feedOk ? crossfadedSource(pcm, loop, feedPos_, feedXw) : 0.0f;
const double shift = baseRatio_ * envFactor;
shiftL_.setShiftRatio(shift);
const double shiftedL = static_cast<double>(shiftL_.process(feedL));
if (stereoOut) shiftR_.setShiftRatio(shift);
// 0..kMaxFeedPerFrame source frames fall due this output frame. All but the LAST are
// written without producing output; the last rides the ordinary 1-in-1-out
// process(), so a rate of exactly 1.0 walks the pre-stretch code path unchanged.
// Crossfaded on the way IN to the shifter, not on the way out: loop the source,
// shift the output.
const std::int64_t due = stretch_.due(stretchRate_);
AudioSample feedL = 0.0f, feedR = 0.0f;
bool fed = false;
for (std::int64_t k = 0; k < due; ++k) {
if (fed) { // an earlier frame of this batch: write-only, no output
shiftL_.writeFrame(feedL);
if (stereoOut) shiftR_.writeFrame(feedR);
}
const std::int64_t q = stretch_.next(loop);
if (q >= feedBound) {
shiftL_.freezeTail(); // idempotent; input ignored while frozen
if (stereoOut) shiftR_.freezeTail();
feedL = feedR = 0.0f;
} else {
const double xw = crossfadeWeight(loop, static_cast<double>(q));
feedL = crossfadedSource(pcm, loop, q, xw);
if (stereoOut) feedR = crossfadedSource(pcmR, loop, q, xw);
}
fed = true;
}
const double shiftedL =
fed ? static_cast<double>(shiftL_.process(feedL))
: static_cast<double>(shiftL_.processNoInput());
outL = shiftedL;
if (stereo) {
if (haveR && shiftR_.configured()) {
if (stereoOut) {
// Genuine stereo (linked lag): channel 1's shifter FOLLOWS channel 0's
// splice decisions via processLinked — one correlation search, one lag, one
// splice schedule for both channels (standard stereo SOLA). An independent
// per-channel search re-drew an inter-channel offset of up to +/-maxLag at
// every splice: stereo image wander at the splice cadence + mono-sum
// combing. Each shifter is still processed EXACTLY ONCE per output frame
// (never twice — that would advance its heads twice and corrupt the state).
// (never twice — that would advance its heads twice and corrupt the state);
// the batch's earlier frames go through writeFrame, which produces none.
// Gated on haveR so a MONO sample never touches shiftR_ — start() only
// primes it for genuinely stereo samples, and a stale un-primed ring must
// not leak a previous note.
if (exhausted) shiftR_.freezeTail();
const AudioSample feedR =
feedOk ? crossfadedSource(pcmR, loop, feedPos_, feedXw) : 0.0f;
shiftR_.setShiftRatio(shift);
outRlocal =
static_cast<double>(shiftR_.processLinked(feedR, shiftL_.lastSplice()));
fed ? static_cast<double>(
shiftR_.processLinked(feedR, shiftL_.lastSplice()))
: static_cast<double>(
shiftR_.processNoInputLinked(shiftL_.lastSplice()));
} else {
// Mono sample in stereo mode (dual-mono): shiftL_ already produced the
// shifted value from the mono feed; mirror it to R. Do NOT call
@@ -498,9 +525,10 @@ private:
outRlocal = shiftedL;
}
}
++feedPos_;
// Preserve advances the read head at the SOURCE rate (duration preserved).
ratio_ = 1.0;
// Preserve advances the read head at the STRETCH rate — the one duration control.
// Everything downstream of it (the loop wrap, the Trigger span, the spline phase)
// therefore stays a source-frame fact and scales by construction.
ratio_ = stretchRate_;
} else {
// VARISPEED: pitch and duration coupled. The read rate carries the repitch; the
// pitch envelope multiplies the ratio for the read-rate bias (unchanged idiom when
@@ -676,9 +704,10 @@ private:
//
// The shifter rings are primed at start() with the first window of the actual upcoming
// source (silence past the end) — output frame 0 is source frame `start`, no ring-fill
// silence, and splices always land in real history. feedPos_ is the integer source frame
// fed to the shifters next; it runs exactly one window ahead of readPos_ under the same
// sustain-loop wrap rule. Once feedPos_ passes the last real frame (Gate: sample end;
// silence, and splices always land in real history. stretch_ is the integer source frame
// fed to the shifters next plus the fractional rate debt; it runs one window ahead of
// readPos_ under the same sustain-loop wrap rule and at the same rate, so the two stay one
// window apart at every stretch. Once it passes the last real frame (Gate: sample end;
// Trigger: playEnd_), the shifters' writers freeze — no padding enters the rings and the
// splice machinery recycles the frozen real tail through the note end (see advanceFrame).
// primeBuf_ is the presized scratch the prime stream is assembled into.
@@ -686,7 +715,8 @@ private:
PitchEnvelope pitchEnv_;
PitchShifter shiftL_;
PitchShifter shiftR_;
std::int64_t feedPos_ = 0;
instrument::engine::StretchCursor stretch_;
double stretchRate_ = 1.0; // Preserve playback rate, clamped and latched at note-on
std::vector<AudioSample> primeBuf_;
// lastOut{L,R}_ track the voice's most recent rendered output. A takeover/steal start()
+169
View File
@@ -23,6 +23,10 @@
// 6. unity + latency contract — asserted bit-exactly: a warm()ed shifter at ratio 1.0 IS a
// clean window delay; a prime()d one has ZERO added latency (out[i] == src[i] to the
// bit) — the GA2 immediate-onset claim.
// 9. time-stretch — the write rate (duration) and the tap rate (pitch) are independent: a
// source fed faster/slower than the output runs moves along the output timeline with its
// pitch untouched, composes with the full transposition range, and never resamples to do
// it. A resampled read is the explicit non-tautology witness in the duration test.
// 8. stereo linked lag (Q-W0 T1-01) — a follower channel driven via processLinked() mirrors
// the master's splice decision (jump/lag/frac/fadeLen AND firing frame) exactly, on
// decorrelated stereo content where an independent per-channel search provably diverges.
@@ -581,6 +585,168 @@ static void testStereoLinkedLagSharedSchedule() {
CHECK(!mirrorDiverged); // applySplice() reproduces splice() bit-identically
}
// --- 9. Time-stretch: the WRITE rate is duration, the TAP rate is pitch, and they are
// independent. Feeding faster/slower than the output runs moves the content along the
// output timeline WITHOUT moving its pitch — no resampling anywhere, which is what
// "WDL_Resampler is not a Preserve engine" asks for. ---
// Drives the shifter with a fractional feed rate the way the Voice does: all but the last
// source frame due on an output frame go through writeFrame (no output), the last through
// process(); an output frame with none due takes processNoInput(). Returns the output plus,
// via `consumed`, how much source it ate.
static std::vector<double> runStretch(const std::vector<AudioSample>& src, std::int64_t w,
double feedRate, double shift, std::size_t outFrames,
std::size_t* consumed) {
PitchShifter ps;
ps.configure(w);
ps.prime(src.data(), w);
ps.setShiftRatio(shift);
ps.setFeedRate(feedRate);
std::size_t pos = static_cast<std::size_t>(w);
double debt = 0.0;
std::vector<double> out(outFrames);
for (std::size_t i = 0; i < outFrames; ++i) {
debt += feedRate;
const int due = static_cast<int>(debt);
debt -= static_cast<double>(due);
AudioSample last = 0.0f;
bool fed = false;
for (int k = 0; k < due; ++k) {
if (fed) ps.writeFrame(last);
last = pos < src.size() ? src[pos] : 0.0f;
++pos;
fed = true;
}
out[i] = static_cast<double>(fed ? ps.process(last) : ps.processNoInput());
}
if (consumed != nullptr) *consumed = pos - static_cast<std::size_t>(w);
return out;
}
// Mean spacing between positive-going zero crossings over [from, to).
static double periodIn(const std::vector<double>& v, std::size_t from, std::size_t to) {
double sum = 0.0;
std::size_t prev = 0, count = 0;
for (std::size_t i = from + 1; i < to; ++i) {
if (v[i - 1] <= 0.0 && v[i] > 0.0) {
if (count > 0) sum += static_cast<double>(i - prev);
prev = i;
++count;
}
}
return count > 1 ? sum / static_cast<double>(count - 1) : 0.0;
}
static void testStretchMovesDurationNotPitch() {
// A source that changes pitch ONCE, at a known source frame: period 200 before it, period
// 100 after. Where that change lands in the OUTPUT is duration; what the two periods
// measure is pitch. A stretcher moves the first and not the second; a resampled read moves
// both, which is exactly the distinction under test.
const std::int64_t w = 2205;
const std::size_t change = 40000; // source frame where the period halves
const std::size_t srcLen = 160000;
std::vector<AudioSample> src(srcLen);
double phase = 0.0;
for (std::size_t i = 0; i < srcLen; ++i) {
phase += 2.0 * kPi / (i < change ? 200.0 : 100.0);
src[i] = static_cast<AudioSample>(std::sin(phase));
}
for (double rate : {0.5, 1.0, 2.0}) {
// Duration: the source is consumed at the feed rate, so the change lands at
// change/rate in the output — the run is sized to reach past it at every rate.
const std::size_t changeOut = static_cast<std::size_t>(change / rate);
const std::size_t outFrames = changeOut + 12000;
std::size_t consumed = 0;
const std::vector<double> out =
runStretch(src, w, rate, /*shift=*/1.0, outFrames, &consumed);
// Pitch: measured well clear of the transition on both sides, and UNCHANGED by the
// rate — 200 before, 100 after, at 0.5x, 1x and 2x alike.
const double before = periodIn(out, changeOut / 4, changeOut / 4 + 6000);
const double after = periodIn(out, changeOut + 2000, changeOut + 8000);
CHECK(approx(before, 200.0, 10.0));
CHECK(approx(after, 100.0, 5.0));
// Non-tautology witness: a RESAMPLED read of the same source at the same rate would
// have produced 200/rate and 100/rate here. At rate != 1 those differ from the
// measurements above by far more than the tolerances, so the assertions genuinely
// separate a stretch from a resample.
if (rate != 1.0) {
CHECK(std::fabs(before - 200.0 / rate) > 20.0);
CHECK(std::fabs(after - 100.0 / rate) > 20.0);
}
// ...and the source really was consumed at the rate (the duration half of the claim).
CHECK(approx(static_cast<double>(consumed),
static_cast<double>(outFrames) * rate, 2.0));
// No dead stretches anywhere, frame 0 included: the stretch path must not reintroduce
// the onset gap prime() exists to close.
std::size_t worstGap = 0, run = 0;
for (std::size_t i = 0; i < outFrames; ++i) {
if (std::fabs(out[i]) < 1e-3) {
++run;
if (run > worstGap) worstGap = run;
} else {
run = 0;
}
}
CHECK(worstGap < 32);
}
}
// The stretch and the transposition compose over the SAME ring, and a feed rate the shift does
// not match is where the splice fade's headroom is tightest (the outgoing tap closes on the
// writer at ratio - feedRate, which setFeedRate exists to tell it). Bounded, finite, gap-free
// across the corners of the engine's rate range crossed with the full transposition range.
static void testStretchAndShiftComposeSafely() {
const std::int64_t w = 2205;
const double f0 = 1.0 / 196.37; // the adversarial non-integer period
const std::size_t srcLen = 400000;
std::vector<AudioSample> src(srcLen);
for (std::size_t i = 0; i < srcLen; ++i) {
src[i] = static_cast<AudioSample>(std::sin(2.0 * kPi * f0 * static_cast<double>(i)));
}
for (double rate : {0.5, 0.75, 1.0, 1.5, 2.0}) {
for (double semis : {-24.0, -12.0, -5.0, 0.0, 7.0, 12.0, 24.0}) {
const double shift = std::pow(2.0, semis / 12.0);
const std::size_t outFrames = 60000;
const std::vector<double> out =
runStretch(src, w, rate, shift, outFrames, nullptr);
std::size_t worstGap = 0, run = 0;
double peak = 0.0;
for (std::size_t i = 0; i < outFrames; ++i) {
CHECK(std::isfinite(out[i]));
const double a = std::fabs(out[i]);
if (a > peak) peak = a;
if (a < 1e-3) {
++run;
if (run > worstGap) worstGap = run;
} else {
run = 0;
}
}
CHECK(worstGap < 32); // continuous: every splice landed in real, aligned history
CHECK(peak < 1.2); // complementary fades: no cancellation, no bulge
CHECK(peak > 0.8); // ...and it played at full level
// Pitch is the TAP's, not the feed's: the observed period is the source period
// divided by the shift, whatever the rate.
const double p = periodIn(out, 20000, 50000);
if (!approx(p, 196.37 / shift, 196.37 / shift * 0.12)) {
std::printf(" rate %.2f semis %.0f: period %.2f want %.2f\n", rate, semis, p,
196.37 / shift);
}
CHECK(approx(p, 196.37 / shift, 196.37 / shift * 0.12));
}
}
}
// The two new entry points on a shifter that was never configured (a Varispeed voice's) —
// neither may touch the empty ring.
static void testStretchEntryPointsOnPassThrough() {
PitchShifter ps;
CHECK(!ps.configured());
ps.writeFrame(0.5f); // no ring to write into
CHECK(ps.processNoInput() == 0.0f); // no input to pass through
CHECK(ps.process(0.25f) == 0.25f); // and the 1:1 path still passes through
}
int main() {
testDurationInvariance();
testUnityRoughlyReproduces();
@@ -590,6 +756,9 @@ int main() {
testUnityBitExactAndLatency();
testFreezeTailContinuousTone();
testStereoLinkedLagSharedSchedule();
testStretchMovesDurationNotPitch();
testStretchAndShiftComposeSafely();
testStretchEntryPointsOnPassThrough();
if (g_fail == 0) {
std::printf("all pitch_shift tests passed\n");
+307
View File
@@ -21,7 +21,10 @@
#include <algorithm>
#include <cmath>
#include <cstdint>
#include <cstdio>
#include <cstring>
#include <ctime>
#include <vector>
using namespace reasampler;
@@ -2880,6 +2883,303 @@ static void testPreserveSubWindowSampleNoZeroPadInRing() {
CHECK(blockPeak(out, 0, frames) > 0.5); // and it genuinely played at full level
}
// ---------------------------------------------------------------------------
// The Preserve read path's stretch generalization: the source is consumed at the playback
// rate while the shifter's read tap runs at the transposition, over ONE delay ring. Only
// their DIFFERENCE reaches the splice machinery.
// ---------------------------------------------------------------------------
// FNV-1a over the raw float bits — an exact-stream witness, not a tolerance.
static std::uint64_t hashStream(const std::vector<AudioSample>& v) {
std::uint64_t h = 1469598103934665603ull;
for (const AudioSample s : v) {
std::uint32_t bits = 0;
std::memcpy(&bits, &s, sizeof(bits));
for (int b = 0; b < 4; ++b) {
h ^= static_cast<std::uint64_t>((bits >> (8 * b)) & 0xffu);
h *= 1099511628211ull;
}
}
return h;
}
// A source with no symmetry a shifter could accidentally satisfy: a sine at a non-integer
// period plus a deterministic pseudo-random dither, so any change in the splice schedule,
// the fed frame sequence or the tap position moves the hash.
static SampleData stretchProbeSample(std::size_t frames, bool stereo) {
SampleData s;
s.frames.resize(frames);
if (stereo) s.framesR.resize(frames);
std::uint32_t lcg = 12345u;
for (std::size_t i = 0; i < frames; ++i) {
lcg = lcg * 1664525u + 1013904223u;
const double n = static_cast<double>(lcg >> 8) / 8388608.0 - 1.0; // [-1,1)
const double t = static_cast<double>(i);
s.frames[i] = static_cast<float>(0.8 * std::sin(2.0 * kPi * t / 196.37) + 0.1 * n);
if (stereo) {
s.framesR[i] =
static_cast<float>(0.8 * std::sin(2.0 * kPi * t / 123.13) - 0.1 * n);
}
}
s.rootNote = 60;
s.sampleRate = 44100;
s.play.adsr = flatAdsr();
s.play.pitchEngine = PitchEngine::Preserve;
return s;
}
// Renders one raw Voice (not through VoiceEngine, which publishes no rate) for `outFrames`.
static void renderVoice(const SampleData& s, int note, double rate, std::int64_t window,
bool stereo, std::vector<AudioSample>& l, std::vector<AudioSample>& r) {
Voice v;
v.presizePreserveShifters(window);
v.start(note, 127, s, /*declickTakeover=*/false, rate);
for (std::size_t i = 0; i < l.size(); ++i) {
if (stereo) {
AudioSample a = 0.0f, b = 0.0f;
v.renderFrameStereo(a, b);
l[i] = a;
r[i] = b;
} else {
l[i] = v.renderFrame();
}
}
}
// --- The null case, asserted against a baseline the SHIPPED engine produced. ---
// The four constants below were captured by running this same function against the
// pre-stretch build (phase-g, before the rate seam existed) and printing the hashes; they are
// therefore a witness that the generalized read path reproduces the shipped Preserve output
// bit for bit at rate 1.0, not a self-consistency check. A change here is a change to what
// every already-saved project sounds like — re-derive the cause before re-baselining.
static void testPreserveUnityRateIsBitIdenticalToTheShippedRead() {
const std::int64_t w = 2205; // the product window at 44.1k
const std::size_t n = 6000;
struct Case {
int note;
bool stereo;
bool loop;
std::uint64_t hashL;
std::uint64_t hashR;
};
const Case cases[] = {
{60, false, false, 16118581538698271917ull, 0ull}, // on root: unity shift
{67, false, false, 17268489061432447375ull, 0ull}, // +7 st: real splices
{55, false, false, 17626155132441637249ull, 0ull}, // -5 st: down-shift
{67, true, true, 116487689553455907ull, 9528575457480122654ull}, // stereo linked + loop
};
for (const Case& c : cases) {
SampleData s = stretchProbeSample(4000, c.stereo);
if (c.loop) {
s.loop.hasLoop = true;
s.loop.start = 1200;
s.loop.end = 3600;
s.loopCrossfadeFrames = 256;
}
std::vector<AudioSample> l(n), r(c.stereo ? n : 0);
renderVoice(s, c.note, /*rate=*/1.0, w, c.stereo, l, r);
const std::uint64_t hl = hashStream(l);
CHECK(hl == c.hashL);
if (hl != c.hashL) std::printf(" note %d L hash %lluull\n", c.note, hl);
if (c.stereo) {
const std::uint64_t hr = hashStream(r);
CHECK(hr == c.hashR);
if (hr != c.hashR) std::printf(" note %d R hash %lluull\n", c.note, hr);
}
}
}
// --- Rate changes DURATION only; the transposition alone sets pitch. ---
static void testPreserveStretchChangesDurationNotPitch() {
// Gate, no loop: the voice's life is exactly how long the source lasts, so the frame at
// which it goes idle IS the note's duration.
const std::int64_t w = 1024;
const std::size_t frames = 24000;
const double srcPeriod = 160.0;
SampleData s;
s.frames.resize(frames);
for (std::size_t i = 0; i < frames; ++i) {
s.frames[i] = static_cast<float>(std::sin(2.0 * kPi * static_cast<double>(i) / srcPeriod));
}
s.rootNote = 60;
s.play.adsr = flatAdsr();
s.play.pitchEngine = PitchEngine::Preserve;
auto run = [&](double rate, PitchEngine engine, std::size_t& lifeFrames) {
SampleData local = s;
local.play.pitchEngine = engine;
Voice v;
v.presizePreserveShifters(w);
v.start(60, 127, local, /*declickTakeover=*/false, rate);
std::vector<AudioSample> out;
out.reserve(frames * 3);
lifeFrames = 0;
for (std::size_t i = 0; i < frames * 3 && v.active(); ++i) {
out.push_back(v.renderFrame());
++lifeFrames;
}
return out;
};
std::size_t lifeUnity = 0, lifeSlow = 0, lifeFast = 0;
const std::vector<AudioSample> unity = run(1.0, PitchEngine::Preserve, lifeUnity);
const std::vector<AudioSample> slow = run(0.5, PitchEngine::Preserve, lifeSlow);
const std::vector<AudioSample> fast = run(2.0, PitchEngine::Preserve, lifeFast);
// Duration scales by 1/rate (the small excess over the source length is the terminal
// declick ring-out Preserve ends on).
CHECK(approx(static_cast<double>(lifeUnity), 24000.0, 200.0));
CHECK(approx(static_cast<double>(lifeSlow), 48000.0, 400.0));
CHECK(approx(static_cast<double>(lifeFast), 12000.0, 200.0));
// ...and the pitch does not move with it. Measured away from the onset and the tail.
auto period = [](const std::vector<AudioSample>& v, std::size_t from, std::size_t to) {
double sum = 0.0;
std::size_t prev = 0, count = 0;
for (std::size_t i = from + 1; i < to && i < v.size(); ++i) {
if (v[i - 1] <= 0.0f && v[i] > 0.0f) {
if (count > 0) sum += static_cast<double>(i - prev);
prev = i;
++count;
}
}
return count > 1 ? sum / static_cast<double>(count - 1) : 0.0;
};
CHECK(approx(period(unity, 2000, 9000), srcPeriod, 8.0));
CHECK(approx(period(slow, 2000, 9000), srcPeriod, 8.0));
CHECK(approx(period(fast, 2000, 9000), srcPeriod, 8.0));
// The non-tautology witness: VARISPEED is the engine that couples them. Reaching the same
// durations there costs exactly the pitch change Preserve refuses to make — so the three
// equal periods above are a property of the stretcher, not of the measurement.
std::size_t lifeVari = 0;
const std::vector<AudioSample> vari = run(0.5, PitchEngine::Varispeed, lifeVari);
CHECK(approx(static_cast<double>(lifeVari), 24000.0, 200.0)); // rate ignored under Varispeed
SampleData down = s;
down.play.pitchEngine = PitchEngine::Varispeed;
Voice vv;
vv.presizePreserveShifters(w);
vv.start(48, 127, down); // -12 st under Varispeed: duration doubles AND pitch halves
std::vector<AudioSample> variDown;
std::size_t variLife = 0;
for (std::size_t i = 0; i < frames * 3 && vv.active(); ++i) {
variDown.push_back(vv.renderFrame());
++variLife;
}
CHECK(approx(static_cast<double>(variLife), 48000.0, 200.0)); // same duration...
CHECK(approx(period(variDown, 2000, 9000), srcPeriod * 2.0, 16.0)); // ...at half pitch
}
// --- The onset is a regression surface: no added latency at ANY rate. ---
static void testPreserveStretchSpeaksOnFrameZeroAtEveryRate() {
const std::int64_t w = 2048;
SampleData s = stretchProbeSample(12000, false);
s.startFrame = 500; // and the first output frame is the START frame, not frame 0
for (double rate : {0.5, 1.0, 2.0}) {
for (int note : {48, 60, 67}) {
Voice v;
v.presizePreserveShifters(w);
v.start(note, 127, s, /*declickTakeover=*/false, rate);
const AudioSample first = v.renderFrame();
// The primed ring parks the tap ON the start frame, so output frame 0 is source
// frame `startFrame` exactly — at every rate and every transposition. A stretcher
// that buffered a window before speaking would fail here, which is the whole point.
CHECK(first == s.frames[500]);
// ...and it keeps speaking: no first-window dip while the schedule settles. The
// 256-frame measuring window spans most of a period even at the lowest note tested
// (-12 st stretches the probe's 196-frame period to 393), so a continuous tone
// peaks well above the floor in every one of them and only a real gap can sink it.
double lo = 1e9;
for (int i = 0; i < 20; ++i) {
double peak = 0.0;
for (int k = 0; k < 256; ++k) {
peak = (std::max)(peak, std::fabs(static_cast<double>(v.renderFrame())));
}
lo = (std::min)(lo, peak);
}
if (!(lo > 0.5)) std::printf(" rate %.2f note %d: lo %.3f\n", rate, note, lo);
CHECK(lo > 0.5);
}
}
}
// --- "Loop the source, shift the output" is unweakened by a stretch. ---
static void testPreserveStretchLoopsTheSourceSpan() {
for (double rate : {0.5, 1.0, 2.0}) {
for (int note : {48, 60, 72}) {
SampleData s;
s.frames.resize(200, 0.0f);
for (int i = 60; i < 120; ++i) s.frames[i] = 0.5f;
s.rootNote = 60;
s.sampleRate = 48000;
s.loop.hasLoop = true;
s.loop.start = 80;
s.loop.end = 120;
s.play.adsr = flatAdsr();
s.play.pitchEngine = PitchEngine::Preserve;
Voice v;
v.presizePreserveShifters(64);
v.start(note, 127, s, /*declickTakeover=*/false, rate);
std::vector<AudioSample> out(4000);
for (std::size_t i = 0; i < out.size(); ++i) out[i] = v.renderFrame();
// The loop is a SOURCE-frame fact, so it keeps the voice alive and at level for as
// long as it is held, whatever the rate consumes it at.
CHECK(v.active());
double sum = 0.0;
for (std::size_t i = out.size() - 200; i < out.size(); ++i) sum += out[i];
CHECK(approx(sum / 200.0, 0.5, 0.05));
}
}
}
// --- The 32-voice measurement gate. Asserts correctness; PRINTS the cost, which is the
// number reported for the algorithm decision (meaningful only in a Release build). ---
static void testPreserveStretchThirtyTwoVoicesHoldUp() {
const std::int64_t w = 2205; // the product window at 44.1k
const std::size_t blockFrames = 44100; // one second of audio
const std::size_t voiceCount = 32;
SampleData s = stretchProbeSample(200000, true);
s.loop.hasLoop = true; // held notes: all 32 sound for the whole run
s.loop.start = 40000;
s.loop.end = 160000;
s.loopCrossfadeFrames = 1024;
// 1.0 is the reference: it is the cost the shipped Preserve read already carries, so the
// two stretched rows are read as a delta against it rather than in isolation.
for (double rate : {1.0, 0.5, 2.0}) {
std::vector<Voice> voices(voiceCount);
for (std::size_t i = 0; i < voiceCount; ++i) {
voices[i].presizePreserveShifters(w);
voices[i].start(48 + static_cast<int>(i), 100, s, /*declickTakeover=*/false, rate);
}
const std::clock_t t0 = std::clock();
double guard = 0.0;
std::size_t sounding = 0;
for (std::size_t f = 0; f < blockFrames; ++f) {
AudioSample l = 0.0f, r = 0.0f;
for (std::size_t i = 0; i < voiceCount; ++i) {
AudioSample a = 0.0f, b = 0.0f;
voices[i].renderFrameStereo(a, b);
l += a;
r += b;
}
guard += static_cast<double>(l) + static_cast<double>(r);
CHECK(std::isfinite(l) && std::isfinite(r));
}
const double secs = static_cast<double>(std::clock() - t0) / CLOCKS_PER_SEC;
for (std::size_t i = 0; i < voiceCount; ++i) {
if (voices[i].active()) ++sounding;
}
CHECK(sounding == voiceCount); // all 32 held the whole second (the loop kept them up)
CHECK(std::fabs(guard) > 0.0); // ...and genuinely produced audio
std::printf(" [measure] 32 stereo Preserve voices @ rate %.2f: %.3f s wall for 1.0 s "
"audio (%.1f%% of one core, %.1f ns/voice/frame)\n",
rate, secs, 100.0 * secs,
secs * 1e9 / (static_cast<double>(blockFrames) *
static_cast<double>(voiceCount)));
}
}
int main() {
testEveryKeyPlaysTheLoadedCapture();
testUnplayableCaptureRefusesEveryNote();
@@ -3006,6 +3306,13 @@ int main() {
testPreservePrimeStopsAtTriggerPlayEnd();
testPreserveSubWindowSampleNoZeroPadInRing();
// The Preserve read path's stretch generalization.
testPreserveUnityRateIsBitIdenticalToTheShippedRead();
testPreserveStretchChangesDurationNotPitch();
testPreserveStretchSpeaksOnFrameZeroAtEveryRate();
testPreserveStretchLoopsTheSourceSpan();
testPreserveStretchThirtyTwoVoicesHoldUp();
if (g_fail == 0) {
std::printf("all sampler_core tests passed\n");
return 0;
+155
View File
@@ -0,0 +1,155 @@
// Standalone tests for reasampler::instrument::engine::StretchCursor — the Preserve read's
// source-feed schedule. No VST3, no REAPER, no vendor, no test framework.
//
// Covers:
// 1. rate 1.0 is EXACTLY one source frame per output frame, forever and with no residue —
// the mechanism behind the "unity is bit-identical to the shipped Preserve read" gate.
// 2. the schedule tracks the rate: over N output frames the cursor consumes N*rate source
// frames to within one, at rates either side of unity and at irrational ones.
// 3. the per-output-frame feed count never exceeds kMaxFeedPerFrame — the bound that makes
// a variable-length feed loop RT-safe.
// 4. the clamp: out-of-range folds to the bounds, unusable input folds to unity (never to a
// silent stall or a quarter-speed surprise).
// 5. the sustain loop wraps the cursor and never lets it leave [start, end) — "loop the
// source" holds at every rate, including one that steps over the loop end.
#include "../src/core/instrument/engine/time_stretch.h"
#include <cmath>
#include <cstdio>
#include <vector>
using namespace reasampler::instrument::engine;
using reasampler::instrument::engine::loop::ResolvedLoop;
static int g_fail = 0;
#define CHECK(cond) do { if(!(cond)) { \
std::printf("FAIL line %d: %s\n", __LINE__, #cond); ++g_fail; } } while(0)
static ResolvedLoop noLoop() { return ResolvedLoop{}; }
static ResolvedLoop loopSpan(std::int64_t start, std::int64_t end) {
ResolvedLoop lp;
lp.active = true;
lp.start = start;
lp.end = end;
lp.length = end - start;
return lp;
}
// --- 1. Unity is exactly one frame per output frame, with no drifting residue. ---
static void testUnityRateFeedsExactlyOneFramePerOutputFrame() {
StretchCursor c;
c.start(100);
const ResolvedLoop lp = noLoop();
for (std::int64_t i = 0; i < 200000; ++i) {
CHECK(c.due(1.0) == 1);
CHECK(c.next(lp) == 100 + i);
}
// No accumulated debt after 200k frames: the source frame the cursor is about to feed is
// exactly the one an un-stretched integer walk would be at. A residue of even one frame
// over a long note would move the shipped Preserve output.
CHECK(c.frame() == 100 + 200000);
}
// --- 2. The schedule tracks the rate. ---
static void testTotalConsumedTracksTheRate() {
const ResolvedLoop lp = noLoop();
// Includes a rate with no exact binary representation, where a naive per-frame rounding
// would drift without bound rather than carrying the residue.
for (double rate : {0.5, 0.75, 1.0, 1.3333333333333333, 2.0, 1.0 / 3.0 + 1.0}) {
StretchCursor c;
c.start(0);
const std::int64_t outFrames = 100000;
for (std::int64_t i = 0; i < outFrames; ++i) {
const std::int64_t due = c.due(rate);
for (std::int64_t k = 0; k < due; ++k) (void)c.next(lp);
}
const double expected = static_cast<double>(outFrames) * clampStretchRate(rate);
CHECK(std::fabs(static_cast<double>(c.frame()) - expected) <= 1.0);
}
}
// --- 3. The feed count is bounded — the RT-safety argument for a variable-length loop. ---
static void testFeedPerOutputFrameIsBounded() {
const ResolvedLoop lp = noLoop();
// Drive at, above and around the ceiling; an unclamped rate would run the caller's loop
// for as many iterations as the rate names.
for (double rate : {kStretchRateMax, kStretchRateMax * 100.0, 3.99, 2.5}) {
StretchCursor c;
c.start(0);
std::int64_t worst = 0;
for (std::int64_t i = 0; i < 20000; ++i) {
const std::int64_t due = c.due(rate);
if (due > worst) worst = due;
for (std::int64_t k = 0; k < due; ++k) (void)c.next(lp);
}
CHECK(worst <= kMaxFeedPerFrame);
CHECK(worst >= 1); // and the bound is not vacuous — frames genuinely fell due
}
}
// --- 4. The clamp. ---
static void testRateClamp() {
CHECK(clampStretchRate(1.0) == 1.0); // exact: the unity read depends on it
CHECK(clampStretchRate(0.5) == 0.5);
CHECK(clampStretchRate(2.0) == 2.0);
CHECK(clampStretchRate(0.001) == kStretchRateMin);
CHECK(clampStretchRate(1000.0) == kStretchRateMax);
// Unusable input plays at speed rather than stalling or quarter-speeding.
CHECK(clampStretchRate(0.0) == 1.0);
CHECK(clampStretchRate(-2.0) == 1.0);
CHECK(clampStretchRate(std::nan("")) == 1.0);
// ...and the cursor honours it rather than looping on the raw value.
StretchCursor c;
c.start(0);
CHECK(c.due(-5.0) == 1); // folded to unity
StretchCursor d;
d.start(0);
CHECK(d.due(50.0) <= kMaxFeedPerFrame);
}
// --- 5. The loop wraps the SOURCE cursor, at every rate. ---
static void testCursorStaysInsideTheLoopSpan() {
const ResolvedLoop lp = loopSpan(1000, 1040); // a 40-frame loop: rate 4 steps 10% of it
for (double rate : {0.5, 1.0, 2.0, 4.0}) {
StretchCursor c;
c.start(1000);
std::int64_t lowest = 1 << 30, highest = -1;
for (std::int64_t i = 0; i < 50000; ++i) {
const std::int64_t due = c.due(rate);
for (std::int64_t k = 0; k < due; ++k) {
const std::int64_t q = c.next(lp);
if (q < lowest) lowest = q;
if (q > highest) highest = q;
}
}
// Never reads outside the span — the "loop the source, shift the output" contract does
// not weaken under a stretch, because the span is a source-frame fact.
CHECK(lowest >= lp.start);
CHECK(highest < lp.end);
CHECK(highest == lp.end - 1); // and it genuinely covered the span
CHECK(lowest == lp.start);
}
// A cursor started BEYOND the loop end (the start-point-past-the-loop case) is pulled in on
// its first take rather than reading off the end.
StretchCursor c;
c.start(5000);
const std::int64_t q = c.next(lp);
CHECK(q >= lp.start && q < lp.end);
}
int main() {
testUnityRateFeedsExactlyOneFramePerOutputFrame();
testTotalConsumedTracksTheRate();
testFeedPerOutputFrameIsBounded();
testRateClamp();
testCursorStaysInsideTheLoopSpan();
if (g_fail == 0) {
std::printf("all time_stretch tests passed\n");
return 0;
}
std::printf("%d time_stretch check(s) failed\n", g_fail);
return 1;
}