Commits on Source 47

  • cznic's avatar
    nuc64: auto deps · 93a5ae04
    cznic authored
    93a5ae04
  • cznic's avatar
    nuc64: auto deps · ff124043
    cznic authored
    ff124043
  • cznic's avatar
    nuc64: auto deps · a137bbd1
    cznic authored
    a137bbd1
  • cznic's avatar
    driver: don't wait for the streaming goroutine in Rows.Close (audit A33) · 58b84fee
    cznic authored
    Rows.Close waited for the row-streaming goroutine to return. That
    goroutine can be parked acquiring the engine's read lock behind another
    connection's still-open write transaction -- frequently the one held by
    the very caller now blocked in Close -- and a sync.RWMutex read lock
    cannot be cancelled, so Close blocked until that transaction ended.
    
    The wait was added by 1026b9c6 to keep a straggling goroutine from
    observing a concurrently torn-down store and panicking on a nil
    dereference. That commit predates the A31 fix, which moved the same
    protection into the engine, where it belongs: DB.Close now takes the
    write lock before nil-ing store/root, and every read path re-checks for
    a closed DB after acquiring the read lock, so a straggler either
    completes before the teardown or fails cleanly with errClosedDB. The
    driver-level wait had become redundant, and it was the sole reason Close
    could block. Drop it, along with the now write-only finished channel.
    
    Contrary to the note in AUDIT.md, A33 does have a deterministic
    reproduction: under GOMAXPROCS(1) the goroutine spawned by Query stays
    runnable but unscheduled until the caller blocks, so a write transaction
    opened on another connection always takes the exclusive lock first and
    parks it. Adds TestAudit_A33_RowsCloseDeadlock, which reproduced the
    deadlock 20/20 times before this change and 0/20 after. The table it
    queries is kept smaller than the 500-row streaming buffer so that a
    goroutine winning the race runs to completion and releases the read lock
    instead of parking on a full buffer while holding it.
    
    Verified with the full four-backend suite and the gated reproduction
    suite, both plain and under -race.
    
    Fixes AUDIT.md A33.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    58b84fee
  • cznic's avatar
    bench: make the benchmark suite runnable and recordable · 532bdb6e
    cznic authored
    ./benchcmp and the bench Makefile target both omitted -vet=off, so on
    current toolchains go test failed the build on this repo's pre-existing
    vet findings and ran zero benchmarks. Neither passed -count either, so
    every log they produced was single-sample, leaving benchstat no way to
    separate a real change from noise. And bench depended on the all target,
    which runs gofmt -w and unconvert -apply among others, i.e. it rewrote
    the source tree before benchmarking it.
    
    Add bench.sh as the runner: -vet=off, -count=6 by default, a full and a
    quick tier, an overridable BENCH regexp, and a header recording the
    commit, host, Go version and run parameters so a log stays meaningful
    later. benchcmp now delegates to it, bench no longer depends on all, and
    benchbase/benchquick wrap the two tiers.
    
    See PERFORMANCE.md B1, B2 and B3.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    532bdb6e
  • cznic's avatar
    bench: cover UPDATE, DELETE, joins, the driver and per-query overhead · e7e6ad02
    cznic authored
    The pre-audit suite benchmarked SELECT, INSERT and cross joins. It had
    no coverage at all of UPDATE, DELETE, GROUP BY and aggregates, DISTINCT,
    ORDER BY top-N, LIMIT/OFFSET, index point and range lookups, equi- and
    outer joins, EXISTS, Compile, per-query fixed overhead, transaction
    overhead, Options.StatementAtomicity, or the database/sql driver path --
    the last being how essentially every user reaches the engine.
    
    Add bench_test.go with those, as <backend>/<shape> sub-benchmarks over
    mem, V1 file and V2 file, so one -bench regexp selects a backend or a
    shape across backends.
    
    Reporting differs from all_test.go deliberately: no b.SetBytes, so no
    MB/s column that actually means record/s, and an explicit rows/op metric
    where the row count is not evident from the name. ns/op always covers
    one statement execution or one fully drained result set; fixture setup
    is never inside the timed region.
    
    Two measurement notes. BenchmarkDelete times BEGIN + DELETE + ROLLBACK
    as a unit: a committed DELETE is not repeatable, and rolling back keeps
    every iteration identical without spending most of the wall clock
    restoring the fixture. The driver fixture indexes t.k so its point
    lookups are dominated by the driver's own per-call cost rather than by a
    full scan.
    
    See PERFORMANCE.md B6.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    e7e6ad02
  • cznic's avatar
    PERFORMANCE: record the audit and a benchstat-ready baseline · d468a5a0
    cznic authored
    perf-baseline-3900x.log is 284 benchmarks x 6 samples, 53 minutes on a
    3900X. It is a plain go test -bench log, so benchstat reads it directly.
    It was recorded from the tree that became the two preceding commits,
    before they were committed, hence the "(dirty tree)" marker in its
    header.
    
    PERFORMANCE.md documents nine findings about the suite itself (B1-B9)
    and seven about the engine (P1-P7). Every plan-shaped claim is
    corroborated with EXPLAIN rather than inferred from a timing:
    
      P1  UPDATE and DELETE are never planned -- EXPLAIN returns the
          statement text, not a plan -- so neither can ever use an index.
          Same predicate, same data, same index: SELECT 204x faster, UPDATE
          and DELETE 1.00x. The largest optimisation opportunity here.
      P2  Every join is a Cartesian product plus a filter: O(n*m), indexes
          ignored, no hash or merge join.
      P4  A V2 empty COMMIT costs 2.2 ms against V1's 1.5 us. Rollback is
          cheap on both, so it is specific to commit, and database/sql
          autocommit puts that floor under every statement.
      P5  ORDER BY ... LIMIT sorts the whole table; top-N saves nothing.
      P3  Recordset.Do re-plans per call, so execute-then-drain plans twice.
      P6  The driver compiles every non-prepared statement: preparing saves
          9.3 us of a 25.2 us round trip.
    
    P7 corrects AUDIT.md. The A3 resolution states that the per-statement
    savepoint "roughly doubles" file-backend write cost, which is why it is
    opt-in there. That was never measured and is too high: +4-14% on V1 and
    +28-67% on V2. AUDIT.md now carries a correction pointing at the
    numbers. Whether the default should change is left open.
    
    Recorded as known gaps rather than silently addressed: the osFile
    backend has no benchmarks, read-only fixtures are rebuilt per b.N ramp
    step (about a third of a run), and the legacy benchmarks keep their
    record/s-as-MB/s convention and their differing timed unit for
    continuity with the historical logs.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    d468a5a0
  • cznic's avatar
    ql: drive UPDATE from an index when the WHERE clause allows (performance P1) · aca61bda
    cznic authored
    UPDATE had no planner at all: EXPLAIN returned the statement text because
    there was nothing to render, and every UPDATE walked the whole table
    chain evaluating its WHERE clause per row, however selective the
    predicate and whatever indexes existed. An index was worth ~200x on a
    point SELECT and exactly nothing on the point UPDATE of the same column.
    
    The 14 interval methods on indexPlan all ended the same way -- read the
    row at handle h, hand (id, data) to the caller -- so factor that emit
    point out into indexPlan.doHandles and leave do a thin wrapper that
    layers row reading on top. SELECT and UPDATE now run the same interval
    code, so an index cannot select differently for one than for the other.
    
    updateStmt.exec then asks indexPlanFor whether an index narrows its WHERE
    clause. That mirrors what whereRset does for a SELECT, including
    splitting an AND into its conjuncts and chaining them through filter, so
    c > lo && c < hi collapses to a single interval rather than to the first
    conjunct alone.
    
    Two properties keep this safe rather than merely fast. The candidate set
    only has to be a superset: the full WHERE clause is re-evaluated for
    every candidate, so covering one conjunct of an AND is useful, and a
    planning error is reported as "no index applies" -- index selection can
    change how long a statement takes, never what it does. And the handles
    are materialized before anything is written: walking the index while
    updating through it would let a row sort ahead of the cursor and be
    updated twice, the Halloween problem, and would mutate the B-tree under
    its own iterator.
    
    DELETE keeps scanning, with the reason recorded on deleteStmt.exec: a
    table's rows form a singly linked chain and unlinking one needs its
    predecessor's handle, which only a scan provides. Making DELETE
    index-driven means giving the chain back-links, an on-disk format change.
    
    EXPLAIN UPDATE now renders the chosen plan. TestUpdateIndexEquivalence
    guards the invariant across three backends and 31 cases -- every index
    shape, NULL three-valued logic, bool indexes, blob payloads and four
    Halloween cases -- that adding an index does not change what an UPDATE
    does: same rows, same values, same RowsAffected, same error.
    
    Measured in isolation against perf-baseline-3900x.log:
    
      Update/mem/point-index    712.95us -> 8.87us    -98.76% (n=6)
      Update/file/point-index    3.874ms -> 1.574ms   -59.39% (n=8)
      Update/file2/point-index   3.651ms -> 2.641ms   -27.66% (n=8)
      Update/mem/point-noindex   681.4us -> 686.8us   ~ (p=0.393, n=10)
    
    The scan path is unchanged, as it must be. The file backends gain less
    than mem because what remains is dominated by the per-statement commit,
    not by the scan that was removed.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    aca61bda
  • cznic's avatar
    PERFORMANCE: record the P1 resolution and a benchmark isolation caveat · 5dc03774
    cznic authored
    Documents how UPDATE was made index-driven, the measured before/after,
    and why DELETE cannot follow without an on-disk format change. P2, the
    Cartesian-product joins, is now the largest remaining opportunity.
    
    Also records a methodology trap hit while measuring this fix. Running
    file-backend benchmarks in the same process as in-RAM ones perturbs the
    latter: page-cache pressure and write-back stalls do not respect
    benchmark boundaries. Measured mixed, the P1 scan path looked like a
    +110% regression on mem and +80%/+169% on the file backends, all at
    p<=0.002. Re-run one backend per process at -count=10 the same
    comparison is +0.8% at p=0.393, i.e. no change. A small p says a
    difference is real, not that it came from the code under test. Note too
    that -bench matches each /-separated element unanchored, so file selects
    file2 as well.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    5dc03774
  • cznic's avatar
    doc: a result set without ORDER BY has unspecified order · 50e770f7
    cznic authored
    QL never stated whether a SELECT without an ORDER BY returns its records
    in a defined order. In practice callers could observe whatever order the
    chosen plan happened to produce -- for a join, the nested loop's (left
    row, right row) input order -- and testdata.ql case 46 encoded exactly
    that.
    
    Leaving it unstated makes plan changes intractable: any faster plan that
    visits rows in a different order becomes a compatibility question rather
    than an optimisation. State the SQL-conventional rule instead. The order
    is an artifact of the plan, the planner may choose differently whenever
    that is faster, and two releases, two storage backends or the mere
    presence of an index may return the same records in different orders.
    
    Also records the corollary that OFFSET and LIMIT select a window of the
    result set, so they are only meaningful together with an ORDER BY that
    makes that window well defined.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    50e770f7
  • cznic's avatar
    ql: merge join for two-way equi-joins (performance P2) · efb99714
    cznic authored
    Every join was a Cartesian product plus a filter: EXPLAIN showed the
    equality applied as a post-filter over the full product, never used to
    drive a lookup, with no hash or merge join anywhere. Joining two n-row
    tables cost Theta(n*m) whatever indexes existed.
    
    mergeJoinDefaultPlan sorts both sides on the join value and walks the two
    sorted streams once, Theta((n+m)*log(n+m)). Sorting uses the same
    backend-provided temporary B-tree that ORDER BY, GROUP BY and DISTINCT
    already use, so it is disk-backed on the file backends rather than held
    in RAM; only the run of rows sharing one join value is buffered.
    
    As with the UPDATE index work, the join is a candidate generator and not
    the decision: the original predicate stays as a filter above it, so
    merging can only narrow the pairs the filter sees. That matters where
    collation order and QL equality disagree -- NaN collates equal to itself
    but NaN == NaN is false -- and keeps no enumeration of such cases
    load-bearing. Rows whose join value is NULL are dropped from both sides,
    since NULL == x is NULL.
    
    Rows come out in join-value order rather than in the nested loop's input
    order, which the preceding commit makes a permitted difference.
    testdata.ql case 46 is a join without ORDER BY and is updated to the new
    order; it was the only golden case that observed it. An earlier version
    restored the old order by materializing the result, which cost 30% of
    the win.
    
    Applies to a two-way cross join whose sides are both plain table scans,
    equating one column of each, with both columns of the same type. The
    type restriction preserves behaviour: QL's == is not defined between
    arbitrary types, and a merge comparing by collation alone would silently
    return no rows where the nested loop reports a type error. Three-way
    joins, joins under an AND, a side already narrowed to an index scan and
    the explicit OUTER JOIN forms keep the nested loop.
    
    TestMergeJoinEquivalence runs 28 queries on three backends with the
    optimisation on and off. How strictly the two runs are compared is set
    per case, since order is only defined under an ORDER BY: exact where the
    statement orders its result or declines the optimisation, same-records
    otherwise, and row count only for LIMIT/OFFSET without an ORDER BY --
    that selects an arbitrary window of an unordered result and so returns
    different records under a different plan. TestMergeJoinApplies pins down
    where the optimisation engages, so equivalence cannot pass by never
    firing.
    
    Measured against perf-baseline-3900x.log, 300x300 rows:
    
      Join/mem/equi-noindex     82.65ms -> 633.8us   -99.23%
      Join/mem/equi-index       83.30ms -> 669.3us   -99.20%
      Join/file2/equi-noindex   89.96ms -> 3.779ms   -95.80%
      Join/file/equi-noindex   129.60ms -> 15.41ms   -88.11%
    
    geomean -96.61%, with allocated bytes down by a similar margin. The gain
    grows with table size. V1 gains least because its CreateTemp builds real
    temporary files.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    efb99714
  • cznic's avatar
    PERFORMANCE: record the P2 resolution · 9b88f1ad
    cznic authored
    Documents the merge join, the numbers, where it applies and where the
    nested loop remains, and the ordering decision that lets it stream its
    output instead of materializing the result.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    9b88f1ad
  • cznic's avatar
    PERFORMANCE: re-record the baseline after the P1 and P2 fixes · 781dfbf0
    cznic authored
    The recorded reference now measures a tree that has index-driven UPDATE
    and the merge join, so later work compares against current behaviour
    rather than against plans that no longer exist. Full tier, -count=6, 284
    benchmarks, 53 minutes, clean tree at 9b88f1ad.
    
    The before-numbers quoted in PERFORMANCE.md P1 and P2 were taken against
    the superseded baseline, which remains recoverable from this file's
    history.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    781dfbf0
  • cznic's avatar
    ql: require modernc.org/internal v1.1.12 (performance P4) · 1f667a06
    cznic authored
    A V2 commit spent most of its time zeroing a memory page it was about to
    unmap. In the profile of an empty V2 commit fsync is ~3% of CPU while
    WAL.commit -> internal/file.Truncate is 92%, and 77% of that is
    runtime.memclrNoHeapPointers -- so this was never about durability.
    
    V2 truncates its journal back to its start on every commit, empty or
    not, and internal/file.Truncate zeroed the tail of the page the old size
    fell in before unmapping that very page. Zeroing is what makes the
    region above the old EOF read as zero once a file grows over it, so it
    is needed when growing; when shrinking, everything above the new size is
    discarded anyway. With the default 1 MiB pages that was a megabyte of
    writes per commit plus the page faults to back them.
    
    Fixed upstream by one condition in internal/file.Truncate, released as
    v1.1.12. Measured here, isolated against v1.1.11, n=8:
    
      Transaction/file2/insert1-rollback   1154.1us -> 120.8us  -89.53%
      Transaction/file2/empty-commit       2146.6us -> 924.0us  -56.95%
      Delete/file2/point-noindex            2.206ms -> 1.109ms  -49.72%
      Update/file2/point-index              2.377ms -> 1.231ms  -48.21%
      StatementAtomicity/file2/on/1row      5.737ms -> 3.101ms  -45.95%
    
    geomean -41.6% over the V2 write paths; empty-rollback is unchanged, as
    it must be, since rollback does not truncate. The full test suite went
    from ~39s to ~24s. V2 is now faster than V1 on the same point UPDATE,
    inverting the usual "V1 for small transactions" guidance for that shape.
    
    A first comparison, a subset run against the full-suite baseline, showed
    three regressions of +265%, +114% and +168% at p<=0.01. Re-run as
    isolated back-to-back A/Bs all three are ~ (p=0.105, 1.000, 0.514).
    Reads do not truncate, which is what made them implausible enough to
    check rather than report.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    1f667a06
  • cznic's avatar
    ql: V2 scratch buffers must be per call, not per storage (audit A49) · a23972b3
    cznic authored
    dbStorage kept four fixed scratch buffers as struct fields -- for a
    record, a B-tree key, a value and a varint -- and reused them for every
    call. DB.rwmu is a read/write lock, so QL runs any number of readers at
    once: two concurrent SELECTs had one copying a record into the shared
    buffer via internal/file.(*file).ReadAt while the other decoded that same
    memory in decode2 -> binary.Uvarint. A torn decode is not merely a race
    detector warning, it is a silently wrong row.
    
    Reproduced with four connections running read-only, fully drained
    SELECTs -- no writes, no prepared statements, no undrained Close. Races
    on ql2 under -race; ql-mem and ql are unaffected. Predates the recent
    work: the repro involves nothing this cycle changed.
    
    Scratch is per call now and the fields are gone; the buffers are 32
    bytes or less. That costs about 7-9% of V2 read throughput
    (Query/file2/scan/all 471 -> 513 us, where/point-noindex 930 -> 994 us,
    both +-1-3%), because the locals escape through the ReadAt interface
    call and so heap-allocate. Routing them through the buffer pool the
    large-record path already uses recovers only 1.15%, which does not pay
    for the indirection or for having to reason about whether decode2
    retains. Correctness first.
    
    TestConcurrentReads covers it on all three backends.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    a23972b3
  • cznic's avatar
    ql: take the write lock before entering the storage on BEGIN (audit A50) · 7303ab74
    cznic authored
    run1 called db.store.BeginTransaction() and only then db.rwmu.Lock().
    Beginning a transaction mutates storage-level state -- the V2 backend
    swaps the File readers read through, the V1 one installs an
    in-transaction page cache -- and readers hold only rwmu.RLock, not
    db.mu, while they iterate, so db.mu gave them no protection. Every read
    in flight raced with the start of a write transaction.
    
    The already-in-a-transaction branch a few lines below has always taken
    the lock in the correct order, so this was an inconsistency inside one
    function rather than a design question. Reorder to match it, and unlock
    again if the storage refuses to begin.
    
    Reproduced with readers querying while other connections open write
    transactions; races on both file backends under -race, in lldb's
    bitFiler on V1 and on dbStorage.File on V2.
    
    TestConcurrentReadsAndBegin covers it on all three backends.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    7303ab74
  • cznic's avatar
    driver: cache compiled statements per connection (performance P6) · bf626304
    cznic authored
    Exec and Query called Compile on every invocation, which for a short
    statement is a large part of the round trip: preparing a statement,
    whose only saving is that compilation, took a third off a point query.
    
    driverConn now keeps a compiled-statement cache, bounded at 128 entries
    so a caller generating unique SQL cannot grow it without limit. No
    locking -- database/sql uses a driver.Conn from one goroutine at a time
    and the cache lives and dies with the connection. Prepare deliberately
    does not use it: its List is already amortized by the caller, so caching
    there would only widen what is shared.
    
    Reusing a compiled List across executions is what a prepared statement
    already does, so this asks nothing new of a statement implementation.
    Measured, n=8:
    
      Driver/exec/insert-tx-rollback   21.79us -> 14.69us  -32.59%
      Driver/exec/update-autocommit    19.59us -> 13.38us  -31.69%
      Driver/query/columns-only        14.38us -> 10.92us  -24.06%
      Driver/query/point               23.63us -> 18.15us  -23.22%
      Driver/query/point-prepared      14.89us -> 14.96us  ~
      Driver/query/scan-all            823.4us -> 826.6us  ~
    
    geomean -19.56%, and the two that should not move do not.
    
    The concurrency tests written for this cache -- one connection reusing a
    List whose previous execution may still be unwinding, and several
    connections running the same text at once -- found two data races that
    had nothing to do with it and predated it. Both are fixed in the two
    preceding commits, A49 and A50.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    bf626304
  • cznic's avatar
    PERFORMANCE: re-record the baseline after P4, P6 and the A49/A50 fixes · 9f99974a
    cznic authored
    The reference now measures a tree with the V2 commit fix, the driver's
    compiled-statement cache, and the two concurrency fixes -- including the
    ~8% the A49 fix costs V2 reads, so that price is visible in the baseline
    rather than only in a commit message. Full tier, -count=6, 284
    benchmarks, 52 minutes, clean tree at bf626304.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    9f99974a
  • cznic's avatar
    driver: concurrency stress tests, and record the pass in AUDIT · 4e480fb1
    cznic authored
    A49 and A50 were both found by accident while benchmarking, which said
    little for how well this area was covered. These drive each backend
    through database/sql from several goroutines at once: every read plan
    shape (index point and range lookups, scan, GROUP BY, DISTINCT, ORDER BY
    with LIMIT, count, merge join), a prepared statement shared between
    goroutines, autocommit writes, explicit transactions both committed and
    rolled back, result sets abandoned undrained, and index and table DDL run
    against a table while it is being read.
    
    Two things keep them from being smoke tests. Table s is never written,
    so readers assert its contents exactly; and each writer owns a disjoint
    key range, so no operation is allowed to fail -- any error is a failure,
    not just a race. The DDL test leans on an invariant worth stating: an
    index changes how a SELECT is answered, never what it returns, so the
    readers assert exact results while an index on the table they are
    reading is repeatedly created and dropped.
    
    Nothing further turned up, at -count=4 and -cpu 1,4,24. AUDIT.md records
    that, and the two properties worth carrying forward: mem has its own
    RWMutex and stayed clean through both bugs while the file backends rely
    entirely on the engine's rwmu -- a read/write lock, so anything a file
    backend mutates while serving a read is unprotected by construction --
    and a store-mutating operation must take rwmu.Lock before entering the
    storage, not after.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    4e480fb1
  • cznic's avatar
    ql: statement atomicity on by default; add NoStatementAtomicity (audit A3) · 02037485
    cznic authored
    A3 is a Tier-1 durable-corruption finding: an updating statement that
    fails partway through leaves the table and its indexes half-changed. The
    fix -- running each updating statement in its own savepoint -- has been
    available since 2026-07-14 but opt-in on the file back ends, on the
    stated grounds that it roughly doubles their write cost. That figure was
    never measured.
    
    Measured, it is wrong in both directions. On V1 statement atomicity is
    close to free: +1.6% on a single-row UPDATE, +0.7% on a single-row
    DELETE, nothing on a thousand-row UPDATE. On V2 it is far worse than
    "doubles" for small statements: +39% single-row UPDATE, +95% single-row
    DELETE, +560% single-row INSERT. The V2 cost is per statement rather than
    per row -- see the new PERFORMANCE.md P8, a nested transaction there
    creates, maps, unmaps and unlinks a temporary file -- so the smaller the
    statement the larger the multiple.
    
    The default flips anyway. Correctness after a failed statement is worth
    paying for, V1 pays essentially nothing, and V2's share is an
    implementation artifact that P8 can remove rather than an inherent price.
    Options.NoStatementAtomicity opts out for a workload that would rather
    have the writes.
    
    A bool's zero value cannot express "on by default", so the old
    StatementAtomicity field is deprecated and ignored rather than
    reinterpreted; every existing call site still compiles, and one that set
    it true still gets what it asked for. This needs no new major version: no
    correct program stops working, some writes get slower.
    
    BenchmarkStatementAtomicity gains the single-row DELETE and INSERT cases,
    which are what exposed the shape of the V2 cost -- the previous two
    UPDATE cases happened to sample where the fixed overhead is diluted.
    TestStatementAtomicity pins the default, the opt-out actually opting out,
    and the deprecated field, on both formats and mem.
    
    doc.go's change list, silent since 2026-07-06, also gains entries for
    this, for the result-ordering decision, for the two data races fixed in
    A49/A50, and for Rows.Close no longer waiting -- which supersedes the
    2026-07-06 entry claiming it does.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    02037485
  • cznic's avatar
    file2: reuse nested transaction journals instead of a temp file each (P8) · 229045db
    cznic authored
    dbStorage.beginTransaction treats nested transactions very differently
    from top level ones. At the top it points the storage at the existing
    WAL; nested -- which is what a savepoint is, and therefore what every
    updating statement opens now that statement atomicity is on by default --
    it created a temporary file, memory mapped it, built a second WAL over
    the first, and pop() then unmapped and unlinked it. A file create, mmap,
    munmap and unlink per statement.
    
    That was the whole of V2's statement-atomicity cost and explains its
    shape: the overhead was per statement, not per row, so a single-row
    INSERT paid +560% while a thousand-row UPDATE paid +11%. V1 has no
    equivalent -- lldb's nested transactions are in memory -- which is why it
    measured at 0-4%.
    
    Journals are interchangeable once their transaction has ended, so pop
    now returns one to a free list and nestedJournal hands it back out,
    truncated first because NewWAL recovers from whatever a journal already
    holds. The free list is bounded by the deepest nesting the DB has
    reached, one entry in the common case, and dropSpares releases it on
    Close.
    
    They stay files rather than becoming cfile.Mem, deliberately: one
    statement's writes can be arbitrarily large, so an in-memory journal
    would trade a bounded disk cost for an unbounded memory one, and
    Options.TempFile exists precisely so an encrypted back end can keep its
    data off the plain filesystem -- memory would quietly bypass that hook.
    
    Isolated A/B, n=8:
    
      StatementAtomicity/file2/on/insert-1row   804.0us -> 354.9us  -55.85%
      StatementAtomicity/file2/on/update-1row    3.377m ->  2.545m  -24.64%
      StatementAtomicity/file2/on/update-all     4.492m ->  4.446m  ~
    
    geomean -24.4%; the full test suite went from 47.8s to 20.8s. The split
    is what the diagnosis predicts: the fixed per-statement cost is gone, the
    part proportional to what the statement writes is not. Measured off vs on
    in one binary, V2 statement atomicity now costs +6.1% on a single-row
    UPDATE, +8.4% on a thousand-row UPDATE, +36% on a single-row DELETE and
    +101% on a single-row INSERT, against +39%, +11%, +95% and +560% before.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    229045db
  • cznic's avatar
    PERFORMANCE: re-record the baseline after the atomicity default and P8 · 547191b6
    cznic authored
    First baseline taken with statement atomicity on by default, so its write
    benchmarks measure a different configuration than every earlier one, not
    just a faster build of the same. Comparisons straddling that commit will
    show the atomicity cost mixed in with whatever else changed.
    
    292 distinct benchmarks now, up from 284: BenchmarkStatementAtomicity
    gained single-row DELETE and INSERT on both formats, which are what
    showed the V2 cost to be per statement rather than per row.
    
    Full tier, -count=6, 1752 samples, 69 minutes, clean tree at 229045db.
    The run is longer than the previous 52 minutes despite P8, because the
    benchmark mix is write-heavy and now pays for atomicity on every
    updating statement; the test suite, which is not, went the other way,
    from 47.8s to 20.8s.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    547191b6
  • cznic's avatar
    ql: re-plan a reused recordset only when the schema changed (performance P3) · 795764e6
    cznic authored
    The A23 fix made DB.do re-plan on every Recordset.Do, so that a recordset
    reused after a later transaction cloned the tables sees current data
    rather than a stale table clone. Correct, but unconditional: the ordinary
    execute-then-drain of one query planned it twice, which on a small query
    is a large part of its cost.
    
    DB.schemaGen counts the times the table set was replaced. beginTransaction
    clones every table and rollback puts the clones back, so a plan built
    before either points at table structs that are no longer live -- which is
    exactly the condition A23 described. A recordset records the generation it
    was planned against and DB.do re-plans only on a mismatch, leaving the
    guarantee unchanged and dropping the redundant work.
    
    A rollback bumps the counter rather than restoring it, so a recordset
    straddling a rolled-back transaction re-plans once unnecessarily. Correct,
    just not optimal, on a path that previously paid a plan every time.
    
    Isolated A/B, n=8:
    
      QueryOverhead/mem/drain-only        709.2ns -> 264.8ns  -62.66%
      QueryOverhead/file2/drain-only     1017.0ns -> 533.9ns  -47.50%
      QueryOverhead/mem/execute+drain      1.626us -> 1.213us -25.43%
      QueryOverhead/file2/execute+drain    2.009us -> 1.562us -22.27%
      QueryOverhead/file/execute+drain     2.286us -> 1.871us -18.18%
      Query/mem/exists                    1513.2us -> 883.1us -41.64%
      Query/mem/where/point-index          3.374us -> 2.396us -28.99%
    
    Query/mem/exists gains most because planning a WHERE EXISTS runs its
    subquery, so re-planning per Do re-ran it. Whole-table scans barely move,
    per-row work dominating there, and a few came out 1-3% slower -- either
    noise or the reused execCtx carrying a populated plan cache where a fresh
    one was empty.
    
    Verified two ways. TestAudit_A23_RecordsetReuse fails with exactly the
    original stale-data symptom when the invalidation is disabled, so it does
    guard this. TestRecordsetReplan covers what a generation counter in
    particular could get wrong: reuse across a ROLLBACK, across CREATE INDEX
    and ALTER TABLE, across a DROP and recreate under the same name, and
    repeated reuse with no change at all.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    795764e6
  • cznic's avatar
    ql: merge join for a join carrying extra predicates (performance P2) · 78f351b7
    cznic authored
    The merge join only engaged for a WHERE clause that was nothing but the
    join predicate. `... WHERE t.k == u.k && t.v > 10` -- arguably the more
    common shape, since one rarely joins without also filtering -- took the
    andand branch of whereRset.planBinOp, which never reached the merge join
    hook, and fell back to the Cartesian product.
    
    Take whichever conjunct equates one column of each side and merge on it,
    keeping the whole clause as a filter above. That is the same
    candidate-generator property the single-predicate case already relies on:
    the merge only narrows the pairs, the clause still decides, so the other
    conjuncts need no special handling.
    
    A side may now also be an index scan, not just a plain table scan. Both
    carry the *table the type check needs and yield the same (id, row) shape,
    and narrowing a side does not stop it being sortable on the join column.
    That lets the two optimisations compose: a conjunct that pushes down
    narrows one side and the merge then reads the narrowed side.
    
    Isolated A/B, n=8:
    
      Join/mem/equi-filtered           78.51m -> 671.7us  -99.14%
      Join/mem/equi-filtered-index     71.89m -> 671.1us  -99.07%
      Join/file2/equi-filtered         88.16m -> 3.903m   -95.57%
      Join/file2/equi-filtered-index   79.87m -> 3.957m   -95.05%
    
    geomean -86.9%; the unfiltered cases are unchanged, as they should be.
    
    The two differential cases that carry an extra conjunct move from exact
    to set comparison: they now merge, so their row order legitimately
    differs from the nested loop's, and only the records must match. Five
    more cases cover the newly reached shapes, including ones whose extra
    conjunct pushes into an index, and TestMergeJoinApplies gains three
    entries so the equivalence test cannot pass by never engaging.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    78f351b7
  • cznic's avatar
    PERFORMANCE: re-record the baseline after P3 and the filtered-join merge · 8fa9377e
    cznic authored
    Both moved query paths substantially: the double plan per query is gone,
    and a join carrying extra predicates now merges instead of building the
    Cartesian product.
    
    298 distinct benchmarks, up from 292: BenchmarkJoin gained equi-filtered
    and equi-filtered-index, the shapes the filtered-join work reached, and
    the latter also exercises a side narrowed to an index scan.
    
    Full tier, -count=6, 1788 samples, 70 minutes, clean tree at 78f351b7.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    8fa9377e
  • cznic's avatar
    ql: narrow DELETE's walk with an index (performance P1, option D) · 42bb5d38
    cznic authored
    DELETE cannot skip its chain walk: a row is unlinked from its
    predecessor, and only a scan knows it. But when an index narrows the
    WHERE clause the rows to delete are known before the walk starts, and a
    row that is not a candidate then needs to yield nothing but its
    successor -- no column read, which also skips expanding its blob chunks,
    no name map, and no WHERE evaluation.
    
    Reuses the indexCandidates helper UPDATE already uses, so there is no
    format change and no compatibility question.
    
    Isolated A/B, n=8:
    
      Delete/mem/point-index      679.1us -> 205.6us  -69.72%
      Delete/file2/point-index     1.983m ->  1.446m  -27.07%
      Delete/file/point-index      2.339m ->  1.854m  -20.75%
    
    Un-indexed cases and whole-table deletes are unchanged, as they must be:
    with no index there are no candidates, and when every row is a candidate
    there is no fast path. Statements that attempt the narrowing and get
    nothing from it pay ~2% for the planning attempt, once per statement.
    
    The earlier note in PERFORMANCE.md -- that a candidate set could at best
    save the WHERE evaluation while the row reads, the dominant cost,
    remained -- was right for the file backends and wrong for mem, where the
    read is nearly free and the per-row interpretation is the cost. Hence
    70% rather than the 20-40% predicted.
    
    updateindex_test.go, whose harness turned out to be statement-agnostic,
    gains 16 DELETE cases across three backends asserting the same property
    it already asserts for UPDATE: adding an index must not change what the
    statement does. Point, range, interval, AND with a non-indexable
    conjunct, OR, no match, no WHERE, first and last row of the chain, NULL
    three-valued logic, bool, string and blob payloads.
    
    PERFORMANCE.md also records what full index-driven DELETE would cost --
    back-links, in the row or beside it, or tombstones -- since insert is a
    prepend doing zero neighbour writes today and all three variants tax it.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    42bb5d38
  • cznic's avatar
    PERFORMANCE: re-record the baseline after the DELETE index narrowing · c067bdf0
    cznic authored
    Full tier, -count=6, clean tree at 42bb5d38.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    c067bdf0
  • cznic's avatar
    ql: merge join for LEFT OUTER JOIN (performance P2) · 70a04216
    cznic authored
    LEFT OUTER JOIN still built the Cartesian product and filtered it, the
    same Theta(n*m) the inner join had before the merge join landed.
    
    The inner join's safety property does not transfer. It is a candidate
    generator with the predicate re-applied by a filter above, but a filter
    above an outer join would discard exactly the NULL-padded rows, since the
    ON clause is NULL for them. So mergeLeftJoinPlan evaluates ON itself, on
    each candidate pair, exactly as the nested loop does. That rebuilds the
    property by another route: collation only brings candidates together, QL
    equality still decides, and where the two disagree the answer stays right
    -- two NaN rows collate equal, become candidates, ON reports false, no
    pair is emitted, and the left row is padded, which is correct.
    
    NULL handling falls out of the same reading. A left row with a NULL join
    value is kept and always padded, matching nothing but still belonging to
    the result; a right row with a NULL join value is dropped, an unmatched
    right row not being part of a left join's result at all.
    
    Scoped to two-way LEFT. RIGHT reorders its output fields and FULL tracks
    matched right rows in a side B-tree, so each is its own implementation
    rather than a variation on this one; both keep the nested loop.
    
    Isolated A/B, n=8:
    
      Join/mem/left-outer     72.06m -> 625.0us  -99.13%
      Join/file2/left-outer   79.60m -> 3.778m   -95.25%
      Join/file/left-outer   113.42m -> 14.41m   -87.29%
    
    Inner joins and the RIGHT/FULL forms are unchanged. file2/equi-index came
    out +129% in the first comparison and is noise: measured alone at n=12 it
    is ~ at p=0.143, with a 368% spread.
    
    15 differential cases across three backends: duplicate keys on both
    sides, keys on one side only, NULL join values either side, an empty
    side, other key types, composition with ORDER BY and count(), and the
    shapes that must decline. One of them found a real bug -- an empty right
    side took an early exit and returned nothing, where every left row must
    come back padded.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    70a04216
  • cznic's avatar
    ql: merge join for RIGHT and FULL OUTER JOIN; fix A51 · a936300a
    cznic authored
    LEFT already merged; RIGHT and FULL still built the Cartesian product.
    All three are now one plan with a flag saying which side keeps its
    unmatched rows, that being the only thing that differs between them:
    LEFT pads and emits unmatched left rows, RIGHT unmatched right rows,
    FULL both. A side's NULL-valued rows are kept only when that side is
    preserved -- a NULL join value matches nothing, but it still belongs to
    the result of a preserved side.
    
    Doing this found a wrong-results bug in the existing FULL OUTER JOIN,
    recorded as AUDIT.md A51. fullJoinDefaultPlan visits the right side only
    from inside the left side's callback and records unmatched right rows
    there, so an empty left side never visits the right side at all and
    returned nothing -- where a FULL OUTER JOIN must return every right row,
    padded, all of them being unmatched by definition. It surfaced because
    the merge implementation returned those rows and the differential test
    flagged the disagreement.
    
    Both plans are fixed. The nested loop still runs every shape the merge
    does not take, so fixing it there was not optional.
    TestFullJoinEmptyLeft guards it with a non-equi ON clause, which declines
    the merge and so exercises fullJoinDefaultPlan; it returns 0 rows instead
    of 5 without the fix.
    
    Isolated A/B, n=8:
    
      Join/mem/right-outer     71.37m -> 622.0us  -99.13%
      Join/mem/full-outer      71.56m -> 616.8us  -99.14%
      Join/file2/right-outer  101.73m ->  12.80m  -87.42%
    
    BenchmarkJoin gains full-outer, which had no coverage. 16 more
    differential cases: RIGHT and FULL both ways round, other key types,
    either side empty, composition with ORDER BY and count(), and the shapes
    that must decline.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    a936300a
  • cznic's avatar
    PERFORMANCE: re-record the baseline after the outer join merges · 0a1ff249
    cznic authored
    301 distinct benchmarks, up from 298: BenchmarkJoin gained full-outer,
    which had no coverage at all -- part of why A51, a FULL OUTER JOIN
    returning nothing when its left side is empty, went unnoticed.
    
    Full tier, -count=6, 1806 samples, 71 minutes, clean tree at a936300a.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    0a1ff249
  • cznic's avatar
    ql: journal nested transactions in 4 kB pages, not 64 kB · 8dcc868f
    cznic authored
    Profiling the pooled nested journal (PERFORMANCE.md P8) put 63% of a
    single-row INSERT in runtime.memmove, 97% of it under the journal's
    WriteAt. The cost is page granularity, not the journal machinery: the
    first write to a page reads the whole page out of the parent, copies it
    into the journal, and copies it back at commit. memPage.setTail writes
    exactly 8 bytes and was 35% of all WAL.WriteAt time -- 64 kB of traffic
    for those 8 bytes.
    
    A nested journal is a temporary file that never outlives its
    transaction, so no existing database can be holding one that a future
    binary has to replay, and its page size is free to change. Matching the
    allocator's own 4 kB page gives, isolated A/B, n=6:
    
        StatementAtomicity/file2/on/insert-1row  243.1us -> 199.5us  -17.96%
        StatementAtomicity/file2/on/delete-1row  1.651ms -> 1.445ms  -12.48%
    
    both p=0.002, while the /off arm -- which opens no nested transaction --
    does not move, which is the check that this is really the nested
    journal. A sweep of 2/4/8/16/64 kB flattens out at -16% between 2 and
    8 kB and falls off above it, so 4 kB is on the plateau.
    
    The same read-modify-write applies to the top-level journal, where it is
    worth more: memmove is 58% of a bulk V2 insert and 60% of a single-row
    INSERT with statement atomicity off, i.e. with no nested journal at all.
    wal2PageLog is left alone regardless, because NewWAL rejects a journal
    whose recorded page size differs from the compiled-in one, so a database
    that crashed under the old value could not replay under a new binary.
    Recorded as P9; the fix belongs upstream.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    8dcc868f
  • cznic's avatar
    ql: journal the top-level V2 transaction in 4 kB pages too · 0a2f63af
    cznic authored
    PERFORMANCE.md P8 found V2's write cost to be page granularity rather
    than journal machinery -- the first write to a page copies the whole
    page into the journal and commit copies it back, so an 8-byte change
    moved 64 kB -- and fixed it for nested journals, where the page size was
    free to change. This does the same for the top-level journal, where it
    was not.
    
    The top-level journal is a named file that outlives a crash, and NewWAL
    required a journal's recorded page size to equal the compiled-in one, so
    a store that crashed under one value could not be recovered by a binary
    built with another. modernc.org/file v1.1.3 replays a journal at the
    size recorded in its own trailer instead, which unfreezes this.
    
    Isolated A/B, n=6, two binaries differing only in wal2PageLog:
    
        StatementAtomicity/file2/off/insert-1row  129.2us -> 83.45us  -35.43%
        StatementAtomicity/file2/on/insert-1row   214.1us -> 160.8us  -24.89%
        StatementAtomicity/file2/on/delete-1row   1.562ms -> 1.379ms  -11.72%
        StatementAtomicity/file2/on/update-1row   2.743ms -> 2.434ms  -11.26%
        StatementAtomicity/file2/on/update-all    4.394ms -> 4.154ms   -5.45%
        InsertFileV21kB* (five shapes)                                     ~
    
    all but the last at p<=0.026. The /off row is the cleanest: no nested
    journal exists there, so it is the top-level journal alone.
    
    The bulk inserts not moving corrects the expectation P9 was recorded
    with. memmove really is 58% of a bulk insert, but that is not
    amplification: a transaction writing a thousand 1 kB rows fills the
    pages it touches, so copying them is the journal doing its job, and
    smaller pages just split the same bytes across more copies.
    Amplification is a ratio, large only when a transaction touches many
    pages and writes little of each.
    
    Compatibility checked end to end rather than argued: a store created by
    a 64 kB binary opens under a 4 kB one, reads back, accepts writes and
    reads back again after a reopen; the reverse works too for a cleanly
    closed store. Data files never carried the page size -- only journals
    do, and a clean close leaves none. An older QL cannot replay a journal
    this one leaves behind after a crash, the usual downgrade limitation.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    0a2f63af
  • cznic's avatar
    PERFORMANCE: correct two statements that had gone stale · 73109d22
    cznic authored
    The baseline section still named commit 58b84fee and a dirty tree; the
    log's own header has said a936300a and clean since it was re-recorded.
    Point at the header instead of restating it, so the two cannot drift
    apart again, and say plainly that a baseline predating a landed
    optimisation re-reports that optimisation -- the same trap as the
    subset-versus-full-suite one already documented.
    
    P2's "where it applies" paragraph still listed LEFT/RIGHT/FULL OUTER
    JOIN as nested-loop, which the two subsections immediately below it
    contradict. Three-or-more-way joins are what actually remains open.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    73109d22
  • cznic's avatar
    PERFORMANCE: re-record the baseline after the WAL page size changes · 572ffc28
    cznic authored
    Recorded at 73109d22, clean tree: 301 benchmarks, 6 samples each. The
    previous baseline predated 8dcc868f and 0a2f63af, so comparing against it
    would have re-reported their 12-35% on V2 writes on top of whatever was
    actually being measured.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    572ffc28
  • cznic's avatar
    ql: bound ORDER BY by the LIMIT above it instead of sorting everything · 39e07732
    cznic authored
    PERFORMANCE.md P5: the LIMIT was applied after a full sort, so a top-10
    cost the same as ordering the whole table. orderByDefaultPlan now takes
    a bound from the LIMIT above it, runs a bounded max-heap over its input
    and inserts only the survivors into the temporary B+Tree. The bound is
    LIMIT + OFFSET, since the skipped rows are ordered too.
    
    Isolated A/B, n=6, two binaries differing only in whether the bound is
    applied:
    
        Query/file/orderby/top10   45.65ms -> 4.431ms  -90.29%
        Query/file2/orderby/top10  5.594ms -> 1.750ms  -68.71%
        Query/mem/orderby/top10    1.146ms -> 634.0us  -44.68%
        Query/file/orderby/page    53.05ms -> 25.05ms  -52.78%
        Query/file2/orderby/page   6.495ms -> 4.109ms  -36.74%
        Query/*/orderby (no LIMIT)                           ~
        QueryScale/mem/orderby/*                             ~
    
    geomean -35.60%, allocated bytes -45% to -74% on the bounded shapes. The
    no-LIMIT benchmarks are the control: nothing can bound them and they do
    not move. They earned their place -- comparing a narrow run against the
    full-suite baseline first showed QueryScale/mem/orderby at -5% to -13%,
    p=0.002, for queries this cannot touch.
    
    The survivors are inserted through the same Set the unbounded path uses,
    deliberately: Set and the iterator may normalize what passes through
    them, and on the memory back end they do, since mem.clone turns idealInt
    into int64. clone is not in the storage interface, so a heap emitting
    rows directly would return differently typed values for a computed
    column.
    
    The result is the same rows, not merely as good a set. The sort key's
    last element is the input sequence number, so no two keys compare equal,
    the ordering is total, and "the first N" is one set rather than a choice
    among ties. Most engines leave that open; this does not, so the spec
    needed nothing said about ties.
    
    Only a static LIMIT and OFFSET count, since both are otherwise evaluated
    per row against that row's fields and a bound guessed at plan time could
    cut rows the real values would keep. topNMaxRows caps the bound at 1000
    because the bounded path holds that many rows in memory where the
    unbounded one spills to a disk-backed B+Tree on the file back ends.
    
    Replacing an evicted heap entry overwrites its backing array rather than
    allocating. Without that, mem/orderby/page -- 510 of 1000 rows kept, the
    worst ratio measured -- was a real +10.09% regression, a temp insert on
    mem being cheap enough that 1000 heap operations cost more than the 490
    inserts they save. In place it is ~. What remains there is memory, not
    time: +8.01% B/op, which is what keeping 510 rows to return 10 costs.
    
    TestTopNEquivalence runs 50 queries on 3 back ends with the bound on and
    off and compares exactly, %#v so a type difference fails: tie-heavy and
    all-ties ordering columns, NULLs at both ends of the collation, DESC,
    multi-column keys, computed and constant columns, OFFSET, over-cap
    bounds, empty results, and composition with DISTINCT, GROUP BY and
    joins. It was checked to have teeth by breaking the optimisation two
    ways -- a comparator fixed ascending fails 16 cases, one blind to the
    sequence number fails 10. TestTopNApplies asserts through EXPLAIN which
    shapes engage, so equivalence cannot pass by never optimising.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    39e07732
  • cznic's avatar
    PERFORMANCE: re-record the baseline after the ORDER BY ... LIMIT bound · c282f1b4
    cznic authored
    Recorded at 39e07732, clean tree: 301 benchmarks, 6 samples each.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    c282f1b4
  • cznic's avatar
    ql: chain joins of three or more record sets instead of multiplying them out · 26b0b638
    cznic authored
    PERFORMANCE.md P2 left k-way joins as a Cartesian product plus a filter.
    chainMergeJoin now merges them left to right -- ((r0 join r1) join r2)
    join ... -- one conjunct per step. This is the largest of the join wins,
    because the plan replaced is Theta(n^k) rather than Theta(n^2), and the
    margin widens with each table added. Isolated A/B, n=6:
    
        JoinChain/mem/chain3     57.06ms -> 128.4us  -99.78%
        JoinChain/mem/chain4     169.8ms -> 104.6us  -99.94%
        JoinChain/file/chain3    86.19ms -> 3.028ms  -96.49%
        JoinChain/file/chain4    248.2ms -> 2.374ms  -99.04%
        JoinChain/file2/chain3   63.89ms -> 3.130ms  -95.10%
        JoinChain/file2/chain4   186.9ms -> 4.371ms  -97.66%
    
    geomean -98.92%, allocated bytes -99.49%. The row counts are far below
    BenchmarkJoin's 300 for the same reason the win is large.
    
    The chain follows the record sets in the order written rather than in
    the order the conjuncts suggest, which is what keeps the output columns
    where the Cartesian product put them: each step appends the added set's
    fields, exactly the product's layout, so nothing above needs a
    projection to undo it. Ordering to suit the predicate would need that
    projection, and ordering well would need cardinalities the planner does
    not have. A join whose middle record set is unconstrained declines and
    keeps the product.
    
    Two things had to generalize. joinColType resolves a field recursively
    through the chain to the base table column it came from, since the type
    test can no longer require both sides to be scans; it doubles as the
    test that a side is one this plan can sort. And a side now reports the
    ids of every record set it covers, not one handle.
    
    That second part is where the differential test earned itself. Storing
    the name-to-id map a joined plan reports works in memory and fails on V2
    with "encode2: unexpected data map[string]interface {}", because
    sortJoinSide materializes into a back end B+Tree and that tree encodes
    what it is given. The ids are flattened to one scalar per record set
    ahead of the row instead, which encodes everywhere and also removes an
    aliasing hazard: a joined plan fills one map and reuses it per row, so
    storing the reference and reading it back gives every buffered row the
    ids of whichever row was emitted last. Neither can arise in a two-way
    join, where both sides are scans whose id is a plain handle.
    
    Tests: 25 differential cases on 3 back ends -- three and four way chains
    in both conjunct orders, id() across every link, empty links,
    composition with ORDER BY, DISTINCT, count and LIMIT, and five shapes
    that must decline -- plus 9 TestMergeJoinApplies assertions. Two
    mutations show the two tests divide the work: mis-keying the id prefix
    fails 7 equivalence cases and no applies case, while leaving the
    intermediate field lists untruncated makes the chain silently decline,
    which every equivalence case survives with correct results and only the
    applies test catches.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    26b0b638
  • cznic's avatar
    PERFORMANCE: re-record the baseline after the chained joins · 97915bc9
    cznic authored
    Recorded at 26b0b638, clean tree: 310 benchmarks, 6 samples each. Up nine
    from the last one, which is BenchmarkJoinChain.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    97915bc9
  • cznic's avatar
    ql: cut a join into the runs its clause does connect, not all or nothing · ed92a87e
    cznic authored
    chainMergeJoin merged three or more record sets only when the clause
    reached across every one of them, and fell back to the full Cartesian
    product otherwise. It now cuts them into the longest consecutive runs
    the clause does join and multiplies only those, so a clause that
    connects some of the record sets gets most of the win instead of none.
    
        FROM t, u, v WHERE t.k == u.k              -> (t join u) x v
        FROM t, u, p, q WHERE t.k == u.k
                          && p.m == q.n            -> (t join u) x (p join q)
    
    Isolated A/B, n=10, benchtime 2s:
    
        JoinChain/mem/partial3          62.18ms -> 1.683ms  -97.29%
        JoinChain/mem/partial4-pairs    185.8ms -> 1.064ms  -99.43%
        JoinChain/file/partial3         87.88ms -> 3.970ms  -95.48%
        JoinChain/file/partial4-pairs   258.6ms -> 9.346ms  -96.39%
        JoinChain/file2/partial3        64.75ms -> 3.504ms  -94.59%
        JoinChain/file2/partial4-pairs  191.6ms -> 31.90ms  -83.35%
    
    Runs stay consecutive and in source order for the same reason a chain
    does: that is what keeps the output columns where the product put them.
    Only a clause joining no two adjacent record sets, FROM t, u, v WHERE
    t.k == v.k, is now left alone.
    
    A single equality reaches this too, which it did not before. It cannot
    chain three record sets, but it can join two and leave the rest to a
    smaller product, and planMergeJoin handed a lone conjunct straight to
    mergeJoinPlan, which requires exactly two sides -- so FROM t, u, v WHERE
    t.k == u.k declined for want of a second conjunct rather than for want of
    a joinable pair. Three of the new assertions caught this.
    
    crossJoinDefaultPlan learned that an operand may cover several record
    sets, since one of them is now a join reporting a map of ids rather than
    a record handle. It gets this the easy way: it reads those ids inside the
    operand's callback, before that operand is asked for another row, so
    copying the entries across is enough -- no snapshot and no flattening,
    unlike sortJoinSide.
    
    file2/partial4-pairs gains least for a reason worth recording: a product
    evaluates its right operand once per left row, and that operand is now a
    merge, so its two temporaries are rebuilt for every left row. Twenty left
    rows is forty temporaries, which is nearly all of V2's 31.9 ms; V1, whose
    CreateTemp is cheaper, does the same shape in 9.3 ms. Not a regression --
    both beat products of 192 ms and 259 ms -- but it caps this shape, and
    materializing a joined right operand would lift it.
    
    Tests: six assertions that used to expect a decline now expect a chain,
    including one from the original two-way work; fourteen new equivalence
    cases covering both partial shapes, ids through a product whose operand
    is a chain, and the clause that still connects nothing; and
    TestMergeJoinChainFields, which exists because mis-offsetting a run's
    field list changes no result and no plan choice -- the rows are right
    either way and nothing downstream reads the list -- so only what the plan
    says about itself can catch it.
    
    One case needed a weaker comparison, and it is the ordering decision
    again: a predicate that is a type error is an error for whichever row
    reaches it first, so t.k == u.k && u.s == p.j reports mismatched types
    about row d under the nested loop and about row b when chained. Same
    error, same reason, different example. cmpErrKind compares the
    parenthesized type description and not the operands, and still fails a
    run that does not error at all, which is the case the type test exists
    for.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    ed92a87e
  • cznic's avatar
    PERFORMANCE: re-record the baseline after the partial chains · de3569b7
    cznic authored
    Recorded at ed92a87e, clean tree: 316 benchmarks, 6 samples each. Up six
    from the last one, which is JoinChain's partial3 and partial4-pairs.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    de3569b7
  • cznic's avatar
    ql: replay a product's joined operand instead of re-running it · 76571b3c
    cznic authored
    A Cartesian product evaluates each operand once per row of the one
    before it. For a scan that costs no more than reading the rows again,
    but since partial chains an operand may itself be a merge join, and it
    rebuilt both of its temporaries every time: twenty left rows meant forty
    temporaries, which on V2 was nearly the whole cost of such a product.
    Those operands are now stored once and replayed. Isolated A/B, n=10,
    benchtime 2s:
    
        JoinChain/file2/partial4-pairs  33.73ms -> 4.878ms  -85.54%
        JoinChain/file/partial4-pairs   10.15ms -> 2.472ms  -75.64%
        JoinChain/mem/partial4-pairs    1.083ms -> 606.2us  -44.00%
    
    geomean -72.98%, allocated bytes -67.66%. PERFORMANCE.md predicted this
    would bring V2 toward V1's number for this shape; it went further, from
    3.32x V1 to 1.97x, since what it removes is forty temporaries on V2
    against V1's cheaper forty.
    
    materializeSide stores rows with the ids flattened ahead of them, the
    same layout and for the same encoding reason as sortJoinSide. Storing
    happens on first use, so an operand to the left that turns out empty
    costs nothing, and only operands that are themselves joins are stored.
    
    The first version carried the bookkeeping -- two slices and a deferred
    drop per do -- through every product, including the ones needing none.
    A regression check over BenchmarkCrossJoin* had no result significant at
    n=1-2, but 14 of 16 point estimates were higher and the smallest case was
    up 21%: a per-do allocation is a real share of 5 us of work, and 14 of 16
    in one direction is a sign test at p~0.002 even when no single row says
    so. do now splits, so a product of plain record sets runs exactly the
    code it ran before; at n=8 those benchmarks are ~ (p=0.105-0.721, the 21%
    now +1.2%).
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    76571b3c
  • cznic's avatar
    ql: use the index for UPDATE and DELETE when the value is a parameter · 01b955a0
    cznic authored
    whereRset.planExpr clones a SELECT's WHERE clause with ctx.arg before
    planning it, so by the time an index is chosen a $1 has become the value
    it stands for. updateStmt.exec and deleteStmt.exec planned the clause as
    written, and indexPlanFor cannot seek on a *parameter node because it
    carries no value -- so WHERE k == 500 used the index and WHERE k == $1
    scanned the whole table.
    
    Both forms evaluate a parameter correctly and the clause is re-checked
    per row either way, so this was never a wrong answer, only about 200x
    the work on a 1000-row table and worse as it grows. Which is to say
    every parameterized UPDATE and DELETE, i.e. the way applications write
    them.
    
    cloneWhere now does the substitution, from both statements' exec and
    from their explain. deleteStmt.explain also gained a plan to render at
    all: it echoed the statement and nothing else, so which plan a DELETE
    took was not observable, which is part of why this went unnoticed.
    
    Found by sqlitebench, added here, which compares ql against SQLite
    through database/sql on the same workloads. Its update-point and
    delete-point numbers made no sense next to BenchmarkUpdate's 8.87us for
    an indexed point UPDATE -- the comparison measured 601us for the same
    work -- and the difference between them was a literal against a
    parameter. No benchmark in this repo would have caught it: they all use
    literals, and the bug lived in the gap between how the suite writes
    queries and how applications do.
    
    sqlitebench is a module of its own so that ql does not acquire a
    dependency on SQLite in order to be compared with it, and the root
    module's ./... does not reach it. It compares against modernc.org/sqlite,
    the pure-Go build -- the right comparison for choosing an embedded
    database for a Go program that must not need cgo, and explicitly not a
    comparison against SQLite the C library. Durability is matched rather
    than assumed: ql fsyncs at every top-level commit, so the comparable
    SQLite setting is its own defaults, with WAL reported separately for what
    it is.
    
    Tests: TestParameterizedIndex asserts the index appears in the plan for
    both forms of both statements, and TestParameterizedIndexResults runs
    each statement both ways and compares the table after, since the plan
    changed and what it produces must not have. Reverting the clone fails
    the first and not the second, which is the signature of a bug about
    speed and not results.
    
    Also records P11: count() reads every column of every row and then
    groups them by distinct row before counting, 3.59 ms against SQLite's
    14.5 us. Left open.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    01b955a0
  • cznic's avatar
    PERFORMANCE: record the SQLite comparison · f9cd171a
    cznic authored
    Ten shapes, six engines, count=10, through database/sql in one process
    against modernc.org/sqlite -- the pure-Go build, which is the comparison
    that matters for choosing an embedded database for a Go program that
    must not need cgo, and explicitly not a comparison against SQLite the C
    library. Durability matched rather than assumed: ql fsyncs every
    top-level commit, so SQLite runs at its own defaults, with WAL reported
    separately as the weaker promise it is.
    
    SQLite is faster on most shapes and by a lot: geomean -78% in memory,
    -84% on disk at matched durability, -94% against WAL. A reader choosing
    on speed alone should choose SQLite and the numbers are written so they
    say that plainly.
    
    Where ql holds up is durable single-row writes: +18.4% committing one
    row per transaction, +22.0% on a point UPDATE, and within 9.2% on a
    point SELECT. That is a shape a lot of application traffic has, and the
    one place ql's much simpler journal costs less than SQLite's.
    
    Some of the gap is ql's to close. update-point read 601us on the first
    run and reads 19.4us now, because building this found P10. count is 248x
    slower than SQLite's because of P11, open. delete-point stays far behind
    because a delete walks the row chain to find the predecessor it unlinks,
    which no index avoids -- the back-link question in P1, declined.
    
    Also recorded: V2 beats V1 by 48% geomean over these shapes and loses to
    it on insert-batch by 108%, so the guidance that V1 suits small
    transactions and V2 large ones is backwards here for everything except
    bulk insert.
    
    run.sh now creates its output directory.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    f9cd171a
  • cznic's avatar
    ql: stream a bare aggregate instead of grouping the whole table first · 5d907dcd
    cznic authored
    PERFORMANCE.md P11: a SELECT that aggregates without a GROUP BY went
    through the grouping plan, which writes every row into a temporary
    B+Tree, chains them, and reads them all back so that one group -- the
    whole record set -- can be walked. That copy exists only to be read once.
    aggregateDefaultPlan evaluates the field expressions over the source rows
    directly. Isolated A/B, n=8:
    
        Query/file2/scan/count  4.456ms -> 997.4us  -77.62%
        Query/file2/scan/sum    4.564ms -> 1.082ms  -76.30%
        Query/file/scan/count   14.12ms -> 4.155ms  -70.58%
        Query/file/scan/sum     14.22ms -> 4.194ms  -70.50%
        Query/mem/scan/count    1.443ms -> 695.6us  -51.78%
        Query/mem/scan/sum      1.567ms -> 780.6us  -50.18%
        Query/*/groupby                                   ~
    
    geomean -67.84%, allocated bytes -65.46%. The GROUP BY benchmarks are the
    control: real groups, old plan, no movement.
    
    The aggregate protocol is untouched and is what makes this possible. Each
    field is evaluated once per row, which is how an aggregate accumulates --
    its running value lives in the name map keyed per call site -- and once
    more with $agg set to read the result back out. Only where the rows come
    from changes.
    
    It does not apply to every such SELECT, and the differential test is why.
    SELECT k, count() FROM t names a column of no particular row, and the two
    plans disagree about which one it gets: k came back 5 grouped and 1
    streamed, because the grouping plan chains each row onto the front of its
    group and so walks them in the reverse of the scan order. Rather than
    redefine what such a query returns -- undefined in SQL, but real in ql
    until now -- rowIndependent keeps the grouping plan for any field whose
    value can depend on which row came last, meaning any column reference or
    id() outside an aggregate's arguments. SELECT sum(v) + count() and SELECT
    count(k + 1) still stream.
    
    Tests: TestAggregateEquivalence runs 38 queries on 3 back ends both ways
    and compares with %#v -- NULLs in every aggregated column, empty inputs,
    every aggregate function, composition with DISTINCT, ORDER BY, LIMIT,
    OFFSET, joins and subqueries, and the GROUP BY cases that must keep the
    old plan. Dropping the rowIndependent guard fails 10 of them.
    TestAggregateApplies asserts through EXPLAIN which shapes engage and that
    none both streams and groups.
    
    What remains: ql still reads every column of every row to count them, so
    count() on V2 is now ~63x SQLite's rather than 248x. Closing that needs
    the scan to skip decoding when no field references a column, which is a
    change to how rows are read rather than how aggregates are planned.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    5d907dcd
  • cznic's avatar
    PERFORMANCE: refresh the SQLite comparison after P10 and P11 · db19af25
    cznic authored
    Both tables re-measured on the current tree, so the numbers are from
    after the two bugs the comparison itself found. count on V2 reads 904us
    where the first run read 4.28ms, and select-scan 1.00ms where it read
    4.26ms, both from P11; update-point in memory reads 20.1us where it read
    601us, from P10.
    
    geomeans move accordingly: -75.7% in memory, -79.1% on disk at matched
    durability, -92.6% against WAL. SQLite is still faster on most shapes and
    by a lot, and the text still says so.
    
    The four write shapes are fsync-bound and were the noisiest rows in the
    full run -- SQLite's update-point came back at +-628% -- so the two ql
    wins were re-measured on their own at n=12, where both arms are +-1-3%:
    +19.5% on insert-txn and +30.0% on update-point. An earlier independent
    run put them at +18.4% and +22.0%, so the direction and rough size
    reproduce even though the exact figure moves.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    db19af25
  • cznic's avatar
    PERFORMANCE: re-record the baseline after P10 and P11 · 9bf26843
    cznic authored
    Recorded at db19af25, clean tree: 316 benchmarks, 6 samples each, the
    same set as the last one since sqlitebench is its own module.
    
    The P11 fix shows where it should: Query/file2/scan/count falls from
    4.57ms to 1.07ms and Query/mem/scan/count from 1.53ms to 716us against
    the previous baseline, which is the -77.6% and -51.8% the isolated A/B
    measured, arrived at independently.
    
    P10 leaves no trace here, which is the point: every benchmark in this
    suite writes its WHERE clause with a literal, so none of them could ever
    have shown it.
    
    Co-Authored-By: default avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    9bf26843
  • cznic's avatar
    nuc64: auto deps · 94f7bc61
    cznic authored
    94f7bc61
Loading
Loading