Skip to content

test: add an OpenSSH regress compatibility harness and score gate #276

Description

@inureyes

Part of the OpenSSH drop-in compatibility epic.

Problem

We have no standing measurement of how compatible bssh actually is. The claim "Drop-in replacement for SSH with compatible command-line syntax" currently sits in README.md with nothing behind it, and the 2026-08-26 measurement put the real figure at 23 of 71 environment-valid tests. Every other sub-issue in this epic needs a number to move, so the instrument has to land first.

Scope

Vendor a harness that runs the OpenSSH regression suite against the bssh binary and scores the result.

  • Fetch and build a pinned openssh-portable release (the measurement used V_10_3_P1) or consume a prebuilt one, and keep the pin in a single declared place so bumps are deliberate.
  • Drive the suite the way openssh-portable's own top-level Makefile does. The t-exec target exports a large env block naming every helper binary (TEST_SSH_SSHD, TEST_SSH_SSHD_SESSION, TEST_SSH_SSHD_AUTH, TEST_SSH_SFTPSERVER and others). Relying on PATH alone is not enough: sshd fails to start and every network test reports a misleading "closed by remote host".
  • Point TEST_SSH_SSH at the bssh binary under test. Leave TEST_SSH_SSHD and the rest on the reference OpenSSH build, since this measures bssh as a client.
  • Enforce a per-test timeout. Several failures manifest as a hang rather than an error.
  • Cross-check every failure by re-running that same test with the reference OpenSSH client, and classify the result as a genuine bssh failure only when the baseline passes. In the initial measurement 8 of 56 failures were environmental (agent-restrict, channel-timeout, connection-timeout, forward-control, forwarding, multiplex, pubkey-priority, scp3).
  • Emit a machine-readable result table (test, verdict, duration, first failure line) and a committed baseline file, so CI fails on regression rather than on absolute score.
  • Record the permanent skip list from the epic as skips with a stated reason, not as failures.

Notes for the implementer

The candidate set used for the baseline was 89 of the suite's 116 shell tests, excluding the pure helpers (test-exec.sh, scp-ssh-wrapper.sh, ssh2putty.sh) and the third-party interop tests (putty-*, dropbear-*, conch-*, ssh-com*). Keep that selection in a data file rather than hardcoding it in a script, since the epic will move tests between the candidate and skip sets as work lands.

Acceptance criteria

  • cargo-invocable or make-invocable target runs the suite against a freshly built bssh and prints the score.
  • The harness runs in GitHub Actions on Linux and macOS, and fails the job when the score drops below the committed baseline.
  • The result table and the baseline are committed artifacts, and a score change shows up as a reviewable diff.
  • Permanent skips are declared with reasons and reported separately from failures.
  • docs/ gains a page describing how to run the harness locally and how to interpret the three verdict classes.

Part of #275

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions