Part of the OpenSSH drop-in compatibility epic.
Problem
We have no standing measurement of how compatible bssh actually is. The claim "Drop-in replacement for SSH with compatible command-line syntax" currently sits in README.md with nothing behind it, and the 2026-08-26 measurement put the real figure at 23 of 71 environment-valid tests. Every other sub-issue in this epic needs a number to move, so the instrument has to land first.
Scope
Vendor a harness that runs the OpenSSH regression suite against the bssh binary and scores the result.
- Fetch and build a pinned
openssh-portable release (the measurement used V_10_3_P1) or consume a prebuilt one, and keep the pin in a single declared place so bumps are deliberate.
- Drive the suite the way openssh-portable's own top-level
Makefile does. The t-exec target exports a large env block naming every helper binary (TEST_SSH_SSHD, TEST_SSH_SSHD_SESSION, TEST_SSH_SSHD_AUTH, TEST_SSH_SFTPSERVER and others). Relying on PATH alone is not enough: sshd fails to start and every network test reports a misleading "closed by remote host".
- Point
TEST_SSH_SSH at the bssh binary under test. Leave TEST_SSH_SSHD and the rest on the reference OpenSSH build, since this measures bssh as a client.
- Enforce a per-test timeout. Several failures manifest as a hang rather than an error.
- Cross-check every failure by re-running that same test with the reference OpenSSH client, and classify the result as a genuine bssh failure only when the baseline passes. In the initial measurement 8 of 56 failures were environmental (
agent-restrict, channel-timeout, connection-timeout, forward-control, forwarding, multiplex, pubkey-priority, scp3).
- Emit a machine-readable result table (test, verdict, duration, first failure line) and a committed baseline file, so CI fails on regression rather than on absolute score.
- Record the permanent skip list from the epic as skips with a stated reason, not as failures.
Notes for the implementer
The candidate set used for the baseline was 89 of the suite's 116 shell tests, excluding the pure helpers (test-exec.sh, scp-ssh-wrapper.sh, ssh2putty.sh) and the third-party interop tests (putty-*, dropbear-*, conch-*, ssh-com*). Keep that selection in a data file rather than hardcoding it in a script, since the epic will move tests between the candidate and skip sets as work lands.
Acceptance criteria
Part of #275
Part of the OpenSSH drop-in compatibility epic.
Problem
We have no standing measurement of how compatible bssh actually is. The claim "Drop-in replacement for SSH with compatible command-line syntax" currently sits in
README.mdwith nothing behind it, and the 2026-08-26 measurement put the real figure at 23 of 71 environment-valid tests. Every other sub-issue in this epic needs a number to move, so the instrument has to land first.Scope
Vendor a harness that runs the OpenSSH regression suite against the bssh binary and scores the result.
openssh-portablerelease (the measurement usedV_10_3_P1) or consume a prebuilt one, and keep the pin in a single declared place so bumps are deliberate.Makefiledoes. Thet-exectarget exports a large env block naming every helper binary (TEST_SSH_SSHD,TEST_SSH_SSHD_SESSION,TEST_SSH_SSHD_AUTH,TEST_SSH_SFTPSERVERand others). Relying onPATHalone is not enough:sshdfails to start and every network test reports a misleading "closed by remote host".TEST_SSH_SSHat the bssh binary under test. LeaveTEST_SSH_SSHDand the rest on the reference OpenSSH build, since this measures bssh as a client.agent-restrict,channel-timeout,connection-timeout,forward-control,forwarding,multiplex,pubkey-priority,scp3).Notes for the implementer
The candidate set used for the baseline was 89 of the suite's 116 shell tests, excluding the pure helpers (
test-exec.sh,scp-ssh-wrapper.sh,ssh2putty.sh) and the third-party interop tests (putty-*,dropbear-*,conch-*,ssh-com*). Keep that selection in a data file rather than hardcoding it in a script, since the epic will move tests between the candidate and skip sets as work lands.Acceptance criteria
cargo-invocable ormake-invocable target runs the suite against a freshly built bssh and prints the score.docs/gains a page describing how to run the harness locally and how to interpret the three verdict classes.Part of #275