• SBBSecho truncates UTF-8 to/from/subject header fields mid-character w

    From Rob Swindell@1:103/705 to GitLab issue in main/sbbs on Thu Oct 1 14:54:32 2026
    open https://gitlab.synchro.net/main/sbbs/-/work_items/1276

    FTS-0001 packed-message headers give the to/from fields 36 bytes and the subject 72 bytes, each NUL-terminated, so 35 and 71 bytes are usable. SMB header fields have no such limit, so when SBBSecho exports a message it cuts these fields to fit. The cut counts bytes, not characters: on a UTF-8 message (`CHRS: UTF-8 4`) a field longer than the limit can end in a partial multi-byte sequence, producing an ill-formed UTF-8 string that readers display as a replacement character.

    Cyrillic, Greek, etc. take 2 bytes per letter in UTF-8, so a 35-byte cut lands mid-character about half the time.

    ## Where

    - Echomail export, `export_echomail()` in `src/sbbs3/sbbsecho.c`: `SAFECOPY(hdr.from, msg.from)`, `SAFECOPY(hdr.to, msg.to)`, `SAFECOPY(hdr.subj, msg.subj)`.
    - Netmail, `create_netmail()`: `SAFECOPY(hdr.to, to)`, `SAFECOPY(hdr.from, from)`, `SAFECOPY(hdr.subj, subject)`.
    - On sub-boards with `SUB_ASCII` set, export truncates first and converts UTF-8 to CP437 afterwards (`utf8_to_cp437_inplace(hdr.to)` etc.). That converts an already-split sequence, and gives away room: in CP437 each of those characters is 1 byte, so converting first and truncating second would keep up to twice as many characters.

    ## Seen in the wild

    The FidoNet UTF-8 echo discussed this in September 2026 ("Testing long header", "FSP-1030 test", "UCS kludges and long headers to base"). FMail truncates the same way when packing from JAM (100-byte fields), and messages from 2:280/5555 arrived here with a `to` field of exactly 35 bytes ending in a lone `0xD0` lead byte. GoldED+ now does the cut itself on a character boundary before the tosser sees the message, but that does not help messages posted locally on a Synchronet system, where the stored field is never cut until SBBSecho packs it.

    ## Suggested fix

    - When `smb_msg_is_utf8()` is true and the field is valid UTF-8, copy with `utf8_strlcpy()` (`src/encode/utf8.c`), which never leaves a partial sequence at the end of the destination. `SAFECOPY_UTF8()` in `sbbs.h` already wraps this pattern, but `sbbsecho.c` does not include `sbbs.h`.
    - In the `SUB_ASCII` path, convert to CP437 into a buffer first, then truncate.

    ## Out of scope / possible follow-up

    - Inbound: SBBSecho imports ill-formed header fields as received. Repairing an upstream cut (dropping a trailing partial sequence) on import could be considered separately.
    - FSP-1030 (preserved as FRL-1021) `UCSTO:` / `UCSFROM:` / `UCSSUBJ:` kludges carry the full field when it had to be cut. SBBSecho neither writes nor reads them. GoldED+ and AmberEdit now emit them, but the proposal was never adopted as a standard.

    -- *Authored by Claude (Claude Code), on behalf of @rswindell*
    --- SBBSecho 3.38-Linux
    * Origin: Vertrauen - [vert/cvs/bbs].synchro.net (1:103/705)
  • From Rob Swindell@1:103/705 to GitLab issue in main/sbbs on Thu Oct 1 15:27:19 2026
    close https://gitlab.synchro.net/main/sbbs/-/work_items/1276
    --- SBBSecho 3.38-Linux
    * Origin: Vertrauen - [vert/cvs/bbs].synchro.net (1:103/705)