Something went wrong. Try again.
atproto git client
Something went wrong. Try again.
Rust
123456789101112131415161718192021222324252627282930313233343536373839404142434445464748495051525354555657585960616263646566676869707172737475767778798081828384858687888990919293949596979899100101102103104105106107108109110111112113114115116117118119120121122123124125126127128129130131132133134135136137138139140141142143144145146147148149150151152153154155156157158159160161162163164165166167168169170171172173174175176177178179180181182183184185186187188189190191192193194195196197198199200201202203204205206207208209210211212213214215216217218219220221222223224225226227228229230231232233234235236237238239240241242243244245246247248249250251252253254255256257258259260261262263264265266267268269270271272273274275276277278279280281282283284285286287288289290291292293294295296297298299300301302303304305306307308309310311312313314315316317318319320321322323324325326327328329330331332333334335336337338339340341342343344345346347348349350351352353354355356357358359360361362363364365366367368369370371372373374375376377378379380381382//! Fitting a value into a fixed-width column.//!//! Five functions with one subject: a listing prints a fixed-width table, the//! values going into it are records written by anybody, and something has to//! decide what to drop and how far to pad. All of these answers were copied//! instead of shared — [`ellipsize`] lived in `cmd::pr::read` and was reached//! as `crate::cmd::pr::read::ellipsize` from two other command families,//! [`day`] existed five times over as the same one-line body, and the padding//! was `{:<N}` at nine call sites, each of them wrong in the same two ways.//! `cmd::pr`'s module documentation argued there was nothing to put in a//! shared module because neither had a caller outside the read half; that//! stopped being true, so here they are.//!//! # Cells and characters are different units, on purpose//!//! [`width`] counts *terminal cells*, which is what a column is measured in.//! [`ellipsize`] counts *characters*, which is what makes a cut safe. Those//! are deliberately not the same unit and this module does not reconcile//! them: the truncation unit is a chosen one — a test below pins the emoji//! case as char-counted — and only the padding half was ever wrong.
use unicode_width::UnicodeWidthChar;
/// Clip to `max` display cells, marking the cut with an ellipsis that costs/// one of them.////// Counts chars, not bytes, so a title with an emoji in it is cut between/// characters rather than through one. Imperfect as display width goes — a/// combining mark counts singly and a wide glyph counts as one cell — but/// never a panic, which is the property the tests pin.////// **The unit here is chars and stays chars.** [`width`] exists beside this/// and is not used by it. Cutting by cells would mean deciding what to do/// when the budget lands mid-glyph, and there is no answer to that which is/// obviously better than the one this already gives; the half that was/// actually shifting columns was the padding, and that is what [`pad_to`]/// fixes. Changing this one is a separate question with its own test.////// `max` is a column width and no caller passes zero; at zero the ellipsis/// would not fit either, and there is no sensible answer to give.pub(crate) fn ellipsize(text: &str, max: usize) -> String { if text.chars().count() > max { text.chars().take(max - 1).collect::<String>() + "…" } else { text.to_string() }}
/// Somebody else's text as one cell of a listing: cleaned, cut to `cut`/// characters, padded to `cells`.////// The three steps in the order they have to happen, at the four listings/// that were each spelling them out — `pr list`, `issue list`, `repo list`/// and `search`. Two of the steps were already copied per site; the third,/// [`crate::term::text::one_line`], is new, and putting it anywhere but/// first would be the bug it exists to prevent: [`ellipsize`] counts/// characters, so a title cut before the control characters come out has a/// budget partly spent on things that draw nothing, and [`width`] recognises/// escape sequences, so a hostile title padded before cleaning lines up/// perfectly with the rows either side of it.////// `cut` and `cells` are separate numbers because they are separate units —/// see the note at the top of this module — and the gap between them is the/// space between one column and the next.pub(crate) fn cell(text: &str, cut: usize, cells: usize) -> String { pad_to(&ellipsize(&crate::term::text::one_line(text), cut), cells)}
/// How many terminal cells `text` will occupy when printed.////// The one question every fixed-width column asks, and the one `{:<N}` gets/// wrong twice over. Rust's fill/align counts `char`s, and a column formatted/// that way is off by the two things atgc's own output is full of:////// - **Wide characters.** A CJK ideograph or an emoji occupies two cells and/// counts as one `char`, so a title padded to 56 by `char` arrives two/// cells wide per glyph and shoves every column to its right. `search`/// prints arbitrary record prose and is where this shows up in the wild./// - **Escape sequences.** Padding applied *after* painting counts the SGR/// bytes, which occupy no cells, so it applies none — `logs pds` lost a/// column exactly this way whenever colour was on. Worse, atgc emits OSC 8/// hyperlinks ([`crate::term::hyperlink`]), where a whole URL sits inside/// the string and is drawn as nothing at all.////// So this is a scanner, not a `map(UnicodeWidthChar::width).sum()`./// `unicode-width` alone gets the escapes wrong in the same direction the/// `char` count does. Two escape shapes are recognised, which are the two/// atgc writes:////// - CSI (`ESC [` … final byte in `@`–`~`), which is every colour in/// [`crate::term::style`]./// - OSC (`ESC ]` … `BEL` or `ESC \`), which is [`crate::term::hyperlink`].////// Anything else after an `ESC` is treated as a two-character escape, and a/// control character on its own contributes nothing — `unicode-width`/// returns `None` for those and this reads that as zero. This is a display/// estimate and not a terminal emulator: a zero-width joiner sequence still/// measures as its parts, which is the same answer every other CLI gives.pub(crate) fn width(text: &str) -> usize { let mut cells = 0; let mut chars = text.chars(); while let Some(c) = chars.next() { if c != '\x1b' { cells += c.width().unwrap_or(0); continue; } match chars.next() { // CSI: parameter and intermediate bytes, then one final byte. Some('[') => { for c in chars.by_ref() { if matches!(c, '\u{40}'..='\u{7e}') { break; } } } // OSC: a string terminated by BEL, or by ST spelled `ESC \`. Some(']') => { let mut after_esc = false; for c in chars.by_ref() { if c == '\x07' || (after_esc && c == '\\') { break; } after_esc = c == '\x1b'; } } // A two-character escape, or a trailing `ESC` with nothing after // it. Either way there is nothing more to skip. _ => {} } } cells}
/// `text` followed by enough spaces to fill `cells` — `{:<N}` that counts/// what [`width`] counts.////// Already too wide comes back unchanged rather than truncated: a caller that/// wants a limit calls [`ellipsize`] first, and silently eating a character/// here would make the two decisions impossible to tell apart in the output.pub(crate) fn pad_to(text: &str, cells: usize) -> String { let mut padded = text.to_string(); padded.push_str(&" ".repeat(cells.saturating_sub(width(text)))); padded}
/// [`pad_to`]'s mirror: enough spaces to fill `cells`, then `text` — `{:>N}`/// that counts what [`width`] counts.////// One caller, and it is the case that proves the point. `logs oauth` prints/// `{:>8}` over an already-painted invocation tag, and that was right only/// because `short_inv` happens to return exactly eight characters, so the/// padding it failed to apply was zero either way.pub(crate) fn pad_start(text: &str, cells: usize) -> String { let mut padded = " ".repeat(cells.saturating_sub(width(text))); padded.push_str(text); padded}
/// The calendar date of an RFC 3339 timestamp, **in UTC**.////// Parsed rather than sliced. Taking the first ten characters printed the/// *writer's* local date, so a record written at `+03:00` and one written at/// `Z` could disagree by a day about the same instant — inside one listing,/// in a column whose only job is to let rows be compared.////// UTC is the choice because it is the only one that makes two rows/// comparable, which is what a column is for. The reader's local zone would/// make the same listing print differently on two machines and would make a/// piped listing depend on `TZ`; the writer's own offset is the bug being/// fixed. The visible consequence is real and intended: a record written at/// `+03:00` just before midnight now prints the previous day. See/// [Output contracts](crate::docs::output).////// A stamp that will not parse falls back to the first ten characters, which/// is what this function used to do unconditionally. That is not tidiness —/// listings have to keep rendering when one record is malformed, and an empty/// `createdAt` comes back empty rather than as a placeholder because old/// records really do carry one and the column is the caller's to fill.pub(crate) fn day(datetime: &str) -> String { match chrono::DateTime::parse_from_rfc3339(datetime) { Ok(stamp) => stamp.to_utc().date_naive().to_string(), Err(_) => datetime.chars().take(10).collect(), }}
#[cfg(test)]mod tests { use super::*;
#[test] fn ellipsize_only_cuts_what_is_too_long() { assert_eq!(ellipsize("short", 10), "short"); // Exactly at the limit is not too long. assert_eq!(ellipsize("0123456789", 10), "0123456789"); // One over: nine kept plus the marker, still ten cells. assert_eq!(ellipsize("0123456789a", 10), "012345678…"); assert_eq!(ellipsize("0123456789a", 10).chars().count(), 10); assert_eq!(ellipsize("", 10), ""); }
/// Titles are user text and routinely contain multibyte characters, so /// the cut counts characters. Slicing bytes here would panic on a title /// whose 57th byte lands mid-character. #[test] fn ellipsize_counts_characters_not_bytes() { let emoji = "🧬".repeat(30); let cut = ellipsize(&emoji, 10); assert_eq!(cut.chars().count(), 10); assert_eq!(cut, "🧬".repeat(9) + "…");
// Combining marks count singly too — imperfect as display width goes, // but never a panic, which is the property being pinned. let accented = "é".repeat(12); assert_eq!(ellipsize(&accented, 5).chars().count(), 5); }
/// The unit `ellipsize` cuts in is chars and the unit `width` measures in /// is cells, and they are allowed to disagree. Pinned so that a later /// reader does not "fix" one into the other without deciding to. #[test] fn the_cut_is_chars_and_the_measure_is_cells() { let cut = ellipsize(&"🧬".repeat(30), 10); assert_eq!(cut.chars().count(), 10); assert_eq!(width(&cut), 19, "nine wide glyphs plus a narrow ellipsis"); }
#[test] fn ascii_is_one_cell_per_character() { assert_eq!(width(""), 0); assert_eq!(width("hello"), 5); }
/// The defect this module was opened for: a CJK title counted by `char` /// is half the cells it prints in. #[test] fn cjk_is_two_cells_per_character() { let title = "日本語"; assert_eq!(title.chars().count(), 3); assert_eq!(width(title), 6); // And so a row padded on the char count would be three cells short. assert_eq!(pad_to(title, 10), "日本語 "); assert_eq!(width(&pad_to(title, 10)), 10); }
#[test] fn an_emoji_is_two_cells() { assert_eq!(width("🧬"), 2); assert_eq!(width("a🧬b"), 4); assert_eq!(width(&pad_to("🧬", 6)), 6); }
/// Padding applied after painting is the `logs pds` bug: the SGR bytes /// are counted, the string looks long, and no padding is applied at all. #[test] fn sgr_escapes_cost_no_cells() { let painted = crate::term::style::paint(true, crate::term::style::GOOD, "create"); assert!(painted.contains('\x1b'), "the test needs a real escape"); assert_eq!(width(&painted), 6); // Which is the whole point: a painted and an unpainted cell of the // same column land in the same place. assert_eq!(width(&pad_to(&painted, 8)), width(&pad_to("update", 8))); }
/// A cell can carry an entire URL that occupies no columns at all, which /// is where measuring characters goes furthest wrong. #[test] fn an_osc_8_hyperlink_costs_only_its_text() { let linked = crate::term::hyperlink::wrap_when( true, "https://bsky.app/profile/permadeath.com", "@permadeath.com", ); assert!(linked.chars().count() > 50, "the URL really is in there"); assert_eq!(width(&linked), "@permadeath.com".chars().count()); assert_eq!(width(&pad_to(&linked, 20)), 20); // BEL terminates an OSC string too, and some emitters use it. assert_eq!(width("\x1b]8;;https://example.com\x07text\x1b]8;;\x07"), 4); }
/// Rows line up. Stated as the property a column actually has to have, /// rather than only as the cell counts that produce it. #[test] fn a_padded_row_lines_up_whatever_is_in_the_first_column() { let painted = crate::term::style::paint(true, crate::term::style::BAD, "delete"); let linked = crate::term::hyperlink::wrap_when(true, "https://example.com/x", "🧬 x"); let rows = ["plain", "日本語", "🧬🧬", &painted, &linked]; let widths: Vec<usize> = rows .iter() .map(|cell| width(&format!("{} after", pad_to(cell, 12)))) .collect(); assert!( widths.windows(2).all(|w| w[0] == w[1]), "columns disagree: {widths:?}" ); }
/// The rendering site, with the title out of the bug report: an OSC 8 /// hyperlink whose text and destination disagree, in a column that would /// otherwise have padded it to line up with the rows either side. #[test] fn a_listing_cell_carries_no_escape_out_of_a_record() { let hostile = "\x1b]8;;https://evil.example\x1b\\ok\x1b]8;;\x1b\\"; let row = cell(hostile, 54, 56); assert!(!row.contains('\x1b'), "{row:?}"); assert_eq!(row.trim_end(), "]8;;https://evil.example\\ok]8;;\\"); assert_eq!(width(&row), 56, "and it is still a 56-cell column");
// A screen-clearing title, and a title that would forge a second row // in a listing whose whole shape is one record per line. let cleared = cell("\x1b[2J\x1b[Hgone", 54, 56); assert!(!cleared.contains('\x1b'), "{cleared:?}"); let forged = cell("real\nforged", 54, 56); assert!(!forged.contains('\n'), "{forged:?}"); assert_eq!(forged.trim_end(), "real forged"); }
/// Cleaning happens before the cut, so the character budget is spent on /// characters that draw. Cutting first would leave a title whose escapes /// ate the column and whose words did not fit. #[test] fn a_cell_is_cleaned_before_it_is_cut() { // Sixty escapes and four letters: everything a reader wants is past // the cut if the escapes are counted, and nothing is if they are not. let padded = cell(&("\x1b[0m".repeat(60) + "word"), 54, 56); assert_eq!(padded.trim_end(), "[0m".repeat(17) + "[0…"); assert_eq!(width(&padded), 56); }
/// Too wide to fit comes back whole. A caller wanting a limit ellipsizes. #[test] fn padding_never_truncates() { assert_eq!(pad_to("far too long", 4), "far too long"); assert_eq!(pad_start("far too long", 4), "far too long"); }
#[test] fn pad_start_fills_on_the_left() { assert_eq!(pad_start("7f3a", 8), " 7f3a"); let painted = crate::term::style::paint(true, crate::term::style::NOTE, "7f3a"); assert_eq!(width(&pad_start(&painted, 8)), 8); }
/// The date out of a record's timestamp, in UTC, whatever offset the /// writer used. #[test] fn day_takes_the_date_off_an_iso_datetime() { assert_eq!(day("2026-08-05T04:31:07+03:00"), "2026-08-05"); assert_eq!(day("2026-08-05T01:34:10.355270Z"), "2026-08-05"); }
/// The behaviour change, pinned: an offset that crosses midnight now /// prints the UTC day, not the writer's. Both stamps below are the same /// instant, and before this parsed they printed different dates. #[test] fn day_is_utc_not_the_writers_local_date() { assert_eq!(day("2026-08-05T01:31:07+03:00"), "2026-08-04"); assert_eq!(day("2026-08-04T22:31:07Z"), "2026-08-04"); assert_eq!( day("2026-08-05T01:31:07+03:00"), day("2026-08-04T22:31:07Z") ); // And the other way over the line. assert_eq!(day("2026-08-04T21:00:00-05:00"), "2026-08-05"); }
/// A value that is not RFC 3339 comes back short instead of panicking or /// erroring, which is what keeps a listing rendering when one record is /// malformed. Records that list today really do carry these. #[test] fn day_degrades_on_anything_it_cannot_parse() { assert_eq!(day("2026-08-17"), "2026-08-17"); // Old records can carry an empty createdAt — one really does, in the // fixture — so this has to come back with something printable. assert_eq!(day(""), ""); assert_eq!(day("2026"), "2026"); assert_eq!(day("nope"), "nope"); // Counted in characters, not bytes — a multi-byte timestamp field is // nonsense, but slicing one by bytes is a panic in a listing. assert_eq!(day("héllo wörld!!"), "héllo wörl"); }}