Keep a catalogue of your own in a spreadsheet: read it from the CSV file the sheet saves, with its JSON header beside it, and write one back - #889
Conversation
…ith its JSON header beside it, and write one back io.read_catalogue reads a .csv file whose header, the catalogue's JSON document without its rows, sits beside it as <name>.phonometry.json (or where header_path says) and declares the delimiter and the decimal mark; nothing about the dialect is guessed. A cell holds one value, bound, range or word in a closed grammar (0.85, ~0.85, <=30, >=5, 0.30..0.50, 0.85±0.05, [AFr5], true or false in a column of flags), the columns are the row class's fields in any unit of their family, basis, the provenance.* columns and columns of your own named x-, and the lines go through the same pass as a JSON document's rows, so every problem is raised at once and placed at its line and its column as a spreadsheet letters them. Text where a number goes is never read as a word, a NaN or zero, a thousands separator is never read, and when every failing number is written with the other decimal mark, or the first line splits at another delimiter, the refusal says which to declare. io.write_catalogue writes a .csv file with a byte order mark and its header, in the delimiter and decimal mark asked for, refuses what one cell cannot hold at the pointer a JSON document would write it at, and puts an apostrophe before any text a spreadsheet would run as a formula, which the reader takes off again. Catalogue.header_sha256 is the hash of the header a CSV file was read with.
…ed table written out as a template to a refusal at its line and column The guide Your own catalogues gains a section, in both languages and in docs/, that writes Arau-Puchades's Tabla 6.1 as a CSV file and its header, reads a tile range typed into a sheet in the closed grammar of cells, shows a misspelt column and a corrupt glyph refused at their lines and columns, and shows Bies's Table 6.2, which credits its rows, refused as a CSV file at its JSON pointer. It cites RFC 4180, says what only a JSON document can hold, and that XLSX is not read. The Files overview, the published catalogues page, the READMEs, the API reference rows of read_catalogue, write_catalogue, Catalogue and CatalogueIssue, the CHANGELOG and the llms files say a catalogue may be a CSV file too.
…the cells' When the header of a CSV file (or a JSON document) stops the read, as a provenance with no date it was consulted does, the issues are sorted as they are when the rows were read: the document's own keys first, then each row by its line.
…how to write it, and a control character in it is named, never shown A range typed with a dash, a plus-or-minus typed as +-, a unit after the number and a number grouped in thousands with a point or a comma, alone or before a fraction, each get their own refusal with the cell to write instead. A figure such as 12,500, which could be either decimal mark, no longer proposes one for the whole file. A control character or a mark that reorders text in a numeric cell is refused by name, and the refusal never carries the cell raw. A column named Key or Name is answered with key or name, and a first line of one column named key no longer proposes a delimiter. The tests now hold the writer to every mark a cell writes in three dialects, the conventions of every published table through the round trip, and each kind of field to the column it makes.
…d names the header limit of a CSV file Table 6.2 of Bies, Hansen and Howard prints one credit, to Beranek and Hidaka (1998), inside the name of its first row; the guide said it credited every row. The guide and the read_catalogue reference now list the 64 KiB limit on the header of a CSV file beside the others, the guide names the refusals that tell a mistyped cell how to write it, the guides index says a catalogue of your own may be a spreadsheet saved as CSV, and the CHANGELOG says both json and csv read the file.
There was a problem hiding this comment.
Sorry @jmrplens, your pull request is larger than the review limit of 150,000 diff characters
|
Warning Review limit reachedNext included review available in 59 minutes. View limit detailsLimit details: You’ve used all 2 included reviews currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: Repository: jmrplens/phonometry/.coderabbit.yaml Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (26)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Numerical conformance1447/1447 checks pass across 97 domains and 477 standards (196 normative designations, 114 further published sources). Used in the tables below is how much of that clause's published tolerance the deviation consumes: 100 % sits exactly on the limit, 5 % uses a twentieth of the allowance, and a dash means the clause states no two-sided tolerance for the quantity, so there is no budget to spend. It is reported and never used to decide a verdict, which is settled at full precision before any rounding. Nothing moved: same 1447 checks, same verdicts, same numbers. Closest to their published limit (top 5) The rows with the least room left, so the ones a change is most likely to push over.
|
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## catalogues/lookups-and-guide #889 +/- ##
================================================================
+ Coverage 96.82% 96.84% +0.01%
================================================================
Files 394 395 +1
Lines 65711 66384 +673
================================================================
+ Hits 63626 64287 +661
- Misses 2085 2097 +12 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
A data sheet is more often typed into a spreadsheet than into JSON, and a spreadsheet saves CSV.
io.read_cataloguenow reads a name ending in.csv(in any case): the rows come from the CSV file, one per line under a first line that names the columns, and the rest of the document from a JSON header beside it, named as the file with.phonometry.jsonafter it, or wherever the newheader_path=says. The header is the catalogue document without its rows plus"csv": {"delimiter": ";", "decimal": ","}, and the dialect is never guessed: a delimiter is a comma, a semicolon or a tab, a decimal mark a point or a comma, and a decimal comma between commas is refused. The calibration sidecar of an audio file takes the same tail, and a sidecar handed over as a header is refused by its schema.Each numeric cell holds one thing, in a closed grammar that writes the same hedges a JSON row does: empty for nothing printed,
0,85,~0,85,<=30(or<30,≤30),>=5(or>5,≥5),0,30..0,50(with~before it when approximate),0,85±0,05(or+/-),[AFr5]for a word where the number would be, andtrueorfalsein a column of flags (in any case, as a spreadsheet saves them). A minus sign may be U+2212, and a thousands separator is never read. The columns arekey, the row class's fields in their own unit or in any other of the same kind (flow_resistivity_kpa_s_m2),basisfor the row as a whole,provenance.page,provenance.printed_table,provenance.laboratory,provenance.accreditation,provenance.report,provenance.test_dateandprovenance.test_standardfor one row's provenance, and columns of your own namedx-. A column the sheet cannot have is refused on the first line with what to do instead: a hedge is written in the cell itself, and what one cell cannot say (several readings, a misprint, a value not to derive, a carried cell, a figure in a unit no family holds, a credit, a basis or a standard for one cell, a bound or a range with a plus-or-minus, an approximate bound) is written in a JSON document.The lines are turned into the rows a JSON document would write and go through the same pass, so a sheet is held to every rule a JSON document is, and every problem is raised at once in one
io.CatalogueError. Each issue is placed back where a person finds it: "tiles.csv, line 2, column D (absorption_coefficient_1000), row 'e400': '0.^G' is not a number; if the datasheet prints this text where the number would be, write it as [0.^G]", or a JSON pointer into the header for a problem of the header.CatalogueIssue.__str__reads a location given as a line and a column as a phrase. A text where a number goes is never read as a word, aNaNor a zero; a cell written with the other decimal mark, a number grouped in thousands with a point, a comma or a space (alone or before a fraction), a range typed with a dash, a plus-or-minus typed as+-, a unit typed after the number, several numbers in one cell, a flag in a numeric column and an open bracket each get their own words and the cell to write instead. A control character or a mark that reorders text in a cell is refused by name, and a refusal never carries the cell raw: it is quoted escaped, and a word is offered between brackets only when it prints. When every number that fails is written with the other decimal mark and none with the declared one, or the first line reads as one column that splits at another delimiter into several columns withkeyamong them, the refusal also says which one to declare, at the header's/csv/decimalor/csv/delimiter; a figure such as12,500, which could be twelve and a half or twelve thousand five hundred, is called ambiguous and proposes no decimal mark. A column named in other letters than its field (Key,Name) is answered with the name it spells. The file is UTF-8 with or without the byte order mark a spreadsheet writes, and any other encoding is refused with the advice to save it as "CSV UTF-8". A CSV file is held to 16 MiB and 50 000 rows, its header to 64 KiB, both checked before a byte is read, andcsv.field_size_limitis left alone. A line with another number of cells than the first line is refused, and lines of empty cells are skipped. A header that stops the read now names its problems before the cells', as the rest of the refusals do.io.write_cataloguewrites a name ending in.csvas the CSV file and its header, in thedelimiter=anddecimal=asked for (a comma and a point by default), with a byte order mark and CRLF line ends. A text that a spreadsheet would run as a formula (one starting with=,+,-,@, a tab or a carriage return) is written after an apostrophe, and the reader takes it off; a text that begins with apostrophes before such a character gets one more, so it reads back exactly, and a name like's-Hertogenboschis left as it is. Numbers are not guarded. What a cell cannot hold is refused before anything is written, each thing at the pointer the JSON document would write it at, and so is an empty text where the field's default says something (an absorption area per nothing), which an empty cell would read as the default. Both files are checked before either is written; each is written beside its name and renamed into place, and the pair is not written as one.io.Catalogue.header_sha256is the SHA-256 of the header a CSV file was read with, besidefile_sha256, now the CSV file's.The guide "Your own catalogues" ("Tus propios catálogos"), in both languages and in
docs/, gains a section that writes Tabla 6.1 of Arau-Puchades as a CSV template, reads a fictitious tile range typed into a sheet, shows a misspelt column and a corrupt glyph refused at their lines and columns, and shows Table 6.2 of Bies, which credits its first row to Beranek and Hidaka (1998), refused as a CSV file with nothing written. It cites RFC 4180 and says that a spreadsheet's own format (XLSX, ODS) is not read. The Files overview, the guides index, the published catalogues page, the READMEs, the API reference rows ofread_catalogue,write_catalogue,CatalogueandCatalogueIssue, the CHANGELOG and the llms files say a catalogue may be a CSV file, and the limits the guide and the reference list include the 64 KiB of a header.What breaks: nothing a caller wrote.
header_path,delimiteranddecimalare new keyword arguments with defaults, andheader_sha256a new field with a default. A.csvname used to be refused withValueError; a name ending in neither.jsonnor.csvstill is, and so areheader_path=beside a JSON document and a dialect beside one. A JSON document that holds a top-levelcsvis refused with its own message instead of as an unknown key. The one line a JSON syntax error prints now reads "x.json, line 1, column 2: this is not JSON: ...".How it was checked. Every table of every published catalogue is written as a CSV file in both dialects and either read back and compared field for field with the rows it came from, the type of every value included, with the table's conventions and its about equal to those of the same rows written as JSON, or refused with nothing written, each issue at a JSON pointer into a member only JSON writes or naming one of the combinations a cell does not hold: 44 of the tables go into a sheet whole, a floor the test holds, and the rest print a credit, several readings or a figure in a unit no family holds. Every mark a cell writes (an upper and a lower bound, a plus-or-minus, an approximate one, an approximate range, a range, a word, a negative value and an exponent) is written in three dialects as the exact cell the grammar reads and read back into the same row. A sheet and its JSON twin read into equal rows, and a published table read from a sheet equals the same table read from JSON. The grammar is held cell by cell to the hedge each form writes, and each refusal to its words and its line and column, the '0.^G' sentence among them. Two guards close the class: every published row class reads from a one-row sheet that fills a number and a whole number, and every field of every row class is held by its kind to what the sheet does with it (a number, a whole number, a flag or a text is a column of its kind, a set or a mapping is the basis column, a mark in the cell or a form only JSON writes, the provenance is narrowed in its own columns), so a kind the CSV front end does not know fails the guard instead of reading as a plausible refusal (the guard caught a bare
provenancecolumn answered with "did you mean 'provenance.page'?"). The physical line of a row after a quoted line break, the formula guard on the provenance columns, and a row with a refused cell not being held to the row contract each have a test of their own. The formula guard reads back exactly for ten texts from=SUM(A1)to's-Hertogenbosch. The printed pages behind the guide: RFC 4180, section 2, rules 1 to 7, and Table 6.2 of Bies, Hansen and Howard, Engineering Noise Control, 5th edition (2017), PDF pages 367 to 369 (printed folios 338 to 340), whose only credit is the one inside the name of its first row.Gates: ruff check and format, mypy over src, scripts, stub/src and tests/static_typing, bandit, the full test suite, every documentation snippet run, the site type check, build and HTML validation, the conformance report with no drift (1447 of 1447 checks), the API reference and llms files regenerated, catalogue data and site reports current, the guides index and i18n parity checks, and the related checks (frozen constants, parameter units, published sources, published catalogues, API reference, em dashes, digit grouping, decimal comma, markdown hazards, errata evidence, control characters, Spanish accents).