tools-update-cron: sync 2026-08-16 — 16 skill(s) updated
This commit is contained in:
+160
-69
@@ -1,105 +1,196 @@
|
||||
---
|
||||
name: xlsx
|
||||
description: "Create, read, edit Excel .xlsx spreadsheets and CSVs."
|
||||
version: 1.0.0
|
||||
author: Anthropic (adapted by Nous Research)
|
||||
license: Proprietary. LICENSE.txt has complete terms
|
||||
description: Create, read, edit Excel .xlsx workbooks and CSVs.
|
||||
version: 1.1.0
|
||||
author: Nous Research
|
||||
license: MIT
|
||||
platforms: [linux, macos, windows]
|
||||
metadata:
|
||||
hermes:
|
||||
tags: [Excel, XLSX, Spreadsheets, Office, Productivity]
|
||||
tags: [excel, spreadsheet, xlsx, csv, openpyxl, productivity]
|
||||
category: productivity
|
||||
related_skills: [docx, pdf, powerpoint]
|
||||
---
|
||||
|
||||
# XLSX Skill
|
||||
# Xlsx Skill
|
||||
|
||||
Create, read, and edit Excel workbooks — formulas, formatting, charts, data cleaning, and format conversion. Every formula-bearing output must be recalculated and error-free before delivery.
|
||||
Work with Excel .xlsx workbooks using Python and openpyxl: build styled
|
||||
multi-sheet workbooks with formulas and charts, inspect or dump existing
|
||||
files, edit cells and structure, and convert to/from CSV. All helper
|
||||
scripts are argparse CLIs that print JSON and use explicit UTF-8 I/O.
|
||||
|
||||
## When to Use
|
||||
|
||||
Use this skill any time a spreadsheet file is the primary input or output: opening, reading, editing, or fixing an existing .xlsx, .xlsm, .xltx, .csv, or .tsv file; creating a new spreadsheet from scratch or from other data; converting between tabular formats; cleaning messy tabular data into a proper spreadsheet. Trigger whenever the user references a spreadsheet file by name or path — even casually. Do NOT trigger when the deliverable is a Word document (`docx` skill), HTML report, standalone script, or Google Sheets API integration. For finance-grade modeling conventions (DCF, LBO, three-statement), the optional `excel-author` skill adds stricter standards on top of this one.
|
||||
- Creating .xlsx reports: multiple sheets, number formats, styling,
|
||||
merged cells, freeze panes, autofilter, conditional formatting,
|
||||
charts, data-validation dropdowns, native Excel tables, defined
|
||||
names, hyperlinks, cell notes, sheet protection.
|
||||
- Reading a workbook: sheet inventory, dumping data as JSON or CSV,
|
||||
listing formulas vs cached values, notes, defined names, tables.
|
||||
- Editing existing files: set cells, append rows, insert/delete
|
||||
rows/columns (reference-aware via `xlsx_restructure.py`),
|
||||
copy/rename sheets, tables, names, notes, protection.
|
||||
- Recalculating formulas headlessly via LibreOffice
|
||||
(`xlsx_recalc.py`).
|
||||
- CSV interop with type inference and non-UTF-8 encodings.
|
||||
- Not for the legacy .xls binary format (use LibreOffice to convert
|
||||
first: `soffice --headless --convert-to xlsx old.xls`).
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- Python 3.10+ with `openpyxl` (`pip install openpyxl`). No other
|
||||
third-party packages are needed; everything else is stdlib.
|
||||
- Optional: LibreOffice (`soffice`) for headless recalculation or
|
||||
format conversion.
|
||||
|
||||
## How to Run
|
||||
|
||||
Run the helper scripts with the `terminal` tool from this skill's
|
||||
`scripts/` directory (every script supports `--help`):
|
||||
|
||||
```bash
|
||||
pip install openpyxl pandas "markitdown[xlsx]"
|
||||
which soffice || sudo apt install -y libreoffice # formula recalculation (scripts/recalc.py)
|
||||
python scripts/xlsx_create.py spec.json report.xlsx # build from JSON spec
|
||||
python scripts/xlsx_read.py report.xlsx --sheets # inventory
|
||||
python scripts/xlsx_read.py report.xlsx --json --sheet Data
|
||||
python scripts/xlsx_read.py report.xlsx --formulas
|
||||
python scripts/xlsx_edit.py report.xlsx --sheet Data --set B2=42 --recalc
|
||||
python scripts/xlsx_restructure.py report.xlsx --sheet Data --insert-rows 3:2
|
||||
python scripts/xlsx_recalc.py report.xlsx
|
||||
python scripts/csv_to_xlsx.py data.csv out.xlsx --encoding utf-8
|
||||
python scripts/xlsx_to_csv.py report.xlsx out.csv --sheet Data
|
||||
```
|
||||
|
||||
macOS: `brew install libreoffice`.
|
||||
Author the JSON spec with `write_file`, inspect script JSON output with
|
||||
`read_file` or directly from stdout.
|
||||
|
||||
## Quick Reference
|
||||
|
||||
| Task | Approach |
|
||||
| Task | Command |
|
||||
|---|---|
|
||||
| **Create** or **edit** with formulas/formatting | `openpyxl` — see gotchas below |
|
||||
| **Bulk data** in or out | `pandas` (`read_excel`, `to_excel`) |
|
||||
| **Quick look** at a sheet | `markitdown file.xlsx` — `## SheetName` per sheet; reads `.xlsm` too. No cell coordinates, so don't plan edits from it. (`read_file` also auto-extracts .xlsx) |
|
||||
| **Read** a model (formulas *and* values) | two `load_workbook` passes — see gotchas |
|
||||
| Create workbook from spec | `xlsx_create.py spec.json out.xlsx` |
|
||||
| Sheet names + dimensions | `xlsx_read.py f.xlsx --sheets` |
|
||||
| Dump sheet as JSON | `xlsx_read.py f.xlsx --json --sheet S` |
|
||||
| Dump sheet as CSV | `xlsx_read.py f.xlsx --csv --out d.csv` |
|
||||
| List formulas + cached values | `xlsx_read.py f.xlsx --formulas` |
|
||||
| Set a cell / formula | `xlsx_edit.py f.xlsx --set "A1==SUM(B:B)"` |
|
||||
| Append a row | `xlsx_edit.py f.xlsx --append '[1,"x",true]'` |
|
||||
| Insert 2 rows, refs NOT shifted | `xlsx_edit.py f.xlsx --insert-rows 3:2` |
|
||||
| Insert 2 rows, refs shifted | `xlsx_restructure.py f.xlsx --insert-rows 3:2` |
|
||||
| Delete a column, refs shifted | `xlsx_restructure.py f.xlsx --delete-cols B` |
|
||||
| Create a native table | `xlsx_edit.py f.xlsx --add-table Sales:A1:C9` |
|
||||
| Append inside a table | `--table-append 'Sales=["West",5]'` |
|
||||
| List tables | `xlsx_edit.py f.xlsx --list-tables` |
|
||||
| Defined names | `--define-name "Rates='Data'!$B$2:$B$9"` / `--delete-name Rates` / `xlsx_read.py f.xlsx --names` |
|
||||
| Hyperlink | `--hyperlink "A1=https://example.com|Docs"` |
|
||||
| Cell note | `--note "B2=Check this|Reviewer"`; read via `xlsx_read.py f.xlsx --notes` |
|
||||
| Protect sheet (see Pitfalls) | `--protect your-password --unlock B2:B9` |
|
||||
| Recalculate via LibreOffice | `xlsx_recalc.py f.xlsx` |
|
||||
| Copy / rename sheet | `--copy-sheet Src:New --rename-sheet Old:New` |
|
||||
| Force recalc on open | `xlsx_edit.py f.xlsx --recalc` |
|
||||
| CSV -> styled xlsx | `csv_to_xlsx.py in.csv out.xlsx` |
|
||||
| xlsx -> CSV | `xlsx_to_csv.py f.xlsx out.csv --encoding utf-8` |
|
||||
|
||||
> Script paths below are relative to this skill's directory.
|
||||
## Procedure
|
||||
|
||||
## Requirements for every output
|
||||
1. **Create**: write a JSON spec (schema documented in
|
||||
`xlsx_create.py --help` and its docstring). Each sheet supports
|
||||
`rows` (scalars or styled cell objects), sparse `cells` overrides,
|
||||
`column_widths`, `row_heights`, `merges`, `freeze_panes`,
|
||||
`autofilter`, `conditional_formats` (cell_is rules and color
|
||||
scales), `charts` (bar/line/pie from cell ranges),
|
||||
`validations` (list dropdowns), `tables` (native Excel tables with
|
||||
a style name), and `protection`. Workbook-level `defined_names`
|
||||
maps names to refs. Cell objects also take `hyperlink` and `note`.
|
||||
Typed values: JSON numbers/bools
|
||||
pass through; dates use `{"value": "2026-01-31", "type": "date"}`.
|
||||
Number formats are Excel format strings: currency `"$#,##0.00"`,
|
||||
percent `"0.0%"`, date `"yyyy-mm-dd"`.
|
||||
2. **Formulas**: set with `"formula": "SUM(B2:B9)"` in the spec or
|
||||
`--set "C1==SUM(A:A)"` in the editor. When writing formulas, add
|
||||
`"full_calc_on_load": true` (spec) or `--recalc` (editor); this sets
|
||||
the workbook's `fullCalcOnLoad` flag so Excel/LibreOffice recompute
|
||||
everything on open. openpyxl itself NEVER evaluates formulas.
|
||||
3. **Read**: `--sheets` for inventory (names, dimensions, merged
|
||||
ranges, chart count, tables, protection, defined names),
|
||||
`--json`/`--csv` for data, `--formulas` to
|
||||
pair each formula string with its cached result, `--notes` for
|
||||
cell comments, `--names` for defined names. Cached results
|
||||
exist only if the file was last saved by a real spreadsheet app;
|
||||
files fresh from openpyxl return `null` there. To materialize
|
||||
results headlessly run `xlsx_recalc.py file.xlsx` (uses
|
||||
LibreOffice; prints `{"recalculated": false, ...}` and exits 0
|
||||
when `soffice` is absent), then reload with `--data-only`.
|
||||
4. **Edit**: `xlsx_edit.py` applies renames/copies first, then
|
||||
structural row/column changes, then `--set`/`--append`. It edits in
|
||||
place unless `--out` is given — copy the file first if you need the
|
||||
original.
|
||||
5. **Restructure**: for insert/delete on sheets that have formulas,
|
||||
merges, tables, or filters, use `xlsx_restructure.py` instead of
|
||||
`xlsx_edit.py`. It rewrites formula references on ALL sheets
|
||||
(absolute `$` refs, ranges, cross-sheet refs), shifts merges,
|
||||
autofilter, freeze panes, validation and conditional-format
|
||||
ranges, table refs, defined names, and row/column dimensions, then
|
||||
prints a JSON report including a `not_shifted` list. Rules and
|
||||
limits: `references/restructuring.md`.
|
||||
6. **CSV interop**: `csv_to_xlsx.py` infers int/float/bool/ISO-date
|
||||
per cell and styles the header row; `xlsx_to_csv.py` writes ISO
|
||||
dates and blank strings for empty cells. Both default to UTF-8 and
|
||||
accept `--encoding` (e.g. `utf-8-sig` for Excel-friendly BOM,
|
||||
`cp1252` for legacy Windows exports).
|
||||
|
||||
- **Professional font** (Arial, Times New Roman) throughout, unless the user says otherwise.
|
||||
- **Zero formula errors.** Never ship while `recalc.py` reports `errors_found`. If you think an error predates you, prove it: load the *original* with `data_only=True` and look at that cell. An error you introduced looks exactly like one you inherited.
|
||||
- **Use formulas, never hardcoded results.** Write `sheet['B10'] = '=SUM(B2:B9)'`, not the Python-computed total. The sheet must recalculate when its inputs change.
|
||||
- **Follow the user's spec literally.** Exact tab names, exact column headers, and the formula they spelled out. A redesign that computes something else fails, however elegant.
|
||||
- **Document every assumption and hardcoded number** where the reader will see it — a cell comment, or an adjacent cell at a table's end. Cite a real source when one exists; when the number came from the user, say so plainly.
|
||||
- **A workbook *you create* for someone to fill in** needs a short legend naming which cells to edit, and one example row of realistic values showing the expected format. Never add such a row to a file you were asked to edit.
|
||||
- **Editing an existing file: match its conventions exactly.** They override every guideline here. Find its designated input cells first — a distinct font color, fill, or shading marks them — write only there, and leave every existing formula untouched.
|
||||
## Converting to PDF
|
||||
|
||||
## Recalculate (mandatory whenever the file contains formulas)
|
||||
|
||||
openpyxl writes formulas as strings with **no cached values**. Until you recalculate, every formula cell reads back as `None` to anything reading cached values — `pandas`, `load_workbook(data_only=True)`, and most previewers.
|
||||
LibreOffice converts headlessly (also works for CSV export of a single
|
||||
sheet):
|
||||
|
||||
```bash
|
||||
python scripts/recalc.py output.xlsx [timeout_seconds] # default 30
|
||||
soffice --headless --convert-to pdf report.xlsx --outdir out/
|
||||
soffice --headless --convert-to csv report.xlsx --outdir out/ # 1st sheet only
|
||||
```
|
||||
|
||||
LibreOffice computes every formula, the file is **rewritten in place**, and you get JSON: `status` (`success` | `errors_found`), `total_formulas`, `total_errors`, and an `error_summary` naming up to 100 cells per error type (`locations_truncated` says how many it withheld — trust `total_errors`, not the length of the list). Fix what it names and run it again. **JSON with an `error` key instead of a `status` means nothing was recalculated**, and only that case exits non-zero — `errors_found` exits 0, so never treat a clean exit as a clean workbook.
|
||||
Only the first sheet lands in a CSV; for other sheets use
|
||||
`xlsx_to_csv.py --sheet NAME`. If `soffice` is missing, install
|
||||
LibreOffice or hand the file to the user unconverted.
|
||||
|
||||
**A green recalc proves your formulas *evaluate*, not that they are *right*.** An off-by-one range or a reference to the wrong row yields a clean, error-free file with wrong numbers. Write 2–3 formulas first and check they pull the values you expect, before building out a grid.
|
||||
## Pitfalls
|
||||
|
||||
**A workbook that links to another file loses those links** if you re-save it with openpyxl and then recalculate. Such a formula reads `='[1]Returns Analysis'!$B$2` — the `[1]` is an index into the workbook's external-reference list, naming a *separate file on disk*, not a sheet. That file is rarely present, so the cell's cached value is the only thing holding its data. openpyxl strips that value on save; LibreOffice then has to resolve the reference for real, fails, writes `#NAME?`, and deletes every link. `recalc.py` refuses to run in that state — copy those cells' values out of the original before you save over them (`--force` overrides, and accepts the loss).
|
||||
|
||||
## Choosing formulas that survive verification
|
||||
|
||||
LibreOffice implements fewer functions than Excel, and one it cannot evaluate becomes a literal `#NAME?` baked into the file you deliver.
|
||||
|
||||
- **Prefer Excel-2007-era functions** — `SUMIFS`, `INDEX`, `MATCH`, `IFERROR`, `SUMPRODUCT` — which need no prefix.
|
||||
- **Six post-2007 functions work, but only with an `_xlfn.` prefix**, because openpyxl writes your formula into the XML verbatim and Excel stores post-2007 names prefixed (its UI hides the prefix): `_xlfn.TEXTJOIN`, `_xlfn.CONCAT`, `_xlfn.IFS`, `_xlfn.SWITCH`, `_xlfn.MAXIFS`, `_xlfn.MINIFS`. Written bare, each yields `#NAME?`.
|
||||
- **Never use `XLOOKUP`, `XMATCH`, `SORT`, `FILTER`, `UNIQUE`, or `SEQUENCE`.** LibreOffice cannot reliably evaluate them; newer builds that do are spilling array functions, and an openpyxl-written file has no spill metadata, so only the top-left cell of the range gets a value — and `recalc.py` reports `total_errors: 0` on the truncated result. Use `INDEX`/`MATCH` for lookups, and sort, filter, and de-duplicate in Python before writing the cells.
|
||||
- A formula LibreOffice could not parse is written back **lowercased** — a quick tell beside a `#NAME?`.
|
||||
|
||||
## openpyxl gotchas
|
||||
|
||||
- **Reading a model takes two loads.** `data_only=True` yields cached values with the formulas gone; the default yields formula strings with no values. One pass cannot give you both.
|
||||
- **`data_only=True` is destructive if you save.** That workbook has no formulas left, so saving replaces every one with a literal — permanently.
|
||||
- **`data_only=True` on a file openpyxl just wrote returns `None` everywhere** — run `recalc.py` first. (A formula whose result is `""` also reads back as `None`.)
|
||||
- **Merged cells: write the top-left anchor only.** Every other cell in the range is a `MergedCell` whose `.value` is read-only.
|
||||
- **`.xlsm` loses its macros unless you pass `keep_vba=True`** to `load_workbook`.
|
||||
- **A sheet name containing a space must be quoted** in a cross-sheet reference: `='Assumptions Inputs'!$B$5`. Unquoted, it evaluates to `#VALUE!`.
|
||||
|
||||
## Financial models
|
||||
|
||||
Unless the user says otherwise, or the existing file already does something else.
|
||||
|
||||
**Color:** blue text (`0,0,255`) for hardcoded inputs and scenario levers · black for formulas · green (`0,128,0`) for links to another sheet · red (`255,0,0`) for links to another file · yellow fill (`255,255,0`) for key assumptions and cells the user should fill in.
|
||||
|
||||
**Numbers:** currency `$#,##0`, with the unit named in the header (`Revenue ($mm)`) · zeros render as `-`, including in percentages (`$#,##0;($#,##0);-`) · negatives in parentheses · percentages `0.0%`, **stored as fractions** (`0.15` renders `15.0%`; storing `15` renders `1500.0%`) · valuation multiples `0.0x` · years as text (`"2024"`, never `2,024`).
|
||||
|
||||
**Structure:** every assumption in its own labeled cell, referenced by the formulas that use it (`=B5*(1+$B$6)`, never `=B5*1.05`) · formulas consistent across every projection period, since a lone edited cell mid-row is the commonest silent error · guard denominators that can be zero.
|
||||
|
||||
For full investment-banking conventions (balance checks, sensitivity tables, named ranges), install the optional skill: `hermes skills install official/finance/excel-author`.
|
||||
- **openpyxl does not calculate.** Formula results are available only
|
||||
via `load_workbook(path, data_only=True)` and only when the file was
|
||||
previously saved by Excel/LibreOffice. Otherwise you get `None`.
|
||||
- **`xlsx_edit.py` insert/delete does not shift references** (raw
|
||||
openpyxl behavior). Use `xlsx_restructure.py`, which does — but even
|
||||
it cannot move chart anchors, images, or conditional-format RULE
|
||||
formulas; read its JSON report's `not_shifted` list and
|
||||
`references/restructuring.md`.
|
||||
- **Sheet protection is NOT security.** `--protect` sets the standard
|
||||
xlsx sheet-protection hash: it signals "don't edit this" to
|
||||
well-behaved apps and nothing more. Anyone can strip it by editing
|
||||
the zip's XML or unchecking it in LibreOffice. Never rely on it for
|
||||
confidentiality or integrity; it does not encrypt anything.
|
||||
- **`data_only=True` then save** silently discards all formulas
|
||||
(cached values replace them). Never save a workbook loaded that way
|
||||
unless that is the goal.
|
||||
- **Loading strips charts/images**: openpyxl does not round-trip
|
||||
charts, so editing a charted workbook and saving drops the charts.
|
||||
Re-add charts after editing, or avoid re-saving charted files.
|
||||
- **CSV locale traps**: always pass explicit encodings (the scripts
|
||||
already do) and remember European CSVs often use `;` delimiters and
|
||||
decimal commas — use `--delimiter ';'` and expect strings like
|
||||
`"12,5"` to stay strings.
|
||||
- **Dates are datetimes**: Excel stores dates as serial numbers;
|
||||
openpyxl returns `datetime`/`date` objects. Dumps here emit ISO
|
||||
strings.
|
||||
- Sheet names are capped at 31 chars and reject `[ ] : * ? / \`.
|
||||
|
||||
## Verification
|
||||
|
||||
1. `python scripts/recalc.py output.xlsx` → `status: success`, `total_errors: 0`.
|
||||
2. Spot-check 2–3 computed cells against expected values (`load_workbook(data_only=True)` *after* recalc).
|
||||
3. `markitdown output.xlsx` — scan for missing sheets, misplaced headers, leftover placeholders.
|
||||
|
||||
## Related skills
|
||||
|
||||
`docx` (Word documents), `pdf` (PDF work), `powerpoint` (decks), optional `excel-author` (finance-grade modeling standards).
|
||||
- After creating: `xlsx_read.py out.xlsx --sheets` and confirm sheet
|
||||
names, dimensions, merged ranges, and chart counts match intent.
|
||||
- Dump data with `--json` and compare against the source values.
|
||||
- After edits: re-dump the touched range; if formulas were written,
|
||||
confirm `--formulas` lists them and that `--recalc` was applied.
|
||||
- After `xlsx_restructure.py`: read its JSON report, then re-run
|
||||
`--formulas` and `--sheets` to confirm references and ranges landed
|
||||
where expected.
|
||||
- For a full visual check, open in LibreOffice:
|
||||
`soffice --headless --convert-to pdf out.xlsx` and inspect the PDF.
|
||||
|
||||
Reference in New Issue
Block a user