feat(utilities): add HoodieTableLayoutAnalyzer for small-file, micro-partition and hot-partition detection - #20152
nsivabalan wants to merge 1 commit into
Conversation
…and JSON output Adds a read-only Spark utility that reports how a Hudi table's bytes and base files are spread across partitions and whether the layout shows the symptoms that slow queries and table services down. What it reports: - table-level base-file size distribution and per-partition file-count distribution, with partition-size skew (CV, Gini, top-N share, outliers) - per-partition rows sorted by size (--enable-partition-stats, --top-n) - per-partition record counts (--include-row-counts), from the metadata table's column stats when present, with per-file Parquet footer reads only for files the column stats do not cover - three detectors (--analyze-table-characteristics): micro-partitioning, small-file pile-up with CLEAN / MODERATE / SEVERE / SKIPPED tiers, and hot partitions by recent ingest commits (compaction and clustering excluded), each with individually tunable thresholds - --output TABLE or JSON; the JSON detector section follows the same envelope as HoodieTableHealthChecker: a status per detector, a one-line summary, findings written as actionable sentences, and effectiveConfigs with effective.* thresholds and observed.* measurements - --props-path to run against several tables in one process The per-table file system view is built once and honors the table's metadata-table setting, so large tables are listed from the metadata table rather than one storage listing per partition. Metadata readers and file system views are closed per table. Also adds a runbook next to the class and a hudi-agent-gateway skill (hudi-table-layout) that runs the tool and interprets its JSON report. The existing TableSizeStats utility and its tests are not modified.
0547485 to
b07e054
Compare
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## master #20152 +/- ##
============================================
- Coverage 80.42% 80.41% -0.01%
- Complexity 34899 35061 +162
============================================
Files 2546 2547 +1
Lines 142931 143828 +897
Branches 17383 17563 +180
============================================
+ Hits 114949 115663 +714
- Misses 20065 20188 +123
- Partials 7917 7977 +60
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
hudi-agent
left a comment
There was a problem hiding this comment.
Thanks for working on this! The PR adds HoodieTableLayoutAnalyzer, a read-only Spark utility. It reports per-partition size and file-count distributions, skew metrics and optional row counts, plus micro-partition, small-file and hot-partition detectors with JSON output, and ships a Claude Code skill that reads the report. The inline comments cover a few things worth double-checking: numFiles gets capped by the histogram's sample size on large tables; failures exit with code 0, which an agent can't detect; compaction and log-compaction commits get counted as ingests; --hoodie-conf is advertised but has no effect; and the detectors interact with the date filter in ways that may be surprising. Please take a look at any inline comments, and this should be ready for a Hudi committer or PMC member to take it from here. Mostly clean and well-documented; one minor nit about a couple of redundant field comments, nothing else rises to the bar for comment.
| if (cfg.tableStats || !cfg.partitionStats) { | ||
| Snapshot ts = tableSizeHist.getSnapshot(); | ||
| System.out.println("Table-level file size distribution:"); | ||
| System.out.printf(" numFiles=%d totalBytes=%s%n", ts.size(), formatBytes(tableTotalBytes)); |
There was a problem hiding this comment.
🤖 ts.size() on a UniformReservoir is capped at the reservoir size, so any table with more than 1M base files prints numFiles=1000000. Could this use tableTotalFiles instead? The same cap affects tableSizeStats.count and the per-partition sizeStats.count (4096) in the JSON, and min/max become sampled too. Since PartitionRow.sizes is kept anyway, exact stats might be cheap.
| try { | ||
| HoodieTableLayoutAnalyzer analyzer = new HoodieTableLayoutAnalyzer(jsc, cfg); | ||
| analyzer.run(); | ||
| } catch (TableNotFoundException e) { |
There was a problem hiding this comment.
🤖 run() wraps everything in a new HoodieException, so this TableNotFoundException branch can never match. And the Throwable branch logs and returns, so spark-submit exits 0 on any failure. For a tool an agent drives, could main exit non-zero on failure (e.g. System.exit(1) after jsc.stop())? Otherwise a failed run looks the same as a run with empty output.
| return timeline.filterCompletedInstants().getInstantsAsStream() | ||
| .filter(i -> { | ||
| String action = i.getAction(); | ||
| if (action.equals(HoodieTimeline.COMMIT_ACTION) || action.equals(HoodieTimeline.DELTA_COMMIT_ACTION)) { |
There was a problem hiding this comment.
🤖 On a Merge-on-Read table, a completed compaction is a commit instant, so it gets counted here as an ingest. The javadoc says table-service commits should be excluded from this count, and clustering already is. Separately, the hot-partition window (line ~819) skips COMPACT/CLUSTER but not LOG_COMPACT. Would WriteOperationType.isTableService(op) work in both places?
| this.jsc = jsc; | ||
| this.cfg = cfg; | ||
|
|
||
| this.props = cfg.propsFilePath == null |
There was a problem hiding this comment.
🤖 props is built here but nothing ever reads it, so --hoodie-conf does nothing, even though the .md lists it as a repeatable key=value Hudi property. For example, --hoodie-conf hoodie.metadata.enable=false would be silently ignored. Could these props go into the HoodieMetadataConfig/metaClient, or could the flag be dropped? Also, --props-path is parsed both as a properties file here and as a list of base paths in run().
| * table age, since new tables legitimately have small partitions. | ||
| */ | ||
| private void runMicroPartitionDetector(List<PartitionRow> rows, TableCharacteristics tc) { | ||
| boolean countTrigger = tc.numPartitions > cfg.microPartitionCountThreshold; |
There was a problem hiding this comment.
🤖 tc.numPartitions is rows.size() after the date filter. So --num-days 7 --analyze-table-characteristics on a table with 50K daily partitions judges the count rule against about 7 partitions and reports CLEAN. Is that intended? The hot-partition scan also ignores the date filter. Maybe use allPartitions.size() for the count rule, or note the scope in the detector summary.
| if (cfg.numDays == 0) { | ||
| log.info("Start date not specified: {}.", startDate); | ||
| } else if (cfg.numDays > 0) { | ||
| endDate = LocalDate.now(); |
There was a problem hiding this comment.
🤖 With --num-days N, the interval is [today - N, today), so today's partition is always left out: --num-days 1 returns only yesterday. The skill maps "the last N days" to this flag, and today's partition is usually the hot one users care about. Is leaving it out intended? It also silently overrides any --end-date the user passed.
| // per partition and would explode driver heap on tables with many partitions. When we | ||
| // do build it, use a 4096-slot reservoir, plenty of headroom for the per-partition | ||
| // file count in practice, while keeping the per-partition cost at ~32 KB. | ||
| Histogram partitionSizeHist = cfg.partitionStats ? new Histogram(new UniformReservoir(4096)) : null; |
There was a problem hiding this comment.
🤖 The file count is already known here, so it might be worth sizing this as new UniformReservoir(Math.max(1, Math.min(4096, baseFiles.size()))). At a fixed 32 KB per partition, a table with 100K partitions needs about 3.2 GB of driver heap just for these histograms. Those are exactly the tables the micro-partition detector targets, and the skill always adds --enable-partition-stats for detector runs.
|
|
||
| // Spark context | ||
| private transient JavaSparkContext jsc; | ||
| // config |
There was a problem hiding this comment.
🤖 nit: // config and // Spark context / // Properties with source, hoodie client, key generator etc. on the fields above just restate the field name/type — could drop these and let the field names speak for themselves.
Describe the issue this Pull Request addresses
Part of #19262 (Hudi-aware agent skills and tools), under the Agentic Lakehouse umbrella #19256;
the overall plan is in #19262 (comment). This is
the "is the data laid out well?" tool; #20151 covers "are table services keeping up?".
Operators have no single way to tell whether a table's layout has gone bad — too many tiny files,
partitions too fine to be useful, or a handful of partitions absorbing every write. The existing
TableSizeStatsprints size histograms to the log, which is useful for a human eyeballing onetable but gives nothing an operator can threshold on or an agent can read.
Summary and Changelog
Adds
HoodieTableLayoutAnalyzer, a Spark utility that walks a table's partitions and reportsper-partition size and file-count distributions, skew metrics, optional row counts, and — under
--analyze-table-characteristics— three detectors with explicit verdicts:CLEAN | FLAGGED | SKIPPED): partition count above a threshold,or partitions holding many files that average well under the target file size. The size rule
only applies once the table is old enough for it to mean something.
CLEAN | MODERATE | SEVERE | SKIPPED): the fraction of qualifying partitionswhose files average under the small-file threshold. Skipped when the table has too few ingest
commits for the signal to be reliable.
commits, with compaction and clustering excluded from the window.
Every threshold is a CLI flag with a documented default. Output is a human-readable table or
--output JSON; the JSON carries, per detector, astatus, a one-linesummary,findings[]written as actionable sentences, and
effectiveConfigs{}whereeffective.*keys are thethresholds actually applied and
observed.*keys are measurements — the same envelope conventionas
HoodieTableHealthCheckerin #20151, so one skill or agent can read both.Row counts (
--include-row-counts) use metadata-table column stats when available and fall back toParquet footers only for files the stats did not cover.
Why a new class rather than extending
TableSizeStats.TableSizeStatsis a long-standingspark-submitentry point whose log output someone may be scraping; changing it to structuredoutput would be a silent behavior change. And "size stats" no longer describes a tool with
detectors, verdicts, and skew analysis. So the analyzer is additive:
TableSizeStatsand its testsare untouched, and whether to deprecate it later is a separate conversation. The analyzer builds
on the same partition-walking and histogram approach but is self-contained; no shared helpers
were extracted.
The per-table
FileSystemViewis built once and honors the table's metadata-table setting, so onlarge object-store tables partitions are listed from the metadata table instead of one storage
listing per partition. Both the metadata reader and the view are closed per table, which matters
under
--props-pathbatch runs over many tables.Docs —
HoodieTableLayoutAnalyzer.md: invocation, every flag, each detector's rule, the JSONschema, and caveats (table age and ingest-commit counts are lower bounds; log files are not
counted toward small-file sizing on Merge-on-Read).
Skill —
hudi-agent-gateway/skills/hudi-table-layout/, following the layout ofhudi-architect(#19380) andhudi-table-health(#20151):SKILL.mddrives collecting the basepath and thresholds, running with
--output JSON, and turning each detector's verdict into arecommendation (clustering, small-file config, partitioning revisit) without acting on the table.
Tests —
TestHoodieTableLayoutAnalyzer, 31 tests onHoodieSparkClientTestBase: each detector'sverdict at and around its thresholds, the skipped cases, hot-partition record accounting
(per-operation counters, not
numWrites), JSON escaping, and the JSON envelope shape.Impact
Purely additive. New class, new test, new runbook, new skill directory.
@PublicAPIClass/@PublicAPIMethodtouched.hoodie.*configs. Thresholds are CLI flags on this tool only.TableSizeStatsis not modified.Risk Level
none
Read-only tool on a new code path, reachable only by explicit invocation.
On size: this is 3.5k lines, of which ~1.1k are tests and ~0.6k are docs and the skill. It is one
self-contained tool; the only clean split would be moving the skill directory to a follow-up.
Happy to do that if reviewers prefer.
Documentation Update
HoodieTableLayoutAnalyzer.mdis added alongside the tool. No new configs. Happy to add a websitepage if reviewers would like one.
Contributor's checklist