Skip to content

feat(utilities): add HoodieTableLayoutAnalyzer for small-file, micro-partition and hot-partition detection - #20152

Open
nsivabalan wants to merge 1 commit into
apache:masterfrom
nsivabalan:layout-analyzer
Open

nsivabalan wants to merge 1 commit into
apache:masterfrom
nsivabalan:layout-analyzer

Conversation

@nsivabalan

Copy link
Copy Markdown
Contributor

Describe the issue this Pull Request addresses

Part of #19262 (Hudi-aware agent skills and tools), under the Agentic Lakehouse umbrella #19256;
the overall plan is in #19262 (comment). This is
the "is the data laid out well?" tool; #20151 covers "are table services keeping up?".

Operators have no single way to tell whether a table's layout has gone bad — too many tiny files,
partitions too fine to be useful, or a handful of partitions absorbing every write. The existing
TableSizeStats prints size histograms to the log, which is useful for a human eyeballing one
table but gives nothing an operator can threshold on or an agent can read.

Summary and Changelog

Adds HoodieTableLayoutAnalyzer, a Spark utility that walks a table's partitions and reports
per-partition size and file-count distributions, skew metrics, optional row counts, and — under
--analyze-table-characteristics — three detectors with explicit verdicts:

  • micro-partition (table-level, CLEAN | FLAGGED | SKIPPED): partition count above a threshold,
    or partitions holding many files that average well under the target file size. The size rule
    only applies once the table is old enough for it to mean something.
  • small-files (CLEAN | MODERATE | SEVERE | SKIPPED): the fraction of qualifying partitions
    whose files average under the small-file threshold. Skipped when the table has too few ingest
    commits for the signal to be reliable.
  • hot-partitions (per-partition): partitions written by at least half of the last N ingest
    commits, with compaction and clustering excluded from the window.

Every threshold is a CLI flag with a documented default. Output is a human-readable table or
--output JSON; the JSON carries, per detector, a status, a one-line summary, findings[]
written as actionable sentences, and effectiveConfigs{} where effective.* keys are the
thresholds actually applied and observed.* keys are measurements — the same envelope convention
as HoodieTableHealthChecker in #20151, so one skill or agent can read both.

Row counts (--include-row-counts) use metadata-table column stats when available and fall back to
Parquet footers only for files the stats did not cover.

Why a new class rather than extending TableSizeStats. TableSizeStats is a long-standing
spark-submit entry point whose log output someone may be scraping; changing it to structured
output would be a silent behavior change. And "size stats" no longer describes a tool with
detectors, verdicts, and skew analysis. So the analyzer is additive: TableSizeStats and its tests
are untouched, and whether to deprecate it later is a separate conversation. The analyzer builds
on the same partition-walking and histogram approach but is self-contained; no shared helpers
were extracted.

The per-table FileSystemView is built once and honors the table's metadata-table setting, so on
large object-store tables partitions are listed from the metadata table instead of one storage
listing per partition. Both the metadata reader and the view are closed per table, which matters
under --props-path batch runs over many tables.

Docs — HoodieTableLayoutAnalyzer.md: invocation, every flag, each detector's rule, the JSON
schema, and caveats (table age and ingest-commit counts are lower bounds; log files are not
counted toward small-file sizing on Merge-on-Read).

Skill — hudi-agent-gateway/skills/hudi-table-layout/, following the layout of
hudi-architect (#19380) and hudi-table-health (#20151): SKILL.md drives collecting the base
path and thresholds, running with --output JSON, and turning each detector's verdict into a
recommendation (clustering, small-file config, partitioning revisit) without acting on the table.

Tests — TestHoodieTableLayoutAnalyzer, 31 tests on HoodieSparkClientTestBase: each detector's
verdict at and around its thresholds, the skipped cases, hot-partition record accounting
(per-operation counters, not numWrites), JSON escaping, and the JSON envelope shape.

Impact

Purely additive. New class, new test, new runbook, new skill directory.

  • No public API changes; no @PublicAPIClass / @PublicAPIMethod touched.
  • No storage format changes. Read-only; never modifies a table.
  • No new hoodie.* configs. Thresholds are CLI flags on this tool only.
  • TableSizeStats is not modified.

Risk Level

none

Read-only tool on a new code path, reachable only by explicit invocation.

On size: this is 3.5k lines, of which ~1.1k are tests and ~0.6k are docs and the skill. It is one
self-contained tool; the only clean split would be moving the skill directory to a follow-up.
Happy to do that if reviewers prefer.

Documentation Update

HoodieTableLayoutAnalyzer.md is added alongside the tool. No new configs. Happy to add a website
page if reviewers would like one.

Contributor's checklist

  • Read through contributor's guide
  • Enough context is provided in the sections above
  • Adequate tests were added if applicable

@github-actions github-actions Bot added the size:XL PR with lines of changes > 1000 label Sep 30, 2026
…and JSON output

Adds a read-only Spark utility that reports how a Hudi table's bytes and
base files are spread across partitions and whether the layout shows the
symptoms that slow queries and table services down.

What it reports:
- table-level base-file size distribution and per-partition file-count
  distribution, with partition-size skew (CV, Gini, top-N share, outliers)
- per-partition rows sorted by size (--enable-partition-stats, --top-n)
- per-partition record counts (--include-row-counts), from the metadata
  table's column stats when present, with per-file Parquet footer reads
  only for files the column stats do not cover
- three detectors (--analyze-table-characteristics): micro-partitioning,
  small-file pile-up with CLEAN / MODERATE / SEVERE / SKIPPED tiers, and
  hot partitions by recent ingest commits (compaction and clustering
  excluded), each with individually tunable thresholds
- --output TABLE or JSON; the JSON detector section follows the same
  envelope as HoodieTableHealthChecker: a status per detector, a one-line
  summary, findings written as actionable sentences, and effectiveConfigs
  with effective.* thresholds and observed.* measurements
- --props-path to run against several tables in one process

The per-table file system view is built once and honors the table's
metadata-table setting, so large tables are listed from the metadata
table rather than one storage listing per partition. Metadata readers and
file system views are closed per table.

Also adds a runbook next to the class and a hudi-agent-gateway skill
(hudi-table-layout) that runs the tool and interprets its JSON report.

The existing TableSizeStats utility and its tests are not modified.
@codecov-commenter

codecov-commenter commented Sep 30, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 79.04125% with 188 lines in your changes missing coverage. Please review.
✅ Project coverage is 80.41%. Comparing base (3ec54a5) to head (b07e054).
⚠️ Report is 1 commits behind head on master.

Files with missing lines Patch % Lines
...ache/hudi/utilities/HoodieTableLayoutAnalyzer.java 79.04% 124 Missing and 64 partials ⚠️
Additional details and impacted files
@@             Coverage Diff              @@
##             master   #20152      +/-   ##
============================================
- Coverage     80.42%   80.41%   -0.01%     
- Complexity    34899    35061     +162     
============================================
  Files          2546     2547       +1     
  Lines        142931   143828     +897     
  Branches      17383    17563     +180     
============================================
+ Hits         114949   115663     +714     
- Misses        20065    20188     +123     
- Partials       7917     7977      +60     
Components Coverage Δ
hudi-common 83.98% <ø> (+0.01%) ⬆️
hudi-client 83.51% <ø> (+0.03%) ⬆️
hudi-flink 85.75% <ø> (+0.01%) ⬆️
hudi-spark-datasource 73.85% <ø> (ø)
hudi-utilities 78.21% <79.04%> (+0.02%) ⬆️
hudi-cli 70.05% <ø> (ø)
hudi-hadoop 70.99% <ø> (+0.01%) ⬆️
hudi-sync 76.02% <ø> (ø)
hudi-io 81.50% <ø> (ø)
hudi-timeline-service 83.06% <ø> (ø)
hudi-cloud 81.00% <ø> (ø)
hudi-kafka-connect 53.20% <ø> (-0.77%) ⬇️
Flag Coverage Δ
common-and-other-modules 52.02% <0.00%> (-0.34%) ⬇️
flink-integration-tests 49.46% <ø> (+<0.01%) ⬆️
hadoop-mr-java-client 43.98% <ø> (-0.01%) ⬇️
integration-tests 13.36% <0.00%> (-0.09%) ⬇️
spark-client-hadoop-common 38.30% <0.00%> (-0.29%) ⬇️
spark-java-tests 52.00% <0.00%> (-0.39%) ⬇️
spark-scala-tests 46.70% <0.00%> (-0.36%) ⬇️
utilities 37.18% <79.04%> (+0.33%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
...ache/hudi/utilities/HoodieTableLayoutAnalyzer.java 79.04% <79.04%> (ø)

... and 12 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@hudi-agent hudi-agent left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ 🤖 This review was generated by an AI agent and may contain mistakes. Please verify any suggestions before applying.

Thanks for working on this! The PR adds HoodieTableLayoutAnalyzer, a read-only Spark utility. It reports per-partition size and file-count distributions, skew metrics and optional row counts, plus micro-partition, small-file and hot-partition detectors with JSON output, and ships a Claude Code skill that reads the report. The inline comments cover a few things worth double-checking: numFiles gets capped by the histogram's sample size on large tables; failures exit with code 0, which an agent can't detect; compaction and log-compaction commits get counted as ingests; --hoodie-conf is advertised but has no effect; and the detectors interact with the date filter in ways that may be surprising. Please take a look at any inline comments, and this should be ready for a Hudi committer or PMC member to take it from here. Mostly clean and well-documented; one minor nit about a couple of redundant field comments, nothing else rises to the bar for comment.

if (cfg.tableStats || !cfg.partitionStats) {
Snapshot ts = tableSizeHist.getSnapshot();
System.out.println("Table-level file size distribution:");
System.out.printf(" numFiles=%d totalBytes=%s%n", ts.size(), formatBytes(tableTotalBytes));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 ts.size() on a UniformReservoir is capped at the reservoir size, so any table with more than 1M base files prints numFiles=1000000. Could this use tableTotalFiles instead? The same cap affects tableSizeStats.count and the per-partition sizeStats.count (4096) in the JSON, and min/max become sampled too. Since PartitionRow.sizes is kept anyway, exact stats might be cheap.

⚠️ AI-generated; verify before applying. React 👍/👎 to flag quality.

try {
HoodieTableLayoutAnalyzer analyzer = new HoodieTableLayoutAnalyzer(jsc, cfg);
analyzer.run();
} catch (TableNotFoundException e) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 run() wraps everything in a new HoodieException, so this TableNotFoundException branch can never match. And the Throwable branch logs and returns, so spark-submit exits 0 on any failure. For a tool an agent drives, could main exit non-zero on failure (e.g. System.exit(1) after jsc.stop())? Otherwise a failed run looks the same as a run with empty output.

⚠️ AI-generated; verify before applying. React 👍/👎 to flag quality.

return timeline.filterCompletedInstants().getInstantsAsStream()
.filter(i -> {
String action = i.getAction();
if (action.equals(HoodieTimeline.COMMIT_ACTION) || action.equals(HoodieTimeline.DELTA_COMMIT_ACTION)) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 On a Merge-on-Read table, a completed compaction is a commit instant, so it gets counted here as an ingest. The javadoc says table-service commits should be excluded from this count, and clustering already is. Separately, the hot-partition window (line ~819) skips COMPACT/CLUSTER but not LOG_COMPACT. Would WriteOperationType.isTableService(op) work in both places?

⚠️ AI-generated; verify before applying. React 👍/👎 to flag quality.

this.jsc = jsc;
this.cfg = cfg;

this.props = cfg.propsFilePath == null

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 props is built here but nothing ever reads it, so --hoodie-conf does nothing, even though the .md lists it as a repeatable key=value Hudi property. For example, --hoodie-conf hoodie.metadata.enable=false would be silently ignored. Could these props go into the HoodieMetadataConfig/metaClient, or could the flag be dropped? Also, --props-path is parsed both as a properties file here and as a list of base paths in run().

⚠️ AI-generated; verify before applying. React 👍/👎 to flag quality.

* table age, since new tables legitimately have small partitions.
*/
private void runMicroPartitionDetector(List<PartitionRow> rows, TableCharacteristics tc) {
boolean countTrigger = tc.numPartitions > cfg.microPartitionCountThreshold;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 tc.numPartitions is rows.size() after the date filter. So --num-days 7 --analyze-table-characteristics on a table with 50K daily partitions judges the count rule against about 7 partitions and reports CLEAN. Is that intended? The hot-partition scan also ignores the date filter. Maybe use allPartitions.size() for the count rule, or note the scope in the detector summary.

⚠️ AI-generated; verify before applying. React 👍/👎 to flag quality.

if (cfg.numDays == 0) {
log.info("Start date not specified: {}.", startDate);
} else if (cfg.numDays > 0) {
endDate = LocalDate.now();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 With --num-days N, the interval is [today - N, today), so today's partition is always left out: --num-days 1 returns only yesterday. The skill maps "the last N days" to this flag, and today's partition is usually the hot one users care about. Is leaving it out intended? It also silently overrides any --end-date the user passed.

⚠️ AI-generated; verify before applying. React 👍/👎 to flag quality.

// per partition and would explode driver heap on tables with many partitions. When we
// do build it, use a 4096-slot reservoir, plenty of headroom for the per-partition
// file count in practice, while keeping the per-partition cost at ~32 KB.
Histogram partitionSizeHist = cfg.partitionStats ? new Histogram(new UniformReservoir(4096)) : null;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 The file count is already known here, so it might be worth sizing this as new UniformReservoir(Math.max(1, Math.min(4096, baseFiles.size()))). At a fixed 32 KB per partition, a table with 100K partitions needs about 3.2 GB of driver heap just for these histograms. Those are exactly the tables the micro-partition detector targets, and the skill always adds --enable-partition-stats for detector runs.

⚠️ AI-generated; verify before applying. React 👍/👎 to flag quality.


// Spark context
private transient JavaSparkContext jsc;
// config

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 nit: // config and // Spark context / // Properties with source, hoodie client, key generator etc. on the fields above just restate the field name/type — could drop these and let the field names speak for themselves.

⚠️ AI-generated; verify before applying. React 👍/👎 to flag quality.

@hudi-bot

Copy link
Copy Markdown
Collaborator

CI report:

Bot commands @hudi-bot supports the following commands:
  • @hudi-bot run azure re-run the last Azure build

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL PR with lines of changes > 1000

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants