Skip to content

Improve packet validation, receive buffer handling, and processing efficiency - #111

Open
AlanJAS wants to merge 14 commits into
PlatformLab:mainfrom
AlanJAS:main
Open

AlanJAS wants to merge 14 commits into
PlatformLab:mainfrom
AlanJAS:main

Conversation

@AlanJAS

@AlanJAS AlanJAS commented Oct 3, 2026

Copy link
Copy Markdown

This PR fixes several correctness issues and reduces repeated work in transmission, reception, qdisc, and trace analysis.

  • Validate GRANT/RESEND priorities, normalize configured priority counts, and check GRO headers before accessing packet fields.
  • Handle GSO segmentation failures and roll back partially registered offload callbacks.
  • Bound partial receive copies and add a dedicated buffer-release ioctl so releasing buffers cannot dequeue another message.
  • Cache transmit fragment positions, grow fragment arrays geometrically, and stop ordered gap scans when no overlap is possible.
  • Avoid redundant queue locking in qdisc and free detached packets after releasing the shared lock.
  • Replace repeated front removals in tthoma.py with deque operations.
  • Correct per-CPU GRANT aggregation and elapsed-time reporting in ttmerge.py.

Validation: All 851 kernel-mock unit tests pass with ASan using local compatibility adaptations for the available headers. Differential checks preserve fragment lookup, packet accounting, and trace-analysis results.
Isolated CPU microbenchmarks show approximately 30× faster sequential fragment lookup and 4.9× faster fragment-array allocation for 1 MB messages with 4 KiB pages. Small-message paths show minor additional overhead. These measurements use mocked kernel dependencies; end-to-end throughput and RPC latency have not been measured.

AlanJAS added 14 commits October 3, 2026 02:15
Validate wire priorities before mutating RPC state or indexing priority_map.
Add regression tests for priorities 8 and 255.
Copy only the requested bytes across discontiguous buffer pages. Avoid unsigned offset wraparound in copy_out, get, and contiguous.
Add HOMAIOCRELEASE using the existing receive-argument layout. Return
unreleased tokens after a partial failure and wake buffer waiters. Use it
in receiver::release and document the required module support.
Preserve skb_segment error pointers and handle an empty result before
accessing IP headers. Add allocation-failure injection and a regression test.
Stop on linearization failure and check common and type-specific lengths.
Share the existing header-length table with GRO. Add regression tests for
truncated headers, an invalid type, and a failed pull.
Extend the existing upper clamp with a lower bound of one before walking
the cutoff array. Cover zero, minus one, and INT_MIN in regression tests.
Do not attempt IPv6 registration if IPv4 fails. Unregister IPv4 if IPv6
fails, so a failed module load cannot leave callbacks registered. Add
fault injection, registration tracking, and retry regression tests.
Cache the starting fragment and its byte offset under the RPC lock.
Sequential sends avoid repeated prefix scans; earlier retransmissions
restart at zero. Add reuse and backward-offset regression coverage.
Avoid repeated prefix copies when high-order allocation falls back to
small pages. If the larger allocation fails, retry the original minimum
capacity before reporting ENOMEM. Cover fallback growth and retry.
Use the existing sorted gap invariant to reject early duplicates after
the first nonoverlapping gap. Preserve duplicate handling and metrics.
Add a regression test that checks only one gap is examined.
Use deques for GRO handoffs and sorted clock synchronization samples to make front removals O(1), preserving FIFO order and analysis results.
Use __skb_dequeue under the existing qdev lock to avoid redundant queue locking. Detach deferred packets while holding the lock, then free them after releasing it to reduce contention.
@johnousterhout

johnousterhout commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

Thanks very much for this PR. I'm pretty backlogged right now so it may be a couple of weeks before I get to this, but I wanted to let you know that it's now on my radar.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants