Skip to content

fix(calendar): preserve weekly and monthly periods in sparse calendars - #2366

Open
Oleg Zholobov (parlorsky) wants to merge 2 commits into
microsoft:mainfrom
parlorsky:olegzh/fix-sparse-calendar-resampling
Open

Oleg Zholobov (parlorsky) wants to merge 2 commits into
microsoft:mainfrom
parlorsky:olegzh/fix-sparse-calendar-resampling

Conversation

@parlorsky

@parlorsky Oleg Zholobov (parlorsky) commented Oct 7, 2026 •

Copy link
Copy Markdown

Description

Fix weekly and monthly calendar resampling when the source calendar has gaps or is already sampled at those frequencies. Identify boundaries by calendar period, then select the first available date in each period.

For 2week/2month, the count advances only for periods present in the input. An entirely missing week or month does not advance it. For example, with the week of February 12 absent, weekly starts of January 29, February 5, February 19, February 26, and March 4 produce 2week samples on January 29, February 19, and March 4. This preserves the existing count behavior and is now documented and tested.

Motivation and Context

The current implementation detects a new week/month only when the weekday/day-of-month decreases. This silently drops periods when there are gaps: resampling 2024-01-01, 2024-01-02, 2024-01-09, 2024-01-10, 2024-01-17 to weeks returns only January 1 instead of January 1, 9, and 17. Weekly inputs that always fall on Monday and monthly inputs that always fall on the first have the same problem when resampling to 2week/2month.

The regression cases cover sparse daily calendars, weekly/monthly inputs, year boundaries, intraday normalization, frequency counts across entirely missing weeks/months, and empty calendars. Four of the original regression cases fail before the boundary fix.

How Has This Been Tested?

  • The upstream tests/test_all_pipeline.py suite: 3 passed for the production fix (training, backtesting, experiment manager). Used a local runner to point TestAutoData at an isolated data directory; the test implementation was unchanged. The follow-up adds only tests and a docstring clarification.
  • pytest tests/misc/test_resam.py tests/misc/test_utils.py -q: 13 passed after adding the missing-period cases.
  • Black and flake8 on the changed files using the project's settings; git diff --check.

Environment: Python 3.12, macOS arm64, pandas 2.3.3, MLflow 3.12.0. The test runs emit existing deprecation/runtime warnings.

Types of changes

  • Fix bugs
  • Add new feature
  • Update documentation

@parlorsky

Copy link
Copy Markdown
Author

@microsoft-github-policy-service agree

Zholobov Oleg (Zholobov Oleg (@parlorsky)) please read the following Contributor License Agreement(CLA). If you agree with the CLA, please reply with the following information.

@microsoft-github-policy-service agree [company="{your company}"]

Options:

  • (default - no company specified) I have sole ownership of intellectual property rights to my Submissions and I am not making Submissions in the course of work for my employer.
@microsoft-github-policy-service agree
  • (when company given) I am making Submissions in the course of work for my employer (or my employer has intellectual property rights in my Submissions by contract or applicable law). I have permission from my employer to make Submissions and enter into this Agreement on behalf of my employer. By signing below, the defined term “You” includes me and my employer.
@microsoft-github-policy-service agree company="Microsoft"

Contributor License Agreement

@microsoft-github-policy-service agree

@arhancanli

Copy link
Copy Markdown

Read the diff (not run). The period-based boundary is the right fix for the dropped weeks, and the sparse-calendar cases cover it. One behaviour I'd like the author to decide on and pin with a test: [:: freq_sam.count] is still applied to the list of weeks that have at least one trading day, not to calendar weeks.

With a business-day calendar where 2024-02-12..16 are closed (a Spring Festival style gap), I ran the new boundary logic in plain pandas (to_period("W") plus ~duplicated(), then [::2]):

period starts: 2024-01-29, 2024-02-05, 2024-02-19, 2024-02-26, 2024-03-04
2week:         2024-01-29, 2024-02-19, 2024-03-04

So consecutive 2week samples are 3 weeks apart across the gap and 2 weeks apart elsewhere, and the phase of every later sample shifts by one week after each empty week. The old code had the same property, so this isn't a regression, but the PR description says it keeps the "frequency-count behavior", and 2month has the same issue when a whole month is missing from a sparse input.

If counting present periods is intended, a test with an empty week between two populated ones (expected output spelled out) documents it. If a fixed grid is intended, the count would have to be applied to the period ordinal, e.g. (_week.asi8 - _week.asi8[0]) % count == 0 combined with the first-per-period mask.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants