TSFM Robustness Benchmark is a systematic testing tool designed to evaluate the engineering robustness of time series foundation models under edge scenarios, such as frequency mismatch, data contamination, and covariate interference. This release includes a systematic evaluation of TimechoAI as the first target model. More models will be integrated in subsequent iterations.
- This project is developed based on Python 3.12+, with core dependencies on
pytest,timecho-ai, andpandas. - The system adopts a clear layered architecture, decoupling business logic, infrastructure utilities, and test execution.
- The middle infrastructure layer natively supports cross-platform execution (Windows / macOS / Linux) and concurrent execution (distributed test scheduling via
pytest-xdist, plus process-level concurrency control). - The test toolkit (
neuraxis_testkit) is packaged as a standalone SDK under thesrc/directory, fully decoupled from business code and ready for reuse across other product lines.
The project follows a standard layered architecture, with the directory structure as follows:
project/
├── config/
│ ├── constants.py # Global business constants
│ └── settings.py # Global environment variable configuration (global path configuration, etc.)
│
├── core/ # Business core common components layer (encapsulates business logic and state management)
│ ├── assertions.py # Business assertions -> depends on neuraxis_testkit.utils.assertions
│ ├── client.py # Low-level client connection (get_timecho_client, etc.)
│ ├── metrics.py # Evaluation metrics calculation
│ ├── models.py # Business data models (request/response)
│ ├── results.py # Business result persistence (CSV/JSON), depends on utils.files
│ ├── resume.py # Policy controller (rate limiting / resume from breakpoint)
│ └── timecho.py # Timecho API client
│
├── env/ # Environment variable configuration directory, loaded by `config/settings.py` via load_dotenv()
│ └── .env.example # Environment variable example file
│
├── src/ # SDK source directory
│ └── neuraxis_testkit/ # Test toolkit
│ ├── log/ # Logging management
│ │ ├── __init__.py # Unified interface exposed externally
│ │ ├── config.py # Variable configuration
│ │ ├── context.py # Context managers (`LogLevelContext`)
│ │ ├── core.py # Core `Logger` class
│ │ ├── decorators.py # Decorators (`log_execution`, `log_time`)
│ │ ├── filters.py # Filters (`ModuleLevelFilter`, `IgnoredLoggerFilter`)
│ │ ├── formatters.py # Formatters (`ColoredFormatter`)
│ │ └── logging.yaml # Logging configuration
│ ├── pytest_infra/ # Pytest infrastructure layer
│ │ ├── __init__.py # Unified interface exposed externally
│ │ ├── collection.py # Dynamic collection (calls manifest_loader)
│ │ ├── fixtures.py # All @pytest.fixture
│ │ ├── hooks.py # All hookimpl (including pytest_addoption / pytest_configure)
│ │ ├── manifest_loader.py # Loads YAML manifests and generates parametrization
│ │ ├── models.py # Test-specific data models (e.g., test case parameters)
│ │ ├── paths.py # NeuraxisPaths + get_paths (pure data + accessors)
│ │ ├── resume.py # Resume-from-breakpoint logic (based on historical results)
│ │ ├── session_manager.py # Test session management (shared resources, locks) SessionFileLock + SessionManager
│ │ ├── test_helpers.py # pytest test-process helpers (glue layer)
│ │ └── test_recorder.py # Test result recorder
│ └── utils/ # Common utilities layer
│ ├── __init__.py # Unified interface exposed externally
│ ├── assertions.py # Generic Assertions Library (Pure Logic)
│ ├── concurrent.py # Concurrency-safe utilities (portalocker wrapper)
│ ├── data_sanitizer.py # Data cleaning and type safety utilities
│ ├── files.py # File operation utilities
│ └── runner.py # Callable execution primitives (process-level timeout + retry)
│
├── tests/ # Time-Series Large Model TestCases
│ └── futureCovs/
│ └── dirtyData/
│ ├── test_dirty.py
│ └── data/ # Test data files (inputs required by test cases)
│ └── test_dirty_s0.csv
│
├── outputs/ # Generated at runtime: logs, results, HTML reports
│ ├── results/ # Business results (CSV/JSON)
│ ├── reports/ # pytest reports (HTML+CSV)
│ └── logs/ # Log files
│ └── tsfm_benchmark_20260824.log # Filename dynamically includes execution date
│
├── conftest.py # Repository-level pytest adaptation entry point
├── pyproject.toml # Project configuration management
├── README.md # Project documentation (English), providing project overview, usage, notes, etc.
└── .python-version
Key file descriptions:
conftest.py: Repository-level pytest adaptation entry point. It declares project-level CLI options such as--project-root,--output-dir,--results-dir,--logs-dir, and sets a custom report header. Common pytest hooks/fixtures are automatically discovered byneuraxis_testkit.pytest_infravia thepytest11entry point, and do not need to be manually bridged inconftest.py.utils/runner.pyonly provides process-level timeout / retry primitives for callables invoked inside a test case. It is not a test entry point; test discovery, execution, and reporting are owned by pytest andpytest_infra/test_recorder.py.pyproject.toml: Project configuration, dependency declarations, pytest configuration, andpytest11plugin entry points.src/neuraxis_testkit/: SDK source directory, which will be split into an independent project later.tests/: Business test cases in the current repository, which will be split into an independent business test repository later.
- Configuration Initialization:
config/settings.pyloads environment variable files (e.g.,.env) from theenv/directory viaload_dotenv(), and exports global paths and runtime configurations. - Model Initialization: Initialize the TimechoAI model using the provided API key.
- Test Execution: Execute the specified test workflow according to the provided command-line arguments.
- Result Output: Output test results to the console or a specified file.
- Python 3.12 or higher
- Virtual environment recommended
python -m venv .venvChoose the corresponding command based on your operating system:
macOS / Linux:
source .venv/bin/activateWindows (CMD):
.venv\Scripts\activate.batWindows (PowerShell):
.venv\Scripts\Activate.ps1Note for Windows PowerShell users: If you encounter a "running scripts is disabled" error, run PowerShell as Administrator and execute:
Set-ExecutionPolicy -ExecutionPolicy RemoteSigned -Scope CurrentUser
deactivateIt is recommended to install the project in editable mode. -e stands for editable, meaning source code changes take effect without reinstallation.
Development/Complete installation (recommended):
python -m pip install -e ".[test]"This command installs:
- Runtime dependencies:
timecho_ai,pandas,requests,pytest,portalocker,python-dotenv,pyyaml - Development dependencies:
pytest-xdist,pytest-html,pytest-cov,pytest-timeout,pytest-mock,pytest-randomly,pre-commit
Install runtime dependencies only:
python -m pip install -e .Windows Platform Note:
The pyproject.toml declares portalocker>=4.3.0, which does not automatically install the portalocker[win32] extension.
If cross-process file locking issues occur on Windows, install it additionally:
python -m pip install "portalocker[win32]"When [project.entry-points.pytest11] in pyproject.toml changes, or when you need to refresh the package metadata without touching other dependencies, run:
python -m pip install -e ".[test]" --force-reinstall --no-depsAfter installation, verify that the pytest plugin is registered successfully:
python -c "import importlib.metadata as m; [print(ep) for ep in m.entry_points(group='pytest11') if 'neuraxis' in ep.name]"Expected output similar to:
EntryPoint(name='neuraxis_testkit_hooks', value='neuraxis_testkit.pytest_infra.hooks', group='pytest11')
EntryPoint(name='neuraxis_testkit_fixtures', value='neuraxis_testkit.pytest_infra.fixtures', group='pytest11')
The project has fully switched to pytest native mode and no longer uses run.py. It is recommended to use python -m pytest consistently to ensure the Python interpreter from the current virtual environment is used.
Run all tests:
python -m pytestRun by module name or keyword:
python -m pytest -k <module_name>Run by file path:
python -m pytest tests/path/to/test_file.pyRun by marker:
python -m pytest -m smoke
python -m pytest -m "not slow"Run concurrently:
python -m pytest -n autoOverride path configurations:
python -m pytest \
--project-root . \
--output-dir ./outputs \
--results-dir ./outputs/results \
--logs-dir ./outputs/logsThe default HTML report path is controlled by addopts in pyproject.toml:
outputs/reports/report.html
The [tool.pytest.ini_options] section in pyproject.toml is configured as follows:
testpaths = ["tests"]- Test file pattern:
test_*.py - Test class pattern:
Test* - Test function pattern:
test_* - Detailed output, HTML report, strict markers, and short traceback enabled by default
- Default failure limit:
--maxfail=5 - CLI logging enabled
- Custom markers registered
- Custom session label and log basename
pytest11 entry points:
[project.entry-points.pytest11]
neuraxis_testkit_hooks = "neuraxis_testkit.pytest_infra.hooks"
neuraxis_testkit_fixtures = "neuraxis_testkit.pytest_infra.fixtures"Therefore, as long as the project is installed via pip install -e . or a regular installation, pytest will automatically load the hooks and fixtures from neuraxis_testkit.pytest_infra.
- Edge scenario detection: Systematically verify the engineering robustness of models against boundary conditions such as complex queries, multi-replica inconsistency, and out-of-order time series writes.
- Defensive architecture validation: Stress-test with strict engineering standards to examine model degradation behavior and recovery capability under non-ideal inputs.
The test results of this framework are limited by specific model versions, data preprocessing strategies, and runtime environments. This tool is intended to provide an objective reference perspective for the engineering defensive architecture design of time series models, rather than an absolute assertion of the final performance of any commercial product.
MIT License