Skip to content

Difficulty reproducing results from raw GAIA dataset #6

Description

@arinagoncharova2005

Good day!

Thank you very much for your excellent work!

I was able to reproduce the results when using the preprocessed data provided in this repository. However, I encountered difficulties when trying to reproduce the full pipeline starting from the raw GAIA dataset to model training.

For raw data preparation, I used the code from the following repository:
https://github.com/tangpan360/TVDiag-data-analysis

Then I ran main.py with reconstruct=True. In this setting, the training completes successfully, but the final prediction metrics are much lower than expected. Here are my results:

2026-03-10 01:01:59 [INFO]: [Root localization] HR@1: 13.845%, HR@2: 27.476%, HR@3: 40.362%, HR@4: 55.272%, HR@5: 66.880%, avg@3: 0.272, MRR@3: 0.250 2026-03-10 01:01:59 [INFO]: [Failure type classification] precision: 45.393%, recall: 46.006%, f1-score: 45.698% 2026-03-10 01:01:59 [INFO]: The average test time is 0.00485848415646944[s] 2026-03-10 01:01:59 [INFO]: The total test time is 4.562116622924805[s]

Has anyone experienced a similar issue when reproducing the pipeline from the raw GAIA dataset?
I would greatly appreciate any suggestions on whether there might be an additional preprocessing step, configuration detail, or dataset version difference (I am using release 1 of the GAIA dataset) that could affect the results.

Thank you in advance for your help!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions