Skip to content

Turn preprints into simulation datasets automatically using LLMs #26

Description

@Essmaw

Context

Currently, manually writing titles and descriptions for simulation datasets is time-consuming and prone to errors. We need a tool that leverages LLMs to extract metadata directly from a preprint PDF to automate the creation of repository-ready records.

Objective

Create a pipeline that takes a preprint as input and outputs:

  1. A Repository-Ready Title & Description: Specifically designed for data repositories (like Zenodo or Materials Cloud), ensuring the dataset is discoverable and well-documented.
  2. A Structured Config File: A formatted file containing the technical keys required for reproduction.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions