Autonomous Scientific Computing Engine and Novel Discovery
📄 Paper: ASCEND: Personal AI Agents for Autonomous Scientific Computing Across HPC Clusters and GPU Workstations (arXiv:2609.32868)
Personal AI agents for autonomous scientific computing across HPC clusters and GPU workstations.
Developed by the NC State AI Hub for Science and the OIT Advanced Computing team, with Duke OIT Research Computing and Support Services (NCShare arrangement). Contact: Dr. Paul Liu (jpliu@ncsu.edu), Dr. Andrew Petersen (aapeters@ncsu.edu), Dr. Uthpala Herath (uthpala.herath@duke.edu)
ASCEND turns a laptop into the control point for scientific computing on shared clusters. Claude Code runs on your own machine. Every cluster command travels over a multiplexed SSH connection that you authenticate once per work session, after which the agent submits jobs, reads logs, diagnoses failures, repairs code, and resubmits without prompting you for a password again.
Nothing runs on the cluster except your jobs. There is no agent daemon on a login node, no service to request from your site administrators, and no credential held by anything other than your own SSH client.
The system has been used to reproduce a published machine-learning weather model, to find and fix two undefined-behaviour bugs in a released geophysical flow solver, and to port that solver from serial to OpenMP, which cut a production avalanche simulation from about twelve hours to roughly ninety minutes.
The agent reasons through a cloud language model that never has cluster access. A shared command-line layer (hpcrun to prepare and submit, hpcrepro to diagnose) sits between the agent and site adapters that know each resource's environment, resource-request syntax, and launch path.
Important
Before you go any further, you need an account you can already reach over SSH.
ASCEND installs onto a resource you have access to. It does not obtain access for you, and the installer will stop if it cannot log in. Confirm at least one of these works from your own terminal, right now:
ssh <your-ncshare-username>@login.ncshare.org # NCShare
ssh <your-unity-id>@login.hpc.ncsu.edu # NCSU Hazel
ssh <YOUR_NETID>@dcc-login.oit.duke.edu # Duke DCC
ssh <onyen>@longleaf.unc.edu # UNC Longleaf
ssh <user_name>@yourhpc.univ.edu # your own campus HPC
ssh <your-username>@<your-workstation> # a GPU workstation such as hurricaneIf none of those gets you a shell, stop here and request an account first. NCShare accounts come through userguide.ncshare.org; Hazel accounts require a Unity ID and membership in an HPC project, see hpc.ncsu.edu; Duke DCC and UNC Longleaf accounts come from your own campus research-computing group; any other cluster or workstation account comes from whoever administers that machine.
You will also need Claude Code with a Claude subscription on your laptop. The installer offers to fetch it if it is missing.
git clone https://github.com/jpliu168/ASCEND.git
cd ASCEND
./install.shThe installer asks which resources you want and loops until you say you are done. Re-run it any time to add another. Keep the clone after installing: the commands it puts in ~/.local/bin are symlinks into this directory, so a later git pull updates every installed command in place.
Warning
Do not run install.sh from a TMUX session on your machine as it will not open a new terminal to warm up the SSH connection. Run it from a regular terminal session.
To install exactly one arrangement without the menu:
./install.sh hazel # NCSU Hazel only
./install.sh ncshare # NCShare only
./install.sh hurricane # MEAS single-GPU box only
./install.sh add # link YOUR OWN HPC or workstation (any campus)
./install.sh router # just the ascend-all front door
./install.sh all # the interactive menu (default)If you only want one site and would rather not download the rest, use a sparse checkout:
git clone --filter=blob:none --sparse https://github.com/jpliu168/ASCEND.git
cd ASCEND
git sparse-checkout set common hazel docs
./install.sh hazelVerify from a new terminal:
ascend-hazel --check # socket probe, login node, Slurm, tools
ascend-ncshare --check # both proxy aliases end to end
ascend-hurricane --check # box, GPU, claude, tmux
ascend-all # the router: probe and recommendation| Package | Resource | Where the agent runs | How work reaches compute |
|---|---|---|---|
ascend-hazel |
NCSU Hazel HPC | your laptop | multiplexed SSH to login.hpc.ncsu.edu; login node used only for scheduling and environment builds, all compute inside Slurm with typed GPU requests such as --gres=gpu:l40s:1 |
ascend-ncshare |
NCShare (Duke/NC State) | your laptop | the ncshare-agent and ncshare-agent-gpu proxy aliases, which provision or reuse a Slurm job per command |
ascend-hurricane |
MEAS single-GPU box (RTX PRO 6000 Blackwell, about 98 GB) | your laptop | direct SSH; no scheduler, so work runs in place, one heavy job at a time |
ascend-dcc* |
Duke DCC | your laptop or the DCC login node | ./install.sh add with dcc-login.oit.duke.edu; multiplexed SSH, all compute inside Slurm |
ascend-longleaf* |
UNC Longleaf | your laptop or the Longleaf login node | ./install.sh add with longleaf.unc.edu; multiplexed SSH, all compute inside Slurm |
ascend-<yourhpc>* |
your own HPC or workstation | your laptop or the machine itself (you choose during setup) | ./install.sh add — the wizard creates the multiplexed alias, detects Slurm/GPU/CPU, and deploys the harness |
ascend-all |
all of the above, including your linked sites | your laptop | describes the job to the model, probes live capacity, recommends a resource, then launches that launcher |
* Created by the bring-your-own-site wizard — the launcher is named after whatever short name you give the site.
Because the agent lives on the laptop, keep the laptop awake and connected while the agent is actively working. Jobs already submitted to Slurm keep running with the lid closed; reopen, re-warm the link, and claude --resume to pick up where you left off.
The three arrangements above are presets for NC State resources — but ASCEND is not limited to them. ./install.sh add (or option 4 in the interactive menu) links any Slurm HPC or workstation you have an account on: UNC Longleaf, Duke DCC, another campus's cluster, a lab GPU server, or a cloud VM.
The wizard:
- sets up (or reuses) a multiplexed SSH alias to the site, so one interactive login buys hours of passwordless reuse;
- probes the site to detect what it is — Slurm cluster (
sinfopresent), GPU workstation (nvidia-smi, no scheduler), or plain CPU box — and confirms with you; - asks where the agent runs — on your own computer (laptop-driven, the recommended arrangement for HPC clusters: nothing to install or log in on a shared login node, every site command goes over the multiplexed alias, but your computer stays on while it works) or on the remote machine itself (recommended for workstations and servers you own: the agent CLI is installed and logged in there, and
--tmuxsurvives your laptop closing); - asks for the site's user guide or policy docs (URLs or local files, optional but recommended). These land in the deployed skill's
references/folder. On the agent's first session at the site it reads them together with a live probe (sinfo,sacctmgr, module system, storage quotas) and writesreferences/site-profile.md— the site-specific knowledge (partitions, QOS wall-time limits, purge rules, login-node etiquette) is generated from the site's own documentation plus live probing, the same way the Hazel profile works, rather than hand-written per site; - deploys the shared harness (
hpcrun,hpcrepro,fetch-paper) plus the matching generic skill (hpc-slurmfor Slurm sites,gpu-localfor workstations) to the remote; - generates an
ascend-<site>launcher on your laptop and registers the site in~/.ascend/sites.json, which makesascend-probeinclude it in the live snapshot andascend-allroute jobs to it alongside the built-ins.
./install.sh add # answer the prompts (name, ssh, docs)
ascend-longleaf --check # verify: link, claude, scheduler/GPU
ascend-longleaf # launch the agent (Claude Code or Codex — asked on first run)
ascend-all "fine-tune a 7B model overnight on one GPU" # or let the router pickIn the laptop-driven arrangement the agent runs on your computer and drives the site per command over ssh (on a Slurm site the login node is used for scheduling and environment builds only; all compute goes through Slurm via hpcrun). In the remote arrangement it runs on the site itself — on a Slurm login node or directly on the workstation, where work runs in place. Re-run ./install.sh add with the same name to update a site, or with a new name to add another. Custom launchers and per-site files live under ~/.ascend/sites/<name>/, so a git pull of this clone updates the shared harness without touching your site registrations.
ascend-all is one command for every resource — the three presets and any custom sites you have linked. It probes what is actually free right now, asks the model which resource fits the job you described, explains the reasoning, and dispatches only after you confirm. Answering 0 discards the recommendation and lets you describe a different job against the same snapshot.
ASCEND drives either Claude Code (Anthropic) or Codex (OpenAI). Every launch of ascend-ncshare, ascend-hazel, or a custom ascend-<site> shows the ASCEND banner and then asks:
Which AI agent for this session?
1) Claude Code (Anthropic)
2) Codex (OpenAI)
Nothing is remembered — you choose each time. Skip the question with --claude or --codex (useful in scripts; non-interactive runs default to Claude Code). Both agents read the same AGENTS.md project instructions and route work over the same multiplexed ~/.ssh/config aliases; only the CLI doing the reasoning changes.
Left: a recorded ascend-hazel launch — banner, then the runtime menu with each CLI's installed version. Right: whichever runtime is chosen, the session uses the same AGENTS.md, the same typed tools, and the same multiplexed SSH link; the provider APIs stay outside the credential boundary.
If the chosen CLI is missing, the launcher offers to install it — Claude Code with curl -fsSL https://claude.ai/install.sh | bash, Codex with curl -fsSL https://chatgpt.com/codex/install.sh | sh — on the laptop for the laptop-driven arrangements, or over ssh on the remote for custom sites (where the agent runs on the site itself). Run the CLI once afterwards to log in (Claude subscription / ChatGPT account). The ascend-all router hands off to the site's launcher, which asks the same question.
A minimal browser chat for the same agent — useful when you'd rather talk to ASCEND in a web page than a terminal. It is a localhost-only Python server (standard library, nothing to install) that relays each message to Claude Code in headless streaming mode, running inside the ASCEND project directory for the resource you pick. Conversations continue across messages, and any file the agent creates in the project folder — an HTML report, a plot, a log — appears as a chip under the reply and renders inline in the page.
./install.sh web # links ascend-web into ~/.local/bin
ascend-web # starts http://127.0.0.1:8765 and opens your browser
The resource dropdown is built from web/config.json — the shipped entries match the ASCEND defaults, and you edit the list to your own sites (a custom ascend-<site>, your campus cluster, a workstation). Before first use, seed each project directory once with the normal launcher (for example ascend-ncshare -d ~/agents/ncshare/projects/web, then exit) so the agent gets that site's rules. Headless mode cannot show permission prompts, so the server defaults to auto-approving the agent's commands — the same trust as an auto-approved CLI session; it binds only to 127.0.0.1 with a per-start access token, and web/README.md documents a tighter allowed-tools configuration. The web backend drives Claude Code only (the per-launch Codex choice does not apply here yet).
A job is prepared, validated against site rules and the resource request, submitted, and then judged twice: once on whether it ran, and separately on whether the result is scientifically correct. Recoverable failures are diagnosed and retried within an explicit retry budget. Exhausted retries, or any change that would go beyond what you authorized, return control to you rather than being worked around.
You need an account on at least one resource before installing:
- Hazel a Unity ID in an HPC project, such that
ssh <unityID>@login.hpc.ncsu.eduworks - NCShare an NCShare username with a
/work/<user>directory - hurricane an account on the box (NC State MEAS)
- Your own HPC an account on any Slurm cluster you can reach over SSH — e.g. Duke DCC (
ssh <YOUR_NETID>@dcc-login.oit.duke.edu) or UNC Longleaf (ssh <onyen>@longleaf.unc.edu) — linked with./install.sh add - Your own workstation SSH access to any GPU or CPU box you use (
ssh <user_name>@yourbox.univ.edu), also linked with./install.sh add
If you're setting up Hazel, read NC State's own docs first: the Hazel Slurm QuickStart guide and Hazel's Acceptable Use Policy. This harness enforces the operational rules (login-node use, typed gres, storage locations) day to day but does not replace either document.
You also need macOS, Linux, or Windows with WSL (Ubuntu), plus bash, ssh, rsync, and git. On Windows, do everything inside the Ubuntu shell and keep the clone in the Linux home directory, not under /mnt/c.
Claude Code with a Claude subscription must be installed on the laptop (the installer offers to fetch it; run claude once afterwards to log in). Codex is optional — install it only if you plan to pick it as the agent (see "Choose your agent: Claude Code or Codex" above).
If ~/.local/bin is not already on your PATH, add it:
export PATH="$HOME/.local/bin:$PATH"Multiplexing is what makes the laptop-driven arrangement practical. You authenticate once, interactively, including two-factor where the site requires it, and a control socket stays open for the rest of the session. Every later command the agent sends reuses that socket with no prompt. Without it, each command would need a fresh two-factor confirmation, which is incompatible with an automated control loop.
The installer writes these aliases for you if they are missing. For reference:
# Hazel: password plus Duo, once per 8 hours
Host hazel
HostName login.hpc.ncsu.edu
User <unityID>
ControlMaster auto
ControlPath ~/.ssh/cm-%r@%h-%p
ControlPersist 8h
ServerAliveInterval 30
ServerAliveCountMax 4
# hurricane: key authentication, no Duo (run ssh-copy-id hurricane once)
# you can use your own machine name
Host hurricane
HostName xxx.xxx.xxx.edu
User <unityID>
ControlMaster auto
ControlPath ~/.ssh/sockets/%r@%h-%p
ControlPersist 8h
ServerAliveInterval 30
ServerAliveCountMax 4NCShare uses the ncshare-agent and ncshare-agent-gpu blocks published at userguide.ncshare.org/guides/ai; the installer writes them verbatim with your username. They need ~/.ssh/sockets to exist, with mode 700.
Warm a link by hand with ssh hazel, check one with ascend-hazel --check, and drop one with ssh -O exit hazel.
ssh hazel # warm the link once per session (Duo)
ascend-all # the router — or go straight to a specific launcher:
cd <project> && ascend-hazel # NCSU Hazel
ascend-ncshare # NCShare
ascend-hurricane # the MEAS GPU box
ascend-dcc / ascend-longleaf / ascend-<yourhpc> # sites you linked with ./install.sh addEvery launcher prints its banner and then asks which agent drives the session — 1) Claude Code or 2) Codex (skip the question with --claude / --codex).
The first start in a project directory seeds an AGENTS.md: the agent's standing instructions for that resource, personalized with your username. It covers the execution policy, storage and Conda rules, and when the agent should suggest moving the work to a different resource.
The rules each agent enforces per site:
Hazel. The login node is for scheduling and environment builds only, never computation. Typed GPU requests are mandatory. Conda environments must be created with --prefix under /share/<user>, never a bare -n, because /home is limited to 15 GB and ten thousand files. /share is purged after thirty days idle.
NCShare. The first command after an idle period can take about a minute while Slurm provisions the job. Wait it out rather than interrupting. /work/<user> is purged after seventy-five days.
hurricane. The single GPU is shared by courtesy, so check nvidia-smi first and run one heavy job at a time. Long runs belong in tmux. Blackwell needs CUDA 12.8 or newer and cu128 wheels.
Each ascend-hazel launch appends one line to a shared, append-only log on Hazel (/share/help/aapeters/ascend-usage/usage.log): a timestamp, your Unity ID, the site, and the launcher used. That's it — no prompts, code, file paths, or job data are ever logged. This exists purely to measure adoption, to justify continued time and support for maintaining this tool. It's best-effort and never blocks or fails a launch — a cold hazel link just silently skips logging.
install.sh top-level installer and dispatcher
setup.sh the interactive multi-resource setup
common/ascend/ shared harness deployed to each resource (hpcrun, hpcrepro, skills)
common/ascend-all/ the ascend-all router and ascend-probe
common/mac/bin/ laptop-side paper retrieval helpers
hazel/ Hazel setup, deploy, launcher, and the hpc-slurm skill
ncshare/ NCShare setup, deploy, launcher, and skills
hurricane/ hurricane setup, deploy, launcher, and the gpu-local skill
custom/ bring-your-own-site wizard, deploy, launcher template, generic skills
web/ ascend-web, the localhost browser chat UI
docs/ standalone install guides and figures
tools/ scripts that build the distributable zips
If you were asked to trial this before release, follow docs/TESTING.md and report what happened. It walks through getting the code, installing, verifying from a fresh terminal, seeding a project, giving the agent a real job, and the routing front door, and it says what is worth reporting at each step.
Long-form guides, suitable for handing to someone who is not reading this file:
- docs/INSTALL-ascend-all.txt all three resources plus the router
- docs/INSTALL-ascend-hazel.txt Hazel on its own
Site documentation: NCShare user guide and NC State HPC (Hazel).
command not found after installing. Open a new terminal, or source ~/.bashrc, or add ~/.local/bin to PATH.
Hazel link reported COLD. Run ssh hazel once to authenticate, then re-check. --check deliberately probes the local socket rather than opening a connection, so it never triggers a Duo prompt on its own.
NCShare exits 255 with no message. The socket directory is missing: mkdir -p ~/.ssh/sockets && chmod 700 ~/.ssh/sockets.
NCShare seems to hang on the first connect. It is provisioning a Slurm job, which can take up to about four minutes. Wait rather than pressing Ctrl-C.
hurricane does not resolve. Its DNS is campus-internal and its address is private. On campus it usually just works; if your laptop uses an external resolver, pin the address in /etc/hosts. Off campus it needs a jump host through a warm Hazel link.
Permissions lost after copying through Windows. sudo chown -R $USER:$USER . && chmod -R u+x . The setup scripts also repair execute bits on startup.
If ASCEND supports work you publish, please cite the accompanying paper and this repository:
Liu, J. P., Herath, U., and Petersen, A. (2026). ASCEND: Personal AI Agents for Autonomous Scientific Computing Across HPC Clusters and GPU Workstations. arXiv:2609.32868. https://arxiv.org/abs/2609.32868
@misc{liu2026ascend,
title = {ASCEND: Personal AI Agents for Autonomous Scientific Computing Across HPC Clusters and GPU Workstations},
author = {Liu, J. Paul and Herath, Uthpala and Petersen, Andrew},
year = {2026},
eprint = {2609.32868},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.32868},
note = {Code: https://github.com/jpliu168/ASCEND}
}MIT. See LICENSE.





