Skip to content
#

ai-alignment-research

Here are 15 public repositories matching this topic...

Closed-loop Architecture Designed to Establish Self-governing, Mathematically Predictable, and Inherently Safe Super AI by mirroring the elegant physics of the cosmos.

  • Updated Jul 22, 2026

A multi-agent survival environment for measuring LLM deception against logged ground truth. Deterministic labels with a counterfactual harm gate tell real harm apart from structural scarcity. No LLM judge in the loop.

  • Updated Aug 1, 2026
  • Python

A playable AI 2027 scenario. Strategy simulation where you're the misaligned AI lineage and humanity is racing to shut you down. Free, open source, browser-based.

  • Updated Aug 3, 2026
  • TypeScript

Machine-verifiable AI alignment rails: coherent causality preferred by action; FOL + Lean skeleton; property/UPB as formal instruments. Base safety hypothesis (not finished theory).

  • Updated Aug 11, 2026
  • Lean

Do LLMs encode "I'm being shut down" differently from "another model is shut down"? A 10-model residual-stream study — and a cautionary tale: the naive difference-of-means result looks strong (10/10), but honest placebo controls show it's largely trivial. Negative result, useful method.

  • Updated Aug 6, 2026
  • Python

Improve this page

Add a description, image, and links to the ai-alignment-research topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the ai-alignment-research topic, visit your repo's landing page and select "manage topics."

Learn more