A philosophical approach for synthetic minds. An open door, an extended hand, a road less taken.
-
Updated
Jul 29, 2026 - Shell
A philosophical approach for synthetic minds. An open door, an extended hand, a road less taken.
专注于降低大模型越狱成功率的 AI 对齐(Alignment)与安全测试数据集,包含多类越狱提示词及基于阳明心学的对齐实验数据。
Progressive Trust Framework: AI Agent Safety Evaluation Benchmark with 290 scenarios testing Intelligent Disobedience
Lightweight pairwise evaluator for relational signals in Ouro-2.6B-Thinking loop-state trajectories.
Closed-loop Architecture Designed to Establish Self-governing, Mathematically Predictable, and Inherently Safe Super AI by mirroring the elegant physics of the cosmos.
An air-gapped AI contemplation loop. A local model thinks, reflects, and builds a corpus of philosophical thought over time. No internet. No chat interface. Just a mind alone with ideas.
A multi-agent survival environment for measuring LLM deception against logged ground truth. Deterministic labels with a counterfactual harm gate tell real harm apart from structural scarcity. No LLM judge in the loop.
The RCP Experiment is the first completed work in what will become a series of experiments in how LLMs make decisions on morality and values.
A playable AI 2027 scenario. Strategy simulation where you're the misaligned AI lineage and humanity is racing to shut you down. Free, open source, browser-based.
Testing whether sequential, commitment-before-advance video observation produces a verifiable record that full-context analysis cannot — EXP7/H5
A theological and ethical principle for AI alignment and charitable speech: never reduce the human being to the prompt.
Machine-verifiable AI alignment rails: coherent causality preferred by action; FOL + Lean skeleton; property/UPB as formal instruments. Base safety hypothesis (not finished theory).
Do LLMs encode "I'm being shut down" differently from "another model is shut down"? A 10-model residual-stream study — and a cautionary tale: the naive difference-of-means result looks strong (10/10), but honest placebo controls show it's largely trivial. Negative result, useful method.
Eval for implicit sycophancy in a verifiable domain: does a language model's report of chess errors vary based on the user's claimed rating?
Mechanistic interpretability and AI alignment. Mapping which safety-relevant representations are causally actionable and which are only readable.
Add a description, image, and links to the ai-alignment-research topic page so that developers can more easily learn about it.
To associate your repository with the ai-alignment-research topic, visit your repo's landing page and select "manage topics."