Skip to content
View aashiqmuhamed's full-sized avatar

Block or report aashiqmuhamed

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Pinned Loading

  1. DynamicSAEGuardrails DynamicSAEGuardrails Public

    Code for the paper "SAEs Can Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs"

    Python 8

  2. neulab/gemini-benchmark neulab/gemini-benchmark Public

    Jupyter Notebook 151 21

  3. transformer-gan transformer-gan Public

    Forked from amazon-science/transformer-gan

    Python

  4. poison-set-selection poison-set-selection Public

    Code for Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks.

    Python 2

  5. defending-against-abliteration defending-against-abliteration Public

    Post-hoc, fine-tuning-free defenses against LLM abliteration / refusal feature ablation.

    Python 1

  6. mole mole Public

    Benchmark for detecting insider threats by AI agents in a simulated frontier AI lab

    Python 1