Skip to content

Repository files navigation

AI Image Generation

AI IMAGE GEN — retro travel-poster banner, itself generated by this repo

CI License: MIT

Local image generation on Apple Silicon (M5 Max, 64 GB), two ways:

  • MFLUX — native MLX, command-line, fastest. Z-Image, FLUX.2 [klein], Qwen-Image, FLUX.1.
  • ComfyUI — node UI for workflows, LoRAs, FLUX.2 [dev] 32B, and the pixel-art pipeline.

Everything is isolated (no system-wide Python) and every tool shares one model store (~/ai-models/) so weights download once. Runs natively on Metal/MPS.

Quick start

./generate.sh "a red panda wearing a tiny top hat, watercolor style"   # MFLUX CLI, FLUX.2 klein
MFLUX_MODEL=z-image-turbo ./generate.sh "an astronaut above Earth, golden hour"   # fastest
./comfyui-start.sh                                                       # ComfyUI @ :8188

Showcase

Everything below came out of this repo — offline, on one laptop, no cloud and no API keys.

Locally generated posters and style studies

Top row + bottom left — one Kyoto travel-poster prompt through FLUX.2 [dev] 32B in ComfyUI; clean, correctly spelled display type is what the big model buys you. Bottom middle + right — one fox prompt swept across styles (watercolor, photo, anime, oil, 3D, ukiyo-e), the six on the right from FLUX.2 [klein] 9B at 6 steps, ~12 s apiece via generate.sh.

One prompt, three styles, one frame

Same prompt, same seed, only the style clause changes — FLUX.2 [klein] then keeps the composition, so the three frames line up and can be masked into each other. Watercolor on the left, the photographic pass through the middle, anime cel on the right:

A fox rendered as watercolor, photo and anime, cross-faded into a single image

for st in "delicate watercolor painting, soft washes, visible paper texture" \
          "cinematic photograph, photorealistic, golden hour, shallow depth of field" \
          "anime cel illustration, clean line art, flat vibrant colors"; do
  MFLUX_MODEL=flux2 MFLUX_SEED=7 MFLUX_STEPS=6 MFLUX_W=1536 MFLUX_H=640 \
    ./generate.sh "a red fox sitting in a forest clearing, tall trees behind, ferns and moss on the ground, centered composition, $st"
done

.venv/bin/python blend-styles.py blend.jpg <the three PNGs> --bounds 0,420,1100,1536 --feather 240

Three renders at ~12 s each, then one composite. --bounds puts the seams left and right of the fox, so the subject itself stays in a single style.

The banner at the top of this README is not stock art either — it is one line of the same CLI, on Qwen-Image-2512, which is the model here that reliably spells display type — FLUX.2 [klein] rendered "AI IMAGEGEGEN" on every seed tried:

MFLUX_MODEL=qwen-2512 MFLUX_W=1536 MFLUX_H=640 MFLUX_SEED=55 ./generate.sh \
  "A flat mid-century travel-poster banner, wide format. Large cream capital letters across the upper third spell exactly the three words: AI IMAGE GEN. No other text anywhere in the image. Behind and below the type, a stylized snow-capped mountain range at sunrise with layered ridges, a large sun disc low on the horizon and radiating sunburst rays. Teal, deep navy and burnt-orange palette, subtle paper grain, crisp vector shapes, high contrast, clean accurate typography, correct spelling."

~2½ min at 20 steps on the M5 Max; the only post-processing was cropping the paper margin off.

Desktop app

gui/ is a small egui/eframe front-end over the very same generate.sh — pick a model, type a prompt, watch the render log stream, see the result. It duplicates no model logic: it sets MFLUX_MODEL / MFLUX_SIZE / MFLUX_STEPS / MFLUX_SEED / MFLUX_MODE / MFLUX_STRENGTH and runs the script.

Drag an image onto the window (or Add…) to use it as a reference: with FLUX.2 [klein] the prompt then becomes an instruction — "put a knitted red scarf on the fox" — and several references can be combined; every other model takes the reference as an img2img starting point instead. Details: reference images.

The window is a split view: model, prompt, options and the gallery of past renders on the left, the picture itself filling the whole right-hand side. ⌘⏎ starts a render, ⌘[ / ⌘] walk the gallery, ⌘B hides the sidebar, and right-clicking any render reveals it in Finder, opens it, copies its path or feeds it back in as a reference — menus and shortcuts.

The desktop app: sidebar with model picker, a reference image in edit mode, prompt box and the gallery of past renders; the finished image fills the right-hand pane above its caption and the live render log

cd gui
cargo run --release      # or ./make-app.sh  ->  dist/AI ImageGen.app

Packaging, the .app bundle and the icon pipeline: gui/README.md.

Documentation

Sorted into setup (provision once) and usage & tips (day-to-day):

docs/setup/

  • installation.md — isolated toolchain (uv, Python 3.12, venvs, ComfyUI, HF login)
  • models.md — downloading models + the shared store (opencode reuse, FLUX.2 dev)
  • models-inventory.md — archived pre-clear snapshot

docs/usage/

Scripts

Script Purpose
generate.sh MFLUX multi-model CLI generation
comfyui-start.sh launch ComfyUI wired to the shared store
download-models.sh fetch/warm models into the shared store
blend-styles.py cross-fade same-seed renders into one image (see Showcase)
env.sh shared env (HF_HOME, store paths) — sourced by the others
throttle.sh optional macOS inbound-bandwidth cap (see gotchas)

Layout

.
├── README.md · env.sh · generate.sh · comfyui-start.sh · download-models.sh
│   throttle.sh · blend-styles.py
├── docs/{setup,usage}/           # documentation
├── docs/assets/                  # README imagery
├── gui/                          # desktop front-end (Rust/egui) + .app bundler
├── .venv/                        # MFLUX venv            (git-ignored)
├── comfyui/                      # ComfyUI clone + venv  (git-ignored)
└── generated/                    # output images         (git-ignored)

~/ai-models/                      # shared model store    (outside the repo)

License

MIT for the code in this repo. It ships no model weights — each model carries its own license, summarised in docs/usage/models-guide.md.

About

Local image generation on Apple Silicon: MFLUX + ComfyUI over one shared model store, with a CLI and a small desktop app.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages