Instant answers for what the model knows. Automatic, cited web search for what it doesn't — because it tells you when it's not sure.
❯ ask "explain git rebase vs merge in 3 lines" Git merge joins two histories with a new commit, keeping both branches. Git rebase replays your commits onto another branch for a linear history. Use merge to preserve context; use rebase for a clean log. ❯ ask "latest stable version of Rust" [ask] searching the web (tavily) — model self-assessed: version may have changed Rust 1.98.1 [1] Sources: [1] Announcing Rust 1.98.1 | Rust Blog https://blog.rust-lang.org/...
A quick question should cost a quick command — not a running process.
I wanted lightweight LLM access from the terminal, the place I already live as a developer. But every option made me stop and open something heavy first — a browser tab, a desktop app idling in the background, an Electron client eating hundreds of MB of RAM, or an assistant locked inside the editor. Just to ask "what's the flag for this?" I had to break focus and leave a resource-hungry tool running.
That felt backwards. So ask is built on one rule: spawn, answer, exit. Each call runs for a second or two, prints the answer, and is gone — a single stdlib-only Python file, zero idle memory, no daemon, no background app. You get the answer in your flow, and your machine goes right back to doing nothing.
Most terminal AI tools pick one extreme. ask fills the empty middle.
Fast, but confidently wrong on anything after the model's training cutoff.
Accurate on fresh facts, but slow and heavy on every single query.
The model self-assesses its own answer and only searches when it admits it might be out of date.
A simple, reliable mechanism — model self-reporting, not fragile phrase-matching.
The model answers your question normally.
It appends a hidden line: NEEDS_WEB: yes/no + a reason.
ask reads that line, strips it from view, and searches only if flagged.
Re-asks grounded in results, prints the answer with numbered sources.
Everything you'd want from a terminal companion — and nothing resident when idle.
No daemon, no server. Each call spawns, runs ~1–2s, prints, and exits.
Groq, Gemini, OpenRouter, Cerebras, or a fully local Ollama.
ask -f big.py:50-120 "..." — send only the lines you need.
cat error.log | ask "what broke?"
Keyless DuckDuckGo by default; Tavily, Serper, or Brave with a key.
Readable sessions you can list, resume, and rename.
The package is aolbeam-ask; the command is ask.
pipx (recommended)
pipx install aolbeam-ask
pip
pip install aolbeam-ask
Homebrew
brew install sinhaKAN-ra/tap/ask
One-line script
curl -fsSL https://raw.githubusercontent.com/sinhaKAN-ra/ask/main/install.sh | bash
Run without installing
uvx aolbeam-ask "your question"
Then get a free key at console.groq.com/keys, export GROQ_API_KEY="gsk_...", and run ask "hello". Full docs in the README.