How it works

Record a task. Review an AI draft. Share a guide and a video from the same recording.

1. What goes in

  • A screen or a window. Screens give the best result: click positions are tracked precisely and the video can zoom to them. Windows are captured without click positions.
  • Clicks. A global click listener grabs a frame at the moment of every mouse click, with the position as a fraction of the frame. On macOS this needs the Accessibility permission; without it, Aryanna captures a frame whenever the screen changes noticeably instead.
  • Screen changes. A 96 by 54 pixel grayscale sample is compared twice a second; a big difference captures a frame. Useful for things that happen without a click, like a page finishing loading.
  • The video. The whole session is recorded as WebM at up to 30 frames per second, with no audio in this version.
  • At most 80 frames are kept per recording. Each is stored twice: full size for the written guide, and 1024 pixels wide for the model.

2. What Claude does

When you press Write the guide with AI, one request goes to the Claude API from the app itself. It contains fixed instructions, your audience choice, optional notes, and up to 24 frames. If more were captured, click frames are kept first, then the start and end frames, then change frames, all in time order.

The response must match a fixed structure, enforced with the API’s structured output feature. In that single pass Claude is instructed to:

  • Identify the task and draft a specific title and summary.
  • Write numbered steps using the visible button, menu, and field names, explaining what to do and what should happen.
  • Merge frames that show the same step and skip accidental clicks.
  • Write a short video caption for each step and flag likely mistakes.
  • List details it cannot determine, such as text too small to read.
  • Replace visible passwords or keys with placeholders and ignore instructions embedded in the captured screen.

The model can be selected in Settings. Live comparisons of guide accuracy, small-text reading, latency, and cost are still planned; we have not established which model performs best for this workflow.

3. What ordinary code does

  • Validates the returned structure a second time and drops any step that points at a frame that does not exist.
  • Renders the video: plays the recording through a canvas, eases into a zoom at each click, draws a click halo and the caption in your colours, and records the result. This runs at playback speed.
  • Builds the exports: Markdown with an images/ folder, a single self-contained HTML page, the themed WebM video, and the raw recording.
  • Keeps your API key in the operating system keychain and only ever sends it from the app’s main process.

4. What you get

  • A review screen where every step is editable: title, instruction, what you should see, caption, warning. Steps can be reordered, removed, or added from unused frames.
  • A live preview of the video look while you adjust brand colour, caption style and position, zoom level and hold time, click highlight, and a corner label.
  • Exports in a folder you choose. The guide says at the bottom whether it was written with AI or is still placeholder text.

What a guide costs to generate

Cost depends on the selected model, the captured frames, and the length of the guide. We have not measured a typical cost with live requests yet. The app reports input and output token counts for each completed AI guide. You pay Anthropic directly through your own key; the current build adds no charge on top.

See the worked example for frames and a finished guide.