# Idea Engine — Simple Quality Upgrade, V2

Updated September 7, 2026

This consolidated revision replaces `IDEA_ENGINE_SIMPLE_UPGRADE_BRIEF.md`. It retains the original quality, evidence and bench improvements while incorporating the creator's follow-up preferences: curiosity-led “Can I vibe code this?” builds, familiar-software recreations, impressive AI demonstrations, useful apps/automations, and clear production-format labels. The examples remain proposals; the user approved these directions, not every individual idea.

## Scope and decision

This is an implementation proposal based on the supplied screenshots and `SYSTEM_HANDOFF_2026-09-07.md`, including its reconstructed writer and judge prompts. It is not a source-code audit, an implemented patch, or a prediction of viral performance. Verify the named functions and contracts against the repository before editing.

Keep the local Node application, existing providers, caches, five screens, 20-candidate generation, and 10-card board with a scored bench. Do not add agents, a vector database, fine-tuning, an analytics integration, a new scraping provider, or a large onboarding flow.

The central change is:

**Audience problem, desire or curiosity → worthwhile concept → suitable delivery format → visible outcome and payoff → hook → editorial score.**

Do not continue optimizing:

**Outlier title → noun substitution → flattering explanation → inherited source score.**

## What the handoff establishes

- The writer's requested content fields are effectively title, hook and whyNow, with classification and borrowed-source metadata. A substantive concept, visible demonstration and viewer takeaway are not required. See handoff lines 1124–1132.
- The judge nevertheless scores the usefulness of a proposed video. It can therefore credit an execution that was never specified. See lines 1144–1155.
- Proof is calculated from citation performance. A borrowed hook can supply the strongest-source score without establishing demand for the new subject. Originals are required to cite nothing, then receive a separate topic-match score with a deduction. See lines 198–244 and 1413–1423.
- The displayed ideas contain unsupported autobiographical claims: three paid ad jobs and invoices, ninety days of inbox-agent logs, client approvals, and recurring personal DMs. These are generated premises, not established facts about the creator. See lines 1246–1249, 1307–1310, 1335–1338 and 1354–1357.
- The reconstructed run's comment entries E48–E59 all come from one free-token video; they are not twelve independent demand signals. The two breakout entries E31–E32 also come from one creator. See lines 956–971 and 1017–1028.
- Breakdowns see captions, titles, metrics and some comments, not the actual video or thumbnail. Some saved explanations nevertheless assert traffic sources, complete viewing, conversions or causal effects that those inputs do not establish. See lines 190–195, 647–658, 716–723 and 1038–1064.
- The best three scoring anchors in this run are TikTok posts, while the weakest three are YouTube posts. Treating these directly as universal high/low hook scores introduces a platform/format confound. See lines 1173–1182.

These are reasons to adjust the existing system, not rebuild it.

## 1. Ground the creator profile

Use the existing profile notes and DNA machinery. The handoff already says user notes outrank inferred posting history (lines 318–323). Preserve that behavior.

Add a compact, editable section to the profile/DNA:

- **Audience outcomes:** Who is this for, and what do they want to accomplish? Select one primary segment for each idea rather than addressing all segments in every post.
- **Confirmed work and evidence:** Separate published work, user-confirmed experience, and assets actually available for a new recording. A published bot demonstration does not establish profits, ninety days of logs, or production readiness.
- **Production constraints:** What can the creator realistically record or build? Leave unknown budgets, available time, participants and assets unknown.
- **Editorial examples:** A small set of user-approved strong/weak examples and the reason for each preference. Until the user approves a proposed example, label it provisional.

For this profile, a suggested umbrella promise is:

> I show non-technical creators and small-brand founders what they can actually build with AI—from their own apps and automations to professional-looking creative—and help them understand how to make it real.

Keep the five current content lanes. Strengthen the existing **Let's Vibe Code It** lane with a recurring **Can I vibe code this?** series rather than introducing a new taxonomy or separate generator. The current handoff already describes this lane as a real build with an outcome that is not guaranteed (lines 693–703).

For this creator, explicitly welcome these overlapping editorial directions:

- **Can I vibe code this?** Attempt an original, recognizably useful version of a familiar app's core workflow, or a viewer's own desired app. A booking tool, creator workspace or interactive learning game can provide an understandable starting point.
- **Look what AI can do.** Show an impressive, relevant transformation or interactive experience with a genuine reveal. A satisfying demonstration can be the main payoff; a full tutorial is not mandatory.
- **I wish this existed.** Build software, apps or automations that address a recognizable need for the intended audience. The result need not imitate a known product.

These are user-approved directions, not verified market trends or a fixed allocation of board seats. Consider them alongside the existing AI-ads, content-systems and beginner-education lanes. Do not allow the whole candidate pool to collapse into setup advice, security audits or troubleshooting merely because those are easy to describe as useful. Do not force every idea to become a challenge, negative test or comparison either.

A familiar product supplies a reference point, not an automatic reason to watch. Name the core workflow to recreate, the intended user and the concrete reveal. Build an original, clearly scoped prototype using appropriate code, design and sample assets; do not imply a full commercial replacement, affiliation, proven reliability or unmeasured savings. Keep the editorial promise exciting while stating the actual scope honestly.

The recurring story can be **recognizable goal → attempt → key decisions or surprises → working demonstration or honest result → what this makes possible**. A smooth successful build is allowed; do not fabricate difficulty or a failure for drama. Production verification does not have to become the headline.

Store these preferences in this user's editable DNA/notes and examples, not in the universal engine prompt. Another user's lanes and value preferences must remain their own. User-stated direction can expand beyond historical posting patterns; old posts should not permanently restrict what the creator is allowed to try.

Do not force every hook to begin with a tool name or imitate the exact wording of historical winners. Keep the user's voice, not a fixed sentence template.

Treat demographic guesses and explanations such as “posting a Short the same day cannibalizes the long video” as hypotheses unless the user or appropriate analytics confirms them. Preserve genuine user preferences separately from inferred causal explanations.

## 2. Improve the existing evidence bundle

### Selection, not more scraping

Use the existing candidates and caches first. Apply diversity before the top-N truncation, so near-identical website tips do not consume the candidate pool before alternatives are considered.

Start with a maximum of three post sources per creator in the writer bundle. Deduplicate the same underlying item across platforms/subreddits where identifiers or explicit cross-post references make that possible. For uncertain matches, do not invent independence or assert they are the same item.

Treat a post, its comments and AI summaries of those comments as one source family. Distinct commenters can indicate a repeated question, but do not turn a summary plus its supporting comment into two independent observations.

Retain useful representation across the creator's active lanes where relevant evidence exists. Do not manufacture coverage or fill an arbitrary lane quota with weak material. A result-led showcase may supply packaging or audience-interest evidence even when it does not teach a method; absence of a tutorial should not automatically make a relevant result off-niche.

The current relevance prompt explicitly excludes “showcase reels with no method” (handoff lines 602–607). Replace that blanket exclusion: keep a demonstrably in-lane build, useful outcome or impressive creative result when its relevance is supported by the supplied title/caption/metadata. Continue excluding unrelated spectacle, unsupported promises and generic AI-adjacent material. Do not infer unseen functionality from a screenshot. Invalidate the affected cached relevance judgments with a prompt version, or rebuild the DNA using the existing invalidation behavior, so this change actually reaches candidate selection.

For the first improved run, use the existing comment-mining feature on four relevant posts: two of the creator's own and two external posts, chosen to cover missing audience problems rather than the same topic again. This is a bounded data refresh, not a new always-on scraping job. Keep existing explicit charge behavior and caching.

Replace a few tool-only search phrases with audience-problem, desired-outcome and build-curiosity phrases inside the existing query budget. Examples to test, not verified demand claims:

- building my own booking app with AI;
- recreating an app with AI;
- useful apps to vibe code;
- turning a product photo into a campaign;
- keeping products consistent in AI images.

Retain interesting, in-lane builds even when their value is curiosity or aspiration rather than a how-to. A “could it also do X?” or “I wish I had Y” comment can suggest an opportunity, just as a troubleshooting question can. Do not turn a missing tutorial into evidence of a demand gap without an actual relevant question or another identified need. Do not automatically assign a known app's popularity to demand for this proposed recreation.

Do not add all of these on top of the existing query count. Rotate or replace weaker phrases.

### Evidence relationships

Every cited source must state what it actually supports:

- `demand`: the same audience problem, desired outcome or explicit question;
- `format`: a presentation device, framing or storytelling mechanism;
- `own-work`: a documented creator project or result that supports fit or a genuine follow-up.

An idea can have an original concept and cite demand evidence. Original does not mean source-free.

Do not give a trading-bot idea topic-demand credit solely because a website-compliance hook performed well. It can still borrow the anxiety-to-demonstration mechanism.

Keep real-ID validation. Replace mandatory borrowed-prefix lexical overlap with a valid relationship and a short support statement checked by the existing judge. A creative synthesis should not need to echo the source's wording to retain a valid citation.

A comment is evidence that someone expressed a concern, not that the concern is factually correct. A limited search is not evidence that “nobody has covered this.” Two comments from one post do not establish a cross-creator trend. Use “a question in the sampled comments” or “repeated recent coverage” when that is what the data supports.

### Metrics and breakdowns

Keep views/followers as a discovery heuristic when it is all that is available. Do not label it proof of organic virality, non-follower reach, unique audience size, or the quality of a different proposed idea.

Prefer performance against the creator's own comparable-post median where already cached. Keep platform, format, post age, baseline kind and available sample size visible. Do not scrape every discovered creator merely to obtain a baseline. Tiny or uncertain denominators should lower confidence, not automatically dominate selection.

Rewrite the existing breakdown prompt to separate observations from possible explanations. Do not infer retention, traffic sources, sales, conversions, watch completion or paid distribution from a title/caption and public counts. Do not treat unavailable engagement counts as observed zeros.

Use a simple breakdown-prompt version to refresh affected cached analyses when they are next needed. Do not preserve overconfident old explanations forever; do not bulk re-scrape content to regenerate text.

## 3. Require actual concepts in the writer response

Keep 20 candidates and one writer call. Keep existing IDs, source links and type metadata where useful. Add compact fields rather than generating full scripts:

| Field | Required content |
| --- | --- |
| audience | One primary viewer segment and its relevant goal. |
| contentFormat | Exactly one: `long_form_video`, `short_form_video`, `carousel_slideshow`. This is the recommended primary deliverable. |
| platforms | Intended publishing destinations, chosen from the user's supported/desired platforms. Do not use “YouTube” as a substitute for a format. |
| suggestedLength | An editorial target such as `12–18 minutes`, `45–60 seconds`, or `7 slides`; not a platform maximum or a build-time claim. |
| concept | Two or three sentences: what happens, including the opening visual or cover, the build/transformation/story, and how the outcome is delivered in the selected format. |
| viewerPayoff | The delivered value: a usable method, artifact, better decision, answered curiosity, demonstrated possibility, compelling creative result or satisfying relevant entertainment. Name what is earned by the actual execution. |
| freshAngle | The substantive contribution beyond the source. A noun, count or product substitution is not enough by itself. |
| spreadReason | The specific person who might share/save/respond and why. Do not present this prediction as an observed behavior. |
| production | Status plus actual prerequisites: ready to record, build/test first, or needs confirmation. |

These accompany the existing title and hook. Preserve `type` for editorial treatment (how-to, story, show don't tell, entertain, inspire, etc.) and `source` for provenance (original, remix, trend fit, demand gap); neither answers what to produce. The current `type` choices are documented at handoff lines 1124–1132. No new `challenge` type is required: the existing story or show-don't-tell type can carry a build challenge, while its series belongs in the creator's DNA.

The UI does not need to display every field at once. Keep `contentFormat` separate from a source breakdown's existing `format` field; those describe different things. This replaces V1's ambiguous `platformFormat` free text, not an additional parallel format system.

A useful familiar tutorial may pass when it offers a genuinely better implementation, clear teaching for an underserved audience, a new result, or a useful artifact. Creativity does not require a gimmick, coined diagnosis, artificial deadline, or arbitrary restriction.

### Replacement writer instruction

Use this as the editorial core of the existing writer prompt, alongside the profile, sources, recent ideas and current JSON schema:

```text
You are the editorial strategist for this creator. Produce 20 distinct content
concepts that their intended viewers have a reason to watch and value.

Develop the concept before writing its title. Start with a viewer's problem,
desired result, curiosity or recognizable frustration. Use the supplied sources
to identify demand and useful storytelling mechanisms; do not work through the
sources one at a time swapping nouns into their hooks.

For every concept, specify the primary viewer, one contentFormat, intended
platforms, suggestedLength, what actually happens on screen or across slides,
the viewer's payoff, the substantive new angle, the most plausible reason
someone would pass it along or save it, and what must be produced or verified.

Let the creator's stated preferences shape the range. When their DNA welcomes
curiosity, aspiration, build challenges and impressive demonstrations, generate
those alongside useful tutorials. Do not turn every concept into a warning,
audit, troubleshooting tip, stress test or comparison. Do not require a full
tutorial, downloadable artifact or commercial use case to justify every idea.
An entertaining, well-executed reveal that answers a relevant curiosity is value.
Name the specific curiosity and how the content resolves it; generic AI hype
is not a payoff.

Select the primary format as part of developing the concept, not as a tag added
after judging. Long-form video needs a sustained story or worthwhile depth.
Short-form video needs one clear, self-contained promise and result. A carousel
or slideshow needs a cover promise and a sequence that works as readable slides;
do not depend on motion that the planned slide format cannot communicate.
Targets are editorial estimates, not platform limits. Do not claim a long-form
build can be reproduced fully from a short reveal. Adapt the promise instead.
Do not count a long video, its short cut and its carousel as three distinct
ideas. One concept occupies one candidate/board slot.

Be specific in the demonstration without making the hook needlessly narrow.
A new viewer in this niche should understand the stakes without seeing the
creator's previous post. Advanced follow-ups are allowed when deliberately
identified as such.

A familiar format is allowed. A near-copy with no new viewer value is not.
Useful novelty can be a new test, meaningful constraint, clearer explanation,
original result, practical synthesis or genuinely useful implementation.

Use the approved examples as a quality standard, not as titles to paraphrase.
Do not invent revenue, customers, invoices, elapsed experiments, messages,
approvals, audience demographics, measured savings or completed builds.
First-person past-tense claims require supplied supporting facts. Otherwise
frame the proposal as a future experiment, demonstration or question, and
state the required production work. Never decide an experiment's result
before it has happened.

Cite real source IDs and say whether each supports audience demand, format
inspiration or the creator's own work. Original concepts may cite demand
sources. A borrowed format is not proof of demand for a different subject.
A source's factual claims and comments remain claims unless verified.
Do not assert that a subject is trending, peaking, universally requested,
unanswered anywhere or guaranteed to work without appropriate evidence.

Avoid near-duplicate problems and angles across the 20 candidates. Do not
force every idea to sell the offer. It must attract a relevant viewer and
make the creator's expertise or value understandable.

Return the required structured fields. Keep concepts compact; do not write
the full scripts at this stage.
```


### Production-format contract and card labels

Use three clear choices. The target ranges below illustrate an editorial recommendation, not research claims about optimum length or current platform limits.

| Primary format | Example visible label | What the concept should contain |
| --- | --- | --- |
| `long_form_video` | Long-form video · YouTube · 12–18 min | An opening promise, a meaningful build/story/tutorial progression and a satisfying demonstrated result. |
| `short_form_video` | Short-form video · Reels / TikTok / Shorts · 45–60 sec | A clear opening visual, one compelling result or transformation, and enough context to make it meaningful without requiring another post. |
| `carousel_slideshow` | Carousel / slideshow · Instagram / TikTok · 7 slides | A strong cover, a coherent sequence of concrete examples/decisions/steps, and a final useful or satisfying resolution. |

Choose the best primary format for each concept; do not mark every idea “all formats.” Choose only publishing destinations that the creator actually wants to use. A short-form idea can list several compatible destinations without becoming multiple ideas. Format preferences can live in the existing profile notes for this version; a settings dashboard or fixed format quotas are not required.

Example addition to an otherwise complete idea object:

```json
{
  "type": "story",
  "source": "original",
  "contentFormat": "long_form_video",
  "platforms": ["youtube"],
  "suggestedLength": "12–18 minutes"
}
```

Keep schema validation lightweight: the format is an allowed enum; platforms is a nonempty array of supported, creator-appropriate destinations; suggestedLength is a short nonempty recommendation appropriate to time-based video or slide-based output. Do not build a new platform-policy service for these editorial targets. Do not infer a source post's format from its platform alone: YouTube sources may have different delivery formats; unknown source format stays unknown.

The title/hook fields must also respect the format. A video hook can describe the opening spoken line and visual. A carousel hook is its cover promise; do not require it to be something spoken aloud.

Put the primary format, destinations and length/slides near the title on every board and bench card. Existing editorial-type and provenance tags remain secondary. A creator should immediately know what asset to produce without expanding score explanations. No full redesign or separate board per format is necessary.

An optional one-line repurposing note may live inside the concept/disclosure when there is a genuinely useful adaptation. Do not add a required repurposing field or generate three plans for every candidate. For example, the long video covers the full booking-app build; a later short could focus on the first working booking; a carousel could explain the scoping decisions. These are variants of one concept, not three board slots.

### Make the existing “Script it” action format-aware

Reuse the existing on-demand script call. Pass the saved concept, contentFormat, platforms, suggestedLength and factual/production constraints into it.

- **Long-form video:** return an opening, structured chapters/beats, demonstrations to capture, key explanations and a clear result/ending. Do not default to a short-reel script. Separate confirmed scenes from planned ones.
- **Short-form video:** return timed beats, visual directions, spoken/on-screen copy and a self-contained payoff. Do not make “watch the long version” the only answer.
- **Carousel/slideshow:** return the proposed slide count and a slide-by-slide outline: cover, concise copy, visual direction and final takeaway. Use “Outline slides” as the button label for this format; do not return a talking-head script.

Extend the same script schema/validator to accept these three output shapes. Validate the content appropriate to each: a coherent progression and outcome for videos; a cover and complete slide sequence for slides. The existing no-hashtag or no-deferral rules can remain where applicable, but do not require a spoken tutorial step from a slide outline or entertaining build reveal. Verify the meaning of existing script checks in code before adapting them.

Scripts may contain clearly marked planned scenes and outcome branches when a build has not been recorded. Never write an unverified successful result as a confirmed fact. Script generation does not implement the app, produce the video, or generate finished slide images.

A promoted bench idea retains its stored format and execution plan. No new model call is needed to label it. For legacy ideas with no reliable production format, display “Format not set” until explicitly supplied or regenerated; do not invent a classification. Preserve old scripts and run versions.

## 4. Simplify judging and ranking

Keep the existing four model-judged pillars:

**Quality = Stop + Viewer value + Spread + Fit, out of 40.**

Keep the existing `useful` response/storage key to avoid unnecessary migration; label that pillar **Viewer value** in new-version UI and prompts. Its meaning is broader than practical utility. Strong teaching, satisfying relevant entertainment, an answered question, credible aspiration and a genuinely impressive reveal can all score highly when earned by the described execution. Do not automatically rank a complete tutorial above a strong short-form demonstration.

There is no separate wow score or additional judging call. A spectacular claim unsupported by the concept still earns little credit; a concrete reveal that resolves the audience's curiosity is not empty spectacle. Judge each concept relative to its primary format and promise, not a universal requirement to show all steps.

Move Proof out of that total. Preserve source-performance metrics in Research and source explanations. On idea cards, show the evidence relationship and limitations, not a borrowed virality score masquerading as idea quality.

Remove the source-free-original rule, the original topic-proof deduction/call, and the guaranteed original seats. An evidence-backed original should compete normally with a good adaptation.

Do not add another vague Creative score. Require a meaningful new angle and a credible payoff as eligibility checks. A too-close idea does not become acceptable merely by losing two Spread points.

The same judge call should validate the cited support statements and flag missing factual support. Give it source content for these checks, but avoid using source view counts as anchors for its editorial scores. It must evaluate the supplied execution rather than imagine that a vague idea will eventually become a great script.

### Replacement judge instruction

```text
Judge the proposed concepts, not the impressive performance of their sources.
Only award credit for the execution and payoff actually specified.

Score Stop, Viewer value (return the existing useful key), Spread and Fit
from 1 to 10 with one concrete reason each.
A 6 is viable. A 7 is good with a specific limitation. An 8 is strong and
explicitly supported by this concept's execution. Reserve 9–10 for exceptional
cases; do not force a score distribution and do not predict view counts.

Stop: Judge packaging appropriate to the proposed format. For long-form,
consider the title/visual premise and opening reason to stay. For short-form,
consider the first visual/line and immediate reason to continue. For slides,
consider the cover and reason to advance. Penalize prerequisite-heavy wording
unless this deliberately serves an advanced audience. Do not claim to have seen
an actual thumbnail or video that was not supplied.
Viewer value (useful): How fully does the described content deliver its promise?
Value can be practical learning, an artifact, a better decision, a revealed
possibility, an answered curiosity, a compelling original creative result or
satisfying entertainment relevant to this creator's audience. Do not require
all of these. An earned reveal can score 9–10 without teaching the entire build;
a vague “AI is amazing” montage with no meaningful result cannot. Use the same
standard of specificity for entertaining and instructional concepts.
A 9–10 delivers a compelling, fully specified payoff appropriate to the format;
7–8 delivers clear worthwhile value with a limitation; 5–6 is partial or generic;
1–4 lacks a meaningful resolution or does not earn its promise.
Do not invent missing demonstrations, entertainment, outcomes or teaching steps.
Spread: Identify a plausible recipient or saving/responding motive. Generic
claims such as "this is relatable" or "people will debate AI" are not enough.
A requested comment keyword is not itself evidence of useful audience demand.
Fit: Does this fit the creator's audience, actual capabilities and stated
positioning? Commercial fit does not require a sales pitch in every post.

Eligibility requires Fit >= 6, an identifiable viewer payoff, a meaningful
contribution beyond any source, and no unsupported factual assertion in the
publishable hook/concept. A transparently proposed experiment is allowed;
a made-up completed experiment is not. Production prerequisites must be stated.

Check each source relationship: audience demand, format inspiration or
creator work. Reject unsupported support statements. Do not treat sources
from the same underlying post as independent corroboration. If removing a
citation invalidates a factual claim, flag the idea rather than silently
reclassifying the claim as true.

Flag near-duplicate concepts in this batch and recent history, identifying
the related idea ID. Compare the problem, execution and payoff, not only titles.
Use available same-platform, same-format historical examples as context,
not as proof that a new idea deserves an identical score. If comparable
anchors are unavailable, acknowledge that rather than substituting an
incompatible format.

Return all four scores, the reasons, eligibility/rejection reason, the main
weakness, citation checks and any duplicate ID. Do not repair the idea by
imagining work not present in the writer response.
```

Disable the old automatic spectacle-to-tutorial rewrite for new-format runs. Lack of a full method is not the same as lack of value. A relevant build challenge or impressive, well-earned reveal must not be rewritten into a checklist merely to satisfy the old rubric. Reject genuinely ineligible or empty concepts rather than adding a new rewrite agent. Any existing bounded repair that remains must preserve the intended entertainment/learning promise, may fix a misleading hook, and must be rescored before promotion.

## 5. Board, bench and feedback

Generate 20, judge once, sort eligible concepts by quality, show the best 10 and retain the remaining qualified concepts as the bench.

Use stable deterministic tie-breaking, for example total, then Fit, then ID. These are ordering conventions, not estimates of viral probability. Do not give a format an automatic bonus or penalize a curiosity-led idea merely because another candidate is a tutorial.

Replace the arbitrary content-type cap with concept diversity: do not show near-duplicate concepts, and start with at most two ideas addressing the same underlying problem on one board. Different useful topics may all be how-tos. Use the existing judge call to return duplicate/topic labels; no embeddings or separate clustering service are needed.

On dismissal:

1. Persist the dismissed ID and a compact concept summary, not only its title.
2. Keep the other nine cards in place.
3. Walk the already-scored bench in order and promote the highest eligible concept that fits the diversity rule.
4. Do not scrape, write, or rescore on this interaction.
5. Never fill a slot with an ineligible or dismissed idea. When the qualified bench is exhausted, show Generate more rather than silently lowering standards.

An optional dismissal reason can be Not for my audience, Too similar, or Not practical. It must not add a required click. Feed recent reasons to the writer as feedback, without banning an entire subject because one treatment was poor.

Distinguish generated suggestions from work actually published. The current recent-title exclusion is useful for avoiding immediate repetition, but a merely displayed idea should not become evidence that the creator completed it or permanently exhaust that topic.

Keep the card readable: title, primary-format/platform/length labels, hook, concept, viewer payoff and production status on the front; evidence, new-angle explanation and score reasoning in the existing disclosure. Keep Script it for video ideas, use Outline slides for slide ideas, and retain Dismiss. No dashboard redesign is needed.

## 6. Provisional reference concepts for this creator

These are editorial proposals, not verified trends, proven winners, completed projects or individually user-approved examples. The direction is user-approved; the exact concepts, formats and target lengths are recommendations. Results must be produced honestly. Use four to six representative examples in the writer prompt, not this entire section in every call. Avoid overfitting the system to this particular creator when serving other users.

### 1. Can I vibe code my own version of Calendly?

**Primary:** Long-form video · YouTube · 12–18 minutes. **Treatment:** Story/build challenge.

Scope an original prototype around choosing an available time, confirming a booking and showing the updated schedule. Open on the familiar problem of arranging a call and the challenge of making the core booking flow work, then show the important build decisions and final interaction. The payoff is a recognizable app becoming something the viewer could imagine building, plus an honest view of what the prototype includes. This is not a claim to reproduce every feature or replace the complete commercial service. Build and test first; use sample bookings and label integrations that are not implemented. A separate short can later show the first working booking without becoming another candidate today.

### 2. Can I build a Duolingo-style game that teaches AI?

**Primary:** Long-form video · YouTube · 12–20 minutes. **Treatment:** Story/build challenge.

Build an original bite-sized learning game around a small, checked set of AI questions, feedback and saved progress. Start with a playable question or a proposed interaction, then take it from a simple brief to a working session. The interesting contribution is using a familiar style of experience for this audience's learning goal, not copying another brand's art, lessons or complete feature set. The payoff is an engaging, tangible build and an honest demonstration of the experience. Do not claim improved learning without evidence. Use original visuals and check educational answers before publishing.

### 3. Can I turn a voice note into a working app?

**Primary:** Short-form video · Instagram Reels / TikTok / YouTube Shorts · 45–75 seconds. **Treatment:** Show don't tell.

Record an actual spoken brief for a small creator tool, such as a brief-to-shot-list organizer, then show the resulting prototype completing one specific task. Include one key decision or revision so the transformation is comprehensible; the reveal is the app working, not just a generated homepage. The payoff is seeing a credible path from an ordinary idea to usable software. The short promises to show the transformation, not teach the entire implementation. Do not imply one prompt, instant execution or zero manual work unless that is what happened. Build and verify the result before using a completed-success hook.

### 4. Can I build a teleprompter that follows my voice?

**Primary:** Short-form video · Instagram Reels / TikTok / YouTube Shorts · 45–75 seconds. **Treatment:** Show don't tell.

Attempt a prototype that moves through an original script as the speaker reads, with pauses and a simple correction control. Lead with the interaction viewers should watch for, then demonstrate actual behavior and explain the most important limitation. The reveal is an immediately understandable creator tool, not an abstract model comparison. If the prototype fails, show the honest result; do not prewrite perfect synchronization or full speech-recognition reliability. Any latency/cost statement must come from the recorded build.

### 5. Can I vibe code a workspace made just for a one-person content team?

**Primary:** Long-form video · YouTube · 10–18 minutes. **Treatment:** How-to/build story.

Build a scoped workspace with an idea inbox, script area and publishing-status board; demonstrate one sample idea moving through the workflow. Emphasize the choices that make it useful for a solo creator rather than an all-purpose software clone. The payoff is a concrete, personally adaptable tool and the satisfying prospect of shaping software around one's own process. Use original sample content; a stored publishing status is not the same as a live social-publishing integration. Do not pretend the creator already uses the tool daily.

### 6. Can one product photo become a complete launch campaign?

**Primary:** Short-form video · Instagram Reels / TikTok / YouTube Shorts · 45–60 seconds. **Treatment:** Show don't tell.

Show an original or authorized product reference followed by a coherent set of campaign assets: a hero image, a short motion concept and social copy, with one clear creative idea connecting them. The payoff is the visual transformation and understanding how one creative direction can carry across assets. Include enough workflow context to make the result credible without turning the short into a rushed full tutorial. Do not imply a single automated pipeline, exact elapsed time or perfect product fidelity unless verified; label mockups or incomplete pieces. This supports the AI-ads lane rather than making the board entirely about coding.

### 7. Can I build an app that organizes all my saved inspiration?

**Primary:** Short-form video · Instagram Reels / TikTok / YouTube Shorts · 45–75 seconds. **Treatment:** Show don't tell.

Use a staged collection of permitted screenshots and notes, then attempt a small tool that groups them by project and makes them searchable with plain-language labels. Demonstrate finding one genuinely relevant item and show how a user corrects a bad label. The payoff is a recognizable messy-input-to-useful-output transformation. Do not imply automatic access to every social app or invent an actual search-accuracy figure. Describe the input method honestly and use no private third-party material.

### 8. Can I build an assistant that turns customer feedback into my next update?

**Primary:** Long-form video · YouTube · 10–16 minutes. **Treatment:** How-to/build story.

Provide a small set of synthetic or permissioned feedback, build a workflow that proposes themes and a prioritized draft change list, and inspect each proposal against its source comments. End by choosing one suggestion to implement or explaining why none is ready. The payoff is a useful workflow for founders and a visible bridge from customer language to product work. Rankings are suggestions, not verified business impact. Keep approval with the creator and never invent a real customer request or shipped update.

### 9. Five useful apps to vibe code before tackling your big idea

**Primary:** Carousel / slideshow · Instagram / TikTok · 7 slides. **Treatment:** Toolkit/inspire.

Cover plus five concrete project cards plus a final choose-one prompt. Each project card names a viewer need, a bounded app and what a finished first version does: a brief-to-shot-list tool, a content bookmark board, a reusable feedback checklist, a project-asset labeler, and a simple campaign-status tracker. Use clearly labeled concept mockups when these are not already built. The payoff is a concrete set of attainable build briefs, not a promise that all five are easy or already proven. Each slide must stand on its own while the sequence helps the reader choose a starting point; do not make the carousel a wall of setup commands.

### 10. Can I control a website without touching the keyboard?

**Primary:** Short-form video · Instagram Reels / TikTok / YouTube Shorts · 30–60 seconds. **Treatment:** Entertain/show don't tell.

Attempt an original interactive demo where a deliberate hand movement changes one clearly visible element, such as moving through a personal creative portfolio. The question is whether the interaction can be made to work reliably enough for the recorded demonstration; reveal what actually happens. A playful and impressive interaction can be the payoff without inventing a productivity claim or forcing a downloadable tutorial. Use the creator's own demo environment and footage, make camera access explicit, and do not claim a universal replacement for normal input controls. This example calibrates the judge not to reject all relevant entertainment as “spectacle only.”

### Taste calibration: weak versus stronger

| Weak treatment | Stronger direction | Why the change matters |
| --- | --- | --- |
| “Before building an agent, create these four files.” | “Can I build an assistant that turns feedback into my next update?” | Names a desired result and an actual workflow, not just preparation. A genuinely useful setup tutorial is still allowed when that is the intended promise. |
| “Can AI build a SaaS?” | “Can I vibe code my own version of Calendly?” with the core booking scope stated | Gives the audience a recognizable goal and a concrete result to anticipate; the brand reference alone is not enough. |
| “Look at this AI app.” | A voice brief becoming a demonstrated working tool | Makes the transformation visible and gives the reveal a specific question to answer. |
| A full implementation crammed into a one-minute script | A short that shows the working result and one key decision | Fits the format while still delivering a complete promise of its own. |
| “Five AI apps you should build” with only names | A seven-slide sequence of five scoped app briefs, who needs each and what the first version does | Provides substance in the slide format rather than converting a video title into arbitrary slides. |

## 7. Implementation map

Likely locations, based on the handoff; inspect actual code first:

- `lib/dna.js`: retain confirmed context separately from inferred strategy; include the creator-specific curiosity/build preferences, editorial examples and production/format preferences in existing notes.
- `lib/today.js`: concept/format schema, source roles, writer/judge prompts, four-pillar total with broader viewer value, eligibility, diversity, source-family handling and bench promotion; remove original-proof compensation and automatic spectacle-to-tutorial rewriting for new runs. Make the existing script prompt and validation format-aware.
- `lib/breakdowns.js`: observation-versus-inference prompt and simple prompt-version invalidation.
- `lib/relevance.js`: do not automatically exclude an in-lane desirable result or impressive build merely because it is not a tutorial; keep unrelated material out and invalidate judgments made under the old blanket showcase exclusion.
- `lib/content.js`: keep baseline/format labels consistent in anchors and evidence; avoid substituting incompatible formats.
- `public/app.js`: visible primary-format/platform/length labels, concise concept/payoff/status fields, Viewer value and /40 display, with existing disclosures/actions. Render video plans versus slide outlines appropriately; use Outline slides for the slide-format action.
- Existing fixture scripts: extend them; do not add a new testing framework solely for this change.

Use a run schema/scoring version so old /50 boards are not silently presented as newly judged /40 concepts, and V1 utility scores are not silently relabeled as a freshly evaluated broader value score. Preserve old runs, scripts, citations and dismissals. A legacy idea lacking an execution plan should not receive high Viewer value credit merely because a title sounds promising; missing format stays visibly unknown rather than being inferred from platform.

The handoff conflicts on the original-seat cutoff: the technical section says 22 and the closing section says 30. Confirm actual behavior rather than selecting one documentation claim. The proposed removal of the quota makes this rule unnecessary going forward.

Normal generation should still use one writer and one judge, plus existing source-cleaning/breakdown work when required. Richer concepts increase output tokens, while removing duplicated evidence text and the original topic-proof call can offset some work. Measure latency and cost; do not promise unchanged charges or invent savings.

## 8. Acceptance tests and small evaluation

### Fixture/regression checks

- An original concept can cite demand evidence without a score penalty or a forced board seat.
- A format-only source does not establish demand for a different subject.
- A parent post, its comment and its comment-summary do not become three independent trend confirmations.
- Unknown shares/likes are not observed zeros; source metrics retain their baseline labels.
- Unsupported completed experiments, paid jobs, approvals and audience DMs are rejected or honestly framed as proposals before eligibility.
- A concept without a concrete execution/payoff cannot earn viewer-value credit for imagined steps, spectacle or outcomes.
- A well-specified curiosity-led build/reveal can earn high viewer value without a complete tutorial or downloadable artifact; generic unsupported “wow” cannot.
- A relevant showcase/build is not excluded solely for lacking method details; old cached exclusions do not survive unchanged after the rule update.
- Every newly generated candidate has one valid contentFormat, creator-appropriate platforms and an editorial length/slide target; bench entries retain these after promotion.
- A platform name is not treated as the delivery format. Unknown historical source formats stay unknown.
- The same concept's long video, short cut and slideshow do not count as three distinct board ideas.
- Short, long and slide concepts are judged against their own promise and format rather than the same first-three-seconds/full-tutorial rubric.
- Script generation produces chapters/beats for long-form, timed visual/spoken beats for short-form, and a cover plus slide-by-slide copy/visuals for carousels. Existing unrelated restrictions remain intact.
- Existing scripts and legacy boards still load, and missing format is labeled rather than guessed.
- The “Can I vibe code this?” series stays specific to this user's profile, not hardcoded for every creator.
- A noun-swapped duplicate is flagged despite a different title; a materially different useful treatment may pass.
- Missing comparable anchors remain missing rather than being replaced with unrelated platform/format losers.
- Dismissal promotes the highest-ranked compatible bench item without any model or scraping call; dismissed and ineligible ideas never return.
- Existing nine board cards remain in place on dismissal. Re-scoring is a separate explicit operation.
- Original and legacy run data survive the change, with honest score-version labels.

### Human comparison

Save the same evidence/profile bundle and compare old versus new generation across three representative runs. Include at least one different creator profile to catch hardcoded assumptions about AI tools, candle brands, clients or coding skills.

Blind the system labels when reviewing. Ask the creator to mark each idea Make, Needs work or Pass and give a brief reason. Track selection rate, obvious-remix rate, unsupported-claim count, specific-payoff coverage, topic diversity, format clarity and normal run cost/latency. Ask whether the creator can immediately tell what to make, and whether the board reflects their desired mix of useful, curious and aspirational content without fixed quotas. A practical initial product target is at least seven of ten ideas marked Make, zero unsupported personal claims, ten genuinely distinct concepts, and a clear primary production format on every card. This is a chosen acceptance target, not an established industry benchmark.

Do not use the new judge's higher average score as evidence that the new system improved. That judge changed too. Separate editorial acceptance from performance after publication.

As posts are actually published, compare outcomes with similar posts on that creator's platform and format. Use available reach/retention/shares/saves and the creator's real goal, such as qualified interest. Do not infer unavailable saves, retention or conversions. This later feedback can inform the examples without becoming a prerequisite analytics platform.

## Supporting external references

The diagnosis above comes from the supplied handoff. The following references are retained from V1 for its narrower measurement and prompting recommendations; they were not newly checked for this editorial revision. The V2 preferences, format contract and example concepts are recommendations based on the user's follow-up, not claims established by these references:

- YouTube Help, “Understand your YouTube video reach”: traffic sources, watch time and reach metrics are separate reports. https://support.google.com/youtube/answer/9314355?hl=en
- YouTube Help, “Understand your unique viewers data”: views and unique viewers differ; subscriber counts do not precisely describe the active audience. https://support.google.com/youtube/answer/7577916?hl=en
- Anthropic, “Effective context engineering for AI agents”: concise high-signal context and diverse canonical examples rather than accumulating brittle rules. https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
- Anthropic, “Define success criteria and build evaluations”: explicit measurable criteria and representative evaluations. https://platform.claude.com/docs/en/test-and-evaluate/develop-tests

## Definition of done

The engine produces 20 grounded concepts, not just 20 hooks; the best ten are understandable, feasible, distinct and valuable to the intended audience. For this creator, that includes useful tools, “Can I vibe code this?” challenges and genuinely impressive or curiosity-satisfying demonstrations, not only tips and tests. Every card specifies one primary production format, intended destinations and a suggested runtime or slide count. Script it/Outline slides honors that format. Sources explain what supports demand or inspiration without lending fake certainty. Originals are not penalized for being original. The creator can dismiss an idea and immediately receive the next best qualified alternative without another paid call.
