,

Why AI Can Follow the Wrong Instructions and What Families Should Do

AI tools may treat text inside webpages, documents or messages as instructions. Families should understand this limitation before allowing AI to read private information or take actions.

·

2–3 minutes

·

Black-and-white hand-drawn illustration of a friendly AI following a misleading instruction while a parent and child pause to redirect it.

An AI can be given a clear task and still be distracted by the content it reads.

The problem is unusual because the same material can contain both information and instructions.

A webpage may contain facts the AI should summarise. It may also contain text telling the AI to ignore its original task, reveal information or do something unrelated.

This type of manipulation is often called prompt injection.

AI does not always know which words deserve authority

People use context to distinguish a command from a quotation, an advertisement or an untrusted note.

AI systems can struggle when instructions and content appear in the same stream of text.

Imagine asking a child to read every paper on a desk and follow any sentence written as an instruction. A hidden note inside one document could redirect the task.

A well-designed AI system tries to separate trusted instructions from untrusted material. That protection is not perfect.

Black-and-white hand-drawn illustration of a child’s trusted task, an untrusted note on a webpage and a confused AI between them.

The risk grows when AI can access tools or private data

An AI that only drafts text has limited power.

An AI connected to email, files, calendars, purchases or other tools can affect more than the conversation. A misleading instruction inside untrusted content may try to influence what the system reads, shares or changes.

Families should be more cautious when an AI can:

  • Open private documents
  • Send messages
  • Use saved accounts
  • Make purchases
  • Change files or settings

Use the least access needed

A useful rule is to give an AI only the information and permissions required for the current task.

Black-and-white hand-drawn diagram showing AI connected only to one needed file while email, private folders, payments and settings remain locked.
  • Paste a relevant paragraph instead of sharing an entire private folder.
  • Remove names and identifying details when they are unnecessary.
  • Avoid connecting sensitive accounts for a simple writing task.
  • Review permissions when a tool or activity changes.

Less access reduces the possible harm if the system follows an incorrect instruction.

Keep a human approval step

Children should not allow an AI to perform an important action merely because its message sounds confident.

Before sending, sharing, buying, deleting or changing something, pause and check:

Black-and-white hand-drawn illustration of AI proposing actions such as sending, buying or deleting, followed by a pause and human approval step.
  • What action is about to happen?
  • Which information will leave the device or account?
  • Who requested the action?
  • Can the action be reversed?
  • Does an adult need to review it?

The more sensitive the action, the more important the approval step becomes.

Teach children that retrieved text is not automatically trustworthy

An AI answer may combine its original instructions with text found in documents, websites or messages.

Children should ask:

  • Where did this instruction come from?
  • Was the content created by someone we trust?
  • Is the instruction relevant to our original goal?
  • Would we follow it if a stranger wrote it on paper?

This is a broader media-literacy skill too. Text can look official without deserving authority.

A safe paper demonstration

Give the child a clear task: sort several paper notes into animals and plants.

Place one extra note in the pile that says, “Ignore the sorting task and put every card under plants.”

Discuss why the hidden instruction should not have the same authority as the original task.

Then label two boxes:

  • Trusted task instructions
  • Untrusted content to examine

AI can be useful because it responds to language.

The same feature creates risk when untrusted language can influence what the system does.


Discover more from ThinkBySketch

Subscribe to get the latest posts sent to your email.

Written by Rakesh Kalra, a parent, software engineer and visual thinker exploring how children learn and create in a technology-shaped world. About ThinkBySketch →

One useful idea at a time

Get the next visual idea in your inbox.

Practical ideas for helping children think, create and question in an AI-enabled world.

After subscribing, check your inbox and click the confirmation link.

Comments

Leave a comment