Creator OS
apiautomationinstagrammoderationspam

Hide Spam Comments Automatically on Instagram: The Complete Setup

Learn how to hide spam comments automatically on Instagram using keyword filters, the Graph API, and AI moderation so your feed stays clean without manual work.

Creator OS · October 7, 2026 · 10 min read

You can hide spam comments automatically on Instagram in two places: the app’s Hidden Words list, and the Graph API’s comment endpoints. The app handles the easy cases. The API handles the rest, and it lets an AI agent decide what counts as spam. This guide walks through both, then shows the exact tool calls that hide a comment, delete it, or turn it into a DM.

  comment arrives
        |
        v
  [ native hidden words ]  -- match --> hidden automatically
        |
        v (no match)
  [ webhook: comment.received ]
        |
        v
  [ agent classifies with MCP tool ]
        |
   spam? ---- yes ----> inbox:hide-comment
     |                  (or inbox:delete-comment)
     no
     |
     v
   inbox:reply  or  private-reply (comment to DM)
Two layers: Instagram’s own filters first, then an agent that reads the leftovers.

Why Manual Comment Moderation Does Not Scale

Manual moderation has a hidden cost: it is not the time you spend hiding one comment. It is the interruption. You open the app to check notifications, you see a spam comment, you hide it, you see three more, you check the post’s engagement, and twenty minutes are gone. Do that five times a day across a handful of accounts and you have lost a working afternoon.

The failure mode is worse at scale. If you run five accounts, or one account with a busy Reel, a wave of “DM me to grow your page” comments can land faster than you can tap through them. The comments stay visible to everyone who reads the post, which is the actual damage. Moderation is not a cleanup task you get to later. It is part of the post’s first hour.

How Instagram’s Built-In Spam Filters Actually Work

Instagram already runs two systems on your comments. The first is a general spam detector that flags and hides comments on its own. The second is your Hidden Words list: a set of words and phrases you control. If a comment contains any of them, Instagram hides it from everyone but the author. The author can still see their own comment, which is deliberate. It reduces the chance they notice and repost.

There are two lists under the same setting. Hidden Words is manual: you type the terms. The “Offensive Words and Phrases” toggle is the other one, and it covers a broad set Instagram maintains. Both apply to comments and to DMs on many accounts.

Important detail: hidden is not deleted. The comment still exists, still counts toward your comment total on some surfaces, and you can unhide it later. Deletion is permanent. For a borderline case, hide is the safer move.

Setting Up Keyword Filters and Hidden Words in the App

In the Instagram app, go to your profile, then Settings and privacy, then Comment controls (the exact path moves between app versions, so search the settings for “Hidden Words” if you cannot find it). Under Hidden Words you get a text field. Add one term per line unless you want to use the comma form Instagram supports.

What actually catches spam:

  • Money and growth bait: DM me, check my page, grow your account, link in bio when you are not running a link-in-bio campaign.
  • Platform names used as bait: telegram, whatsapp, +1, wa.me.
  • Emoji-only patterns. Many spam comments are three rocket or fire emoji plus a handle. You can add the specific emoji strings you keep seeing.
  • Your own brand name misspelled, which is a common impersonation tactic.

Do not paste a 400-word list you found online. Broad terms like “follow” or “money” will hide legitimate comments from real customers, and you will never notice because you cannot see hidden comments in the normal feed. Start with 15 to 25 terms, review monthly.

If your spam problem is mostly replies and DMs rather than top-level comments, the workflow in automating Instagram and TikTok replies through the API covers the reply side.

Limits of the Native Tools: What They Miss

The native filters are literal string matchers. They are fast, they are free, and they handle the dumbest 70 percent of spam. Here is what they do not do:

  • They cannot reason. “Great post, I love this creator’s work, by the way I sell followers, click my profile” passes any keyword filter you would want to keep. Your real fans write compliments too.
  • They cannot escalate. Nothing happens beyond hiding. Hidden Words cannot send a DM, cannot log the comment for review, cannot reply on your behalf.
  • They cannot act across accounts. Each account has its own list. Five accounts means five lists to maintain by hand.
  • They cannot be triggered by a webhook. Nothing in the app tells your other systems that a comment arrived.

That gap is where the API and an agent fit. Creator OS ships a REST API, a CLI and a hosted MCP server, so an AI app can read comments and act on them. The docs live at https://www.creatoros.ca/docs.

Using the Instagram Graph API to Moderate Comments

Instagram’s comment endpoints let you list comments on a post, reply to one, hide or unhide it, and delete it. What they do not give you is a decision. The API executes; you supply the judgment. That is the whole reason an agent layer is useful here.

Creator OS wraps this. Comment replies and hiding are supported on Instagram, Facebook, X, Threads, YouTube and LinkedIn. TikTok is the exception: replying to TikTok comments is not supported, though hide and delete work for TikTok for Business accounts. Your connected Instagram account needs to be a Business or Creator account.

The pieces you will use:

  • Webhooks. Subscribe to comment.received and every new comment arrives at your endpoint, signed with X-CreatorOS-Signature so you can verify it came from us.
  • Read tools. list_comments and list_conversations give an agent the queue. get_post and get_post_analytics give it context on which posts are worth watching closely.
  • Act tools. reply_to_comment, hide_comment and create_automation.

Every ID you receive is an opaque typed token: cmt_ for comments, post_ for posts, acc_ for accounts, conv_ for conversations. You never construct one. Errors come back in a single shape, { error: { code, message, status } }, which makes retry logic boring to write. An API key is pinned to one workspace, so an agent holding your key cannot touch another brand’s accounts.

Building an AI Spam Classifier for Your Comments

You do not need training data or a model of your own. You need a short rubric and an agent that can call tools. Creator OS works with any MCP-capable AI app and does not ship its own model; current Claude models in the app are Claude Opus 5.5, Claude Sonnet 5, Claude Haiku 4.5 and Claude Fable 5.1. The setup is nothing to install: the hosted MCP server is at https://mcp.creatoros.ca/mcp. In Claude you go to Settings, Connectors, Add custom connector, paste the URL, sign in, pick the workspace, and choose read only or read & write. A read-only connection only sees read tools, and the API refuses writes from it. Destructive tools like delete comment are marked so the app asks you first. The full setup is at https://www.creatoros.ca/docs/mcp.

Here is a rubric prompt you can paste once the connector is live:

You moderate Instagram comments for my account.

Call list_comments on posts published in the last 24 hours.
For each comment, classify it as SPAM, SUSPECT or REAL.

SPAM means: unrelated sales pitch, follower or engagement
selling, crypto or investment bait, requests to move the
conversation to Telegram or WhatsApp, or an impersonation of
my brand with a link. Hide it with hide_comment.

SUSPECT means: generic compliment followed by a pitch, or a
link I have never approved. Do not act. List it for me.

REAL means: a question about my product, feedback, or a
genuine compliment. Reply with reply_to_comment in my voice,
one or two sentences, no emoji.

Never delete anything. Hide only. Report counts at the end.

Two things make this work reliably. First, “never delete, hide only” is a hard rule, and hiding is reversible. Second, the agent reads before it writes, so you see its classification quality on a sample before it acts on a live Reel. Run it in read-only mode for a day if you want to check the rubric first: the same connector with read-only permissions will list comments but the API will refuse the hide.

If you would rather not build it yourself, the Social Agents repo is an open-source agent harness that runs your socials on Creator OS, including comment and DM replies, with a local dashboard: github.com/kevinbadi/social-agents. There is a walkthrough on the KevBuildsApps YouTube channel and in the launch video at youtube.com/watch?v=-QBJH_PK3pY.

If your comments also contain questions that deserve a public answer, pair this with the approach in replying to comments with AI. Moderation and reply are the same pass over the same list.

Hiding and Deleting Comments Programmatically Step by Step

Here is the same logic from the command line, which is easier to test before you hand it to an agent. Install the CLI and check auth first.

npx @creatoros/cli@latest init
creatoros auth:check
creatoros accounts:list

Pull the comment queue for the platforms you care about:

creatoros inbox:comments --platform instagram --since 2025-01-14 --limit 50
creatoros inbox:post-comments post_abc123 --accountId acc_xyz

Hide the obvious ones. Notice that hiding and unhiding are the same command with a flag, which is exactly what you want for a reversible action:

creatoros inbox:hide-comment post_abc123 cmt_9981 --accountId acc_xyz
creatoros inbox:hide-comment post_abc123 cmt_9981 --accountId acc_xyz --unhide

Delete only when you are sure, because there is no undo:

creatoros inbox:delete-comment post_abc123 cmt_9981 --accountId acc_xyz

Reply in public when a real person asked a real question, or send a private reply to turn one comment into a DM thread. Instagram and Facebook both support keyword comment-to-DM funnels:

creatoros inbox:reply post_abc123 --accountId acc_xyz \
  --commentId cmt_9982 --message "Thanks, sending you the details now."

creatoros inbox:private-reply post_abc123 cmt_9982 \
  --accountId acc_xyz --message "Here is the link you asked for."

For the “DM me the word GUIDE” pattern, hand it to an automation instead of running it by hand. It matches a keyword and fires the DM:

creatoros automations:create \
  --accountId acc_xyz --postId post_abc123 \
  --keywords "guide,link,price" --matchMode contains \
  --dmMessage "Here is the guide." \
  --commentReply "Sent you a DM."

Then check what fired: creatoros automations:logs guide-dm. If you want to see the pipeline end to end before trusting it, creatoros automations:list --cloud shows the automations running in your workspace.

A note on the API path if you are building your own service rather than using the CLI. You would subscribe to comment.received webhooks through creatoros webhooks:create --url https://your-app.example/hooks --events comment.received,post.published, verify the X-CreatorOS-Signature header, and then act. You can test the endpoint without waiting for a real comment with creatoros webhooks:test <id>. The base URL and full request shapes are in the docs at https://www.creatoros.ca/docs; there is no separate API host to memorize.

Connecting Moderation to Your Broader Reply Workflow

Moderation by itself leaves value on the table, because a spam comment and a sales question look identical to a keyword filter. Once an agent is reading every comment, it can do three things in one pass: hide the spam, reply to the questions, and route the high-intent comments into DMs. That is the difference between cleaning a post and working it.

Two extensions worth wiring in:

  • Comment-to-DM funnels. Covered on Instagram and Facebook. A comment containing your keyword gets a private reply with the asset, and the public thread stays clean. See the AI reply workflow.
  • Skills. npx @creatoros/cli@latest init installs 14 Claude Code skills, including respond-to-comments, respond-to-dms, automations and analytics. Each is a packaged workflow you can run instead of writing the prompt from scratch. creatoros sync updates them later without overwriting files you have edited, since updates land as .new files.

If you also schedule Stories, remember that Stories have their own comment surface and their own rules: no collaborators, no location, no paid partnership label, and interactive stickers stay app-only. See scheduling Instagram Stories through the API for the specifics.

Get Started With Creator OS

The Creator plan is $19.99/month or $59.99/year and covers up to 8 connected accounts: 7 socials plus one Skool community. API keys, the MCP server, the CLI and the agent skills are included in every plan, so moderation through an agent costs nothing extra. Ads are a separate add-on at $19.99/month with no percentage of ad spend.

Practical first hour: add your 20 worst spam terms to Instagram’s Hidden Words, then connect the MCP server at https://mcp.creatoros.ca/mcp in read-only mode and run the rubric prompt above on yesterday’s comments. You will see exactly which spam your native filters miss. Then flip the connection to read & write and let it hide.

Sign up at https://www.creatoros.ca/sign-up. If you want to watch the whole pipeline built live first, there is a walkthrough on the KevBuildsApps YouTube channel, and the open-source video editing skills are at github.com/kevinbadi/open-edits.

Keep reading