How to block "print" pages from indexing

dreadmichael

I have a fairly large FAQ section and every article has a "print" button. Unfortunately, this is creating a page for every article which is muddying up the index - especially on my own site using Google Custom Search.

Can you recommend a way to block this from happening?

Example Article:

http://www.knottyboy.com/lore/idx.php/11/183/Maintenance-of-Mature-Locks-6-months-/article/How-do-I-get-sand-out-of-my-dreads.html

Example "Print" page:

http://www.knottyboy.com/lore/article.php?id=052&action=print

SEODinosaur

you can block in .robot text, every page that ends in action=print

dreadmichael

That would be great. Do you mind giving me an example?

jennita

Rather than using robots.txt I'd use a noindex,follow tag instead to the page. This code goes into the tag for each print page. And it will ensure that the pages don't get indexed but that the links are followed.

SEODinosaur

Theres more then one way to skin a chicken.

SEODinosaur

Try This.

User-agent: *

Disallow: /*&action=print

SEODinosaur

http://www.seomoz.org/learn-seo/robotstxt

NakulGoyal

I actually remember Lore from a while ago. It's an interesting, easy to use FAQ CMS.

Anyways, I would also recommend implementing Canonical Tags for any possible duplicate content issues. So whether it's the print or the web version, each one of them will contain a canonical tag pointing to the web url of that article in the section of your website.

rel="canonical" href="http://www.knottyboy.com/lore/idx.php/11/183/Maintenance-of-Mature-Locks-6-months-/article/How-do-I-get-sand-out-of-my-dreads.html" />

dreadmichael

Thanks Donnie. Much appreciated!

jennita

True but using robots.txt does not keep them out of the index. Only using "noindex" will do that.

dreadmichael

Ya it is actually really useful. Unfortunately they are out of business now - so I'm hacking it on my own.

I will take your advice. I've shamefully never used rel= canonical before - so now is a good time to start.

NakulGoyal

Yes, it's strongly recommended. It should be fairly simple to populate this tag with the "full" URL of the article based on the article ID. This approach will not only help you get rid of the duplicate content issue, but a canonical tag essentially works like a 301 redirect. So from all search engine perspective you are 301'ing your print pages to the real web urls without redirecting the actual user's who are browsing the print pages if they need to.

Dr-Pete

I have to agree with Jen - Robots.txt isn't great for getting indexed pages out. It's good for prevention, but tends to be unreliable as a cure. META NOINDEX is probably more reliable.

One trick - DON'T nofollow the print links, at least not yet. You need Google to crawl and read the NOINDEX tags. Once the ?print pages are de-indexed, you could nofollow the links, too.

SEODinosaur

Yes, but Rel=Canonical does not block a page it only tells google which page to follow out of two pages.The question was how to block, not how to tell google which link to follow. I believe you gave credit to the wrong answer.

http://en.wikipedia.org/wiki/Canonical_link_element

This is not fair. lol

SEODinosaur

But the spiders still run on the page and read the canonical link, however with the robot text the spiders will not.

SEODinosaur

Although you are correct... there is still more then one way to skin a chicken.

SEODinosaur

Your welcome : )

dreadmichael

You are right Donnie. I've "good answered" you too.

I've gone ahead and updated my robots.txt file. As soon as I am able, I will use no indexon the page, no follow on the links, and rel=canonical.

This is just what I needed, a quick fix until I can make a more permanent solution.

Dr-Pete

Rel-canonical, in practice, does essentially de-index the non-canonical version. Technically, it's not a de-indexation method, but it works that way.

jennita

Josh, please read my and Dr. Pete's comments below. Don't nofollow the links, but do use the meta noindex,follow on the page.

Welcome to the Q&A Forum

Browse the forum for helpful insights and fresh discussions about all things SEO.

How to block "print" pages from indexing

Products

Moz Solutions

Free SEO Tools

Resources

About Moz

Why Moz

Get Involved