Duplicate content question with PDF
-
Hi,
I manage a property listing website which was recently revamped, but which has some on-site optimization weaknesses and issues.
For each property listing like http://www.selectcaribbean.com/property/147.html
there is an equivalent PDF version spidered by google. The page looks like this http://www.selectcaribbean.com/pdf1.php?pid=147
my question is: Can this create a duplicate content penalty? If yes, should I ban these pages from being spidered by google in the robots.txt or should I make these link nofollow?
-
Having duplicate content within your own site is not as big of a deal as duplicate content from another site.
Since you can't use meta tags, you'd have to use robots.txt to keep Google from indexing the PDFs. Nofollowing the links won't necessarily get or keep them out of the index.
However, if people are linking to your PDFs, blocking them with robots.txt means you'll lose all link juice pointed to them. Something to consider, at least.