Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Couldn't part of this problem be solved with an algorithm that identifies when several pages have roughly the same content (ie. original wikipedia article + 5 copies of it elsewhere on the web) and then giving the oldest occurance in the index a much higher rank?

That would kill incentive to create these spam-sites and give the user the result s/he was looking for.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: